Using Gemini 3.5 Flash on Vercel AI Gateway
Implementation guide · Gemini 3.5 · Google DeepMind
Serverless
Vercel AI Gateway exposes Gemini 3.5 Flash through model ID google/gemini-3.5-flash. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-06-29. Next refresh: weekly.
Quick Start
- 1
- 2Use the Vercel AI Gateway SDK or REST API to call
google/gemini-3.5-flash— see the documentation for request format. - 3
Code Examples
Install
pip install openaiAPI key
AI_GATEWAY_API_KEYModel ID
google/gemini-3.5-flashcreator/model-name e.g. kwaipilot/kat-coder-pro-v2
import os
from openai import OpenAI
client = OpenAI(
api_key=os.environ["AI_GATEWAY_API_KEY"],
base_url="https://ai-gateway.vercel.sh/v1"
)
response = client.chat.completions.create(
model="google/gemini-3.5-flash",
messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)Pricing on Vercel AI Gateway
| Type | Price (per 1M) |
|---|---|
| Input tokens | $1.50 |
| Output tokens | $9.00 |
| Query | $14.00 |
Capabilities
VisionMultimodalReasoningJSON / Tool useStructured OutputsCode ExecutionPrompt CachingBatch APIAudio
About Gemini 3.5 Flash
Gemini 3.5 Flash is Google DeepMind's generally available Flash model for sustained frontier-level performance on agentic and coding tasks. It supports multimodal inputs, native thinking, tool and function calling, structured outputs, code execution, search grounding, batch processing, and long contexts up to 1M tokens.
Model Specs
Released2026-05-19
Context1.05m
ArchitectureDecoder Only
Knowledge cutoff2025-01