Gemini 1.5 Flash on Google Vertex AI (Extended Context) on GCP Vertex AI
Gemini 1.5 · Google DeepMind
Serverless
Last refreshed 2026-06-15. Next refresh: weekly.
Why use Gemini 1.5 Flash on Google Vertex AI (Extended Context) on GCP Vertex AI?
GCP Vertex AI offers Gemini 1.5 Flash on Google Vertex AI (Extended Context) with pay-as-you-go pricing at $0.07/1M input tokens. Vertex AI is Google Cloud's managed AI platform, offering access to Gemini models and hundreds of partner models alongside tools for fine-tuning, grounding, vector search, and end-to-end MLOps pipelines.
Input / 1M
$0.070
Output / 1M
$0.21
Cache
Not sourced
Batch
Not sourced
Setup recipe
Python + curlInstall
pip install google-cloud-aiplatformAuth
export GOOGLE_CLOUD_PROJECT=...Call
import os
import vertexai
from vertexai.generative_models import GenerativeModel
vertexai.init(project=os.environ["GOOGLE_CLOUD_PROJECT"], location="us-central1")Model ID
vertex-gemini-1.5-flash-extendedRequest example
import os
import vertexai
from vertexai.generative_models import GenerativeModel
# Reads GOOGLE_CLOUD_PROJECT from env; authenticates via Application Default Credentials
vertexai.init(project=os.environ["GOOGLE_CLOUD_PROJECT"], location="us-central1")
model = GenerativeModel("vertex-gemini-1.5-flash-extended")
response = model.generate_content("Hello")
print(response.text)Gotchas
- For Google-published models use the model name directly, e.g. "gemini-2.0-flash-001". For third-party publishers (Anthropic, Meta, etc.) use the full publisher path, e.g. "publishers/anthropic/models/claude-3-5-sonnet-v2@20241022".
- The examples expect GOOGLE_CLOUD_PROJECT; rename it only if your application config maps the new variable.
Pricing
| Type | Price (per 1M) |
|---|---|
| Input tokens | $0.07 |
| Output tokens | $0.21 |
Capabilities
VisionMultimodalStructured Outputs
About Gemini 1.5 Flash on Google Vertex AI (Extended Context)
Gemini 1.5 Flash on Google Vertex AI (Extended Context) is Google DeepMind's Gemini 1.5 model with multimodal text and image input. It offers a 1M-token context window.
Model Specs
Released2024-02-15
Context1m
ArchitectureDecoder Only
Knowledge cutoff2024-05