Using Vertex AI Multimodal Embeddings on GCP Vertex AI
Implementation guide · Vertex AI Multimodal Embeddings · Google DeepMind
GCP Vertex AI exposes Vertex AI Multimodal Embeddings through model ID multimodalembedding. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-07-01. Next refresh: weekly.
Quick Start
- 1
- 2Use the GCP Vertex AI SDK or REST API to call
multimodalembedding— see the documentation for request format. - 3
Code Examples
pip install google-cloud-aiplatformGOOGLE_CLOUD_PROJECTmultimodalembeddingFor Google-published models use the model name directly, e.g. "gemini-2.0-flash-001". For third-party publishers (Anthropic, Meta, etc.) use the full publisher path, e.g. "publishers/anthropic/models/claude-3-5-sonnet-v2@20241022".
import os
import vertexai
from vertexai.generative_models import GenerativeModel
# Reads GOOGLE_CLOUD_PROJECT from env; authenticates via Application Default Credentials
vertexai.init(project=os.environ["GOOGLE_CLOUD_PROJECT"], location="us-central1")
model = GenerativeModel("multimodalembedding")
response = model.generate_content("Hello")
print(response.text)Pricing on GCP Vertex AI
| Type | Price (per 1M) |
|---|---|
| Input tokens | Free |
| Output tokens | Free |
Capabilities
About Vertex AI Multimodal Embeddings
Vertex AI Multimodal Embeddings is Google Cloud's foundation embedding model for image, text, and video inputs. It supports cross-modal retrieval and semantic search through the Vertex AI Multimodal Embeddings API.