Gemini 1.0 Pro Vision
About
Gemini 1.0 Pro Vision is a multimodal large language model crafted by Google, excelling in tasks involving both visual and textual data. It boasts advanced capabilities in visual understanding, classification, and summarization, enabling the creation of content from images and videos. The model adeptly processes a range of visual and textual inputs, such as photographs, documents, and infographics, and is capable of generating image descriptions and object identification. Moreover, it supports zero-shot, one-shot, and few-shot learning, enhancing its adaptability to diverse applications. Despite its powerful features, Gemini 1.0 Pro Vision is slated for deprecation, with a removal date set for April 9, 2025, prompting users to transition to updated models like Gemini 1.5 Pro and Gemini 1.5 Flash 15.
Capabilities
Providers(1)
| Provider | Input (per 1M) | Output (per 1M) | Type | |
|---|---|---|---|---|
| GCP Vertex AI | $0.5 | $1.5 | Serverless |