GCP Vertex AI exposes Vicuna 7B through model ID vicuna-7b. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-06-15. Next refresh: weekly.
Quick Start
- 1
- 2
Code Examples
pip install google-cloud-aiplatformGOOGLE_CLOUD_PROJECTvicuna-7bFor Google-published models use the model name directly, e.g. "gemini-2.0-flash-001". For third-party publishers (Anthropic, Meta, etc.) use the full publisher path, e.g. "publishers/anthropic/models/claude-3-5-sonnet-v2@20241022".
import os
import vertexai
from vertexai.generative_models import GenerativeModel
# Reads GOOGLE_CLOUD_PROJECT from env; authenticates via Application Default Credentials
vertexai.init(project=os.environ["GOOGLE_CLOUD_PROJECT"], location="us-central1")
model = GenerativeModel("vicuna-7b")
response = model.generate_content("Hello")
print(response.text)Pricing on GCP Vertex AI
Capabilities
About Vicuna 7B
Vicuna-7B is an open-source language model crafted by LMSYS, fine-tuning the LLaMA model using around 125,000 user conversations from ShareGPT. It's designed for natural and fluent dialogues, effectively addressing a wide array of queries and generating text on diverse subjects. However, while it performs well, it may sometimes produce incorrect or biased responses due to its training limitations. Aimed primarily at research, it comes in various versions and quantizations to cater to different computational needs. Although helpful and polite, its performance is slightly lower compared to larger models like Vicuna-13B or Vicuna-33B 125.