Using Vicuna 13B 16K on GCP Vertex AI
Implementation guide · Vicuna · LMSYS Org
GCP Vertex AI exposes Vicuna 13B 16K through model ID vicuna-13b-16k. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.
Last refreshed 2026-06-15. Next refresh: weekly.
Quick Start
- 1
- 2Use the GCP Vertex AI SDK or REST API to call
vicuna-13b-16k— see the documentation for request format.
Code Examples
pip install google-cloud-aiplatformGOOGLE_CLOUD_PROJECTvicuna-13b-16kFor Google-published models use the model name directly, e.g. "gemini-2.0-flash-001". For third-party publishers (Anthropic, Meta, etc.) use the full publisher path, e.g. "publishers/anthropic/models/claude-3-5-sonnet-v2@20241022".
import os
import vertexai
from vertexai.generative_models import GenerativeModel
# Reads GOOGLE_CLOUD_PROJECT from env; authenticates via Application Default Credentials
vertexai.init(project=os.environ["GOOGLE_CLOUD_PROJECT"], location="us-central1")
model = GenerativeModel("vicuna-13b-16k")
response = model.generate_content("Hello")
print(response.text)Pricing on GCP Vertex AI
Capabilities
About Vicuna 13B 16K
The Vicuna 13B v1.5 16K model by LMSYS is an advanced conversational AI built on the transformer architecture, fine-tuned from Llama 2. It features 13 billion parameters and handles up to 16,000 tokens in context, making it suitable for diverse tasks like chatbots and content generation. Trained on 125,000 conversations from ShareGPT, it excels in generating coherent text and engaging dialogue but may struggle with domain-specific knowledge and nuanced context. Despite potential biases, it represents a significant step forward in conversational AI.