Using Qwen3-Coder-480B-A35B-Instruct on GCP Vertex AI

Implementation guide · Qwen3-Coder · Alibaba

ServerlessOpen Source

GCP Vertex AI exposes Qwen3-Coder-480B-A35B-Instruct through model ID qwen3-coder-480b-a35b-instruct. Use the setup steps, sourced pricing, capabilities, and official provider links below to validate this route before deployment.

Last refreshed 2026-06-19. Next refresh: weekly.

Quick Start

  1. 1
    Create an account at GCP Vertex AI and generate an API key.
  2. 2
    Use the GCP Vertex AI SDK or REST API to call qwen3-coder-480b-a35b-instruct — see the documentation for request format.
  3. 3
    You'll be billed $0.22/1M input, $1.80/1M output tokens. See full pricing.

Code Examples

Install
pip install google-cloud-aiplatform
API key
GOOGLE_CLOUD_PROJECT
Model ID
qwen3-coder-480b-a35b-instruct

For Google-published models use the model name directly, e.g. "gemini-2.0-flash-001". For third-party publishers (Anthropic, Meta, etc.) use the full publisher path, e.g. "publishers/anthropic/models/claude-3-5-sonnet-v2@20241022".

import os
import vertexai
from vertexai.generative_models import GenerativeModel

# Reads GOOGLE_CLOUD_PROJECT from env; authenticates via Application Default Credentials
vertexai.init(project=os.environ["GOOGLE_CLOUD_PROJECT"], location="us-central1")
model = GenerativeModel("qwen3-coder-480b-a35b-instruct")
response = model.generate_content("Hello")
print(response.text)

Pricing on GCP Vertex AI

TypePrice (per 1M)
Input tokens$0.22
Output tokens$1.80

Capabilities

JSON / Tool useStructured OutputsCode Execution

About Qwen3-Coder-480B-A35B-Instruct

Qwen3-Coder-480B-A35B-Instruct is Alibaba's flagship open-source code generation and agentic model, released July 22, 2025 under the Apache 2.0 license. The model has 480 billion total parameters with 35 billion active parameters per token, organized across 62 transformer layers with 160 specialized expert networks and 8 experts activated per token. It uses Grouped Query Attention (GQA) with 96 query heads and 8 key-value heads and supports a native context window of 262,144 tokens, extendable to 1 million tokens via YaRN position scaling.

Model Specs

Released2025-07-22
Parameters480B total, 35B active
Context262k
ArchitectureMixture of Experts

Provider

GCP Vertex AI

Google Cloud Platform (GCP)

Mountain View, California, United States