Last refreshed 2026-06-15. Next refresh: weekly.
Why use Qwen3.5-397B-A17B on Together AI?
Together AI offers Qwen3.5-397B-A17B with pay-as-you-go pricing at $0.60/1M input tokens. Together AI is a platform for running open-source and proprietary LLMs with fast serverless and dedicated endpoints at competitive inference pricing.
Compare Qwen3.5-397B-A17B across 4 providers to find the best fit for your use caseInput / 1M
$0.60
Output / 1M
$3.60
Cache
Not sourced
Batch
Not sourced
Setup recipe
Python + curlInstall
pip install togetherAuth
export TOGETHER_API_KEY=...Call
from together import Together
client = Together() # reads TOGETHER_API_KEY from env
response = client.chat.completions.create(
model="Qwen/Qwen3.5-397B-A17B",Model ID
Qwen/Qwen3.5-397B-A17BRequest example
from together import Together
client = Together() # reads TOGETHER_API_KEY from env
response = client.chat.completions.create(
model="Qwen/Qwen3.5-397B-A17B",
messages=[{"role": "user", "content": "Hello"}]
)
print(response.choices[0].message.content)Gotchas
- Use provider model ID "Qwen/Qwen3.5-397B-A17B", not the LLMReference slug "qwen3.5-397b-a17b".
- Together uses "organization/model-name" format, e.g. "meta-llama/Llama-4-Scout-17B-16E-Instruct" or "Qwen/QwQ-32B". See the Together model catalog for the exact ID.
- The examples expect TOGETHER_API_KEY; rename it only if your application config maps the new variable.
Compare Qwen3.5-397B-A17B Across Providers
| Provider | Input (per 1M) | Output (per 1M) |
|---|---|---|
| OpenRouter | $0.39 | $2.34 |
| Together AI | $0.60 | $3.60 |
| Alibaba Cloud PAI-EAS | $0.39 | $2.34 |
| Novita AI | $0.60 | $3.60 |
Pricing
| Type | Price (per 1M) |
|---|---|
| Input tokens | $0.60 |
| Output tokens | $3.60 |
Capabilities
VisionMultimodalReasoningJSON / Tool useStructured Outputs
About Qwen3.5-397B-A17B
Alibaba's largest Qwen3.5 model, featuring a Mixture-of-Experts architecture with 397B total parameters and 17B active per token (using 512 total experts with 10 routed + 1 shared active). Supports 201 languages with a native 262K token context window extensible to 1M tokens via YaRN. Includes a thinking/reasoning mode, tool calling with MCP integration, and unified vision-language capabilities through early fusion training.
Get Started
Model Specs
Released2026-02-16
Parameters397B
Context262k
ArchitectureMixture of Experts