Mixtral Models by MistralAI
Last refreshed 2026-06-01. Next refresh: weekly.
Details
Capabilities
About
The Mixtral family of large language models (LLMs), developed by Mistral AI, offers a groundbreaking approach in open-source AI through a sparse mixture-of-experts (SMoE) architecture. This innovative design allows the models to manage a significant number of parameters while ensuring efficient inference speed by activating only a subset of parameters for each token. Such architecture enables Mixtral models to deliver performance on par with much larger models, standing out in various benchmarks and outperforming competitors like Llama 2, and even equaling the prowess of closed-source models such as GPT-3.5. These models are multilingual, supporting languages such as English, French, Italian, German, and Spanish, and excel in domains like code generation. Instruction-tuned versions like Mixtral-8x7B-Instruct-v0.1 cater to applications requiring robust instruction-following and chat capabilities. The Mixtral family provides versatile models of differing sizes, addressing diverse computational and application requirements.
Decision facts
- Best fit
- JSON / Tool usecoding
- Capability starting point
- Mixtral 8x22B Instruct v0.3 with 64k context and JSON / Tool use
- Lowest tracked input
- Mixtral 8x7B · $0.15/1M · Mistral AI Studio
- Closest related family
- Agents-A1
Current Variants
Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.
Use when the workload needs 64k context and JSON / Tool use.
Use when the workload needs 33k context and 56B parameters.
| Model | Use when | Released | Signals | Status |
|---|---|---|---|---|
| Mixtral 8x22B Instruct v0.3 | Use when the workload needs 64k context and JSON / Tool use. | 2024-07 | 64k contextJSON / Tool use | Current |
| Mixtral 8x22B Instruct v0.1 | Use when the workload needs 64k context. | 2024-04 | 64k context | Current |
| Mixtral 8x22B v0.1 | Use when the workload needs 64k context. | 2024-04 | 64k context | Current |
| Mixtral 8x7B | Use when the workload needs 32k context. | 2023-12 | 32k context | Current |
| Mixtral 8x7B Instruct v0.1 | Use when the workload needs 33k context and 56B parameters. | 2023-12 | 33k context56B parameters | Current |
Release Timeline
3 release groupsSpecifications(5 models)
| Model | Released | Context | Parameters | JSON / Tool use |
|---|---|---|---|---|
| Mixtral 8x22B Instruct v0.3 | 2024-07 | 64k | 8x22B | Yes |
| Mixtral 8x22B Instruct v0.1 | 2024-04 | 64k | 8x22B | No |
| Mixtral 8x22B v0.1 | 2024-04 | 64k | 8x22B | No |
| Mixtral 8x7B | 2023-12 | 32k | 8x7B | No |
| Mixtral 8x7B Instruct v0.1 | 2023-12 | 33k | 56B | No |
Available From(21 providers)
Pricing
Popular comparisons in this family
- Mistral Large 2 vs Mixtral 8x7B392
- Llama 3 70B Instruct vs Mixtral 8x7B206
- Llama 3.1 70B Instruct vs Mixtral 8x7B100
- Mixtral 8x7B vs Qwen3.6-35B-A3B89
- Llama 3 8B Instruct vs Mixtral 8x7B68
- Llama 2 13B Chat vs Mixtral 8x7B53
- Mixtral 8x7B vs Qwen3.6-27B53
- DeepSeek V3 vs Mixtral 8x7B46
- Mixtral 8x7B vs Phi-3 Mini 4k45
- Mistral Nemotron vs Mixtral 8x7B39


