LLM Reference

Mixtral 8x22B v0.1

Released
2024-04-17
Last refreshed
2026-07-11
Status
Researched 153d ago
Open sourceCommercial use: permittedCodingClassification

Mixtral 8x22B v0.1 is worth evaluating for coding and classification when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding and classification
  • Workloads that can use a 64k context window
  • Buyers comparing 4 tracked provider routes

Do not use it for

  • Vision or document-understanding workloads
  • Strict JSON or tool-calling flows
Specifications
Family
Mixtral
Released
2024-04-17
Context
64k
Parameters
8x22B
Architecture
Mixture of Experts
Knowledge cutoff
2024-01
Specialization
general
Openness
Open source
License
Apache 2.0OSI-approvedCommercial use: permitted
Weights
Unknown
Code
Unknown
Training
Fine-tuned
Created by

Enterprise AI solutions for trust and transparency.

Paris, France
Founded 2023
Website
Pricing
Output / 1M
$0.650
Input / 1M
$0.650

Cheapest of 8 routes · DeepInfra

About

The Mixtral 8x22B v0.1 is a pretrained generative Sparse Mixture of Experts (MoE) model created by Mistral AI [1][2][4]. It utilizes a specialized architecture where different sub-models, termed "experts," manage distinct input segments, enhancing both efficiency and performance relative to traditional large language models [2][10][12]. This model features an impressive 176 billion parameters and supports a context length of 65,000 tokens [10][13]. It excels in text generation, completion, and question answering, outperforming models like LLaMA 2 70B on various benchmarks [4][5][7]. Nonetheless, as a base model, it lacks inherent moderation capabilities, potentially generating inappropriate or harmful content without filtration [2][4][10].

Top use-case fit: coding, agents, and build tasks

Coding

Q/$ B

1 relevant benchmark in the decision map.

Classification

Q/$ C

2 relevant benchmarks in the decision map.

Provider price ladder

Compare all 8

Compare API pricing across 4 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
DeepInfra$0.650$0.650
Serverless
Fireworks AI$1.20$1.20
ServerlessProvisioned
Together AI$1.20$1.20
Serverless
Microsoft Foundry$2.00$6.00
Provisioned

Available via routers & gateways(14)

Capabilities

No model capability flags are currently sourced.

Benchmark peer barsfor Coding

Benchmark scores(5)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
Google-Proof Q&A60.1diamondObserved 2026-03-06research
HellaSwag93.810-shotObserved 2026-03-06research
HumanEval86.2pass@1Observed 2026-03-06research
Massive Multitask Language Understanding84.55-shotObserved 2026-03-06research
AI2 Reasoning Challenge86.0Observed 2026-05-28Source

Migration checks

No linked migration route is available for this model yet.

Compare Mixtral 8x22B v0.1 with other models