LLM Reference

Mixtral Models by MistralAI

MistralAIApache 2.0Open sourceHighlightOpen Source
5 models2023–2024Up to 64k ctxFrom $0.15/1M input

Details

ResearcherMistralAI
LicenseApache 2.0OSI-approved
Commercial useCommercial use: permitted
Models5
Released2023–2024
Max context64k

Capabilities

Function Calling1 of 5 models

About

The Mixtral family of large language models (LLMs), developed by Mistral AI, offers a groundbreaking approach in open-source AI through a sparse mixture-of-experts (SMoE) architecture. This innovative design allows the models to manage a significant number of parameters while ensuring efficient inference speed by activating only a subset of parameters for each token. Such architecture enables Mixtral models to deliver performance on par with much larger models, standing out in various benchmarks and outperforming competitors like Llama 2, and even equaling the prowess of closed-source models such as GPT-3.5. These models are multilingual, supporting languages such as English, French, Italian, German, and Spanish, and excel in domains like code generation. Instruction-tuned versions like Mixtral-8x7B-Instruct-v0.1 cater to applications requiring robust instruction-following and chat capabilities. The Mixtral family provides versatile models of differing sizes, addressing diverse computational and application requirements.

Current Variants

Use-when guidance is based on each model's tracked capabilities, context window, release date, and replacement status.

5 in view

Use when the workload needs 64k context and function calling.

2024-0764k contextfunction calling

Use when the workload needs 64k context.

2024-0464k context

Use when the workload needs 64k context.

2024-0464k context

Use when the workload needs 32k context.

2023-1232k context

Use when the workload needs 33k context and 56B parameters.

2023-1233k context56B parameters

Release Timeline

3 release groups
2024-07
1 current
Mixtral 8x22B Instruct v0.3
64k contextfunction calling
Current
2024-04
2 current
Current
Current
2023-12
2 current
Mixtral 8x7B
32k context
Current
Mixtral 8x7B Instruct v0.1
33k context56B parameters
Current

Specifications(5 models)

Mixtral model specifications comparison
ModelReleasedContextParametersFn Calling
Mixtral 8x22B Instruct v0.32024-0764k8x22BYes
Mixtral 8x22B v0.12024-0464k8x22BNo
Mixtral 8x22B Instruct v0.12024-0464k8x22BNo
Mixtral 8x7B2023-1232k8x7BNo
Mixtral 8x7B Instruct v0.12023-1233k56BNo

Pricing

Popular comparisons in this family

Frequently Asked Questions

What is Mixtral used for?
Mixtral is used for agent workflows and tool use and coding. The family description and listed model capabilities point to those workloads as the best fit.
How does Mixtral compare to Ministral?
Mixtral by MistralAI is strongest where you need agent workflows and tool use, while Ministral by MistralAI is the closest related family to check for vision and multimodal work. Mixtral has 5 listed variants and reaches up to 64k context, while Ministral reaches up to 32k context, so compare the specs and pricing tables before choosing a production model.
Which Mixtral model should I use?
For the lowest listed input price, start with Mixtral 8x7B through Mistral AI Studio at $0.15/1M input tokens. For the most capable/latest local choice, evaluate Mixtral 8x22B Instruct v0.3 with 64k context and function calling.