Llama 3 70B
- HellaSwag
- 92.4%
- Output (from)
- $2.75 / 1M
Last refreshed 2026-09-21. Next refresh: weekly.
Best LLMs for text classification, routing, and moderation in 2026. Covers extraction, safety labeling, and structured output tasks.
Verdict
Gemini 3 Pro is the runner-up; compare HellaSwag against MMLU-Pro.
Classification picks take the strongest score across MMLU-Pro, MMLU, and lighter classification benchmarks, then recency.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Gemini 3.1 Pro Preview PreviewVisionTools Signal used: MMLU 98% | $2.00 | $12.00 | |
| 2 | Llama 3.1 405B Signal used: HellaSwag 95.8% | — | — | |
| 3 | DeepSeek V3 Tools Signal used: HellaSwag 95.7% | $0.10 | $0.28 | |
| 4 | Qwen2.5-72B-Instruct Signal used: HellaSwag 95.6% | $0.18 | $0.28 | |
| 5 | Llama 3.1 70B Instruct Signal used: HellaSwag 94.2% | $0.40 | $0.40 | |
| 6 | Mistral Large 2 VisionTools Signal used: HellaSwag 93.8% | $0.48 | $2.40 | |
| 7 | Mixtral 8x22B v0.1 Signal used: HellaSwag 93.8% | $0.65 | $0.65 | |
| 8 | Falcon 180B Signal used: HellaSwag 92.7% | — | — | |
| 9 | Gemma 2 27B Signal used: HellaSwag 92.6% | $0.08 | $0.24 | |
| 10 | GPT-5.5 ReasoningVisionTools Signal used: MMLU 92.4% | $5.00 | $30.00 | |
| 11 | Llama 3 70B Signal used: HellaSwag 92.4% | $0.65 | $2.75 | |
| 12 | Qwen2-7B Signal used: HellaSwag 92% | $0.05 | $0.15 | |
| 13 | Gemini 3 Pro VisionTools Signal used: MMLU-Pro 91.8% | $1.25 | $5.00 | |
| 14 | Mistral NeMo Instruct (2407) Signal used: HellaSwag 91.8% | $0.02 | $0.04 | |
| 15 | DeepSeek Coder V2 Lite Signal used: HellaSwag 91.4% | $0.50 | $0.50 | |
| 16 | Claude Opus 4.6 ReasoningVisionTools Signal used: MMLU 91.1% | $5.00 | $25.00 | |
| 17 | Llama 3 8B Instruct Signal used: HellaSwag 91.1% | $0.02 | $0.04 | |
| 18 | Mixtral 8x7B Signal used: HellaSwag 90.9% | $0.15 | $0.20 | |
| 19 | Phi-3 Small 128K Signal used: HellaSwag 90.8% | $0.35 | $1.05 | |
| 20 | Mistral 7B Instruct v0.3 Tools Signal used: HellaSwag 90.2% | $0.20 | $0.20 |
Claude Opus 4.6 is Anthropic's Claude 4.6 model with multimodal text and image input and an optional reasoning mode. It offers a 1M-token context window and scores 80.8 on SWE-bench Verified.
91.1%
MMLU
The Llama 3 8B Instruct model, released on April 18, 2024, is Meta's latest instruction-following language model with 8 billion parameters. It utilizes an auto-regressive transformer architecture with Grouped-Query Attention for improved scalability. Trained on over 15 trillion tokens and fine-tuned with 10 million human-annotated examples, it excels in dialogue and conversational tasks. The model outperforms its predecessors on industry benchmarks, scoring 68.4 on MMLU (5-shot). Designed for commercial and research applications, it prioritizes safety and helpfulness, making it suitable for chatbots, virtual assistants, and other interactive AI applications. For more details, visit the Hugging Face page [1].
91.1%
HellaSwag
Mixtral 8x7B, developed by Mistral AI, features a cutting-edge Mixture of Experts (MoE) architecture, utilizing eight experts with seven billion parameters each, yielding a total of 46.7 billion parameters. This architecture activates only two experts per token, allowing for efficient processing and a 6x faster inference rate compared to Llama 2 70B. The model excels in performance, surpassing Llama 2 70B and competing with GPT-3.5 on numerous benchmarks. It supports multiple languages and can handle context up to 32,000 tokens, enhancing understanding of lengthy text. Designed for diverse tasks, it is strong in code generation and available under a permissive Apache 2.0 license, promoting community engagement. Compatible with various optimization tools, its weights are easily deployable, with Mistral AI continuing to improve its capabilities through performance optimizations and fine-tuning efforts.
90.9%
HellaSwag