LLM Reference

Best LLMs for Classification (2026)

Last refreshed 2026-09-21. Next refresh: weekly.

Best LLMs for text classification, routing, and moderation in 2026. Covers extraction, safety labeling, and structured output tasks.

Verdict

Use Llama 3 70B for classification today.

Gemini 3 Pro is the runner-up; compare HellaSwag against MMLU-Pro.

Researched 3d agoWhy this pickMethodology

How we rank

Classification picks take the strongest score across MMLU-Pro, MMLU, and lighter classification benchmarks, then recency.

  1. Eligibility — Models tagged for the classification decision task (routing, moderation, extraction, or classification benchmarks).
  2. Primary ranking — Maximum score among MMLU-Pro, MMLU, HellaSwag, BoolQ, ANLI, ToxiGen coverage, then newer release.
  3. Podium freshness — Shortlist cards require `lastResearched` within 60 days and a tracked public output price. Older or unpriced models remain in the full leaderboard with a visible “Verify pricing” reminder when research is older than 45 days.
  4. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. Pricing — Lowest tracked commercial token pricing.
#ModelInput $/1MOutput $/1M
1Gemini 3.1 Pro Preview
PreviewVisionTools

Signal used: MMLU 98%

$2.00$12.00
2Llama 3.1 405B

Signal used: HellaSwag 95.8%

——
3DeepSeek V3
Tools

Signal used: HellaSwag 95.7%

$0.10$0.28
4Qwen2.5-72B-Instruct

Signal used: HellaSwag 95.6%

$0.18$0.28
5Llama 3.1 70B Instruct

Signal used: HellaSwag 94.2%

$0.40$0.40
6Mistral Large 2
VisionTools

Signal used: HellaSwag 93.8%

$0.48$2.40
7Mixtral 8x22B v0.1

Signal used: HellaSwag 93.8%

$0.65$0.65
8Falcon 180B

Signal used: HellaSwag 92.7%

——
9Gemma 2 27B

Signal used: HellaSwag 92.6%

$0.08$0.24
10GPT-5.5
ReasoningVisionTools

Signal used: MMLU 92.4%

$5.00$30.00
11Llama 3 70B

Signal used: HellaSwag 92.4%

$0.65$2.75
12Qwen2-7B

Signal used: HellaSwag 92%

$0.05$0.15
13Gemini 3 Pro
VisionTools

Signal used: MMLU-Pro 91.8%

$1.25$5.00
14Mistral NeMo Instruct (2407)

Signal used: HellaSwag 91.8%

$0.02$0.04
15DeepSeek Coder V2 Lite

Signal used: HellaSwag 91.4%

$0.50$0.50
16Claude Opus 4.6
ReasoningVisionTools

Signal used: MMLU 91.1%

$5.00$25.00
17Llama 3 8B Instruct

Signal used: HellaSwag 91.1%

$0.02$0.04
18Mixtral 8x7B

Signal used: HellaSwag 90.9%

$0.15$0.20
19Phi-3 Small 128K

Signal used: HellaSwag 90.8%

$0.35$1.05
20Mistral 7B Instruct v0.3
Tools

Signal used: HellaSwag 90.2%

$0.20$0.20

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4Claude Opus 4.6

    Claude Opus 4.6 is Anthropic's Claude 4.6 model with multimodal text and image input and an optional reasoning mode. It offers a 1M-token context window and scores 80.8 on SWE-bench Verified.

    91.1%

    MMLU

  • #5Llama 3 8B Instruct

    The Llama 3 8B Instruct model, released on April 18, 2024, is Meta's latest instruction-following language model with 8 billion parameters. It utilizes an auto-regressive transformer architecture with Grouped-Query Attention for improved scalability. Trained on over 15 trillion tokens and fine-tuned with 10 million human-annotated examples, it excels in dialogue and conversational tasks. The model outperforms its predecessors on industry benchmarks, scoring 68.4 on MMLU (5-shot). Designed for commercial and research applications, it prioritizes safety and helpfulness, making it suitable for chatbots, virtual assistants, and other interactive AI applications. For more details, visit the Hugging Face page [1].

    91.1%

    HellaSwag

  • #6Mixtral 8x7B

    Mixtral 8x7B, developed by Mistral AI, features a cutting-edge Mixture of Experts (MoE) architecture, utilizing eight experts with seven billion parameters each, yielding a total of 46.7 billion parameters. This architecture activates only two experts per token, allowing for efficient processing and a 6x faster inference rate compared to Llama 2 70B. The model excels in performance, surpassing Llama 2 70B and competing with GPT-3.5 on numerous benchmarks. It supports multiple languages and can handle context up to 32,000 tokens, enhancing understanding of lengthy text. Designed for diverse tasks, it is strong in code generation and available under a permissive Apache 2.0 license, promoting community engagement. Compatible with various optimization tools, its weights are easily deployable, with Mistral AI continuing to improve its capabilities through performance optimizations and fine-tuning efforts.

    90.9%

    HellaSwag