LLM Reference

Best LLMs for Classification (2026)

Last refreshed 2026-07-20. Next refresh: weekly.

Best LLMs for text classification, routing, and moderation in 2026. Covers extraction, safety labeling, and structured output tasks.

Verdict

Use GPT-5.5 for classification today.

DeepSeek V4 Pro is the runner-up, 2 points back on MMLU.

Researched 58d agoWhy this pickMethodology

How we rank

Classification picks take the strongest score across MMLU-Pro, MMLU, and lighter classification benchmarks, then recency.

  1. EligibilityModels tagged for the classification decision task (routing, moderation, extraction, or classification benchmarks).
  2. Primary rankingMaximum score among MMLU-Pro, MMLU, HellaSwag, BoolQ, ANLI, ToxiGen coverage, then newer release.
  3. Podium freshnessShortlist cards require `lastResearched` within 60 days and a tracked public output price. Older or unpriced models remain in the full leaderboard with a visible “Verify pricing” reminder when research is older than 45 days.
  4. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. PricingLowest tracked commercial token pricing.
#ModelInput $/1MOutput $/1M
1Gemini 3.1 Pro Preview
PreviewVisionTools

Signal used: MMLU 98%

$2.00$12.00
2Llama 3.1 405B

Signal used: HellaSwag 95.8%

3DeepSeek V3
Tools

Signal used: HellaSwag 95.7%

$0.10$0.28
4Qwen2.5-72B-Instruct

Signal used: HellaSwag 95.6%

$0.18$0.28
5Llama 3.1 70B Instruct

Signal used: HellaSwag 94.2%

$0.40$0.40
6Mistral Large 2
VisionTools

Signal used: HellaSwag 93.8%

$0.48$2.40
7Mixtral 8x22B v0.1

Signal used: HellaSwag 93.8%

$0.65$0.65
8Falcon 180B

Signal used: HellaSwag 92.7%

9Gemma 2 27B

Signal used: HellaSwag 92.6%

$0.08$0.24
10GPT-5.5
ReasoningVisionTools

Signal used: MMLU 92.4%

$5.00$30.00
11Llama 3 70B

Signal used: HellaSwag 92.4%

$0.65$2.75
12Qwen2-7B

Signal used: HellaSwag 92%

$0.05$0.15
13Gemini 3 Pro
VisionTools

Signal used: MMLU-Pro 91.8%

$1.25$5.00
14Mistral NeMo Instruct (2407)

Signal used: HellaSwag 91.8%

$0.02$0.04
15DeepSeek Coder V2 Lite

Signal used: HellaSwag 91.4%

$0.50$0.50
16Claude Opus 4.6
ReasoningVisionTools

Signal used: MMLU 91.1%

$5.00$25.00
17Llama 3 8B Instruct

Signal used: HellaSwag 91.1%

$0.02$0.04
18Mixtral 8x7B

Signal used: HellaSwag 90.9%

$0.15$0.20
19Phi-3 Small 128K

Signal used: HellaSwag 90.8%

$0.35$1.05
20Mistral 7B Instruct v0.3
Tools

Signal used: HellaSwag 90.2%

$0.20$0.20

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.

  • #4Xiaomi MiMo-V2.5-Pro

    Xiaomi's April 22, 2026 public-beta flagship in the MiMo-V2.5 series. The official Xiaomi MiMo page describes MiMo-V2.5-Pro as its most capable model to date, focused on general agentic capability, complex software engineering, long-horizon tasks, and ultra-long-context instruction following. OpenRouter lists it as text-to-text with 1,048,576 token context, 131,072 max completion tokens, reasoning controls, tool use, and response_format support. Xiaomi says the V2.5 series will be open-sourced soon, but no public weights/license were verified at research time.

    89.4%

    MMLU

  • #5Claude Sonnet 4.6

    Claude Sonnet 4.6 is Anthropic's best combination of speed and intelligence. Proprietary decoder-only model with 1M-token context, 64K max output, multimodal vision, extended thinking, and function calling. Available via Anthropic API, AWS Bedrock, GCP Vertex AI, and OpenRouter at $3/1M input and $15/1M output tokens.

    89.3%

    MMLU

  • #6Qwen2.5-7B-Instruct

    Instruction-tuned 7B variant combining strong reasoning with real-time inference on single GPUs, ideal for developer tools and vision applications.

    89.3%

    HellaSwag

Frequently asked questions

Which LLM is best for classification?

GPT-5.5 is the current LLMReference top pick for classification. The verdict uses the stored category signal MMLU: 92.4%. Output pricing starts at $30.00 per 1M tokens. Review the linked model and provider pages before production use because availability and pricing can change.

How does GPT-5.5 compare to DeepSeek V4 Pro for classification?

GPT-5.5 leads DeepSeek V4 Pro in the visible shortlist on MMLU: 92.4% versus 90.1%. The pricing cards show GPT-5.5: output pricing starts at $30.00 per 1m tokens and DeepSeek V4 Pro: output pricing starts at $0.87 per 1m tokens.

How does LLMReference rank LLMs for classification?

LLMReference ranks LLMs for classification from stored model, benchmark, freshness, and pricing data. The current methodology summary is: Classification picks take the strongest score across MMLU-Pro, MMLU, and lighter classification benchmarks, then recency.

How often is this list updated?

The LLM rankings on this page are updated daily as new benchmark scores, provider availability, and pricing data are tracked. The "as of" date at the top of the page shows the most recent refresh.

How do you decide which models appear in the top 3?

The podium picks are driven by the primary benchmark signal for this category (shown in the Methodology section), filtered to non-deprecated models with confirmed API availability. In ties, we prefer the more recently released model.

Are preview or beta models included?

Preview models appear in the "Watch list" section but are not in the main ranked podium unless the category explicitly allows it (e.g., /best/coding and /best/agents, where preview models often lead benchmarks).

Can I compare two specific models head-to-head?

Yes — use the Compare tool at llmreference.com/compare for a side-by-side breakdown of context window, pricing, benchmarks, and provider availability.

Is the pricing data real-time?

Pricing is tracked from provider documentation and updated regularly. It reflects the best available public data, not live API quotes — always verify before billing.