LongCat-2.0
- MMLU-Pro
- —
- Output (from)
- $2.95 / 1M
Last refreshed 2026-07-13. Next refresh: weekly.
The best small LLMs under 10B parameters in 2026 — fast, cheap, and deployable on-device or at the edge with strong benchmark scores.
Verdict
T5Gemma is the runner-up: — vs — on MMLU-Pro.
Single-source resultMiniMax M2.7 scored 80.4% on MMLU-Pro, more than five points above the next GA score (56.0%). We dropped it one GA rank until another source corroborates the result.
Small models (≤10B active parameters) rank on MMLU-Pro, then GPQA Diamond, MMLU, and HellaSwag.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Granite 4.1 8B MMLU-Pro: 55.99% | $0.05 | $0.10 | |
| 2 | MiniMax M2.7 ReasoningTools MMLU-Pro: 80.43% | $0.28 | $1.20 | |
| 3 | Phi-4 Mini MMLU-Pro: 52.8% | $0.90 | $0.90 | |
| 4 | Gemma 2 9B MMLU-Pro: 52.08% | $0.06 | $0.18 | |
| 5 | LFM2.5 8B A1B ReasoningTools MMLU-Pro: 50.5% | — | — | |
| 6 | Granite 4.1 3B MMLU-Pro: 49.83% | — | — | |
| 7 | Phi-3 Mini 4k MMLU-Pro: 45.66% | $0.05 | $0.25 | |
| 8 | LFM2.5 1.2B Instruct Tools MMLU-Pro: 44.35% | — | — | |
| 9 | Llama 3.1 8B Instruct MMLU-Pro: 44.25% | $0.02 | $0.05 | |
| 10 | Llama 3 8B Instruct MMLU-Pro: 40.5% | $0.02 | $0.04 | |
| 11 | Llama 3.2 3B Instruct MMLU-Pro: 34.7% | $0.03 | $0.05 | |
| 12 | Llama 3.2 1B Instruct MMLU-Pro: 20% | $0.03 | $0.10 | |
| 13 | Qwen3-8B MMLU-Pro: — | $0.04 | $0.14 | |
| 14 | Qwen2-7B MMLU-Pro: — | $0.05 | $0.15 | |
| 15 | Gemma 7B Instruct MMLU-Pro: — | $0.05 | $0.07 | |
| 16 | OpenChat 3.5 (0106) MMLU-Pro: — | $0.07 | $0.07 | |
| 17 | Starling LM 7B Beta MMLU-Pro: — | — | — | |
| 18 | Zephyr 7B Beta MMLU-Pro: — | $0.05 | $0.20 | |
| 19 | Qwen2.5-7B-Instruct MMLU-Pro: — | $0.03 | $0.03 | |
| 20 | Aya 23 8B MMLU-Pro: — | — | — |
Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
Mistral Voxtral Mini 3B 2507 is MistralAI's Voxtral model. It was released 2025-07-01.
—
MMLU-Pro
IBM Granite Vision 3.3 2B improves on Granite Vision 3.2 with a new SigLIP2 vision encoder, higher-quality training data, and experimental capabilities: image segmentation, doctags generation, and multi-page support (up to 8 pages). Enhanced safety compared to 3.2. Architecture: SigLIP2 + 2-layer MLP (GELU) + Granite 3.1 2B Instruct. Benchmarks: DocVQA 0.91, TextVQA 0.80, OCRBench 0.79, InfoVQA 0.68. Apache 2.0.
—
MMLU-Pro
Qwen3 Reranker 8B is Alibaba's multilingual reranking model from the Qwen3 generation, designed for retrieval-augmented generation pipelines. Open-sourced under Apache 2.0. Achieves 81.22 on MTEB-Code and 72.94 on MMTEB-R. Released alongside Qwen3 Embedding series June 2025.
—
MMLU-Pro
Side-by-side comparison of the top picks by price, benchmark, and API access.
LongCat-2.0 is the current LLMReference top pick for small-model deployment. The verdict uses the stored category signal MMLU-Pro: —. Output pricing starts at $2.95 per 1M tokens. Review the linked model and provider pages before production use because availability and pricing can change.
LongCat-2.0 leads T5Gemma in the visible shortlist on MMLU-Pro: — versus —. The pricing cards show LongCat-2.0: output pricing starts at $2.95 per 1m tokens and T5Gemma: output pricing is listed as free.
LLMReference ranks LLMs for small-model deployment from stored model, benchmark, freshness, and pricing data. The current methodology summary is: Small models (≤10B active parameters) rank on MMLU-Pro, then GPQA Diamond, MMLU, and HellaSwag.
The LLM rankings on this page are updated daily as new benchmark scores, provider availability, and pricing data are tracked. The "as of" date at the top of the page shows the most recent refresh.
The podium picks are driven by the primary benchmark signal for this category (shown in the Methodology section), filtered to non-deprecated models with confirmed API availability. In ties, we prefer the more recently released model.
Preview models appear in the "Watch list" section but are not in the main ranked podium unless the category explicitly allows it (e.g., /best/coding and /best/agents, where preview models often lead benchmarks).
Yes — use the Compare tool at llmreference.com/compare for a side-by-side breakdown of context window, pricing, benchmarks, and provider availability.
Pricing is tracked from provider documentation and updated regularly. It reflects the best available public data, not live API quotes — always verify before billing.