Best Small Language Models Under 10B Parameters (2026)

Last refreshed 2026-09-30. Next refresh: weekly.

The best small LLMs under 10B parameters in 2026 — fast, cheap, and deployable on-device or at the edge with strong benchmark scores.

Verdict

Use Phi-3 Mini 4k for small-model deployments today.

Qwen3.8-Max is the runner-up: 45.66% vs — on MMLU-Pro.

Researched 9d agoWhy this pickMethodology
Single-source resultMiniMax M2.7 scored 80.4% on MMLU-Pro, more than five points above the next GA score (56.0%). We dropped it one GA rank until another source corroborates the result.

How we rank

Small models (≤10B active parameters) rank on MMLU-Pro, then GPQA Diamond, MMLU, and HellaSwag.

  1. Eligibility — Non-deprecated models with ≤10B parameters (billions-only parser).
  2. Primary ranking — MMLU-Pro, then GPQA Diamond, then MMLU, then HellaSwag, then newer release.
  3. Podium freshness — Shortlist cards require `lastResearched` within 60 days and a tracked public output price. Stale or unpriced SKUs stay in the table with a “Verify pricing” badge once research is past 45 days.
  4. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. Pricing — SLMs often win on unit economics — compare the provider ladder before picking.
#ModelInput $/1MOutput $/1M
1Granite 4.1 8B

MMLU-Pro: 55.99%

$0.05$0.10
2MiniMax M2.7
ReasoningTools

MMLU-Pro: 80.43%

$0.28$1.20
3Phi-4 Mini

MMLU-Pro: 52.8%

$0.90$0.90
4Gemma 2 9B

MMLU-Pro: 52.08%

$0.06$0.18
5LFM2.5 8B A1B
ReasoningTools

MMLU-Pro: 50.5%

——
6Granite 4.1 3B

MMLU-Pro: 49.83%

——
7Phi-3 Mini 4k

MMLU-Pro: 45.66%

$0.05$0.25
8LFM2.5 1.2B Instruct
Tools

MMLU-Pro: 44.35%

——
9Llama 3.1 8B Instruct

MMLU-Pro: 44.25%

$0.02$0.05
10Llama 3 8B Instruct

MMLU-Pro: 40.5%

$0.02$0.04
11Llama 3.2 3B Instruct

MMLU-Pro: 34.7%

$0.03$0.05
12Llama 3.2 1B Instruct

MMLU-Pro: 20%

$0.03$0.10
13Qwen3.8-Max
ReasoningVisionTools

MMLU-Pro: —

$2.00$6.00
14Qwen3-8B

MMLU-Pro: —

$0.04$0.14
15Qwen2-7B

MMLU-Pro: —

$0.05$0.15
16Gemma 7B Instruct

MMLU-Pro: —

$0.05$0.07
17OpenChat 3.5 (0106)

MMLU-Pro: —

$0.07$0.07
18Starling LM 7B Beta

MMLU-Pro: —

——
19Zephyr 7B Beta

MMLU-Pro: —

$0.05$0.20
20Qwen2.5-7B-Instruct

MMLU-Pro: —

$0.03$0.03

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4Qwen2.5-7B-Instruct

    Instruction-tuned 7B variant combining strong reasoning with real-time inference on single GPUs, ideal for developer tools and vision applications.

    —

    MMLU-Pro

  • Aya-23-8B is a multilingual large language model developed by Cohere For AI, featuring 8 billion parameters. As an instruction-fine-tuned model, it is adept at following instructions and is optimized for text generation and understanding. The model employs a decoder-only Transformer architecture, utilizing enhancements like parallel attention and feed-forward layers for efficiency. It supports 23 languages, including Arabic, Chinese, English, and French, and is proficient in tasks such as machine translation, chatbot interactions, and text summarization. Despite its capabilities, its performance might vary across languages, particularly those with less linguistic resources, and it has a context length limit of 8192 tokens. Training involved diverse data sources like human annotations and synthetic datasets to bolster its multilingual proficiency.

    —

    MMLU-Pro

  • #6Qwen2-1.5B

    Qwen2-1.5B is Alibaba's Qwen2 model. It scores 44.2 on the GPQA benchmark.

    —

    MMLU-Pro