LLM Reference

Best Small Language Models Under 10B Parameters (2026)

Last refreshed 2026-09-01. Next refresh: weekly.

The best small LLMs under 10B parameters in 2026 — fast, cheap, and deployable on-device or at the edge with strong benchmark scores.

Verdict

Use Qwen3.8-Max for small-model deployments today.

Jina Reranker v3.5 is the runner-up: — vs — on MMLU-Pro.

Researched 11d agoWhy this pickMethodology
Single-source resultMiniMax M2.7 scored 80.4% on MMLU-Pro, more than five points above the next GA score (56.0%). We dropped it one GA rank until another source corroborates the result.

How we rank

Small models (≤10B active parameters) rank on MMLU-Pro, then GPQA Diamond, MMLU, and HellaSwag.

  1. EligibilityNon-deprecated models with ≤10B parameters (billions-only parser).
  2. Primary rankingMMLU-Pro, then GPQA Diamond, then MMLU, then HellaSwag, then newer release.
  3. Podium freshnessShortlist cards require `lastResearched` within 60 days and a tracked public output price. Stale or unpriced SKUs stay in the table with a “Verify pricing” badge once research is past 45 days.
  4. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. PricingSLMs often win on unit economics — compare the provider ladder before picking.
#ModelInput $/1MOutput $/1M
1Granite 4.1 8B

MMLU-Pro: 55.99%

$0.05$0.10
2MiniMax M2.7
ReasoningTools

MMLU-Pro: 80.43%

$0.28$1.20
3Phi-4 Mini

MMLU-Pro: 52.8%

$0.90$0.90
4Gemma 2 9B

MMLU-Pro: 52.08%

$0.06$0.18
5LFM2.5 8B A1B
ReasoningTools

MMLU-Pro: 50.5%

6Granite 4.1 3B

MMLU-Pro: 49.83%

7Phi-3 Mini 4k

MMLU-Pro: 45.66%

$0.05$0.25
8LFM2.5 1.2B Instruct
Tools

MMLU-Pro: 44.35%

9Llama 3.1 8B Instruct

MMLU-Pro: 44.25%

$0.02$0.05
10Llama 3 8B Instruct

MMLU-Pro: 40.5%

$0.02$0.04
11Llama 3.2 3B Instruct

MMLU-Pro: 34.7%

$0.03$0.05
12Llama 3.2 1B Instruct

MMLU-Pro: 20%

$0.03$0.10
13Qwen3.8-Max
ReasoningVisionTools

MMLU-Pro:

$2.00$6.00
14Qwen3-8B

MMLU-Pro:

$0.04$0.14
15Qwen2-7B

MMLU-Pro:

$0.05$0.15
16Gemma 7B Instruct

MMLU-Pro:

$0.05$0.07
17OpenChat 3.5 (0106)

MMLU-Pro:

$0.07$0.07
18Starling LM 7B Beta

MMLU-Pro:

19Zephyr 7B Beta

MMLU-Pro:

$0.05$0.20
20Qwen2.5-7B-Instruct

MMLU-Pro:

$0.03$0.03

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4Ornith-1.0 9B

    Ornith-1.0 9B Dense is the smallest public variant of DeepReinforce's Ornith 1.0 family. Designed for edge deployment and single-GPU setups (~19 GB in BF16 on a single 80GB GPU). Post-trained on a Qwen 3.5 or Gemma 4 base using self-scaffolding RL for agentic coding tasks. The model learns to build and refine its own execution harness during RL rather than using a fixed scaffold. Supports OpenAI-compatible tool calls. Context: 262,144 tokens confirmed by the official Hugging Face model card and config.

    MMLU-Pro

  • #5Higgs Audio v3 TTS

    Higgs Audio v3 TTS is Boson AI's 4B-parameter text-to-speech model released June 4, 2026. It supports 102 languages (85 at production quality with WER/CER <5%), zero-shot voice cloning, and inline control tokens for emotion (21 types), style (singing/shouting/whispering), sound effects, and prosody. Audio output is 24kHz MP3 or PCM. Open weights available under a non-commercial license; hosted API is in free public preview.

    MMLU-Pro

  • #6MOSS-TTS-v1.5

    MOSS-TTS-v1.5 is an open-weight multilingual text-to-speech model from MOSI AI and the OpenMOSS team. The 8B-parameter MossTTSDelay model supports zero-shot voice cloning, long-form speech generation, explicit pause control with [pause X.Ys] markers, and language-tagged multilingual synthesis across 31 languages. Version 1.5 improves on MOSS-TTS v1.0 with stronger multilingual synthesis, more stable voice cloning, better long-reference short-text handling, and punctuation-driven prosody. The model weights are available on Hugging Face under Apache 2.0; no hosted token-priced API route is confirmed in the June 2026 research handoff.

    MMLU-Pro