Phi-3 Mini 4k
- MMLU-Pro
- 45.66%
- Output (from)
- $0.250 / 1M
Last refreshed 2026-09-30. Next refresh: weekly.
The best small LLMs under 10B parameters in 2026 — fast, cheap, and deployable on-device or at the edge with strong benchmark scores.
Verdict
Qwen3.8-Max is the runner-up: 45.66% vs — on MMLU-Pro.
Small models (≤10B active parameters) rank on MMLU-Pro, then GPQA Diamond, MMLU, and HellaSwag.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Granite 4.1 8B MMLU-Pro: 55.99% | $0.05 | $0.10 | |
| 2 | MiniMax M2.7 ReasoningTools MMLU-Pro: 80.43% | $0.28 | $1.20 | |
| 3 | Phi-4 Mini MMLU-Pro: 52.8% | $0.90 | $0.90 | |
| 4 | Gemma 2 9B MMLU-Pro: 52.08% | $0.06 | $0.18 | |
| 5 | LFM2.5 8B A1B ReasoningTools MMLU-Pro: 50.5% | — | — | |
| 6 | Granite 4.1 3B MMLU-Pro: 49.83% | — | — | |
| 7 | Phi-3 Mini 4k MMLU-Pro: 45.66% | $0.05 | $0.25 | |
| 8 | LFM2.5 1.2B Instruct Tools MMLU-Pro: 44.35% | — | — | |
| 9 | Llama 3.1 8B Instruct MMLU-Pro: 44.25% | $0.02 | $0.05 | |
| 10 | Llama 3 8B Instruct MMLU-Pro: 40.5% | $0.02 | $0.04 | |
| 11 | Llama 3.2 3B Instruct MMLU-Pro: 34.7% | $0.03 | $0.05 | |
| 12 | Llama 3.2 1B Instruct MMLU-Pro: 20% | $0.03 | $0.10 | |
| 13 | Qwen3.8-Max ReasoningVisionTools MMLU-Pro: — | $2.00 | $6.00 | |
| 14 | Qwen3-8B MMLU-Pro: — | $0.04 | $0.14 | |
| 15 | Qwen2-7B MMLU-Pro: — | $0.05 | $0.15 | |
| 16 | Gemma 7B Instruct MMLU-Pro: — | $0.05 | $0.07 | |
| 17 | OpenChat 3.5 (0106) MMLU-Pro: — | $0.07 | $0.07 | |
| 18 | Starling LM 7B Beta MMLU-Pro: — | — | — | |
| 19 | Zephyr 7B Beta MMLU-Pro: — | $0.05 | $0.20 | |
| 20 | Qwen2.5-7B-Instruct MMLU-Pro: — | $0.03 | $0.03 |
Instruction-tuned 7B variant combining strong reasoning with real-time inference on single GPUs, ideal for developer tools and vision applications.
—
MMLU-Pro
Aya-23-8B is a multilingual large language model developed by Cohere For AI, featuring 8 billion parameters. As an instruction-fine-tuned model, it is adept at following instructions and is optimized for text generation and understanding. The model employs a decoder-only Transformer architecture, utilizing enhancements like parallel attention and feed-forward layers for efficiency. It supports 23 languages, including Arabic, Chinese, English, and French, and is proficient in tasks such as machine translation, chatbot interactions, and text summarization. Despite its capabilities, its performance might vary across languages, particularly those with less linguistic resources, and it has a context length limit of 8192 tokens. Training involved diverse data sources like human annotations and synthetic datasets to bolster its multilingual proficiency.
—
MMLU-Pro
Qwen2-1.5B is Alibaba's Qwen2 model. It scores 44.2 on the GPQA benchmark.
—
MMLU-Pro