Qwen3.8-Max
- MMLU-Pro
- —
- Output (from)
- $6.00 / 1M
Last refreshed 2026-09-01. Next refresh: weekly.
The best small LLMs under 10B parameters in 2026 — fast, cheap, and deployable on-device or at the edge with strong benchmark scores.
Verdict
Jina Reranker v3.5 is the runner-up: — vs — on MMLU-Pro.
Small models (≤10B active parameters) rank on MMLU-Pro, then GPQA Diamond, MMLU, and HellaSwag.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Granite 4.1 8B MMLU-Pro: 55.99% | $0.05 | $0.10 | |
| 2 | MiniMax M2.7 ReasoningTools MMLU-Pro: 80.43% | $0.28 | $1.20 | |
| 3 | Phi-4 Mini MMLU-Pro: 52.8% | $0.90 | $0.90 | |
| 4 | Gemma 2 9B MMLU-Pro: 52.08% | $0.06 | $0.18 | |
| 5 | LFM2.5 8B A1B ReasoningTools MMLU-Pro: 50.5% | — | — | |
| 6 | Granite 4.1 3B MMLU-Pro: 49.83% | — | — | |
| 7 | Phi-3 Mini 4k MMLU-Pro: 45.66% | $0.05 | $0.25 | |
| 8 | LFM2.5 1.2B Instruct Tools MMLU-Pro: 44.35% | — | — | |
| 9 | Llama 3.1 8B Instruct MMLU-Pro: 44.25% | $0.02 | $0.05 | |
| 10 | Llama 3 8B Instruct MMLU-Pro: 40.5% | $0.02 | $0.04 | |
| 11 | Llama 3.2 3B Instruct MMLU-Pro: 34.7% | $0.03 | $0.05 | |
| 12 | Llama 3.2 1B Instruct MMLU-Pro: 20% | $0.03 | $0.10 | |
| 13 | Qwen3.8-Max ReasoningVisionTools MMLU-Pro: — | $2.00 | $6.00 | |
| 14 | Qwen3-8B MMLU-Pro: — | $0.04 | $0.14 | |
| 15 | Qwen2-7B MMLU-Pro: — | $0.05 | $0.15 | |
| 16 | Gemma 7B Instruct MMLU-Pro: — | $0.05 | $0.07 | |
| 17 | OpenChat 3.5 (0106) MMLU-Pro: — | $0.07 | $0.07 | |
| 18 | Starling LM 7B Beta MMLU-Pro: — | — | — | |
| 19 | Zephyr 7B Beta MMLU-Pro: — | $0.05 | $0.20 | |
| 20 | Qwen2.5-7B-Instruct MMLU-Pro: — | $0.03 | $0.03 |
Ornith-1.0 9B Dense is the smallest public variant of DeepReinforce's Ornith 1.0 family. Designed for edge deployment and single-GPU setups (~19 GB in BF16 on a single 80GB GPU). Post-trained on a Qwen 3.5 or Gemma 4 base using self-scaffolding RL for agentic coding tasks. The model learns to build and refine its own execution harness during RL rather than using a fixed scaffold. Supports OpenAI-compatible tool calls. Context: 262,144 tokens confirmed by the official Hugging Face model card and config.
—
MMLU-Pro
Higgs Audio v3 TTS is Boson AI's 4B-parameter text-to-speech model released June 4, 2026. It supports 102 languages (85 at production quality with WER/CER <5%), zero-shot voice cloning, and inline control tokens for emotion (21 types), style (singing/shouting/whispering), sound effects, and prosody. Audio output is 24kHz MP3 or PCM. Open weights available under a non-commercial license; hosted API is in free public preview.
—
MMLU-Pro
MOSS-TTS-v1.5 is an open-weight multilingual text-to-speech model from MOSI AI and the OpenMOSS team. The 8B-parameter MossTTSDelay model supports zero-shot voice cloning, long-form speech generation, explicit pause control with [pause X.Ys] markers, and language-tagged multilingual synthesis across 31 languages. Version 1.5 improves on MOSS-TTS v1.0 with stronger multilingual synthesis, more stable voice cloning, better long-reference short-text handling, and punctuation-driven prosody. The model weights are available on Hugging Face under Apache 2.0; no hosted token-priced API route is confirmed in the June 2026 research handoff.
—
MMLU-Pro