LFM2.5 8B A1B
- MMLU-Pro
- 50.5%
- Free basis
- Self-host no token fee
Last refreshed 2026-09-01. Next refresh: weekly.
Free LLMs you can use right now: zero-cost hosted tiers first, then open-weight models you can self-host. Updated as free tiers change.
Verdict
LFM2.5 1.2B Instruct is the runner-up, 6 points back on MMLU-Pro.
Free leaders prioritize models with a current zero-dollar hosted tier, then fall back to open-weight models you can self-host without token fees.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Gemma 4 31B IT VisionTools Capability signal: MMLU-Pro 85.2% | Free | Free | |
| 2 | Gemma 4 26B A4B IT VisionTools Capability signal: MMLU-Pro 82.6% | Free | Free | |
| 3 | Nemotron 3 Nano Omni Capability signal: MMLU-Pro 71.8% | Free | Free | |
| 4 | Gemma 4 E4B IT Tools Capability signal: MMLU-Pro 69.4% | Free | Free | |
| 5 | Gemma 4 E2B IT Tools Capability signal: MMLU-Pro 60% | Free | Free | |
| 6 | gpt-oss-120b Tools Capability signal: GPQA Diamond 78.2% | Free | Free | |
| 7 | gpt-oss-20b Tools Capability signal: GPQA Diamond 68.8% | Free | Free | |
| 8 | CoBuddy ReasoningTools Capability signal: — | Free | Free | |
| 9 | Laguna XS.2 Capability signal: — | Free | Free | |
| 10 | Laguna M.1 Capability signal: — | Free | Free | |
| 11 | T5Gemma Tools Capability signal: — | Free | Free | |
| 12 | Gemma 3 Capability signal: — | Free | Free | |
| 13 | Gemma 3n Capability signal: — | Free | Free | |
| 14 | MedGemma VisionTools Capability signal: — | Free | Free | |
| 15 | MedSigLIP VisionTools Capability signal: — | Free | Free | |
| 16 | Qwen3.6-Plus VisionTools Capability signal: MMLU-Pro 88.5% | $0.33 | $1.95 | |
| 17 | Qwen3.6 Max Preview PreviewReasoningVisionTools Capability signal: MMLU-Pro 88.5% | $1.04 | $6.24 | |
| 18 | Qwen3.5-397B-A17B ReasoningVisionTools Capability signal: MMLU-Pro 87.8% | $0.39 | $2.34 | |
| 19 | DeepSeek V4 Pro ReasoningTools Capability signal: MMLU-Pro 87.5% | $0.43 | $0.87 | |
| 20 | Qwen3.5-122B-A10B ReasoningVisionTools Capability signal: MMLU-Pro 86.7% | $0.26 | $2.08 |
MiniMax M3 is MiniMax's current API flagship (released June 1, 2026) with MiniMax Sparse Attention for economical 1M-token context, native multimodality, and agentic coding. Benchmark rows in LLMReference separate vendor-reported SWE-bench Pro, Terminal-Bench 2.1, MCP-Atlas, and BrowseComp scores (source: minimax.io) from third-party rows — inspect evaluator, variant, and source on each benchmark cell before comparing to leaderboard claims.
92.9%
GPQA Diamond
Qwen3.8-Max is a 2.4-trillion-parameter MoE flagship delivering a comprehensive leap in coding and professional work. It autonomously codes and delivers complete projects spanning 10+ days, handles hundreds of specialized tasks across legal, financial, design, and other professional domains, and produces production-grade results end-to-end in a single conversation. Native visual understanding runs through the full cycle of planning, execution, and verification, enabling deep semantic analysis of ultra-long documents and extended video content. The public Hugging Face checkpoint (Qwen3.8-2.4T-A95B) is text-only under the Qwen3.8-Max License; the hosted qwen3.8-max API adds vision, optional non-thinking mode, 1M context, tools, structured outputs, and prompt caching on the same model ID.
92.6%
GPQA Diamond
Qwen3.6-Max builds upon Qwen3-Max and Qwen3.6-Plus to deliver enhanced vibe coding capabilities, more efficient coding agent execution, and significantly improved front-end development skills. Its long-tail knowledge retention has been further upgraded, making it Alibaba's most capable multimodal flagship as of April 2026.
91.8%
GPQA Diamond