LLM Reference

Best Free LLMs You Can Use Right Now (2026)

Last refreshed 2026-09-01. Next refresh: weekly.

Free LLMs you can use right now: zero-cost hosted tiers first, then open-weight models you can self-host. Updated as free tiers change.

Verdict

Use LFM2.5 8B A1B for free-tier work today.

LFM2.5 1.2B Instruct is the runner-up, 6 points back on MMLU-Pro.

Researched 54d agoWhy this pickMethodology

How we rank

Free leaders prioritize models with a current zero-dollar hosted tier, then fall back to open-weight models you can self-host without token fees.

  1. EligibilityModels qualify if at least one current public/self-serve provider row lists both input and output token pricing as $0, or if the model is open-weight/self-host-capable under the same license flags used by /best/open-source.
  2. Primary rankingZero-cost hosted availability comes first. Within the hosted-free and open-weight fallback groups, rows sort by MMLU-Pro, then GPQA Diamond, then MMLU, then newer release. Podium cards keep hosted $0 routes ahead of open-weight fallback rows; a fallback card is labeled as self-host/no-token-fee rather than paid hosted output.
  3. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. Open-weight fallbackOpen-weight fallback rows represent local or self-hosted use where the model itself has no token fee; you still need to budget your own compute.
  5. Related viewUse /best/open-source when the decision is license and self-host posture first; this free page answers what you can use at zero dollars right now.
#ModelInput $/1MOutput $/1M
1Gemma 4 31B IT
VisionTools

Capability signal: MMLU-Pro 85.2%

FreeFree
2Gemma 4 26B A4B IT
VisionTools

Capability signal: MMLU-Pro 82.6%

FreeFree
3Nemotron 3 Nano Omni

Capability signal: MMLU-Pro 71.8%

FreeFree
4Gemma 4 E4B IT
Tools

Capability signal: MMLU-Pro 69.4%

FreeFree
5Gemma 4 E2B IT
Tools

Capability signal: MMLU-Pro 60%

FreeFree
6gpt-oss-120b
Tools

Capability signal: GPQA Diamond 78.2%

FreeFree
7gpt-oss-20b
Tools

Capability signal: GPQA Diamond 68.8%

FreeFree
8CoBuddy
ReasoningTools

Capability signal:

FreeFree
9Laguna XS.2

Capability signal:

FreeFree
10Laguna M.1

Capability signal:

FreeFree
11T5Gemma
Tools

Capability signal:

FreeFree
12Gemma 3

Capability signal:

FreeFree
13Gemma 3n

Capability signal:

FreeFree
14MedGemma
VisionTools

Capability signal:

FreeFree
15MedSigLIP
VisionTools

Capability signal:

FreeFree
16Qwen3.6-Plus
VisionTools

Capability signal: MMLU-Pro 88.5%

$0.33$1.95
17Qwen3.6 Max Preview
PreviewReasoningVisionTools

Capability signal: MMLU-Pro 88.5%

$1.04$6.24
18Qwen3.5-397B-A17B
ReasoningVisionTools

Capability signal: MMLU-Pro 87.8%

$0.39$2.34
19DeepSeek V4 Pro
ReasoningTools

Capability signal: MMLU-Pro 87.5%

$0.43$0.87
20Qwen3.5-122B-A10B
ReasoningVisionTools

Capability signal: MMLU-Pro 86.7%

$0.26$2.08

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4MiniMax M3

    MiniMax M3 is MiniMax's current API flagship (released June 1, 2026) with MiniMax Sparse Attention for economical 1M-token context, native multimodality, and agentic coding. Benchmark rows in LLMReference separate vendor-reported SWE-bench Pro, Terminal-Bench 2.1, MCP-Atlas, and BrowseComp scores (source: minimax.io) from third-party rows — inspect evaluator, variant, and source on each benchmark cell before comparing to leaderboard claims.

    92.9%

    GPQA Diamond

  • Qwen3.8-Max is a 2.4-trillion-parameter MoE flagship delivering a comprehensive leap in coding and professional work. It autonomously codes and delivers complete projects spanning 10+ days, handles hundreds of specialized tasks across legal, financial, design, and other professional domains, and produces production-grade results end-to-end in a single conversation. Native visual understanding runs through the full cycle of planning, execution, and verification, enabling deep semantic analysis of ultra-long documents and extended video content. The public Hugging Face checkpoint (Qwen3.8-2.4T-A95B) is text-only under the Qwen3.8-Max License; the hosted qwen3.8-max API adds vision, optional non-thinking mode, 1M context, tools, structured outputs, and prompt caching on the same model ID.

    92.6%

    GPQA Diamond

  • #6Qwen3.6-Max

    Qwen3.6-Max builds upon Qwen3-Max and Qwen3.6-Plus to deliver enhanced vibe coding capabilities, more efficient coding agent execution, and significantly improved front-end development skills. Its long-tail knowledge retention has been further upgraded, making it Alibaba's most capable multimodal flagship as of April 2026.

    91.8%

    GPQA Diamond