LLM Reference

Best Open Source LLMs (2026)

Last refreshed 2026-09-01. Next refresh: weekly.

The best open-weight LLMs in 2026, ranked by benchmark scores. Run locally, self-host, or deploy on your own infra — no API key required.

Release watch. Nemotron-Labs-Diffusion from NVIDIA, including 3B, 8B, and 14B diffusion language model variants.

Verdict

Use Kimi K3 for self-hosted open-weight use today.

Qwen3.8-Max is the runner-up, 0.9 points back on GPQA Diamond.

Researched 11d agoWhy this pickMethodology

How we rank

Open-weight boards emphasize GPQA Diamond (harder to game than broad MMLU), then MMLU, then recency.

  1. EligibilityModels marked with supported open licenses/flags in seed data.
  2. Primary rankingGPQA Diamond, then MMLU, then newer release.
  3. Variant collapseWe keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  4. PricingHosted open-weight pricing still varies wildly — use the rate card column for apples-to-apples.
#ModelInput $/1MOutput $/1M
1Kimi K3
ReasoningVisionTools

GPQA Diamond: 93.5%

$3.00$15.00
2MiniMax M3
ReasoningVisionTools

GPQA Diamond: 92.9%

$0.30$1.20
3Qwen3.8-Max
ReasoningVisionTools

GPQA Diamond: 92.6%

$2.00$6.00
4Qwen3.6-Max
Vision

GPQA Diamond: 91.8%

5Qwen3.8-Flash-Next
ReasoningVision

GPQA Diamond: 91.7%

6GLM-5.2
ReasoningTools

GPQA Diamond: 91.2%

$1.40$4.40
7Kimi K2.6
ReasoningVisionTools

GPQA Diamond: 90.5%

$0.73$3.40
8Qwen3.6-Plus
VisionTools

GPQA Diamond: 90.4%

$0.33$1.95
9DeepSeek V4 Pro
ReasoningTools

GPQA Diamond: 90.1%

$0.43$0.87
10Qwen3.5-397B-A17B
ReasoningVisionTools

GPQA Diamond: 89.3%

$0.39$2.34
11Qwen3.8-27B
ReasoningVisionTools

GPQA Diamond: 89.2%

$0.50$3.00
12Trinity-Large-Thinking
ReasoningTools

GPQA Diamond: 89.2%

$0.22$0.85
13Qwen3.5-Plus
Vision

GPQA Diamond: 88.4%

$0.30$1.80
14Ring-2.6-1T
ReasoningTools

GPQA Diamond: 88.27%

$0.07$0.63
15DeepSeek V4 Flash
ReasoningTools

GPQA Diamond: 88.1%

$0.06$0.12
16Qwen3.6-27B
ReasoningVisionTools

GPQA Diamond: 87.8%

$0.32$3.20
17DeepSeek V3 0324

GPQA Diamond: 87.6%

$0.27$1.12
18MiniMax M2.7
ReasoningTools

GPQA Diamond: 87.4%Tied within margin

$0.28$1.20
19Hunyuan Hy3 Preview
PreviewReasoningTools

GPQA Diamond: 87.2%

$0.07$0.26
20GLM-5.1
ReasoningTools

GPQA Diamond: 86.2%

$1.05$3.50

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • #4Trinity-Large-Thinking

    Arcee AI's flagship 400B sparse MoE reasoning model with 13B active parameters per token. Trained on 20T tokens with a STEM-focused curriculum. Designed for agentic workflows, chain-of-thought reasoning, and long-context tasks up to 256K tokens (BF16 API). Open-source under Apache 2.0. Available via Arcee AI API.

    89.2%

    GPQA Diamond

  • #5Qwen3.5-Plus

    Qwen3.5-Plus is the flagship commercial API model of the Qwen3.5 native vision-language series, delivering outstanding performance comparable to state-of-the-art models with significant leaps in both pure-text and multimodal capabilities compared to the Qwen3 series.

    88.4%

    GPQA Diamond

  • #6Ring-2.6-1T

    Ring-2.6-1T is InclusionAI's MIT-licensed trillion-parameter MoE reasoning model for agent workflows, engineering tasks, scientific analysis, and enterprise automation. It supports high and xhigh reasoning effort modes and entered OpenRouter's Programming top 10 in the 2026-05-18 audit.

    88.27%

    GPQA Diamond