Cheapest LLM APIs You Can Call Right Now (2026)

Last refreshed 2026-09-30. Next refresh: weekly.

The cheapest LLM APIs you can call today, ranked by input price with a quality score beside each so you see the trade-off.

Ranking rule. Use this page when token price is the first constraint and quality still matters. The rows below exclude zero-dollar tiers and surface a quality watermark beside tracked input prices.
Related intent. Need no-cost options? Compare the free-model leaderboard separately from paid API pricing.

Verdict

Use Granite 4.0 H Micro for low-cost API calls today.

Ling 3.0 Flash is the runner-up: $0.017 vs $0.021 on Input $/1M.

Researched todayWhy this pickMethodology

How we rank

Cheapest LLM APIs stay a strict price board, with a quality watermark so low-cost rows do not hide weak benchmark coverage.

  1. Eligibility — Chat-completion style APIs with positive priced tiers (we exclude zero-dollar rows here because `/best/free` covers that intent).
  2. Primary ranking — Ascending input $/1M tokens using consolidated provider data.
  3. Quality watermark — Rows show the first available MMLU or GPQA Diamond score as a capability check, but the score does not change cheap-page order.
  4. Variant collapse — We keep one row per model family (`familySlug` + parameter tier). When headline scores tie within ±0.5 pt (±10 Elo on Chatbot Arena), we pick the canonical SKU by lowest tracked input price, then GA over preview or limited access, then newest `release`. A folded sibling within the benchmark noise band can show a "Tied within margin" chip on that score cell.
  5. Coverage caveat — Ultra-cheap models may lack frontier benchmarks — pair this list with `/best/coding` or another task board for quality guardrails.
#ModelInput $/1MOutput $/1M
1Ling-2.6-Flash
Tools

Quality watermark: —

$0.01$0.03
2Granite 4.0 H Micro

Quality watermark: —

$0.02$0.11
3Aleph Alpha Luminous Base

Quality watermark: —

$0.02$0.06
4Together AI - Gemma 3n-e4B
Tools

Quality watermark: —

$0.02$0.04
5Gemma 3n 4B (free)

Quality watermark: —

$0.02$0.04
6Llama 3 8B Instruct

Quality watermark: MMLU 76.9%

$0.02$0.04
7Llama 3.1 8B Instruct

Quality watermark: —

$0.02$0.05
8Mistral NeMo (2407)

Quality watermark: —

$0.02$0.03
9Ling 3.0 Flash
ReasoningTools

Quality watermark: —

$0.02$0.06
10Llama 3.2 1B Instruct

Quality watermark: MMLU 49.3%

$0.03$0.10
11ERNIE Lite Pro

Quality watermark: —

$0.03$0.06
12GLiNER2.5-Decide

Quality watermark: —

$0.03Free
13Granite 3.3 8B Instruct
Tools

Quality watermark: —

$0.03$0.25
14KAT Coder Pro V1
Tools

Quality watermark: —

$0.03$1.20
15LFM2-24B-A2B
Tools

Quality watermark: —

$0.03$0.12
16Llama 3.2 3B Instruct

Quality watermark: —

$0.03$0.05
17Qwen2.5-7B-Instruct

Quality watermark: MMLU 81.2%

$0.03$0.03
18Amazon Nova Micro

Quality watermark: —

$0.04$0.14
19AutoGLM Phone 9B Multilingual
VisionTools

Quality watermark: —

$0.04$0.14
20Nova Micro

Quality watermark: —

$0.04$0.14

Honorable mentions

Next seats in this ranking. Lines below are from each model's stored description in LLMReference seed data—spot-check the model page before relying on a capability claim.
  • IBM Granite 3.3 8B with improved reasoning capabilities. Part of IBM's enterprise-focused Granite model family optimized for instruction following.

    $0.030

    Input $/1M

  • #5KAT Coder Pro V1

    KAT-Coder-Pro V1 is Kwaipilot's first high-performance agentic coding model, achieving 73.4% on SWE-Bench Verified. It uses a Mixture-of-Experts architecture with approximately 72B active parameters. Context: 256K tokens. Proprietary, available via Vercel AI Gateway and OpenRouter.

    $0.030

    Input $/1M

  • #6LFM2-24B-A2B

    LFM2-24B-A2B is the largest model in LiquidAI's LFM-2 hybrid-architecture family, designed for efficient on-device deployment with 24B total parameters and 2B active per token.

    $0.030

    Input $/1M