Granite 4.0 H Micro
- Input $/1M
- $0.017
- Output (from)
- $0.112 / 1M
Last refreshed 2026-09-30. Next refresh: weekly.
The cheapest LLM APIs you can call today, ranked by input price with a quality score beside each so you see the trade-off.
Verdict
Ling 3.0 Flash is the runner-up: $0.017 vs $0.021 on Input $/1M.
Cheapest LLM APIs stay a strict price board, with a quality watermark so low-cost rows do not hide weak benchmark coverage.
| # | Model | Input $/1M | Output $/1M | |
|---|---|---|---|---|
| 1 | Ling-2.6-Flash Tools Quality watermark: — | $0.01 | $0.03 | |
| 2 | Granite 4.0 H Micro Quality watermark: — | $0.02 | $0.11 | |
| 3 | Aleph Alpha Luminous Base Quality watermark: — | $0.02 | $0.06 | |
| 4 | Together AI - Gemma 3n-e4B Tools Quality watermark: — | $0.02 | $0.04 | |
| 5 | Gemma 3n 4B (free) Quality watermark: — | $0.02 | $0.04 | |
| 6 | Llama 3 8B Instruct Quality watermark: MMLU 76.9% | $0.02 | $0.04 | |
| 7 | Llama 3.1 8B Instruct Quality watermark: — | $0.02 | $0.05 | |
| 8 | Mistral NeMo (2407) Quality watermark: — | $0.02 | $0.03 | |
| 9 | Ling 3.0 Flash ReasoningTools Quality watermark: — | $0.02 | $0.06 | |
| 10 | Llama 3.2 1B Instruct Quality watermark: MMLU 49.3% | $0.03 | $0.10 | |
| 11 | ERNIE Lite Pro Quality watermark: — | $0.03 | $0.06 | |
| 12 | GLiNER2.5-Decide Quality watermark: — | $0.03 | Free | |
| 13 | Granite 3.3 8B Instruct Tools Quality watermark: — | $0.03 | $0.25 | |
| 14 | KAT Coder Pro V1 Tools Quality watermark: — | $0.03 | $1.20 | |
| 15 | LFM2-24B-A2B Tools Quality watermark: — | $0.03 | $0.12 | |
| 16 | Llama 3.2 3B Instruct Quality watermark: — | $0.03 | $0.05 | |
| 17 | Qwen2.5-7B-Instruct Quality watermark: MMLU 81.2% | $0.03 | $0.03 | |
| 18 | Amazon Nova Micro Quality watermark: — | $0.04 | $0.14 | |
| 19 | AutoGLM Phone 9B Multilingual VisionTools Quality watermark: — | $0.04 | $0.14 | |
| 20 | Nova Micro Quality watermark: — | $0.04 | $0.14 |
IBM Granite 3.3 8B with improved reasoning capabilities. Part of IBM's enterprise-focused Granite model family optimized for instruction following.
$0.030
Input $/1M
KAT-Coder-Pro V1 is Kwaipilot's first high-performance agentic coding model, achieving 73.4% on SWE-Bench Verified. It uses a Mixture-of-Experts architecture with approximately 72B active parameters. Context: 256K tokens. Proprietary, available via Vercel AI Gateway and OpenRouter.
$0.030
Input $/1M
LFM2-24B-A2B is the largest model in LiquidAI's LFM-2 hybrid-architecture family, designed for efficient on-device deployment with 24B total parameters and 2B active per token.
$0.030
Input $/1M