LLM Reference

Granite 4.1 8B

Released
2026-04-29
Last refreshed
2026-06-15
Status
Researched 128d ago
Open sourceCommercial use: permittedCodingAgentsLong contextClassificationJSON / Tool use

Granite 4.1 8B is worth evaluating for coding, agents, and long context when its provider route and context window match the workload.

Use it for

  • Teams evaluating coding, agents, and long context
  • Workloads that can use a 131k context window
  • Buyers comparing 1 tracked provider route

Do not use it for

  • Vision or document-understanding workloads
  • Strict JSON or tool-calling flows
Specifications
Released
2026-04-29
Context
131k
Parameters
8B
Architecture
Decoder Only
Openness
Open source
License
Apache 2.0OSI-approvedCommercial use: permitted
Weights
Available
Code
Unknown
Created by

Creating reliable and adaptable AI solutions

Armonk, New York, United States
Founded 1945
Website
Pricing
Output / 1M
$0.100
Input / 1M
$0.050

Cheapest of 1 route · OpenRouter

About

IBM Granite 4.1 8B is a dense decoder-only transformer instruct model with 40 layers, 4096 embedding size, GQA (32 attention heads, 8 KV heads). Supports multilingual dialog (12 languages), code with FIM, tool-calling/function-calling, RAG, and summarization. Trained on NVIDIA GB200 NVL72 cluster. Apache 2.0. Benchmarks: MMLU 73.84, HumanEval 85.37, GSM8K 92.49, BFCL v3 68.27.

Top use-case fit: coding, agents, and build tasks

Coding

Q/$ A

2 relevant benchmarks in the decision map.

Agents

Q/$ A

1 relevant benchmark in the decision map.

Long context

Included by capability and metadata signals in the decision map.

Provider price ladder

Compare API pricing across 1 providers for input and output tokens, batch, and cached reads when available.

ProviderInput / 1MOutput / 1MRoute
OpenRouter$0.050$0.100
Serverless

Capabilities

No model capability flags are currently sourced.

Benchmark peer barsfor Coding

Benchmark scores(7)

Scores are benchmark-specific and are direction-aware: the same numeric gap can mean very different outcomes across suites. Use the leaderboard context and this model's provider route to decide whether the winning margin is meaningful for your workload.
BenchmarkScoreVersionEvaluationSource
Berkeley Function Calling Leaderboard v368.3BFCL v3 (accuracy)Observed 2026-06-07Source
BigCodeBench35.0BigCodeBench (pass@1)Observed 2026-06-07Source
Google-Proof Q&A42.00-shot CoT, instruct model (accuracy)Observed 2026-06-07Source
Grade School Math 8K92.58-shot, instruct model (accuracy)Observed 2026-06-07Source
HumanEval87.2pass@1, instruct model (pass@1)Observed 2026-06-07Source
Massive Multitask Language Understanding73.85-shot, instruct model (accuracy)Observed 2026-06-07Source
MMLU PRO56.05-shot CoT, instruct model (accuracy)Observed 2026-06-07Source

Migration checks

No linked migration route is available for this model yet.