LLM Reference
IFEvalactiveGeneral

IFEval: Instruction-Following Evaluation

Metric: Instruction Following Score (higher is better)Introduced: 2023

IFEval measures LLM ability to follow verifiable formatting instructions (e.g., output length, keyword inclusion, capitalization) with exact-match checking.

Models ranked

12

tracked on this benchmark

Score band

94.8 – 38.5

best → lowest tracked

Snapshot trend

+3.87

Apr 29 → Jun 26 · 1 models

Leaderboard

Tracked models ranked by Instruction Following Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Agents-A1
Version: IFEval accuracy, vendor-reported on Agents-A1 model cardHarness: Not recordedEvaluator: Not recordedObserved: Jun 26, 2026Confidence: Not recordedSource

Notes: Official Intern Science model card reports Agents-A1 at 94.82 on IFEval. Evaluator/source: Intern Science self-reported model-card/project benchmark table. Public harness details are not fully reproducible from the seed row. Confidence: medium for reported value, low-medium for reproducibility. Recommended seed value: 94.82.

94.8
2
Qwen3.5-397B-A17B
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 19, 2026Confidence: Not recordedSource
92.6
3
GPT-5.5
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 29, 2026Confidence: Not recordedSource
92.1
4
LFM2.5 8B A1B
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: May 28, 2026Confidence: Not recordedSource
91.8
5
Kimi K2.6
Version: Not recordedHarness: Not recordedEvaluator: Not recordedObserved: Apr 29, 2026Confidence: Not recordedSource
89.8
6
Llama 3 70B Instruct
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
77.8
7
Gemma 2 9B Instruct
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
65.5
8
Llama 3 8B Instruct
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
59.5
9
Qwen2-7B-Instruct
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
57.8
10
Phi-3 Mini 4k
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
45.0
11
Gemma 7B Instruct
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
42.6
12
Mistral 7B Instruct v0.3
Version: v2Harness: Not recordedEvaluator: Not recordedObserved: Apr 14, 2026Confidence: Not recordedSource
38.5

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Related benchmarks

Last reviewed: Apr 15, 2026