LLM Reference
active

GAOKAO

Metric: Accuracy (higher is better)Introduced: 2023

Evaluation benchmark using questions from China's college entrance examination (Gaokao) across math, language arts, science, and social studies.

Models ranked

11

tracked on this benchmark

Score band

72.2 – 21.0

best → lowest tracked

Snapshot trend

need ≥2 snapshots

Leaderboard

Tracked models ranked by Accuracy (higher is better).

Compare candidates
#Model variant and provenanceScore
1
GPT-4
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
72.2
2
Gemini 1.0 Pro
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
57.9
3
ERNIE Bot
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
56.6
4
GPT-3.5 Turbo
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
53.2
5
Baichuan 2 13B Chat
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
43.9
6
ChatGLM2-6B
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
42.7
7
Baichuan 2 7B Chat
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
40.5
8
ChatGLM-6B
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
30.8
9
Baichuan 2 7B
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
27.2
10
LLaMA 7B
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
21.1
11
Vicuna 7B
Version: zero-shot, objective-accuracyHarness: Not recordedEvaluator: Not recordedObserved: May 20, 2023Confidence: Not recordedSource
21.0

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Last reviewed: May 1, 2026