JevBench
Benchmark Heaven modality board for typed-decision / SystemOne-class models. Official frozen revisions published at benchmarkheaven.com/api/jevbench/<version>. Not affiliated with TypeSafe AI.
Models ranked
7
tracked on this benchmark
Score band
71.4 – 42.2
best → lowest tracked
Snapshot trend
-8.39
Oct 2 → Oct 5 · 2 models
Leaderboard
Tracked models ranked by Score (higher is better).
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 1 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Individual publisher; sole notability signal is this board.
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 3 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Measured as 'wity-1 (Wity, reasoning auto)' — the operator's main row on its hosted API (API default is reasoning off; off-row 40.53 and always-row 70.97 are listed but unranked). BH gives no per-row measurement date (last_measured_on null, measurement_date_status 'unknown'); timestamp = artifact scoring date ('scored 2026-10-05 09:06:16 UTC' in source_note).
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 4 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Card discloses public JevBench results influenced checkpoint selection.
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 6 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Measured as 'Winnow-12B Q8' (Q8_0 GGUF) on an evaluator-owned RTX6000 pod. Card's author-run JevBench public-subset numbers are not this score.
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 9 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Measured as 'Sage 1.3.0' (model pin levanto-sage-v1.3), hosted API, text only. Supplementary A3 run (300 sealed + P300) scored 5 Oct 2026; scorer source sha256 7ac71f7459bcb1a2062ce818c8b23dc38962793293d944abc787c06066bd2d40.
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 13 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Individual publisher (Chinh Nguyen); family quyet included in wave 1. Capability rank 37 of 92.
Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 14 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Interfaze first-party LoRA adapter on Qwen3.5-4B; BH pin interfaze-ai/lev@7bdc748d + Abhinavexists/lev@cf104b69; 1475/1500 answered OK.
How to read this benchmark
This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.
Trust this score when
- There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
- The model list covers the same version family you can actually deploy today.
- Top candidates overlap with your required routing and feature requirements.
Be cautious when
- There is only one benchmark snapshot or the dataset appears stale.
- The benchmark metric direction is opposite of your decision objective.
- The score difference between options is narrow and likely within implementation variance.
Last reviewed: Oct 6, 2026