active

JevBench

Metric: Score (higher is better)

Benchmark Heaven modality board for typed-decision / SystemOne-class models. Official frozen revisions published at benchmarkheaven.com/api/jevbench/<version>. Not affiliated with TypeSafe AI.

Models ranked

7

tracked on this benchmark

Score band

71.4 – 42.2

best → lowest tracked

Snapshot trend

-8.39

Oct 2 → Oct 5 · 2 models

Leaderboard

Tracked models ranked by Score (higher is better).

Compare candidates
#Model variant and provenanceScore
1
Quyet-1.0-Large
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 4, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 1 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Individual publisher; sole notability signal is this board.

71.4
2
wity-1
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 5, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 3 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Measured as 'wity-1 (Wity, reasoning auto)' — the operator's main row on its hosted API (API default is reasoning off; off-row 40.53 and always-row 70.97 are listed but unranked). BH gives no per-row measurement date (last_measured_on null, measurement_date_status 'unknown'); timestamp = artifact scoring date ('scored 2026-10-05 09:06:16 UTC' in source_note).

70.9
3
Torchcast Decision 12B
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 4, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 4 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Card discloses public JevBench results influenced checkpoint selection.

69.9
4
Winnow-12B
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 2, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 6 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Measured as 'Winnow-12B Q8' (Q8_0 GGUF) on an evaluator-owned RTX6000 pod. Card's author-run JevBench public-subset numbers are not this score.

68.9
5
Levanto Sage
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 5, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 9 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Measured as 'Sage 1.3.0' (model pin levanto-sage-v1.3), hosted API, text only. Supplementary A3 run (300 sealed + P300) scored 5 Oct 2026; scorer source sha256 7ac71f7459bcb1a2062ce818c8b23dc38962793293d944abc787c06066bd2d40.

50.1
6
Quyet-1.0-Medium
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 4, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 13 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Individual publisher (Chinh Nguyen); family quyet included in wave 1. Capability rank 37 of 92.

43.6
7
lev
Version: v1.6.0Harness: Not recordedEvaluator: Not recordedObserved: Oct 4, 2026Confidence: Not recordedSource

Notes: Benchmark Heaven official JevBench v1.6.0 composite (rank 14 of 92). Frozen artifact sha256 b8560f6e00c45d5dfdf254954d2176b015aad279c741873c0ed7345673354866. Interfaze first-party LoRA adapter on Qwen3.5-4B; BH pin interfaze-ai/lev@7bdc748d + Abhinavexists/lev@cf104b69; 1475/1500 answered OK.

42.2

How to read this benchmark

This benchmark scores models where higher is better. Use scores for directional filtering and shortlisting, not universal quality ranking; then validate pricing, context window, provider availability, and fit for your workload.

Trust this score when

  • There is a fresh timestamped snapshot (or multiple snapshots) for this benchmark.
  • The model list covers the same version family you can actually deploy today.
  • Top candidates overlap with your required routing and feature requirements.

Be cautious when

  • There is only one benchmark snapshot or the dataset appears stale.
  • The benchmark metric direction is opposite of your decision objective.
  • The score difference between options is narrow and likely within implementation variance.

Last reviewed: Oct 6, 2026

Resources