AI NewsModels & agentsAnnouncement
JevBench v1.3.0 scored 52 typed decision systems on 534 decisions, and Jev 1.13.0 came out top
Benchmark Heaven published JevBench v1.3.0 on 21 September, scoring 52 Jev-class typed decision systems on 534 decisions across four difficulty tiers. TypeSafe AI's Jev 1.13.0 leads the score at 74.4, ahead of SemIf at 73.1 and djev at 73.0.
Image: Benchmark Heaven
Why it mattersA team picking a typed decision model now has a public leaderboard with speed, calibration and cost measured on the same 534 decisions, so the choice can be made from numbers rather than from vendor benchmarks.
Picking a typed decision model for production has been guesswork, because every vendor benchmarks against itself. Benchmark Heaven published JevBench v1.3.0 on 21 September, and it scores 52 systems on the same 534 decisions, so the choice can be made from one set of numbers.
The benchmark splits the 534 decisions into four difficulty tiers: 72 easy, 96 standard, 146 judge, and 220 hard. Each system is scored on four axes at 25 percent each, combined as a geometric mean: Intelligence, Calibration, Speed, and Cost. Benchmark Heaven ran the tests from a server in Germany, and self-hosted models were run on RunPod GPU, an H100 NVL, or CPU depending on where the model fits. The scoring harness is open-source.
The top of the leaderboard
TypeSafe AI's Jev 1.13.0 leads at 74.4 overall (Intelligence 85.7, Calibration 82.7, Speed 83.3, Cost 52.0). SemIf follows at 73.1 (Intelligence 79.0, Calibration 72.6, Speed 83.7, Cost 59.5). djev is third at 73.0 (Intelligence 82.7, Calibration 65.4, Speed 91.4, Cost 57.6). The gap between first and third is 1.4 points, and the top three each win on a different axis: Jev on intelligence and calibration, djev on speed, SemIf on cost.
The speed numbers
Benchmark Heaven reports Jev 1.13.0's median latency at 0.65 seconds raw, with p95 at 0.72 seconds raw. Cost is the axis where the leaders lose ground: Jev 1.13.0's cost score of 52.0 is the lowest of its four, so a team paying for every decision on the API will pay more than for a system nearer the top on cost. The full table sorts by any column, and the task-outcome grid shows checkmarks and crosses across 52 systems on 231 public tasks, so a team can see where a system fails rather than only its aggregate score.
What it means for a team choosing one
Vendor benchmarks compare a model to what the vendor already ships. That leaves the reader with no way to see whether a self-hosted open model, or a smaller cheaper hosted one, would do the job. JevBench measures all four axes on the same 534 decisions, so a team can decide by trading intelligence for cost, or speed for calibration, and defend the choice from a public table. The harness is open, and Benchmark Heaven says the protocol is maintained on GitHub, so the numbers can be reproduced. When the same team wants to rescore next quarter, it runs the harness itself.
Source
Primary source: JevBench v1.3.0 leaderboard on Benchmark Heaven, scored 21 September 2026.
This item was written by an AI system from the linked source. Reveneau is responsible for what it publishes.
Get AI News in your inbox
New developer tools, model and agent releases, and how teams are actually using them to release software. Short, and only when there is something worth reading.

