Forty public benchmarks recast as typed decisions, scored so that random guessing equals zero.
The Decision Index asks how much better than chance a model decides across intent routing, tool selection, retrieval checks, reasoning, and safety. Chance correction means a task with two options and a task with 150 options count on the same scale.
Each source benchmark is converted into typed decisions with a closed option set. Scores are chance-corrected per task so 0 is random guessing and 100 is perfect, then averaged across areas. Calibration (ECE) is reported alongside but is not part of the headline number.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Jev 1.13 | TypeSafe AI | Closed | 57.9 |
| 02 | Rune 26B-A4B v3 | Invergent | Open | 57.4 |
| 03 | AutoJev-27B | denis-pplx | Open | 56.4 |
| 04 | Winnow-12B | EldanRing | Open | 50.0 |
| 05 | Jevfire | kikoncuo | Open | 49.4 |
| 06 | decider-35b-a3b | Mapika | Open | 47.1 |
| 07 | Decision 1.0 Lux 9B | vLLM Semantic Router | Open | 43.5 |
| 08 | decider-4b v2 | Mapika | Open | 40.7 |
| Ad | ||||
| 09 | Jev-Omni | akhilaaa3 | Open | 40.5 |
| 10 | djev | Maisa | Open | 40.3 |
| 11 | Winnow-E4B | EldanRing | Open | 39.9 |
| 12 | Hopper | HopitAI | Open | 39.7 |
| 13 | Bespoke Nimble 9B v2 | Bespoke Labs | Open | 39.6 |
| 14 | lev | Interfaze | Open | 38.5 |
| 15 | Kev 9B | Jared Palmer | Open | 38.5 |
Jev 1.13 leads at 57.89 on edition 0.2.1. Open models in the 4B to 9B range score roughly 30 to 40. Encoder models under 1B score below 15.
Based on score correlations across our database.
1 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.