The main community board for System One models. One composite of intelligence, calibration, speed, and cost.
JevBench scores decision models on four axes: how often they are right (intelligence), how honest their probabilities are (calibration), how fast they answer, and what each decision costs. The composite rewards models that are good enough and cheap enough to run inside a hot path, which is where System One models are used.
Every model answers the same public and sealed items. Each axis is scaled to 0-100 and combined into one composite. Self-hosted latency is adjusted so local and API models compare fairly. The sealed set limits overfitting to published items.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | decider-4b v2 | Mapika | Open | 64.1 |
| 02 | Jev 1.13 | TypeSafe AI | Closed | 63.3 |
| 03 | Cygnet | blockbrain | Open | 61.8 |
| 04 | Hopper | HopitAI | Open | 59.4 |
| 05 | Winnow-12B | EldanRing | Open | 55.6 |
| 06 | reflex 4B | kshetrajna12 | Open | 54.0 |
| 07 | djev | Maisa | Open | 52.2 |
| 08 | Jev-Omni | akhilaaa3 | Open | 51.3 |
| Ad | ||||
| 09 | Metask-Jev-4B | Wayfind (Metask AI) | Open | 47.8 |
| 10 | SemIf | TheoLeeCJ | Open | 47.7 |
| 11 | Jobe | MantisShrimpdev | Open | 46.9 |
| 12 | system-one-open | mithalouni | Open | 45.1 |
| 13 | spark-s1-4b v6 | Abhishek Rai (Nokast) | Open | 44.6 |
| 14 | Malkuth-4B | newfull5 | Open | 44.5 |
| 15 | decider-35b-a3b | Mapika | Open | 41.2 |
3 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.
As of 24 September 2026, decider-4b v2 is first at 64.13 and Jev 1.13 is second at 63.29. Jev still leads on the intelligence axis.
The composite includes speed and cost. A frontier LLM can top the intelligence axis and still land far down the composite because each decision is slower and more expensive.
Based on score correlations across our database.