Bounded decisions taken from real software: routing complaints, picking clauses, selecting tools, reading receipts.
Decision Bench tests the decisions a product team would actually hand to a model: which queue a complaint goes to, which contract clause applies, which tool to call, or what a receipt says. It reports accuracy, calibration, latency, and cost for each model.
Every row has a closed option set and one correct answer. Accuracy is the share of rows answered correctly. Calibration is reported as ECE and Brier score, latency as the median per row over the network, and cost per 1,000 rows.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Jev 1.13 | TypeSafe AI | Closed | 92.4% |
| 02 | Sage | Levanto | Closed | 92.1% |
| 03 | Tev1-4B Experimental | Together AI | Open | 85.4% |
| 04 | Laya | Convai Innovations | Open | 52.8% |
2 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.
Not enough scored models yet.
Jev 1.13 scores 92.4% at a 439 ms median. The best LLMs score 93% to 94% but cost more per decision.
Based on score correlations across our database.