Pick one intent out of 77, scored on how much better the probabilities are than guessing base rates.
The Choice primitive picks one option from a list. Jevals tests it on banking support messages where the model must choose the right intent out of 77. This is the closest public proxy for ticket routing and intent classification.
Each item has a human gold label. Models return a probability for every option, and the Decision Score measures skill over the label prior using the Brier score (or ranked probability score for ordered scales). 100 is perfect, 0 matches guessing the label base rates, and negative means worse than guessing. Each run is repeated five times.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Jev 1.13 | TypeSafe AI | Closed | 67.8 |
No models in this category.
Not enough scored models yet.
It is skill over the label prior. Jev 1.13 scores 67.78, Gemini 3.8 Flash scores 74.11 as of 18 September 2026.
Based on score correlations across our database.