Rate how helpful a response is on an ordered scale. The hardest primitive: most models barely beat guessing.
Score is the graded primitive: the model rates content on an ordered scale and returns a probability for each level. Jevals tests it on human helpfulness ratings from HelpSteer2. It is a proxy for lead scoring, quality review, and ranking.
Each item has a human gold label. Models return a probability for every option, and the Decision Score measures skill over the label prior using the Brier score (or ranked probability score for ordered scales). 100 is perfect, 0 matches guessing the label base rates, and negative means worse than guessing. Each run is repeated five times.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Jev 1.13 | TypeSafe AI | Closed | 9.2 |
No models in this category.
Not enough scored models yet.
Human helpfulness ratings are noisy and hard to predict. Jev 1.13 leads at 9.2, and several LLMs score below zero, which means worse than guessing the base rates.
Based on score correlations across our database.