Yes/no questions on medical research abstracts, scored on the quality of the probability, not just the answer.
Noul is the yes/no primitive: the model returns one probability between 0 and 1. Jevals tests it on PubMedQA, where each question asks whether a research abstract supports a claim. It is a proxy for verification, compliance, and guardrail checks.
Each item has a human gold label. Models return a probability for every option, and the Decision Score measures skill over the label prior using the Brier score (or ranked probability score for ordered scales). 100 is perfect, 0 matches guessing the label base rates, and negative means worse than guessing. Each run is repeated five times.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Jev 1.13 | TypeSafe AI | Closed | 69.0 |
No models in this category.
Not enough scored models yet.
Jev 1.13 scores 69.03 and Gemini 3.8 Flash leads at 72.98 as of 18 September 2026.
Based on score correlations across our database.