Vendor-run workflows for routing, scoring, and moderation. The numbers you will see in every Jev launch post.
TypeSafe runs Jev and several LLMs on four workflows that look like production use: routing, scoring, and moderation-style judgments. It reports accuracy next to cost and time per case.
Accuracy is the mean across the four workflows. Labels come from a consensus of frontier LLMs rather than human annotators.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Jev 1.13 | TypeSafe AI | Closed | 67.8% |
No models in this category.
Not enough scored models yet.
Use them for speed and cost, which are easy to verify. For accuracy, check an independent board like Decision Bench or Jevals as well.
Based on score correlations across our database.