Twenty-five hundred expert-written questions designed to be unsolvable by any current AI system, across every academic field.
HLE is the hardest broad-coverage benchmark in public use. The questions were crowdsourced from a thousand subject experts and explicitly filtered to defeat frontier models at the time of release. About 14% are multimodal, requiring image understanding. HLE measures how close a model is to the ceiling of human expert knowledge, and how much further the field still has to go.
Questions are short-answer or multiple choice. Scoring is exact-match for short-answer items and accuracy for multiple choice. Many questions include an image or diagram, so a fair score requires a multimodal model.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Claude Fable 5 | Anthropic | Closed | 55.5 |
| 02 | GPT-5.6 Sol | OpenAI | Closed | 49.5 |
| 03 | Claude Opus 4.8 | Anthropic | Closed | 48.7 |
| 04 | Gemini 3.1 Pro Preview | Closed | 47.0 | |
| 05 | GPT-5.5 | OpenAI | Closed | 45.8 |
| 06 | GPT-5.4 | OpenAI | Closed | 43.7 |
| 07 | GPT-5.4 High | OpenAI | Closed | 43.7 |
| 08 | GPT-5.6 Terra | OpenAI | Closed | 42.9 |
| 09 | Grok 4.5 | xAI | Closed | 42.7 |
| 10 | Gemini 3.5 Flash | Closed | 42.7 | |
| 11 | Claude Opus 4.7 Thinking | Anthropic | Closed | 42.3 |
| 12 | Claude Opus 4.7 | Anthropic | Closed | 42.3 |
| 13 | GLM-5.2 | Z.ai | Open | 41.1 |
| 14 | Qwen 3.7 Max | Alibaba | Closed | 40.5 |
| 15 | Claude Opus 4.6 (Thinking) | Anthropic | Closed | 39.9 |
82 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.
The authors built it to be a benchmark that humanity might run out of room to keep designing. The questions are at or beyond the level of a top expert in each field, which makes it useful even as models improve dramatically.
Most strong open-weight models score under 10%. Frontier closed models in 2026 are between 20% and 35%. Even the best models are far from human-expert performance, which is the explicit design goal.
About 14% of items have images. Text-only models can still be evaluated on the remaining 86%, but the official score assumes full multimodal capability. Compare like with like when reading leaderboards.
Based on score correlations across our database.