Artificial Analysis composite score that blends a dozen reasoning, coding, and math benchmarks into a single number.
The Intelligence Index is Artificial Analysis’s way of giving every LLM a single comparable score across reasoning, coding, math, and long-context ability. They run a dozen public benchmarks under the same prompt and scoring settings, normalize each one to a 0–100 range, and average them with fixed weights. It is most useful as a quick way to spot the top tier without studying every sub-benchmark.
Artificial Analysis runs each model on the underlying evaluations using a shared prompting harness, normalizes each per-benchmark score, and aggregates them with published weights. The composite is recomputed whenever a new model lands or a benchmark is added.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Claude Fable 5 | Anthropic | Closed | 62.1 |
| 02 | GPT-5.6 Sol | OpenAI | Closed | 60.9 |
| 03 | Claude Opus 4.8 | Anthropic | Closed | 57.3 |
| 04 | GPT-5.6 Terra | OpenAI | Closed | 56.6 |
| 05 | GPT-5.5 | OpenAI | Closed | 56.3 |
| 06 | Grok 4.5 | xAI | Closed | 55.8 |
| 07 | Claude Opus 4.7 Thinking | Anthropic | Closed | 55.0 |
| 08 |
Our average pulls in every benchmark we track, including Arena scores and HF leaderboards. The AA Intelligence Index is Artificial Analysis’s own composite over their evaluation suite. Use it as a second opinion, not a replacement.
Yes. Every benchmark inside the Intelligence Index has its own deep-dive page, so you can drill into MMLU-Pro, GPQA Diamond, LiveCodeBench, and the others independently.
Based on score correlations across our database.
| Claude Opus 4.7 |
| Anthropic |
| Closed |
| 55.0 |
| 09 | GPT-5.4 | OpenAI | Closed | 53.1 |
| 10 | GPT-5.4 High | OpenAI | Closed | 53.1 |
| 11 | GLM-5.2 | Z.ai | Open | 52.6 |
| 12 | GPT-5.6 Luna | OpenAI | Closed | 52.3 |
| 13 | Gemini 3.5 Flash | Closed | 52.0 |
| 14 | Gemini 3.1 Pro Preview | Closed | 47.7 |
| 15 | Qwen 3.7 Max | Alibaba | Closed | 46.7 |
85 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.