Artificial Analysis composite score that blends a dozen reasoning, coding, and math benchmarks into a single number.
The Intelligence Index is Artificial Analysis’s way of giving every LLM a single comparable score across reasoning, coding, math, and long-context ability. They run a dozen public benchmarks under the same prompt and scoring settings, normalize each one to a 0–100 range, and average them with fixed weights. It is most useful as a quick way to spot the top tier without studying every sub-benchmark.
Artificial Analysis runs each model on the underlying evaluations using a shared prompting harness, normalizes each per-benchmark score, and aggregates them with published weights. The composite is recomputed whenever a new model lands or a benchmark is added.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Claude Opus 5.5 | Anthropic | Closed | 57.6 |
| 02 | Claude Sonnet 5.5 | Anthropic | Closed | 56.0 |
| 03 | Gemini 4 Argon | Closed | 52.6 | |
| 04 | GPT-6.1 Sol | OpenAI | Closed | 51.8 |
| 05 | Claude Fable 5 | Anthropic | Closed | 49.6 |
| 06 | GPT-6 Sol | OpenAI | Closed | 47.6 |
| 07 | GPT-5.6 Sol | OpenAI | Closed | 47.0 |
| 08 | Qwen3.8-Max | Alibaba | Open | 45.4 |
| Ad | ||||
| 09 | Kimi K3 | Moonshot AI | Open | 43.6 |
| 10 | GPT-5.6 Terra | OpenAI | Closed | 42.1 |
| 11 | Claude Opus 4.8 | Anthropic | Closed | 41.8 |
| 12 | Ling-3.1-flash | Ant Group | Closed | 41.1 |
| 13 | Claude Opus 4.7 Thinking | Anthropic | Closed | 40.7 |
| 14 | Claude Opus 4.7 | Anthropic | Closed | 40.7 |
| 15 | DeepSeek-V4.1-Flash | deepseek-ai | Open | 39.5 |
Our average pulls in every benchmark we track, including Arena scores and HF leaderboards. The AA Intelligence Index is Artificial Analysis’s own composite over their evaluation suite. Use it as a second opinion, not a replacement.
Yes. Every benchmark inside the Intelligence Index has its own deep-dive page, so you can drill into MMLU-Pro, GPQA Diamond, LiveCodeBench, and the others independently.
Based on score correlations across our database.
93 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.