Tests whether a model obeys precise formatting, length, and constraint rules.
IFBench measures whether a model actually follows precise instructions. Each prompt has verifiable rules: respond in exactly four sentences, include the word "elephant" twice, output JSON with these exact keys. The benchmark cares about obedience, not creativity, which makes it a strong signal for production reliability.
Each prompt has one or more programmatic verifiers. The model’s output is checked rule-by-rule. The score is the fraction of rules satisfied, averaged across prompts.
| # | Model | Lab | Source | Score |
|---|---|---|---|---|
| 01 | Grok 4.20 Beta 0309 Reasoning | xAI | Closed | 82.9 |
| 02 | MiniMax M3 | MiniMax | Open | 82.9 |
| 03 | Grok 4.3 | xAI | Closed | 81.3 |
| 04 | Grok 4.3 beta | xAI | Closed | 81.3 |
| 05 | Qwen 3.7 Max | Alibaba | Closed | 80.5 |
| 06 | MiMo-V2.5-Pro | Xiaomi | Closed | 79.9 |
| 07 | DeepSeek-V4-Flash | DeepSeek | Open | 79.2 |
| 08 | Qwen3.5-397B-A17B | Alibaba | Open | 78.8 |
| Ad | ||||
| 09 | Gemini 3 Flash (Thinking Minimal) | Closed | 78.0 | |
| 10 | Gemini 3.1 Flash Lite Preview | Closed | 77.2 | |
| 11 | Gemini 3.1 Pro Preview | Closed | 77.1 | |
| 12 | Qwen3.6 Max Preview | Alibaba | Closed | 76.6 |
| 13 | DeepSeek-V4-Pro | DeepSeek | Open | 76.5 |
| 14 | Gemini 3.5 Flash | Closed | 76.3 | |
| 15 | GLM-5.1 | Z.ai | Open | 76.3 |
Real applications need predictable output: JSON with specific keys, summaries with specific lengths, responses in specific languages. A high IFBench score means the model will hold up when the prompt has hard constraints.
Based on score correlations across our database.
74 model(s) with undisclosed parameter counts not shown. Most closed-source labs do not publish model size.