Model rankingsIndices are sourced from Artificial Analysis. WMQ = 50% agentic + 40% coding + 10% speed.Read methodology →
Models ranked by AA benchmark indices. All metric values carry an AA prefix; a — means the index is not disclosed for that model. Scores computed 2026-06-23.
| # | Provider | Confidence | ||
|---|---|---|---|---|
| 1 | Claude Fable 5 | Anthropic | 73.9 | Verified |
| 2 | GPT-5.5 | OpenAI | 72.4 | Verified |
| 3 | Claude Opus 4.8 | Anthropic | 71.9 | Verified |
| 4 | Claude Opus 4.7 | Anthropic | 71.2 | Verified |
| 5 | GPT-5.4 | OpenAI | 69 | Verified |
| 6 | Gemini 3.5 Flash | 68.1 | Verified | |
| 7 | Gemini 3.1 Pro Preview | 66.9 | Verified | |
| 8 | Claude Sonnet 4.6 | Anthropic | 61.7 | Verified |
| 9 | Kimi K2.7 Code | Kimi | 59.7 | Verified |
| 10 | GPT-5.4 mini | OpenAI | 55.5 | Verified |