Model efficiency index
Models ranked by Intel · t/s ÷ $/100 tasks — a composite that rewards intelligence, throughput, and cost-efficiency simultaneously. Sources: Artificial Analysis (intel, speed, verbosity) + OpenRouter API pricing, verified 2026-06-20.
| # | Model | Lab | Intel | t/s | $/100T | Intel/$100T | Intel·t/s/$100T | Efficiency |
|---|---|---|---|---|---|---|---|---|
| 1 | GPT-OSS 120B | US | 24 | 344 | $0.687 | 34.9 | 12,017 | |
| 2 | Mimo V2.5~ | CN | 40 | 82 | $0.659 | 60.7 | 4,978 | |
| 3 | DS V4 Flash | CN | 40 | 108 | $0.879 | 45.5 | 4,915 | |
| 4 | MiniMax M2.5~ | CN | 34 | 209 | $1.560 | 21.8 | 4,555 | |
| 5 | Gemma 4 31B | US | 29 | 35 | $0.531 | 54.6 | 1,911 | |
| 6 | Grok 4.3 | US | 38 | 135 | $4.435 | 8.6 | 1,157 | |
| 7 | DS V4 Pro | CN | 44 | 91 | $3.525 | 12.5 | 1,136 | |
| 8 | MiMo-V2.5-Pro | CN | 42 | 53 | $2.046 | 20.5 | 1,088 | |
| 9 | MiniMax-M2.7 | CN | 38 | 49 | $1.990 | 19.1 | 936 | |
| 10 | MiniMax-M3 | CN | 44 | 57 | $3.108 | 14.2 | 807 | |
| 11 | GPT-5.4 Nano | US | 38 | 162 | $9.021 | 4.2 | 682 | |
| 12 | GLM-5~ | CN | 40 | 77 | $5.448 | 7.3 | 565 | |
| 13 | GLM-5.1 | CN | 40 | 90 | $8.841 | 4.5 | 407 | |
| 14 | Haiku 4.5~ | US | 24 | 89 | $5.730 | 4.2 | 373 | |
| 15 | Kimi K2.5~ | CN | 38 | 55 | $6.450 | 5.9 | 324 | |
| 16 | Kimi K2.6 | CN | 43 | 45 | $12.811 | 3.4 | 151 |
~ = Tok/Task estimated; cost approximate. US = US lab, CN = Chinese lab. Intel = AA Intelligence Index · t/s = AA median output speed · OR pricing.
Methodology
Raw inputs
- Intel
- Artificial Analysis Intelligence Index — composite benchmark score
- t/s
- Artificial Analysis median output speed in tokens per second
- Tok/Task
- AA "Output Tokens per II Task" — median output tokens per benchmark task
- OR In$/1M
- OpenRouter input token price (cache-miss / fresh)
- OR CH$/1M
- OpenRouter cache-hit input price (— = no discount offered)
- OR Out$/1M
- OpenRouter output token price
Cost model — $/Task
Assumes a representative agentic request of 10,000 input tokens, split 70/30 fresh vs. cached:
$/Task = ( 7,000 × OR_In$/1M + 3,000 × OR_CH$/1M ← OR_In if no cache + Tok/Task × OR_Out$/1M ) ÷ 1,000,000
$/100T = $/Task × 100, so dollar amounts are human-readable. No model on OpenRouter charges for cache writes.
Intel / $100T
Intel ÷ $/100T
Intelligence value per $100 spent. Analogous to miles-per-gallon — higher means more intelligence per dollar. Does not account for speed.
Intel · t/s / $100T
(Intel × t/s) ÷ $/100T
The composite score. Rewards models that are simultaneously smart, fast, and cheap. Doubling speed at constant cost and intelligence doubles the score. Doubling cost halves it. Ranking by this metric is the primary sort.
Key findings
- →GPT-OSS 120B ranks #1 on the composite by a 2.4× margin, driven entirely by its 344 t/s throughput — fastest in the set. Trade-off: lowest intelligence score (24).
- →Mimo V2.5 wins the pure Intel/$100T race and holds #2 on the composite. Balanced across all three dimensions: intel 40, 82 t/s, $0.659/100T.
- →Gemma 4 31B is the cheapest model ($0.531/100T) and ranks #2 on Intel/$100T, but falls to #5 on the composite — 35 t/s is the anchor.
- →Cliff at rank 4→5: MiniMax M2.5 (rank 4, 4,555) is more than double Gemma (rank 5, 1,912). Practical panel candidates live in the top 4.
- →Haiku 4.5 is the worst US-lab model on both metrics — premium-priced at $5.73/100T with the same intel as GPT-OSS 120B but ¼ the throughput.