CodeSmart

Model efficiency index

Models ranked by Intel · t/s ÷ $/100 tasks — a composite that rewards intelligence, throughput, and cost-efficiency simultaneously. Sources: Artificial Analysis (intel, speed, verbosity) + OpenRouter API pricing, verified 2026-06-20.

#ModelLabIntelt/s$/100TIntel/$100TIntel·t/s/$100TEfficiency
1GPT-OSS 120BUS24344$0.68734.912,017
2Mimo V2.5~CN4082$0.65960.74,978
3DS V4 FlashCN40108$0.87945.54,915
4MiniMax M2.5~CN34209$1.56021.84,555
5Gemma 4 31BUS2935$0.53154.61,911
6Grok 4.3US38135$4.4358.61,157
7DS V4 ProCN4491$3.52512.51,136
8MiMo-V2.5-ProCN4253$2.04620.51,088
9MiniMax-M2.7CN3849$1.99019.1936
10MiniMax-M3CN4457$3.10814.2807
11GPT-5.4 NanoUS38162$9.0214.2682
12GLM-5~CN4077$5.4487.3565
13GLM-5.1CN4090$8.8414.5407
14Haiku 4.5~US2489$5.7304.2373
15Kimi K2.5~CN3855$6.4505.9324
16Kimi K2.6CN4345$12.8113.4151

~ = Tok/Task estimated; cost approximate. US = US lab, CN = Chinese lab. Intel = AA Intelligence Index · t/s = AA median output speed · OR pricing.

Methodology

Raw inputs

Intel
Artificial Analysis Intelligence Index — composite benchmark score
t/s
Artificial Analysis median output speed in tokens per second
Tok/Task
AA "Output Tokens per II Task" — median output tokens per benchmark task
OR In$/1M
OpenRouter input token price (cache-miss / fresh)
OR CH$/1M
OpenRouter cache-hit input price (— = no discount offered)
OR Out$/1M
OpenRouter output token price

Cost model — $/Task

Assumes a representative agentic request of 10,000 input tokens, split 70/30 fresh vs. cached:

$/Task =
  ( 7,000 × OR_In$/1M
  + 3,000 × OR_CH$/1M   ← OR_In if no cache
  + Tok/Task × OR_Out$/1M
  ) ÷ 1,000,000

$/100T = $/Task × 100, so dollar amounts are human-readable. No model on OpenRouter charges for cache writes.

Intel / $100T

Intel ÷ $/100T

Intelligence value per $100 spent. Analogous to miles-per-gallon — higher means more intelligence per dollar. Does not account for speed.

Intel · t/s / $100T

(Intel × t/s) ÷ $/100T

The composite score. Rewards models that are simultaneously smart, fast, and cheap. Doubling speed at constant cost and intelligence doubles the score. Doubling cost halves it. Ranking by this metric is the primary sort.

Key findings

  • GPT-OSS 120B ranks #1 on the composite by a 2.4× margin, driven entirely by its 344 t/s throughput — fastest in the set. Trade-off: lowest intelligence score (24).
  • Mimo V2.5 wins the pure Intel/$100T race and holds #2 on the composite. Balanced across all three dimensions: intel 40, 82 t/s, $0.659/100T.
  • Gemma 4 31B is the cheapest model ($0.531/100T) and ranks #2 on Intel/$100T, but falls to #5 on the composite — 35 t/s is the anchor.
  • Cliff at rank 4→5: MiniMax M2.5 (rank 4, 4,555) is more than double Gemma (rank 5, 1,912). Practical panel candidates live in the top 4.
  • Haiku 4.5 is the worst US-lab model on both metrics — premium-priced at $5.73/100T with the same intel as GPT-OSS 120B but ¼ the throughput.