REASON
AI thinking xray 13 models
Projects / Reasoning benchmark
PROJ-004 RESEARCH PUBLISHEDJun - Aug 2026

Reasoning benchmark

Reasoning is 80% of a token budget. This project documents how it is spent.

#reasoning#benchmark#tokenomics
/ Key Findings
60-80%
of tokenspend is consumed by reasoning
13
models tested, single-turn
25-40x
spread in efficiency across panel of models
2x
median variance on billed tokens on identical input
/ background
Learning objective: find out what reasoning looks like across a spread of models, what reasoning costs and how the models behave on everyday tasks

The rise of the reasoning model has brought about a massive increase in tokenspend. For most people it was u model upgrade in their favourite chatbot but for companies reasoning and agentic AI has made the costs skyrocket. Most companies pay the bill, but few have a through understanding of what it covers and if they are paying for. Reasoning is hidden to the average user and when the value is unclear it is hard to compare and consider alternatives. That is why I dedided to map the thinking of 13 different models - to show the elephant inside the snake.

/ stack
grouped by what it's for
Harness
PythonYAMLOpenRouterPinned model versionsJSONL
Analysis
PythonLLM as a judge
Datasite
Static html
/ keep going
2 more projects