PROJ-004◇ RESEARCH● PUBLISHEDMay 2026
Token Language Tax
Danish costs 50% more than English to say the same thing. This project measures the surcharge across 7 models and 11 languages
#Tokenomics#Sovereign AI#Benchmark
/ KEY FINDINGS
+55%
average output tax across 11 non-English languages
7x11
models tested across languages, input and output
+40 → +72%
spread in output tax depending on model choice alone
$1.68
cost of entire experiment
/ background
“Learning objective: find out whether the token language tax that research has measured on input also applies on output, where the money actually is and what it looks like across most used models”
Token billing has become the default pricing model for AI, and almost everyone treats the price per token as the price. It is only half the equation. The other half is how many tokens your text turns into, and that depends on the language you write in. English is what tokenizers were built for, and every other language pays a surcharge that cannot be optimised away. Research had already measured this on the input side. Input is the cheap side. So I set out to measure the expensive one.
/ artifacts
/ related writing
/ keep going
2 more projects