yield point
AI / LLM Internalsintermediateupdated 2026-08-25

Tokenization and the price of a script

The same sentence costs several times more in one language than another, and roughly a third of that gap was decided by UTF-8 in 2003, before anyone trained a tokenizer.

Everything a model reads and writes is tokens, and tokens are what you are billed for, what the context window is measured in, and what latency is proportional to. So it matters that the number of them is not a property of the text. It is a property of the tokenizer, and tokenizers are trained.

Watch the squares get coarser as merges are learned. Then switch the script to CJK: the same word count, the same frequency shape, three times the bytes, and a vocabulary that was mostly spent on somebody else's language.

Tokens billed, against the raw byte count1,261 / 1,261
a whole worda piece of onea single byte
Tokens per word
5.25
Bytes per token
1.00
Vocabulary
256
UTF-8 cost
1 byte/char
Billing
text it trained on
Merges spent
0 / 300
tick 0 / 420
Break it
Do

The training budget. Vocabulary size is exactly 256 bytes plus this.

How much of the trainer's corpus was written in the script being billed. The rest is Latin.

How top-heavy the word distribution is. Natural language sits near 1.

Safety properties
  • ✓
    Every token decodes back to exactly the bytes it replaced

    held on every tick so far

  • ✓
    Each merge application removes exactly one token

    held on every tick so far

Measured vs theory
MetricSimulatedFormulaError
vocabSize256.000556.00054.0%

It starts as bytes

Byte-pair encoding begins with an alphabet of 256 byte values and nothing else.[1]paperNeural Machine Translation of Rare Words with Subword UnitsSennrich, R. et al., ACL 2016, 2016 Then it repeats one step: find the most frequent adjacent pair in the corpus, and give that pair a new id. Three hundred merges later there are 556 ids, and common sequences have collapsed into single tokens.

That is the whole algorithm. It is running in the widget above: each tick finds the most frequent pair in a training corpus and merges it, then replays that merge onto the text being billed, which is exactly what encoding does.

Drag Merges learned to 0. Every square goes red, one per byte, and tokens per word is just the byte count. That is the floor, and it is where every tokenizer starts.

The first multiplier was fixed in 2003

UTF-8 encodes an ASCII character in one byte, most accented Latin and Cyrillic in two, and CJK and Devanagari in three.[3]RFCRFC 3629: UTF-8, a transformation format of ISO 10646Yergeau, F., 2003 This is not a modelling decision or a vendor’s choice. It is the text encoding the whole industry standardised on, and it was settled long before anyone was billing per token.

So before a tokenizer exists at all, the same sentence in Hindi is roughly three times as many bytes as the same sentence in English. Set Merges learned to 0 and switch scripts to watch that with nothing else in the way.

The second multiplier is the training corpus

Merges are spent on whatever is frequent in the training data. A script that made up a tenth of the corpus wins roughly a tenth of the vocabulary, and everything it sends after that is billed near the raw byte rate.

This is measured, and it is large. Petrov and colleagues found the same content costing up to 15 times more tokens across languages in commercial tokenizers, and traced the effect straight through to API cost, context window, and latency.[2]paperLanguage Model Tokenizers Introduce Unfairness Between LanguagesPetrov, A. et al., NeurIPS 2023, 2023

Move the Share of training text in that script slider with CJK selected. Tokens per word falls sharply. That is real, and it is the argument for training your tokenizer on the distribution you actually serve.[4]paperSentencePiece: A simple and language independent subword tokenizerKudo, T. & Richardson, J., EMNLP 2018, 2018

The part that does not move

Identical settings on both sides: same merge budget, same share of the training corpus, same words drawn from the same frequency shape. The only difference is what UTF-8 charges per character, and it is the difference that survives everything else.

Latin script

One byte per character, and the alphabet the merge table was mostly built from.

Tokens billed, against the raw byte count387 / 1,250
a whole worda piece of onea single byte
Tokens per word
1.61
Bytes per token
3.23
Vocabulary
556
UTF-8 cost
1 byte/char
Billing
text it trained on
Merges spent
300 / 300

CJK script

Three bytes per character, spent before any tokenizer gets a say.

Tokens billed, against the raw byte count1,136 / 3,624
a whole worda piece of onea single byte
Tokens per word
4.73
Bytes per token
3.19
Vocabulary
556
UTF-8 cost
3 bytes/char
Billing
text it trained on
Merges spent
300 / 300
Measured at tick 310 - both sides, same seed, same inputs
MeasureLatin scriptCJK scriptGap
Tokens per word1.614.732.9×
Bytes per token3.233.19same
Vocabulary size556.0556.0same
tick 310
Break it
Do

The training budget. Vocabulary size is exactly 256 bytes plus this.

How much of the trainer's corpus was written in the script being billed. The rest is Latin.

How top-heavy the word distribution is. Natural language sits near 1.

Now set that slider all the way to 1 and the merge budget to its maximum, on both sides. This is the most favourable case that exists: a tokenizer trained exclusively on the script being billed, with every merge you can buy.

The gap does not close. It stays around threefold.

Look at the bytes per token row while you do it. It is the same on both sides. The tokenizer is not doing a worse job on CJK - it is compressing just as hard, getting just as many bytes into each token. There is simply three times as much to compress, and a merge can only ever join two things into one.

The distinction is worth holding onto, because the two halves have different fixes. Training share is a decision somebody made and can remake. The encoding is not a decision anybody is going to remake.

What defeats a vocabulary

Press Bill words it never trained on. Same script, same word lengths, same frequency shape, a lexicon the trainer never saw - and the cost jumps. Nothing about the text is unusual. It is simply not the text the merges were bought for.

Press Bill it numbers. Digits are the sharpest version of the problem, because character frequency and sequence frequency come apart completely: every digit is common, so pairs of digits are learned almost immediately, but no particular number is frequent enough to earn an id. The string 84713 occurs once in a corpus and never again.

What to do with this

Measure your own text, not a benchmark. Fertility is a property of the pairing between a tokenizer and a distribution. Run your real traffic through the tokenizer you are actually billed by and count. The number you get is the one that matters, and it will not match a published average.[5]source codetiktoken - the BPE tokenizer used by OpenAI models

Budget prompts in tokens, never in characters. The folk rule of four characters per token is a fit to English prose. It holds well enough there and fails badly everywhere else - in CJK, in code, in anything with long identifiers. A truncation rule written in characters silently sends a different amount of text per language.

Treat a fertility gap as a pricing fact, not a bug report. If your product serves users in a script the tokenizer was not trained on, those users get a smaller context window and a larger bill for the same work. That is worth knowing before someone asks why, and worth pricing for rather than explaining away.

The dial

You gainA vocabulary that turns the text you actually send into a few tokens per word
You payEverything else pays close to the raw byte rate, and one part of that bill cannot be trained away

The line - when someone asks

A tokenizer starts from bytes and learns merges, so its whole benefit is a refund on text that looks like its training corpus. Two independent things set the bill: UTF-8 spends one byte on an ASCII character and three on CJK, and the merge budget goes to whatever was frequent in training, so a script that was a tenth of the corpus wins a tenth of the vocabulary. Retraining fixes the second and nothing fixes the first, because a merge joins two symbols into one and there are three times as many symbols to begin with.

Recall

Loading…

Where are you with this?

Saved on this device. Sign in to keep it across devices.

Sources

Primary

Secondary