How BPE Tokenization Works
After reading this you will be able to predict roughly how many tokens a piece of text costs, explain why byte-pair encoding produces the tokens it does, and know why a model cannot count the letters in "strawberry".
What a tokenizer does and why it matters
A language model never sees your text as characters. Before any math happens, the text is chopped into tokens: short chunks of bytes, each mapped to an integer ID. The model only ever works with those integers. The tokenizer is the fixed lookup table that turns text into IDs and back.
Byte-pair encoding (BPE) is the most common way to build that table. The idea is short: start with single characters, then repeatedly glue together the most frequent adjacent pair into one new symbol. Do that a few thousand times and you get a vocabulary where common fragments like th, ing and the are single tokens, while rare strings stay in small pieces.
Here is the hook. Type strawberry into a typical tokenizer and it becomes something like str + aw + berry, three tokens, three opaque IDs. Ask the model how many r characters the word has and it has no reliable way to answer, because the individual letters were merged away before the model saw anything. The count is not a reasoning failure so much as a consequence of tokenization.
When to think about tokens, and when not to
Reach for a token count whenever money or context length is involved. API pricing is per million tokens, not per word or per character, so the token count is the number on your invoice. Context windows are measured in tokens too, so knowing that a document is 4,000 tokens rather than 3,000 words tells you whether it fits.
You can usually ignore tokenization for casual English prose, where the rule of thumb 4 \text{ chars} \approx 1 \text{ token} holds well enough. Start paying attention when your text is unusual: source code, JSON, non-English languages, long numbers, emoji, or repeated whitespace. Those cost far more tokens than their character count suggests, and the visualizer shows exactly where the cost lands.
The tokenizer in this tool is small and trained on a tiny corpus, so its exact merges differ from GPT-family or Llama tokenizers. The mechanism is identical. Treat the counts here as an illustration of behavior, not as a substitute for the real tokenizer of a specific model.
The algorithm and its one rule
BPE training is a loop with a single decision rule. Given the current corpus split into symbols, count every adjacent pair, find the most frequent one, and merge it everywhere. Record that merge. Repeat until the vocabulary reaches the target size.
Here (a, b) is an adjacent pair of current symbols and \text{count}(a,b) is how many times that pair appears across the whole corpus. The winning pair (a,b)^* becomes a new symbol ab, and the vocabulary grows by one. The choice is greedy: it takes the best pair right now and never reconsiders.
Two consequences follow directly. First, early merges are the most common fragments in the language, because frequency drives every step. Second, the merge list is ordered. When you later tokenize new text, you apply the merges in the same order they were learned, so the vocabulary size acts like a dial: more merges means longer, fewer tokens.
A worked example you can reproduce
Take a corpus small enough to check by hand. Suppose after splitting into words and counting, you have these four word types with their frequencies:
| Word | Count | Initial symbols |
|---|---|---|
| low | 5 | l o w |
| lower | 2 | l o w e r |
| newest | 6 | n e w e s t |
| widest | 3 | w i d e s t |
Three merges by hand
- Count adjacent pairs weighted by word frequency. The pair
e sappears innewest(6) andwidest(3), giving 9. The pairs talso appears in both, giving 9. The pairl oappears inlow(5) andlower(2), giving 7. Break the 9-tie by pickinge s. - Merge
e sintoes. Nownewestisn e w es tandwidestisw i d es t. Recount. The paires tnow scores 6 + 3 = 9, the highest. Merge it intoest. - Recount again. The pair
l oscores 5 + 2 = 7, now the winner. Merge it intolo. Your first three merges arees,est,lo.
Notice that est was built from a previous merge. That is why the list is ordered: the merge for es t can only fire after e s already exists.
The default corpus in the tool mixes English and code, so its first merges look like th, in, en, t and =-adjacent pairs from the code. Drag the vocabulary slider and watch the list rebuild in exactly this frequency order.
Reading the token count and the chars-per-token ratio
The most useful single number the tool reports is chars per token: total characters divided by total tokens. Higher is better, because it means each token carries more text and therefore costs less.
For ordinary English trained on a matching corpus, expect roughly 3.5 to 4.5. The sentence the quick brown fox is 19 characters. If a well-trained tokenizer turns it into 5 tokens ( the, quick, brown, fox plus a leading fragment), the ratio is 19 / 5 = 3.8. Now feed it dense JSON like {"id":42,"ok":true}. The same 19 characters might split into 11 tokens because braces, quotes and colons rarely merge, giving 19 / 11 \approx 1.7. Same length, more than double the token cost.
How vocabulary size changes the count
The comparison panel tokenizes the same input at vocabulary 300, 1000 and 2000. As the vocabulary grows, more merges are available, so common runs of characters collapse into single tokens and the count falls. The drop is fast at first and then flattens, because the early merges catch the highest-frequency pairs and each later merge helps less.
The chars-per-token ratio moves the opposite way. At 58 tokens the ratio is 200 / 58 \approx 3.4; at 34 tokens it is 200 / 34 \approx 5.9 for text that matches the corpus. Production tokenizers use vocabularies of 30,000 to 130,000 or more, far past the flat part of this curve, so their ratios for in-distribution English sit near 4.
Turning tokens into money
API providers bill per million tokens. Multiply your token count by the price and divide by a million.
Say a request sends 39 tokens at 2.50 dollars per million input tokens. The cost is 39 / 1{,}000{,}000 \times 2.50 = 0.0000975 dollars, about a hundredth of a cent. That looks negligible until you multiply by volume. At one million such requests per day the input alone is 39 \times 10^6 / 10^6 \times 2.50 = 97.50 dollars per day, roughly 2,925 dollars per month. If your text is dense JSON that doubles the token count, so does the bill.
Common mistakes
The errors below all come from treating tokens as if they were words or characters.
- Counting words instead of tokens
- Words and tokens diverge fast. A 750-word page of English is roughly 1,000 tokens, but the same page of code or JSON can be 2,000 or more. Budget in tokens.
- Ignoring the leading space
theandthe(with a leading space) are different tokens with different IDs. The tool marks the space with a visible ␣. This is why concatenating strings without spaces can change token counts unexpectedly.- Expecting the model to count characters
- Once
strawberryis three token IDs, the letters are gone. Character-level questions ask the model to recover information the tokenizer discarded. - Assuming all text costs the same per character
- Text far from the training corpus falls back to single characters, so the same meaning in another language or format can cost three to five times more tokens.
Do not size a context window using a word count and the 4-chars-per-token rule for anything but plain English. For code, tables, or non-Latin scripts, tokenize the real text. Underestimating here truncates prompts silently and inflates cost.
Related tools
Once text is tokens, the model predicts the next one. The Next-Token Sampling Playground shows the actual next-token distribution of a small model and lets you reshape it with temperature, top-k and top-p, so you can see what happens after tokenization. When you are ready to turn token counts into a budget across models and context sizes, the LLM API Cost & Context Planner estimates spend per day and month, including RAG context and prompt caching.
Frequently asked questions
Why does the model get the number of r's in strawberry wrong?
Because it never sees the letters. The tokenizer merges strawberry into a couple of subword tokens before the model runs, and each token is one integer ID. Counting characters would require information the tokenization already threw away.
How many tokens is a typical English word?
Short common words are usually one token, often with a leading space attached, like the. Longer or rarer words split into two or three. Across a page of English the average lands near 4 characters per token, so a 5-character word averages a little over one token.
Why does code cost more tokens than prose of the same length?
Code is full of punctuation, indentation and identifiers that appear rarely in the training corpus, so few adjacent pairs merge. Symbols like {, }, : and repeated spaces often stay as single-character tokens, pushing the chars-per-token ratio down toward 1.7 or lower.
Does a bigger vocabulary always mean fewer tokens?
For text similar to the training corpus, yes, but with diminishing returns. In the example above, going from vocabulary 300 to 1000 cut tokens from 58 to 39, while 1000 to 2000 only reached 34. Text unlike the corpus barely improves at any vocabulary size.
Is this tokenizer the same one GPT or Llama uses?
No. This is a small BPE tokenizer trained on a tiny local corpus to show the mechanism. Real models use much larger vocabularies and corpora, so the exact splits differ. The training loop and the merge rule are the same.