LLM Tokenizer Visualizer
Large language models never read letters or words — they read tokens produced by byte-pair encoding (BPE): start from single characters and repeatedly merge the most frequent adjacent pair until the vocabulary is full. Here the whole pipeline runs live: edit the training corpus (mixed English and code by default), drag the vocabulary slider, and watch the merge list rebuild in order of frequency. Then type anything and see it as colored token chips with a token count and chars-per-token ratio. A comparison panel tokenizes the same input at vocabulary 300, 1000 and 2000 to show how counts shrink as merges accumulate, and a cost row turns the token count into dollars at an editable price per million tokens. It explains at a glance why models miscount the letters in "strawberry", and why text unlike the training corpus costs more tokens.
Runs 100% in your browser — models are trained and computed locally on your device.
Read the full guide to this tool
Notes
- Merges are learned greedily: at every step the single most frequent adjacent symbol pair in the corpus becomes one new vocabulary entry, so early merges are common fragments like "th" or "in".
- A model sees "strawberry" as a couple of opaque token IDs, not eleven letters — counting the r’s requires knowledge the tokenization already threw away.
- Whitespace is part of tokens: " the" (with a leading space) and "the" are different tokens, which is why the chips here carry a visible ␣ marker.
- Text that looks unlike the training corpus — other languages, rare symbols, dense code — falls back to short fragments or single characters, so the same meaning costs several times more tokens and therefore more money.
- Runs 100% in your browser — models are trained and computed locally on your device.