Zipf & Heaps' Law Explorer
Zipf’s law is one of the eeriest regularities in language: rank the words of almost any text by frequency and the 2nd word appears about half as often as the 1st, the 3rd a third as often — frequency falls as 1/rank, a straight line on a log-log plot. Paste or upload any text — an article, a novel, your own writing — and see its rank-frequency plot with the fitted exponent next to the ideal Zipf line. Then switch to Heaps’ law and watch the vocabulary curve: how many distinct words have appeared after n words of reading, growing as a power law that never quite flattens out.
Runs 100% in your browser — simulations are computed locally on your device.
Read the full guide to this tool
Notes
- Zipf’s law holds across languages, authors and centuries; nobody fully agrees why. Explanations range from Mandelbrot’s information-theoretic optimality to preferential attachment — even random typing produces Zipf-like curves.
- Heaps’ law says vocabulary grows as V ≈ K·n^β with β typically 0.4–0.6: new words keep arriving forever, just ever more slowly — the reason a dictionary is never finished.
- Typically 40–60% of a book’s distinct words appear exactly once (hapax legomena) — a direct consequence of the Zipf tail.
- The fitted exponent is a least-squares slope on the log-log points; real texts bend away from a perfect line at both the top ranks and the rare-word tail.
- Runs 100% in your browser — simulations are computed locally on your device.