Zipf and Heaps' Law, Explained
After reading this you will be able to rank the words of any text, check whether their frequencies fall as one over rank, read the fitted exponent off a log-log plot, and predict how fast the vocabulary grows as the text gets longer.
What these two laws say
Take any long text. Count how often each distinct word appears. Sort the words from most common to least common and give them ranks: rank 1 is the most frequent word, rank 2 the next, and so on. Zipf's law says the frequency of the word at rank r is roughly proportional to 1/r. The second word appears about half as often as the first, the third about a third as often, the tenth about a tenth as often.
Here is the hook. In a typical English novel the word "the" might make up about 6% of all tokens. Zipf predicts "of" (rank 2) at about 3%, "and" (rank 3) at about 2%, "a" (rank 4) at about 1.5%. Different authors, different centuries, different languages: the same falling staircase appears. On a log-log plot, where both rank and frequency are plotted on logarithmic axes, that 1/r curve becomes a straight line with slope close to -1.
Heaps' law is the companion. As you read more words, you keep meeting words you have not seen before, but the rate of new words slows down. The count of distinct words V after reading n tokens grows as a power of n with an exponent below 1. The two laws are not independent: a Zipf tail full of rare words is exactly what forces vocabulary to keep growing.
When these laws apply, and when they do not
Both laws describe natural word streams: prose, speech transcripts, articles, chat logs. They hold best when you have a lot of tokens, say 10,000 or more, and when the text is genuine language rather than a table of numbers or a list of unique identifiers.
They break down in predictable ways. A very short text has too few counts for a stable slope. A text stuffed with repeated boilerplate (a legal template, a song with a heavy chorus) will bulge at the top ranks. A random list of serial numbers has no Zipf structure at all, because every "word" appears exactly once. Treat the fitted exponent as a description of the text in front of you, not as a law of physics you are testing.
The single most abused habit is fitting a straight line through all the log-log points and reporting one slope. Real texts bend at both ends: the top few ranks sit below the line and the rare-word tail flattens. A slope of -1.07 reported to three digits hides that curvature. Look at the plot, not only the number.
The formulas and what each symbol means
The clean form of Zipf's law introduces one exponent, usually written s, close to 1.
Here f(r) is the frequency (or probability) of the word at rank r, s is the Zipf exponent, and C is a normalising constant so the frequencies sum to the total token count. When s = 1 you get the pure 1/r shape. Take logs of both sides and the reason for the log-log plot appears:
That is a straight line in \log r with slope -s. Fit a least-squares line to the points (\log r, \log f) and the negative of its slope is your estimate of s.
Heaps' law has its own two parameters.
V(n) is the number of distinct words seen after n total tokens, K is a constant (often between 10 and 100 for English), and \beta is the growth exponent, typically 0.4 to 0.6. Because \beta \lt 1, the derivative K\beta n^{\beta - 1} keeps shrinking but never reaches zero: new words keep arriving forever, just more slowly.
The two exponents are linked. If the Zipf exponent is s then the Heaps exponent is approximately \beta \approx 1/s when s \gt 1, and \beta \approx 1 when s \le 1. A steeper Zipf tail means fewer rare words and therefore slower vocabulary growth.
Reproducing the demo text
The demo loads a public-domain novel of about 50,000 tokens. Below are the counts for the top five ranks that the explorer reports, alongside what a pure 1/r law predicts from the rank-1 count.
| Rank | Word | Count | Predicted C/r | Ratio |
|---|---|---|---|---|
| 1 | the | 3000 | 3000 | 1.00 |
| 2 | of | 1580 | 1500 | 1.05 |
| 3 | and | 1490 | 1000 | 1.49 |
| 4 | to | 950 | 750 | 1.27 |
| 5 | a | 760 | 600 | 1.27 |
- Set the constant from rank 1: C = f(1) = 3000.
- Predict rank 2 as 3000/2 = 1500. Observed 1580, so the ratio is
1.05, close to 1. - Predict rank 3 as 3000/3 = 1000. Observed 1490. The word "and" runs high, which is common in narrative prose.
- Fit the line \log f = \log C - s \log r across all ranks. The explorer reports a slope near
-1.06, so s \approx 1.06. - Predict Heaps from Zipf: \beta \approx 1/1.06 \approx 0.94 is the loose upper bound, but measured vocabulary growth on this text gives \beta \approx 0.52, inside the usual 0.4 to 0.6 band. The gap is the curvature warning in action: the simple 1/s link is only exact for an ideal infinite Zipf tail.
Reading the two plots
On the rank-frequency plot, three regions matter. The head (ranks 1 to about 10) often sits a little below the fitted line because a handful of function words are slightly less dominant than a perfect law demands. The middle is the cleanest straight stretch; this is where the slope you trust comes from. The tail is a cloud of words that each appear once or twice. It flattens into a horizontal band because you cannot have a count below 1.
That tail is not noise to be discarded. Between 40% and 60% of the distinct words in a typical book appear exactly once. Those are the hapax legomena. They are the direct cause of Heaps' law: every fresh rare word is a new entry in the vocabulary. If you removed them, both the tail and the vocabulary growth would collapse.
On the Heaps plot, vocabulary rises steeply at first, then bends toward a gentler climb. It never goes flat. Plot it on log-log axes and it straightens into a line of slope \beta. If you double the text from 50,000 to 100,000 tokens and \beta = 0.52, the vocabulary grows by a factor of 2^{0.52} \approx 1.43, so about 43% more distinct words, not twice as many.
Common mistakes
Watch for these when you read your own results.
- Tokenising badly
- If you keep punctuation attached, "dog." and "dog," and "dog" become three words. Lowercasing and stripping punctuation before counting changes the top ranks and can shift the slope by 0.1 or more.
- Fitting the whole line
- The tail's flattening pulls a global least-squares slope toward zero. Many analysts fit only the middle ranks, or use a maximum-likelihood estimator for a power law instead of a line fit.
- Confusing frequency with probability
- The shape is identical either way, since dividing every count by n only shifts the log-log line vertically by \log n. The slope -s does not change.
- Reading meaning into the exponent
- A slope near
-1is the null expectation, not evidence of anything special. Even random typing with a space key produces Zipf-like curves, so a good fit is weak proof of deep structure.
Why the law appears at all
No single explanation has won. Mandelbrot derived a Zipf-like law from minimising the average cost of encoding messages: if common ideas get short cheap words, the frequency-length trade-off forces a power law. A different family of arguments uses preferential attachment, where a word already used often is more likely to be reused, which grows heavy-tailed distributions in many systems. Most unsettling, a monkey hitting random keys including a space bar produces a rank-frequency curve that looks Zipfian, because word length and frequency line up by pure combinatorics.
The honest position: Zipf's law is a robust statistical signature that many generating processes share. Seeing it in a text tells you the text behaves like natural language. It does not by itself pick out which mechanism produced it. That ambiguity is why the law is famous and still argued over.
Related tools
Zipf's law is one power law among many. To see another one measured straight off a log-log slope, try the Box-Counting Dimension Lab, which reads a fractal's dimension from exactly the same kind of straight-line fit. For power laws born from random reuse, the drunkard paths in the Random Walk Explorer include heavy-tailed Lévy flights. If you want to watch text itself get squeezed by exploiting the very frequency imbalance Zipf describes, the Compression Playground shows Huffman coding give short codes to frequent tokens. And for the pure statistics of averages settling down as data piles up, the Law of Large Numbers demo is the cleanest place to start.
Frequently asked questions
How much text do I need for a stable slope?
Aim for at least 10,000 tokens. Below about 2,000 the top ranks jump around and the slope can swing by 0.3 between two similar passages. At 50,000 tokens the middle-rank slope typically settles to within about 0.05.
Why is the Zipf exponent often above 1 rather than exactly 1?
The pure 1/r law is an idealisation. Real English texts usually fit somewhere between 1.0 and 1.1 in the middle ranks, and the exact value depends on tokenisation, genre and how much of the tail you include in the fit.
What are hapax legomena and why do so many exist?
They are words that appear exactly once in the text. In most books they are 40% to 60% of the distinct vocabulary. They exist because the Zipf tail is so long: there are far more rare words than common ones, and the rarest bucket, count of one, is always the most crowded.
Does Heaps' law mean vocabulary never stops growing?
Yes, in the model. Because \beta \lt 1, the curve K n^{\beta} keeps rising without bound, just ever more slowly. In practice a language has a finite word stock, so real growth eventually curves below the power law, but for any text you will actually read it keeps climbing.
Can I compare two authors by their exponents?
Cautiously. A difference of 0.02 in s is within the noise of tokenisation choices. A difference of 0.15 or more, computed the same way on texts of similar length, may reflect a real stylistic gap, such as a richer or poorer use of rare words.