The Shannon Guessing Game, Explained

After reading this you will understand why guessing the next letter of a text measures the entropy of English, how the game turns your guesses into Shannon's upper and lower bounds, and how to read the roughly 1 bit per letter that good play produces.

What the game measures

English is redundant. If you read the quick brown fo you already know the next letter is x with near certainty. That predictability is not a vague impression: it is a number, measured in bits per letter, and you can extract it by playing a game.

In 1951 Claude Shannon published "Prediction and Entropy of Printed English." His idea was simple and clever. A person who has read a lot of English carries an internal model of the language. Ask that person to guess the next letter of a hidden text, one letter at a time, and count how many guesses each letter takes. A perfect predictor needs few guesses where the language is predictable and many where it is not. The distribution of those guess counts is a fingerprint of the language's entropy.

This tool is that experiment, run in your browser. A text is hidden. You guess letter by letter. Every guess you make is a data point, and from the tally the tool computes bounds on the entropy of English: the same calculation that gave Shannon his famous estimate near 1.3 bits per letter, well below the 4.75 bits you would need if all 27 symbols were equally likely.

When to use it, and when not

Use the game when you want to feel entropy rather than read about it. Playing 50 letters teaches more about redundancy than any table. It is also a genuine measurement: paste a passage in a style you care about, guess it honestly, and you get an entropy estimate for that style of writing.

Do not treat the number as a universal constant. Your result depends on you (how well you predict), on the passage (a legal contract is more predictable than a poem), and on how many letters you guessed. Fewer than about 40 letters gives noise, not a measurement. And the game measures the entropy of English as modeled by your brain. A weak guesser produces a high estimate because the method reports an upper bound tied to your skill, not the true minimum.

Shannon's subjects worked with a 27-symbol alphabet: the 26 letters A to Z plus space. This tool normalizes any text you paste the same way, folding case and stripping punctuation and digits, so your numbers are comparable to his.

The formula and the intuition

Let q_i be the fraction of letters you guessed correctly on exactly your i-th attempt. If you got 30 of 50 letters on the first try, then q_1 = 0.60. These fractions sum to 1 across all guess ranks.

Shannon's upper bound is just the entropy of that guess-count distribution:

H_{\text{upper}} = -\sum_{i=1}^{27} q_i \log_2 q_i

Here q_i is the fraction of letters that took i guesses, and the sum runs over every possible rank from 1 to 27. The logarithm base 2 makes the answer come out in bits. This is an upper bound because a real guessing sequence can never be more efficient than the ideal code built from its own statistics.

The lower bound is subtler. Shannon showed it as:

H_{\text{lower}} = \sum_{i=1}^{27} i \cdot (q_i - q_{i+1}) \cdot \log_2 i

Each term weighs the rank i by the drop q_i - q_{i+1} from one rank to the next, scaled by \log_2 i. The lower bound uses only the ranking, not the exact probabilities, so it says: however clever the guesser's model is, the entropy cannot fall below this. Take q_{28} = 0 so the last term is well defined. The true entropy of English sits between these two bounds.

Reproducing the built-in passage

Suppose you play the default passage and, over 50 guessed letters, your attempts fall out like this. These counts are the kind of distribution a fluent reader produces.

Guess-count tally over 50 letters
Rank iLettersqᵢqᵢ log₂ qᵢ
1390.78-0.2792
260.12-0.3671
330.06-0.2435
410.02-0.1129
510.02-0.1129
  1. Convert counts to fractions: 39/50 = 0.78, 6/50 = 0.12, and so on.
  2. Sum the last column for the upper bound: 0.2792 + 0.3671 + 0.2435 + 0.1129 + 0.1129 = 1.1156. Flip the sign convention (the terms are already negative), so H_{\text{upper}} \approx 1.12 bits per letter.
  3. For the lower bound, use the drops. With q_5 = 0.02 and q_6 = 0: term at i=1 is 1·(0.78−0.12)·log₂1 = 0 (since log₂1 = 0). Term at i=2 is 2·(0.12−0.06)·1 = 0.12. Term at 3 is 3·(0.06−0.02)·1.585 = 0.1902. Term at 4 is 4·(0.02−0.02)·2 = 0. Term at 5 is 5·(0.02−0)·2.322 = 0.2322.
  4. Sum: 0 + 0.12 + 0.1902 + 0 + 0.2322 = 0.5424, so H_{\text{lower}} \approx 0.54 bits per letter.

Your estimate for this passage: between 0.54 and 1.12 bits per letter. That straddles Shannon's classic range and confirms English carries a lot of redundancy. If all 27 symbols were equally likely, both bounds would sit at 4.75 bits.

Seeing the distribution

The shape that matters is how sharply the guess counts fall off. A steep drop, most letters guessed on the first try, means low entropy. A flat spread across many ranks means the guesser is often uncertain, and the entropy climbs.

Most letters (78 percent) fall on the first guess. The tall first bar is what pulls the entropy down toward 1 bit.

As the fraction of first-guess correct letters rises from 0.20 to 0.95, both entropy bounds fall. At 0.20 first-guess accuracy the upper bound is near 3.5 bits; at 0.95 it drops below 0.5 bits. The remaining probability is spread evenly across ranks 2 to 6.

Reading and interpreting your result

Read the two bounds as a range, not a point. If the tool reports 0.6 to 1.3 bits, that is the interval Shannon's own human subjects produced, and it means the passage carries about 1 bit of genuine unpredictability per letter. The remaining 4.75 - 1 = 3.75 bits are redundancy: structure you could reconstruct from context.

A wide gap between the bounds (say 0.5 to 2.0) usually means you have not guessed enough letters, or the passage mixes very predictable stretches with surprising ones. A narrow gap on 100+ letters is a trustworthy measurement.

Entropy
The average number of bits needed to specify the next letter given everything before it. Lower means more predictable.
Redundancy
The gap between maximum entropy (4.75 bits for 27 symbols) and measured entropy. English runs about 75 percent redundant.
Upper bound
Entropy of your guess-count distribution. Reflects your skill; a weak guesser inflates it.
Lower bound
Uses only the ranking of guesses. The true entropy cannot be below it.

Common mistakes

Do not quote a number from 10 or 20 letters. Short runs produce wild swings: a single lucky stretch can push the upper bound below 0.5 bits, and one hard word can double it. Guess at least 40 to 50 letters before trusting anything.

A second mistake is comparing your result to Shannon's 1.3 bits as if it were a target you should hit. His figure came from specific subjects on specific 27-symbol text. Your passage and your prediction skill both move the number. The point is the method, not matching a historical constant.

Third, remember what normalization does. Case is folded and punctuation dropped, so Don't! becomes dont, four letters. Restoring punctuation would add entropy the game does not measure. The 27-symbol scope is a deliberate simplification, not an oversight.

Related tools

Entropy is one lens on text. To count what you are working with, use the Word & Character Counter. To see which letters and words carry the predictable mass, the Word Frequency & N-gram Analyzer shows the bigrams and trigrams that make guessing easy. Redundancy and readability are cousins, so the Readability Scorer is a useful companion. If you want to inspect exactly which codepoints survive normalization, the Unicode Inspector shows every character. And for a different fingerprint of style, the Stylometry & Authorship Comparator compares two texts with Burrows' Delta.

Frequently asked questions

Why is the entropy of English about 1 bit per letter?

Because English is highly constrained. Spelling rules, common words, and grammar mean most letters are largely determined by context. Shannon measured 0.6 to 1.3 bits with human guessers, far below the 4.75 bits of a 27-symbol random source.

Why two bounds instead of one number?

A guessing experiment cannot pin entropy exactly. The upper bound comes from the full guess distribution, the lower bound from the ranking alone. The true entropy of the language sits between them, and better data narrows the gap.

Does my typing skill or knowledge affect the result?

Yes. The game measures the entropy of English relative to your predictive model. A fluent guesser gets a lower, more accurate estimate. A distracted or unfamiliar guesser inflates the upper bound.

Can I test a language other than English?

You can paste any text, but the 27-symbol A to Z plus space alphabet fits English and other Latin-script languages loosely. Accented letters get stripped, which distorts the count for languages that use them heavily.

How many letters should I guess?

At least 40 to 50 for a rough figure, and 100 or more for a stable one. Watch the gap between the bounds shrink as you add letters.