The Base Rate Fallacy, Explained

After reading this you will be able to take a test's accuracy and a condition's rarity and compute the only number that matters to you: the chance you are actually sick given a positive result. You will also see why that number is so often much smaller than people expect.

What it is and one example that stings

Suppose a screening test for a disease is described as "99% accurate." You take it. It comes back positive. What is the chance you have the disease?

Most people, including most doctors, answer around 99%. The real answer, for a disease that affects 1 in 1000 people, is close to 9%. Ninety-one out of every hundred positive results are false alarms.

The mistake has a name: the base rate fallacy. It happens when you read a test's accuracy as if it were your personal probability of being sick, and quietly ignore how rare the condition is in the first place. The rarity is the base rate. The visualizer draws 1000 dots, tests every one of them, and recolours them so you can count the false positives directly instead of trusting your gut.

The three numbers that decide everything

You need exactly three inputs, and it helps to keep them straight because two describe the test and one describes the world.

Prevalence (base rate)
The fraction of the population that actually has the condition, before any test is run. Written p. A prevalence of 1 in 1000 means p = 0.001.
Sensitivity
The probability the test says positive when the person is sick. This is the true positive rate. A sensitivity of 0.99 means the test catches 99 of every 100 sick people and misses 1.
Specificity
The probability the test says negative when the person is healthy. This is the true negative rate. A specificity of 0.99 means the test correctly clears 99 of every 100 healthy people and falsely flags 1.

Notice that sensitivity and specificity say nothing about how common the disease is. They are properties of the test measured in a lab. The prevalence is what varies from one setting to another, and it is the number people forget.

"99% accurate" is ambiguous on purpose. It usually means sensitivity and specificity are both around 0.99, but a single accuracy figure hides which errors the test makes. Always ask for the two rates separately.

The formula and why it flips your intuition

The clean way to state the answer is Bayes' theorem written in counts. Take a population of size N. The number of sick people who test positive (true positives) is

\text{TP} = N \cdot p \cdot \text{sens}

Here p is prevalence and \text{sens} is sensitivity. The number of healthy people who test positive by mistake (false positives) is

\text{FP} = N \cdot (1 - p) \cdot (1 - \text{spec})

where \text{spec} is specificity, so 1 - \text{spec} is the false positive rate. The probability you want, the chance you are sick given a positive test, is the true positives divided by all positives:

P(\text{sick} \mid +) = \frac{\text{TP}}{\text{TP} + \text{FP}}

The intuition sits in the two products. The true positives scale with p, the small number. The false positives scale with 1 - p, the large number. When the disease is rare, the healthy majority is enormous, and even a tiny false positive rate applied to that huge group produces more false positives than there are true positives. The healthy majority drowns the signal.

Worked example: the demo defaults

1 in 1000 prevalence, 99% sensitivity, 99% specificity

These are the tool's default settings. Run the demo and follow the arithmetic on a population of N = 1000.

  1. Sick people: 1000 × 0.001 = 1. On average exactly 1 of the 1000 dots has the disease.
  2. True positives: 1 × 0.99 = 0.99. The test almost certainly catches that 1 person, so round to 1 true positive.
  3. Healthy people: 1000 × 0.999 = 999.
  4. False positives: 999 × (1 - 0.99) = 999 × 0.01 = 9.99, about 10 healthy people flagged by mistake.
  5. Positive results total: 1 + 10 = 11.
  6. Chance of being sick given a positive: 1 / 11 = 0.0909, about 9.1%.

So of the 11 people who test positive, only 1 is sick. Ten are healthy and terrified for no reason. A 99% accurate test delivered a 9% verdict, and every step above is one you can count with the recoloured dots.

Out of 11 positive results, one comes from the single sick person and ten come from the 999 healthy people. The false positives dominate.

How the answer moves as the disease gets more common

The 9% figure is not a property of the test. Hold sensitivity and specificity at 0.99 and slide the prevalence up. As sick people become common, true positives multiply and start to outweigh the fixed 1% trickle of false positives.

P(sick | positive) with sensitivity and specificity fixed at 0.99
PrevalenceTrue positives per 1000False positives per 1000P(sick | positive)
1 in 10000.999.990.090
1 in 1009.99.90.500
1 in 2049.59.50.839
1 in 249550.990

Read the middle row carefully. At a prevalence of 1 in 100, the true and false positives are almost exactly equal (9.9 versus 9.9), so a positive result is a genuine coin flip. The disease has to be common (roughly as common as the test's error rate) before a positive result carries the weight people instinctively give it.

With sensitivity and specificity both fixed at 0.99, the posterior probability P(sick | positive) rises from 0.090 at a prevalence of 1 in 1000, to 0.500 at 1 in 100, to 0.839 at 1 in 20, to 0.990 at 1 in 2. The rarer the condition, the more a positive result is dominated by false alarms.

Natural frequencies: the trick that fixes the intuition

The reason the tool draws dots and phrases the answer as "1 of 11 people" is not decoration. Gerd Gigerenzer and colleagues found that when the same problem is posed in percentages, most physicians get it wrong, but when it is posed in natural frequencies (counts out of a fixed population), the majority get it right.

Compare the two phrasings of the identical problem:

  • Percentage version: prevalence 0.1%, sensitivity 99%, false positive rate 1%. What is P(sick | positive)? Hard.
  • Frequency version: 1 of every 1000 people is sick. That person almost certainly tests positive. Of the other 999, about 10 test positive anyway. So about 11 test positive and 1 is sick. Easy.

Same numbers, different cognitive load. The frequency tree in the tool traces the whole population down the sick and healthy branches, then splits each branch by test outcome, so the four boxes at the bottom (true positive, false positive, true negative, false negative) add back to 1000.

The curve rises steeply from rare to common. The marked line at p = 0.01 is where a positive is exactly a coin flip (50%).

Common mistakes

The most frequent error is treating sensitivity as the answer. A test that catches 99% of sick people does not mean a positive result is 99% likely to be sick. Those are different conditional probabilities: P(positive | sick) versus P(sick | positive). Swapping them is the base rate fallacy in one sentence.

Three more traps to avoid:

  • Confusing accuracy with specificity. A test can be 99% "accurate" overall on a rare disease simply by calling everyone healthy, since it will be right 99.9% of the time. Overall accuracy is a useless summary when classes are imbalanced.
  • Applying a screening prevalence to a symptomatic patient. If someone already has symptoms, their prevalence is not the population base rate; it is much higher, so their post-test probability is much higher too. Use the right base rate for the right person.
  • Forgetting that a negative can also mislead. With the defaults, a negative result is almost certainly correct because the false negatives are rare, but for a high-prevalence condition with modest sensitivity, a negative result is not the reassurance it seems.

The same math dooms mass screening

Push the prevalence to the extreme and the arithmetic gets brutal. Airport security screens for something like 1 in a million travellers carrying a threat. Even a scanner that is 99% accurate on both rates produces, per million people:

  • True threats caught: 1 × 0.99 ≈ 1.
  • False alarms: 999999 × 0.01 ≈ 10000.

So each genuine catch comes buried under roughly 10,000 innocent people flagged. That is not a flaw in the scanner; it is the base rate fallacy operating at scale. Screening for very rare events with an imperfect test is inherently a false-alarm machine, which is why such systems need a cheap, accurate second stage to sort the flagged pile.

Related tools

The base rate fallacy is one of a family of probability surprises worth playing with. The Monty Hall Simulator is another problem where conditional probability defeats intuition. For the machinery underneath, the tree of outcomes generalises to the Law of Large Numbers, which explains why the false positive count settles near its expected value across many people. If you care about how selection bends probabilities, see Berkson's Paradox and Simpson's Paradox Visualizer. And for how easy it is to conjure a false positive by hunting, try the p-Hacking Simulator and the A/B Test Peeking Simulator.

Frequently asked questions

Why is a 99% accurate test only 9% likely to be right about me?

Because "99% accurate" describes what the test does to sick and healthy people separately, not what a positive result means for you. If only 1 in 1000 people is sick, the 999 healthy people generate about 10 false positives while the sick group generates only 1 true positive, so a positive is sick only 1 time in 11, about 9%.

Does taking the test again help?

Yes, if the errors are independent. After a first positive, your probability of being sick jumps from 0.001 to about 0.09. Use 0.09 as the new base rate for a second independent test: TP becomes 0.09 × 0.99 = 0.089 and FP becomes 0.91 × 0.01 = 0.0091, giving about 0.907, roughly 91%. A second positive is far more convincing than the first.

What is the difference between sensitivity and specificity?

Sensitivity is how well the test detects sick people: P(positive | sick). Specificity is how well it clears healthy people: P(negative | healthy). A test can be high on one and low on the other, and the two errors matter differently depending on whether the disease is rare or common.

Is the base rate fallacy only about medicine?

No. It applies to any rare event flagged by an imperfect test: airport screening, spam filters, fraud detection, DNA database matches, forecasting rare failures. Whenever the true positives are outnumbered by false positives, the same 1-divided-by-11 arithmetic applies.

What prevalence makes a positive result a coin flip?

When true positives equal false positives. With sensitivity and specificity both at 0.99, that happens at prevalence 0.01 (1 in 100), where each side contributes about 9.9 people per 1000 and P(sick | positive) is exactly 0.5.