Regression to the Mean, Explained

After reading this you can predict how far the top performers of one season will slide in the next, compute the exact fall-back from a single correlation, and spot the false lessons that regression to the mean has taught people for a century.

What it is and one hook example

Every measured performance is part signal and part noise. Call the stable part skill and the fresh part luck. Regression to the mean is the fact that when you select the extreme performers, you have selected people with unusually good luck as well as skill, and luck does not repeat. Next time, the skill stays and the luck resets, so the group drifts back toward the average.

Here is the hook. A rookie hits .330 and wins an award. Next year he hits .290 and the press calls it a sophomore slump. His skill did not collapse. His first year mixed real ability with a hot streak, and the hot streak was never going to come back on command. The decline is not a story about pressure or ego. It is arithmetic.

This simulator gives each of N players a fixed true skill, then draws a fresh random luck term each season and adds it to skill to make an observed score. Nothing about the players changes between seasons. Only the luck is redrawn. That single mechanism reproduces the slump, the fading fund manager, and the flight instructor who wrongly concluded that punishment works.

When the effect appears, and when it does not

Regression to the mean shows up whenever you do two things at once: you select on an extreme of a noisy measurement, and you measure a second time. Both halves are required.

If you pick a random player and check next season, there is no systematic drift: a random player has average luck by definition. The drift is entirely a product of the selection. Pick the top 10% and they fall. Pick the bottom 10% and they climb. Pick the middle and almost nothing happens.

The effect also disappears at the extremes of the noise. If a score is pure skill with no luck, the top group scored high because they are genuinely the best, and they will score high again. If a score is pure luck with no skill, the top group was a coin-flip fluke and the next measurement is a fresh coin flip, so they regress all the way to the mean. Real cases sit in between.

Regression is not a force that pulls scores toward the center. No player is dragged down. The population average of skill is unchanged. What changes is which players you are looking at: you chose them for luck they no longer have.

The formula and why it is a share of variance

Model each observed score as skill plus luck, with both centered so the population mean is zero:

X = S + L, \quad Y = S + L'

Here S is the fixed skill (same both seasons), L and L' are independent luck draws, X is season 1 and Y is season 2. Let skill have variance \sigma^2_{\text{skill}} and luck have variance \sigma^2_{\text{luck}}. The correlation between the two seasons is

r = \frac{\sigma^2_{\text{skill}}}{\sigma^2_{\text{skill}} + \sigma^2_{\text{luck}}}

That is the share of the total variance that is skill. It runs from 0 (all luck) to 1 (all skill). The between-season correlation is not some separate number to memorize: it is the reliability of the measurement.

Now the payoff. If a group of players averaged some distance d above the population mean in season 1, their expected season-2 average is only r \cdot d above the mean. They fall back by the fraction (1-r) of their lead:

\text{expected fall-back} = (1 - r)\, d

All skill (r = 1) means zero fall-back. All luck (r = 0) means complete fall-back to the mean. Everything in between splits the difference in exact proportion.

A worked example that reproduces the demo

The demo defaults give skill and luck equal spread. Take \sigma_{\text{skill}} = 10 and \sigma_{\text{luck}} = 10, so both variances are 100. That is a clean, checkable case.

Top 10% with equal skill and luck

  1. Compute the correlation: r = \frac{100}{100 + 100} = 0.5. Half the variance is skill.
  2. Find the season-1 average of the top 10%. Observed scores have variance 100 + 100 = 200, so standard deviation \sqrt{200} \approx 14.14. For a normal distribution the mean of the top 10% sits about 1.755 standard deviations above the mean. So d = 1.755 \times 14.14 \approx 24.8 points above average.
  3. Apply the fall-back formula: expected season-2 average = r \cdot d = 0.5 \times 24.8 \approx 12.4 points above the mean.
  4. The top group loses (1 - r)\,d = 0.5 \times 24.8 \approx 12.4 points, exactly half its lead, in one season.

Nobody in that group got worse at the task. Their skill average is fixed. They started 24.8 above because they had good skill and good luck, and only the skill part, worth 12.4, carries over.

The season-1 lead of the top group is the same in each case. What differs is how much survives: the surviving lead equals r times the original, so at r = 0.5 exactly half remains.

Reading the scatter and the readout

The main plot puts season 1 on one axis and season 2 on the other, one dot per player. A cloud of dots stretched into an oval tells the whole story. The long axis of the oval is skill, shared by both seasons. The width perpendicular to it is luck, different each season.

The best-fit line through the cloud has slope r, not slope 1. That slope less than one is the regression. Highlight the top 10% of season 1 (a vertical strip on the right) and look at where those dots sit vertically: their average height is below their average horizontal position, pulled toward the center by the factor r. The bottom 10% (a strip on the left) sits above its horizontal position by the same logic.

Watch the readout for r as you move the sliders. When you widen the luck spread, the oval fattens, r drops, and the highlighted top group collapses toward the mean. When you widen the skill spread instead, the oval stretches along its long axis, r rises, and the top group barely moves.

Common mistakes

The costly error is to explain the drift with a story. The flight instructors in Kahneman's account praised smooth landings and saw the next landing get worse, scolded rough landings and saw the next one improve. They concluded punishment beats praise. Both patterns were pure regression: an unusually smooth landing is followed on average by a more ordinary one whatever you say afterward, and an unusually rough one is followed by a better one. The instructors read causation into arithmetic and taught themselves the wrong lesson.

A second mistake is treating the fall-back as evidence of a real decline. A fund that beat the market by a wide margin one year, then merely matched it the next, has not necessarily lost its edge. If manager skill explains only r = 0.2 of return variance, an outperformer keeps just 20% of the gap on average. The other 80% was luck that did not recur.

A third mistake is designing a before-and-after study without a control group. Give a remedial program to the worst-scoring students and their scores rise. Some of that rise is the program and some is regression, because the worst scorers were partly unlucky. Without a control group selected the same way, you cannot tell the two apart, and you will overstate the program. The Berkson's paradox and Simpson's paradox visualizers show two more ways selection quietly bends a conclusion.

Explore the effect yourself

With skill spread 10 and luck spread 10 the between-season correlation is r = 0.5, and the top 10% of season 1 (about 24.8 points above the mean) is expected to average only 12.4 points above the mean in season 2, a fall-back of half its lead. Raising the luck spread lowers r and increases the fall-back; raising the skill spread does the opposite.

Related tools on this site

Regression to the mean is one member of a family of results about noise, selection and averaging. To see averages themselves settle onto a stable value, try the Law of Large Numbers, and to watch those averages become bell-shaped no matter the source distribution, the Central Limit Theorem demo. The bell curve building itself from independent bounces is the Galton board.

To build intuition for the correlation r that drives the fall-back, guess it from scatterplots in Guess the Correlation. For the ways selection and repeated testing manufacture false signals, see the p-hacking simulator, the A/B test peeking simulator, and the base rate visualizer. And for randomness as a constructive tool rather than a nuisance, the Monte Carlo playground throws random samples to estimate quantities you cannot compute directly.

Frequently asked questions

Does regression to the mean mean everyone becomes average over time?

No. The population spread of skill is fixed and does not shrink. Regression only describes what happens to a group you selected on an extreme. The overall distribution of scores stays the same width season after season. Top players keep being replaced by different top players, whose luck is fresh.

Is the sophomore slump real or a statistical artefact?

It is an artefact of selecting on the standout rookies. Award winners were the extreme of a noisy first year, so their second year regresses. The same math predicts that the worst rookies improve, which nobody writes headlines about. Both moves are the same effect from opposite ends.

How is regression to the mean different from ordinary linear regression?

They share a name and a slope but describe different things. Linear regression fits a line to data. Regression to the mean is the observation, first made by Galton, that the fitted slope from one noisy measurement to a later one is less than one, so predicted values are pulled toward the mean. The slope in that specific setup equals the correlation r.

If skill never changes, why not just average many seasons?

You should. Averaging k independent seasons cuts the luck variance by a factor of k while leaving skill untouched, which raises the effective correlation. That is why long track records are more trustworthy than single-season peaks: the luck averages out and the skill remains.

Can I compute the fall-back without simulating?

Yes. Get the correlation r from the skill and luck variances, measure how far your selected group sits above the mean in units you like, then multiply that distance by r to get the expected follow-up and by (1-r) to get the expected drop. With r = 0.5 and a lead of 24.8, the drop is 12.4.