The Iterated Prisoner's Dilemma, Explained
After reading this you can score a prisoner's dilemma match by hand, predict which strategies survive a round-robin, and explain why generous strategies beat unforgiving ones once you add noise.
What the tournament is, and one game to fix ideas
Two players each choose, in secret, to cooperate (C) or defect (D). The payoffs are fixed by four numbers. If both cooperate they each score the reward R = 3. If both defect they each score the punishment P = 1. If one defects while the other cooperates, the defector gets the temptation T = 5 and the cooperator gets the sucker payoff S = 0.
Play once and defection wins outright. Whatever the other player does, defecting pays more: 5 > 3 if they cooperate, 1 > 0 if they defect. That is the dilemma. Mutual cooperation (3 each) beats mutual defection (1 each), yet each player is pulled toward the move that wrecks it.
The interesting version is iterated: the same two players meet for many rounds and remember what happened. Now a defection can be punished next round, and cooperation can be rewarded. Robert Axelrod ran two tournaments in 1980 where submitted strategies played long matches against each other. A four-line strategy called Tit for Tat won both. This tool rebuilds that tournament and then goes further: it lets the winners breed.
When this model helps, and when it misleads
Use the iterated prisoner's dilemma to reason about repeated interactions where both sides can help or exploit each other and both remember: trade partners, feuding neighbours, firms in a price war, animals grooming each other. The model earns its keep by showing that cooperation needs no altruism, no central enforcer, and no communication. It can emerge from selfish players who simply expect to meet again.
Do not read the payoff numbers as money or utility in any real case. They are a ranking, and the model only requires the order T \gt R \gt P \gt S plus 2R \gt T + S (so that alternating exploitation beats steady cooperation for neither side). Real conflicts also have more than two moves, imperfect memory, side deals, and reputations that travel. Treat the tournament as a clean thought experiment that isolates one mechanism, not as a forecast.
The condition 2R \gt T + S matters. With the defaults, 2(3) = 6 > 5 + 0 = 5, so two players who take turns exploiting each other (average 2.5 per round) do worse than two steady cooperators (3 each). If you edit the payoffs so that T + S > 2R, the whole logic of the tournament changes.
The strategies on the roster
- Always Cooperate
- Plays C every round. Generous and exploitable.
- Always Defect
- Plays D every round. Safe and predatory.
- Tit for Tat
- Cooperates first, then copies the opponent's last move. Nice, retaliatory, forgiving, clear.
- Generous Tit for Tat
- Like Tit for Tat, but forgives a defection with some fixed probability, breaking chains of mutual punishment.
- Grudger
- Cooperates until the opponent defects once, then defects forever.
- Pavlov
- Win-stay, lose-shift. Repeats its last move after a good payoff (
TorR), switches after a bad one (PorS). - Random
- Cooperates with probability
0.5each round. - Prober
- Opens with a defection to test the opponent, then plays like Tit for Tat unless it spots a pushover to exploit.
Scoring a match, with the formula
A match is a sequence of rounds. Each round produces one payoff for each player from the table below. The match score is the sum over all rounds.
Here V_i is player i's total, n is the number of rounds, a_i^t is player i's move in round t, and g is the payoff function: g(C,C)=3, g(D,C)=5, g(C,D)=0, g(D,D)=1.
A useful shortcut for a long match is the average per round. Two Tit for Tat players lock into mutual cooperation from round one, so each averages 3.0. Two Always Defect players average 1.0. Tit for Tat against Always Defect cooperates once (scoring 0), then defects forever (scoring 1), so over n rounds it averages (0 + (n-1)\cdot 1)/n, which is 0.99 over 100 rounds. The Always Defect side scores 5 once then 1 thereafter, averaging 1.04. It wins that pairing by a hair, but only a hair.
A worked round-robin on the demo roster
Run the demo defaults: the eight strategies above, payoffs 5/3/1/0, 100 rounds per match, no noise, round-robin mode. Every strategy plays every other and also a copy of itself. Below are per-match totals for the four cleanest pairings, computed with the rules above over 100 rounds.
Four exact match totals
- Tit for Tat vs Tit for Tat. Both cooperate every round. Each scores
100 × 3 = 300. - Always Defect vs Always Cooperate. The defector scores
100 × 5 = 500; the cooperator scores100 × 0 = 0. This is the largest exploitation in the tournament. - Tit for Tat vs Always Defect. Tit for Tat plays C in round 1 (scores
0), then D for rounds 2 to 100 (scores99 × 1 = 99), total99. Always Defect scores5in round 1 then99 × 1 = 99, total104. - Grudger vs Always Cooperate. Neither ever triggers the grudge, so both cooperate throughout:
300each.
The pattern is the point. A nice strategy never loses a pairing by more than 5 points, yet it collects the full 300 against every other nice strategy. Always Defect wins several pairings by 5 points each, but earns only 1 per round against every retaliator it meets. Add up the whole table and Tit for Tat's steady 300s outrun Always Defect's scattered wins.
From leaderboard to evolution
Round-robin ranks strategies once. Evolutionary mode asks a harder question: if the population is a mix of strategies and the successful ones reproduce, what mix survives? Start each strategy with an equal share of the population. Each generation, replay the tournament, let each strategy's average score set its fitness, then update shares by the replicator equation.
Here x_i^t is strategy i's share of the population in generation t, f_i is its average payoff against the current population mix, and \bar{f} is the population-wide average payoff. A strategy that scores above the average grows; one below the average shrinks. Shares always sum to 1.
This makes success frequency-dependent. When cooperators are common, Always Defect scores far above average and its share climbs fast. But as it climbs, it starts meeting itself, scoring only 1 per round, and its fitness collapses. Meanwhile retaliators, sheltered from exploitation after round one, keep scoring near 3 among themselves. Always Defect eats the naive cooperators, then starves. The retaliators inherit the wreckage.
Watch the stacked area chart in two phases. Early generations: Always Defect and Prober swell while Always Cooperate crashes. Later generations: with easy prey gone, the defectors shrink and the Tit-for-Tat family fills the space. The crossover is the whole story.
Why noise changes the winner
Turn the noise slider up and each move flips with some small probability. At 1% noise, roughly one move in a hundred is the opposite of what the strategy intended. This is where plain Tit for Tat breaks.
Two Tit for Tat players cooperating happily hit one accidental defection. Player A now defects to punish, so B punishes back, so A punishes again. They lock into alternating defection: D,C,D,C,..., each averaging (5+0)/2 = 2.5 per round instead of 3. A single error costs them the difference for the rest of the match. Two Grudgers are worse: one accidental defection triggers permanent mutual defection at 1 per round.
Generous Tit for Tat fixes this. By forgiving a defection with probability around 0.1, it breaks the echo. When A accidentally defects and B retaliates, A sometimes cooperates anyway, and the pair snaps back to mutual cooperation. Under noise the generous strategy climbs the leaderboard while the unforgiving ones fall. The interactive below lets you watch the crossover.
Common mistakes when reading the results
Do not conclude that Always Defect is a bad strategy because it lost. It wins every single pairing it plays. It loses the tournament only because it wins by tiny margins (5 points) while collecting the worst possible ongoing score (1 per round) against everyone who retaliates.
Three traps catch new readers. First, judging a strategy by whether it beats its opponent head to head. Tit for Tat never beats anyone: it ties or loses every match, yet it wins the tournament on totals. The prisoner's dilemma is not zero-sum, so your score against a partner matters more than the gap between you.
Second, forgetting that evolutionary results depend on the starting mix. Always Defect only starves after it has eliminated the naive cooperators. Seed the population with no Always Cooperate and the defector's early boom shrinks. Change the roster and you change who inherits the population.
Third, treating a short match like a long one. With only 5 rounds, the round-1 defection by Always Defect against Tit for Tat is a fifth of the match, and defectors do relatively better. Long matches (100+ rounds) let cooperation compound and are where the classic result holds.
Related simulations on this site
The prisoner's dilemma is one of several models where simple local rules produce a surprising collective outcome. If cooperation over a shared resource interests you, the Tragedy of the Commons shows a stock collapsing under selfish harvesting and recovering under quotas. For agents sorting by a mild local preference, see Schelling's Segregation Model. The Pareto Frontier Evolution tool uses the same replicator idea to breed designs, and Opinion Spread traces how a rule on a network settles into consensus or division. For a population model with boom-and-bust dynamics, try the Predator–Prey Simulator.
Frequently asked questions
Why does Tit for Tat win if it never beats anyone?
Because the game is not zero-sum. Tit for Tat ties every nice player at 300 and loses to Always Defect by only 5 points. Always Defect wins several matches by 5 but scores just 1 per round against every retaliator. Summed across all pairings, steady 3-per-round cooperation outweighs a handful of 5-point wins.
What is the best number of rounds to use?
Use at least 100 rounds for the classic behaviour. In short matches the opening moves dominate and defectors gain an edge. The strategies also work best when neither player knows exactly which round is the last, so a fixed long match is a fair stand-in.
Does adding noise ever help a strategy?
Noise never helps unforgiving strategies. It helps generous ones relative to the field, because forgiveness recovers from accidental defections that would otherwise start a vendetta. At 1% noise a plain Tit-for-Tat self-match drifts toward 2.5 per round, while a generous self-match stays near 2.9.
Why does Always Defect grow and then shrink in evolutionary mode?
Its fitness is frequency-dependent. When cooperators are common it scores far above average (up to 5 per round against Always Cooperate) and its share climbs. As it grows it plays itself more often, scoring only 1 per round, so its fitness drops below average and its share falls back.
What do the four payoff numbers have to satisfy?
Two conditions. The ranking T \gt R \gt P \gt S makes it a prisoner's dilemma at all. The extra condition 2R \gt T + S makes steady cooperation beat taking turns to exploit. With the defaults, 5 > 3 > 1 > 0 and 6 > 5, so both hold.