Spaced Repetition, Scheduled Three Ways
After reading this you will understand how SM-2, FSRS and half-life regression each decide when to review a flashcard, why those decisions produce different daily workloads and retention, and how to pick a retention target that fits your time.
What the simulator does
Three copies of one flashcard deck run at the same time. Each copy uses a different scheduler: SM-2 (the classic Anki algorithm), an FSRS-style rule, and a Duolingo-style half-life regression. The same simulated student answers all three decks on the same simulated days, so any difference you see in workload or retention comes from scheduling alone.
Here is the hook. Suppose you have a 500-card deck and you set your retention target to 90%. After one simulated year, all three schedulers keep roughly 88 to 91% of reviewed cards retrievable. But the daily review counts differ by a factor of two or more. One scheduler asks for 40 reviews a day, another for 22. Same deck, same student, same memory model. The gap is entirely how each algorithm spaces its reviews.
The memory model underneath
Every card in every deck follows the same two-quantity model. The first is stability S, measured in days: how slowly the memory decays. The second is retrievability R: the probability you recall the card right now. They are linked by an exponential forgetting curve.
Here t is the number of days since the last review and S is the current stability. When t = S, retrievability is \exp(-1) \approx 0.368, so stability is the time it takes recall to fall to about 37%. This is the modern restatement of the curve Hermann Ebbinghaus measured on himself in 1885, memorizing nonsense syllables and re-testing at fixed delays. His savings dropped fast at first and then leveled off, exactly the shape of a decaying exponential.
A successful review multiplies stability. A lapse (a failed review) collapses it back toward a small value. The spacing effect falls straight out of this: reviewing a card while R is still high (say 0.98) barely raises S, while reviewing after R has dropped to 0.85 raises it much more. That is why cramming, which repeats cards while they are fresh, wastes effort.
How the three schedulers differ
All three want long intervals with few lapses. They disagree on how to estimate the next interval.
- SM-2 (Anki classic)
- Each card carries an ease factor, starting at 2.5. On a correct answer the next interval is the current interval times the ease factor, so intervals grow geometrically: 1 day, 6 days, then roughly 15, 37, 92 days. A lapse resets the interval to 1 day and drops the ease factor by 0.2. SM-2 never measures retrievability directly. It just multiplies.
- FSRS-style
- This rule tracks S explicitly and reviews each card exactly when its retrievability falls to your target. If your target is 0.90, it solves 0.90 = \exp(-t/S) for t, giving t = -S \ln(0.90) \approx 0.105\,S. Higher stability means a longer wait, and the target directly sets how often you see each card.
- Half-life regression (Duolingo-style)
- Half-life is the time for recall to fall to 0.5, which equals S \ln 2 \approx 0.693\,S. This rule keeps a running half-life estimate, doubles it on a success, and collapses it on a lapse. It is coarse but cheap, which is why a system serving millions of learners used a version of it.
These are simplified caricatures for building intuition, not the production algorithms. Real FSRS fits about 17 parameters to your review history. Treat the numbers here as illustrative of the mechanics, not as benchmarks.
The retention target, turned into an interval
The single most instructive control is the retention target, because it converts directly into workload through the forgetting curve. For the FSRS-style rule, the interval after a review is set by solving the curve for the target R_{\text{target}}.
Read this carefully, because it drives everything. If a card has stability S = 30 days:
- Target 0.80: t = -30 \ln 0.80 = 30 \times 0.223 = 6.7 days.
- Target 0.90: t = -30 \ln 0.90 = 30 \times 0.105 = 3.2 days.
- Target 0.95: t = -30 \ln 0.95 = 30 \times 0.0513 = 1.5 days.
Going from 90% to 95% retention more than doubles the review frequency for the same card. That is the workload-retention trade-off in one line: the interval shrinks as the logarithm of your target, so the last few percent of retention cost the most reviews.
A worked year with the demo deck
500 cards, 90% target, one simulated year
Load the demo (field defaults: a 500-card deck, retention target 0.90, one year of daily study). Follow the arithmetic for a single card that never lapses, then scale up.
- A new card starts at stability S = 1 day. First interval at target 0.90: t = -1 \times \ln 0.90 = 0.105, rounded up to 1 day.
- A correct review roughly triples stability in this model, so S goes 1, 3, 9, 27, 81, 243 days over six successful reviews.
- At S = 27 the FSRS interval is -27 \ln 0.90 = 2.84 days short of that, so intervals lag stability. Once S = 200 the interval reaches 21 days, the maturity cutoff.
- A card needs about 6 to 7 successful reviews to mature. Across 500 cards introduced over the year, the deck reaches roughly 320 mature cards by day 365.
Now the workload. SM-2 multiplies by a fixed 2.5 ease and reaches long intervals fastest, so it settles near 22 reviews a day. The FSRS-style rule pins each card to exactly 90% retrievability, spending about 30 reviews a day but holding measured retention tightest. Half-life regression doubles on success, which is aggressive, so it front-loads reviews then thins out to about 25 a day, with slightly more lapses because doubling sometimes overshoots.
| Scheduler | Reviews/day | Observed retention | Mature cards |
|---|---|---|---|
| SM-2 | 22 | 0.885 | 305 |
| FSRS-style | 30 | 0.902 | 322 |
| Half-life regression | 25 | 0.874 | 298 |
Reading and interpreting the results
Three numbers matter after a run.
Reviews per day is your time cost. At 10 seconds per card, 30 reviews is 5 minutes; 22 reviews is under 4. Small per-card savings add up across months and across a 5,000-card deck.
Observed retention is the fraction of due reviews answered correctly. Compare it to your target. If FSRS targets 0.90 and observes 0.902, its model is well calibrated. If SM-2 observes 0.885 against a nominal 0.90, its fixed multiplier is slightly too aggressive for this student, which is the known reason FSRS was built.
Mature cards counts cards whose interval has reached 21 days. This is your real progress: cards you will keep for months, not cards you saw once yesterday. A scheduler that shows high daily reviews but few mature cards is churning, not building.
Judge a scheduler by mature cards per review, not by raw retention. A deck that hits 92% retention but matures only 250 cards is spending your time on repetition you did not need.
Common mistakes
Do not set retention to 0.99 hoping to "really learn" the material. The interval formula punishes you: at 0.99 the interval is -S \ln 0.99 = 0.01\,S, one hundredth of stability, so a 100-day card is reviewed every day. You will drown in reviews and mature almost nothing.
A second mistake is confusing stability with retrievability. A card can have huge stability (S = 200 days) and still show low retrievability if you have not seen it in a year. Stability is durability; retrievability is your recall right now.
A third is treating these caricatures as tuned to you. The multipliers here are fixed. Real algorithms fit your history, so your actual mileage differs. Use the simulator to feel the shape of the trade-off, not to predict your exact Anki load.
A fourth is ignoring lapses. Every failed review collapses stability and restarts the interval ladder. A scheduler with slightly higher retention but far fewer lapses often wins on total workload, because lapses are expensive: they throw away weeks of interval growth.
Related tools
To estimate the daily cost and backlog of your own real deck, use the Flashcard Workload Planner. To see the single-card forgetting curve flatten as you add reviews, open the Forgetting Curve Visualizer. To watch a tutoring system infer mastery from right and wrong answers rather than from scheduling, try the Knowledge Tracing Simulator.
Frequently asked questions
What retention target should I use?
90% is the common sweet spot. It holds recall high while keeping the interval at -S \ln 0.90 = 0.105\,S, ten times longer per card than 0.99 would allow. Drop to 0.85 if your deck is huge and time is short; the interval grows to 0.163\,S, cutting daily reviews by about a third for a few points of recall.
Why does FSRS-style cost more reviews than SM-2 here?
Because it hits the target precisely. SM-2's fixed 2.5 multiplier lets intervals run slightly long, so it drifts below target (0.885 versus 0.90) and saves reviews by accepting more forgetting. FSRS spends those reviews to stay calibrated.
Is half-life the same as stability?
They are proportional. Half-life is the time to reach 0.5 recall, which is S \ln 2 \approx 0.693\,S. Stability is the time to reach 0.368 recall. A card with S = 30 days has a half-life of about 20.8 days.
Does the spacing effect really come from the model?
Yes. In this model a success multiplies stability more when retrievability is lower. Review at R = 0.98 and stability barely moves; review at R = 0.85 and it jumps. Cramming keeps R near 1.0, so it wastes the multiplier.
Do my reviews leave my device?
No. The simulator runs entirely in your browser. Nothing you type or practice is sent anywhere.