Bayesian Knowledge Tracing, Explained

After reading this you will be able to compute, by hand, how a tutoring system updates its estimate of your mastery after each right or wrong answer, and why guessing and slipping can push that estimate the wrong way.

What Bayesian Knowledge Tracing is

Bayesian Knowledge Tracing (BKT) is the model that mastery-based tutoring systems use to decide when you have learned a skill. Corbett and Anderson introduced it in 1994 for the Cognitive Tutors used in algebra classrooms. The same idea drives Khan Academy-style mastery bars, ASSISTments, and many adaptive homework platforms.

The model tracks one hidden fact about you: do you know the skill or not? You cannot observe that directly. You can only observe answers. A student who knows a skill can still slip and get it wrong. A student who does not know it can still guess and get it right. BKT reconciles the noisy answers with the hidden state using Bayes' rule, and reports a single number: the probability that you have mastered the skill right now.

Here is the hook. Suppose the system currently believes there is a 60% chance you know a skill, and the questions are four-option multiple choice with a 25% guess rate. You answer three in a row correctly. Your mastery estimate climbs, but not to 100%: it might land near 97%. Three correct answers under a 25% guess rate still leave a small chance you got lucky each time. That gap is the whole reason some tutors demand a streak before they let you move on.

When to use this model, and when not

BKT fits skills that are roughly binary and stable: a specific procedure like adding fractions with unlike denominators, or a single vocabulary item. It assumes that once you learn the skill you do not forget it inside the session, and that every question on that skill has the same difficulty.

Those assumptions break in predictable ways. If the skill decays over days, BKT will not model the loss, and you want the Forgetting curve visualizer or a spaced schedule instead. If one skill really bundles three sub-skills, a single mastery number hides which part you are missing. And if your questions are easy multiple choice with a high guess rate, BKT cannot separate knowledge from luck, so the estimate crawls upward slowly no matter how good the update math is.

BKT models learning within a topic, not retention across weeks. For scheduling reviews so you do not forget, pair it with the Spaced repetition simulator, which handles decay explicitly.

The four parameters and the update rule

BKT has exactly four parameters. Two describe the hidden state, two describe the noisy observation.

p(L_0)
Prior probability you already know the skill before the first question.
p(T)
Transition, or learning rate: the chance that, on any question you did not know, you learn the skill by doing it.
p(G)
Guess: the chance of a correct answer when you do not know the skill. For four-option multiple choice this is near 0.25.
p(S)
Slip: the chance of a wrong answer when you do know the skill. Careless errors, misread question, typo.

Each answer triggers two steps. First, condition on the answer using Bayes' rule. After a correct answer, the posterior probability you knew the skill is:

p(L_n \mid \text{correct}) = \frac{p(L_n) \cdot (1 - p(S))}{p(L_n) \cdot (1 - p(S)) + (1 - p(L_n)) \cdot p(G)}

Here p(L_n) is the current mastery estimate. The numerator is the chance you knew it and did not slip. The denominator adds the chance you did not know it and guessed right. After a wrong answer, swap the roles:

p(L_n \mid \text{wrong}) = \frac{p(L_n) \cdot p(S)}{p(L_n) \cdot p(S) + (1 - p(L_n)) \cdot (1 - p(G))}

Second, apply learning. Even a student who did not know the skill may have learned it by attempting the question, so the estimate carried into the next question is:

p(L_{n+1}) = p(L_n \mid \text{obs}) + (1 - p(L_n \mid \text{obs})) \cdot p(T)

The conditioned estimate is nudged toward 1 by the learning rate p(T). That is why mastery never quite falls back after a wrong answer the way it would with pure Bayes: learning pulls it up every step.

A worked example with the demo defaults

The demo uses p(L_0) = 0.3, p(T) = 0.15, p(G) = 0.25, p(S) = 0.10. Suppose the answer sequence is correct, correct, wrong, correct.

Four answers, step by step

  1. Start at 0.3. Answer 1 is correct. Numerator: 0.3 x 0.9 = 0.27. Denominator: 0.27 + 0.7 x 0.25 = 0.445. Conditioned: 0.6067. Apply learning: 0.6067 + 0.3933 x 0.15 = 0.6657.
  2. Answer 2 is correct. Numerator: 0.6657 x 0.9 = 0.5991. Denominator: 0.5991 + 0.3343 x 0.25 = 0.6827. Conditioned: 0.8776. Apply learning: 0.8776 + 0.1224 x 0.15 = 0.8960.
  3. Answer 3 is wrong. Numerator: 0.8960 x 0.10 = 0.0896. Denominator: 0.0896 + 0.1040 x 0.75 = 0.1676. Conditioned: 0.5346. Apply learning: 0.5346 + 0.4654 x 0.15 = 0.6044.
  4. Answer 4 is correct. Numerator: 0.6044 x 0.9 = 0.5440. Denominator: 0.5440 + 0.3956 x 0.25 = 0.6429. Conditioned: 0.8462. Apply learning: 0.8462 + 0.1538 x 0.15 = 0.8693.

The estimate runs 0.30, 0.6657, 0.8960, 0.6044, 0.8693. Notice step 3: one wrong answer cut the estimate from 0.896 to 0.604, a drop of nearly 0.30. A single slip with a 10% slip rate carries a lot of weight, because a well-mastered student rarely slips.

The estimate climbs on correct answers, drops sharply on the wrong answer, then recovers. It never reaches 1.0 because guessing and slipping keep uncertainty alive.

Explore how guessing changes everything

The most instructive knob is the guess rate. Raise it and correct answers stop meaning much, so mastery climbs slowly and a streak of three right answers no longer proves you learned anything.

With p(G) = 0.25 and p(S) = 0.10, five correct answers in a row starting from p(L_0) = 0.3 give estimates 0.666, 0.896, 0.966, 0.988, 0.996. Raise p(G) to 0.50 and the same five correct answers give 0.507, 0.669, 0.796, 0.882, 0.933: markedly slower, because a correct answer is now only slightly more likely from a knower than a guesser.

Reading and interpreting the estimate

The output is a probability, not a grade. A mastery estimate of 0.87 means the model puts an 87% chance on you knowing the skill given everything you have answered. Most tutors set a threshold, often 0.95, and let you advance once the estimate crosses it. In the demo defaults, a clean streak from 0.3 reaches roughly 0.966 on the third correct answer, which is exactly why "three in a row" is a common rule.

Two systems with the same threshold can behave very differently. With p(G) = 0.25 you cross 0.95 on the third correct answer. With p(G) = 0.50 you need six correct answers to cross it (0.933 on the fifth, above 0.95 on the sixth). The threshold did not change. The information per answer did.

If you design questions, cutting the guess rate helps far more than tuning the model. Moving from four options to a free-response format drops p(G) from about 0.25 to near 0.05, and mastery detection speeds up sharply.

Common mistakes

The most frequent error is reading the estimate as a test score. It is a belief about a hidden state, and it depends heavily on the four parameters you chose. Set p(S) = 0 and a single wrong answer will crash the estimate toward zero, because the model now insists that knowers never miss.

A second mistake is fitting the parameters loosely. If p(G) + p(S) \ge 1, the model becomes degenerate: a correct answer can lower your mastery estimate, which is nonsense. Beck and Chang showed in 2007 that BKT parameters are not always identifiable, meaning different parameter sets can fit the same data equally well, so trust fitted values only with enough responses per skill (hundreds, not dozens).

A third mistake is expecting BKT to remember forgetting. Standard BKT has no decay term. If a week passes between sessions, the model still reports the mastery it had at the end of the last session. Handle retention with a scheduling model rather than by twisting BKT.

Related tools on this site

BKT answers "have you learned it yet." Two neighboring questions need other tools. To decide when to review so you do not forget, use the Spaced repetition simulator. To estimate the daily review load a deck creates, and how a break builds a backlog, use the Flashcard workload planner.

Frequently asked questions

Why does mastery never reach 100 percent?

Because guessing and slipping keep a residual chance you got lucky or careless. With p(G) = 0.25 and p(S) = 0.10, even after eight correct answers the estimate sits near 0.999, not exactly 1.

Why did one wrong answer drop my estimate so much?

A low slip rate makes wrong answers strong evidence. In the worked example, p(S) = 0.10 means a knower slips only 10% of the time, so a wrong answer pulled the estimate from 0.896 down to 0.604.

How many correct answers does mastery usually require?

With typical defaults and a 0.95 threshold, three correct answers in a row is common. Raising the guess rate or lowering the learning rate pushes that number up. There is no universal answer; it falls out of the four parameters.

Can the estimate go down after a correct answer?

Only if the parameters are broken, specifically when p(G) + p(S) \ge 1. With sensible values (p(G) \lt 0.5, p(S) \lt 0.5) a correct answer always raises the estimate.

Does BKT account for forgetting between days?

No. Standard BKT assumes no forgetting once learned. For decay over days, use a spaced repetition or forgetting-curve model instead.