Burrows' Delta and Stylometry, Explained

After reading this, you will understand how Burrows' Delta turns two texts into style fingerprints, how it converts those into a single distance number, and when that number is trustworthy enough to say two texts might share an author.

What stylometry actually measures

Stylometry studies the countable habits of writing. Not what an author says, but how often they reach for small structural words without thinking. Consider two questions: how often does a text use the, and how often does it use whale? The second answer tells you the topic. The first tells you something about the person, because writers use the, of, and, to, and that at rates that stay remarkably stable across everything they write and vary noticeably between people.

Burrows' Delta, introduced by John Burrows in 2002, is the standard first tool for authorship attribution. It ignores content words entirely and compares the frequencies of the most common function words. The output is a distance: a small number means the two style profiles are close, a large number means they are far apart. This tool then checks that distance against how much each text differs from itself, and gives a same-author or different-author reading.

The hook: function words are hard to fake because you rarely notice using them. You can change your vocabulary to disguise a topic, but keeping your rate of the steady at, say, 6.2 percent while consciously suppressing it is exhausting and usually leaves other markers untouched.

When to use it, and when not to

Delta answers one narrow question well: given two texts, is their style close enough that a single author is plausible? It works best on running prose of the same broad kind: two essays, two chapters, two blog posts.

It struggles in four common situations. Translation replaces the original author's function words with the translator's. Heavy editing smooths one text toward a house style. Genre shifts change function word rates on their own, so a writer's fiction and their technical manual will look different even though it is one person. And short texts are simply noisy. Below roughly 1,500 words per side, the frequency of the can swing by a full percentage point from sampling alone.

Delta never proves authorship. A close result says the styles are consistent with one author, not that they are the same person. A jury-grade claim needs a candidate pool, controls, and a domain expert. Treat the verdict here as a screening step.

The formula and the intuition behind it

Start by choosing the n most frequent words across your reference material (this tool uses a fixed list of common English function words). For each word i, measure its relative frequency in a text: occurrences divided by total words.

Raw frequencies are not comparable across words, because the at 6 percent and that at 0.8 percent live on different scales. Burrows fixes this with a z-score. For word i, take its mean frequency \mu_i and standard deviation \sigma_i across your corpus, then standardize each text's frequency:

z_i = \frac{f_i - \mu_i}{\sigma_i}

Here f_i is the word's frequency in one text, \mu_i is its average frequency, and \sigma_i is how much that frequency varies. A z-score of +1.5 means this text uses the word 1.5 standard deviations more often than average. Now every word is on the same scale.

Delta is then the mean absolute difference between the two texts' z-scores:

\Delta_{AB} = \frac{1}{n} \sum_{i=1}^{n} \left| z_i^{A} - z_i^{B} \right|

The symbols: n is the number of words in the list, z_i^{A} is the z-score of word i in text A, and z_i^{B} is the same for text B. The vertical bars mean absolute value, so over-use and under-use both count as distance. Averaging keeps the scale interpretable: a Delta near 0 means the profiles match word for word, and each unit is roughly one standard deviation of disagreement per word.

A worked example with three short words

Standardizing and comparing two texts

To keep the arithmetic checkable, use three words instead of the full list. Suppose across your corpus these are the means and standard deviations, and these are the measured frequencies in texts A and B.

Frequencies (percent of all words) and corpus statistics
WordText AText BMean μStd σ
the6.205.906.000.40
of3.103.803.500.30
and2.902.602.700.20
  1. z-scores for A: the = (6.20 − 6.00)/0.40 = +0.50; of = (3.10 − 3.50)/0.30 = −1.333; and = (2.90 − 2.70)/0.20 = +1.00.
  2. z-scores for B: the = (5.90 − 6.00)/0.40 = −0.25; of = (3.80 − 3.50)/0.30 = +1.00; and = (2.60 − 2.70)/0.20 = −0.50.
  3. Absolute differences: the = |0.50 − (−0.25)| = 0.75; of = |−1.333 − 1.00| = 2.333; and = |1.00 − (−0.50)| = 1.50.
  4. Delta = (0.75 + 2.333 + 1.50) / 3 = 1.528.

A Delta of 1.528 is large. The two texts disagree by about 1.5 standard deviations per word, driven mostly by of, where A is well below average and B is above it. On three words this is only illustrative, but the mechanism is exactly what the tool runs on its full word list.

Reading the verdict against internal variation

A raw Delta of 1.528 means nothing on its own, because Delta values depend on the word list, the corpus, and the text lengths. The trick this tool uses is a self-comparison baseline. Split text A into two halves and compute the Delta between them: that is how much A differs from itself. Do the same for B. Call these the internal distances.

Now compare. If A differs from B by no more than A differs from its own two halves, the styles are indistinguishable at the resolution you have. If the A-to-B distance clearly exceeds both internal distances, the styles are different.

A same-author pattern: the A-to-B bar (0.68) sits between the two internal bars (0.62 and 0.71), so the texts differ from each other no more than each differs from itself.

In the chart above the cross-text distance of 0.68 falls inside the internal range of 0.62 to 0.71. That is the signature of a single author. If the A-to-B bar were at 1.40 while the internal bars stayed near 0.65, you would read that as two authors.

Imagine three bars: text A's internal Delta fixed at 0.65, text B's internal Delta fixed at 0.70, and a movable A-to-B Delta. When the A-to-B bar is below about 0.70 the verdict reads "same author plausible"; as you drag it past the internal bars toward 1.5 the verdict flips to "different authors", with the gap widening.

Common mistakes that ruin a comparison

The most frequent error is feeding in texts that are too short. At 500 words, a single repeated that shifts its frequency by 0.2 percent, which can move a z-score by half a unit. Aim for 1,500 words or more per side, and treat anything shorter as suggestive at best.

The second is mixing genres. A novelist's dialogue-heavy chapter uses the far less than their descriptive chapter. Compare like with like. The third is forgetting about translation and editing: both replace the author's function-word habits with someone else's, so a genuine match can read as a mismatch, and two edited house-style pieces can read as a false match.

The fourth is over-reading a single number. Delta ranks candidates; it does not certify. If you have three possible authors, compute Delta against all three and look at the ordering, not just whether one crosses a threshold.

Related tools on this site

Delta rests on word counts, so the counting tools pair naturally with it. The Word Frequency & N-gram Analyzer shows you the raw function-word rates that feed the z-scores. The Word & Character Counter confirms each text clears the 1,500-word floor. To see whether two drafts differ in wording rather than style, the Text Diff Checker compares them line by line. For a different lens on writing habits, the Readability Scorer measures sentence and word length instead of function-word rates, and the Repeated Phrase Finder surfaces phrasing tics that Delta ignores.

Frequently asked questions

Does a low Delta prove the same author wrote both texts?

No. It shows the two style profiles are close enough that one author is plausible, given your texts and word list. Proof of authorship needs a controlled candidate pool and expert judgement. Read the verdict as a screen, not a conclusion.

How many words do the texts need?

Aim for at least 1,500 words each. Below that, function-word frequencies swing enough from sampling alone that a genuine match and a mismatch can look the same. The longer both texts are, the more stable every z-score becomes.

Why ignore content words like names and topics?

Content words track subject matter, which any author can change at will and which two texts on different topics will never share. Function words like the and of reflect unconscious habit, so they separate authors rather than topics.

Can someone fool Delta on purpose?

With effort, partly. Consciously changing function-word rates is possible but hard to sustain across thousands of words, and disguising one marker usually leaves others intact. Translation and heavy editing fool it far more easily, because they rewrite the function words without anyone trying to.

Why compare against each text's internal variation?

Raw Delta values are not calibrated across word lists and text lengths, so a value of 0.9 means nothing by itself. Splitting each text in half gives a same-author baseline for that specific text. If A differs from B no more than A differs from itself, the styles are indistinguishable.