The Repeated Phrase Finder, Explained

After reading this you will know how repeated-phrase detection works, how the counts are computed, why only maximal repeats are shown, and how to read the results without chasing phrases that were never a problem.

What the tool does, with one example

The Repeated Phrase Finder scans your text for runs of words that appear verbatim more than once, then ranks them by how many times each run recurs. A run can be as short as three words or as long as a whole recycled sentence.

Here is the kind of thing it catches. Suppose a chapter contains these three sentences, spread across twenty pages:

At the end of the day, the numbers have to add up.
But at the end of the day, readers want a story.
And at the end of the day, that is what sells.

Each sentence reads fine on its own. Across the manuscript, at the end of the day appears 3 times. That is a five-word phrase repeated 3 times, so it accounts for 15 words of recycled text. The tool surfaces it because your eye rarely does: the repeats are too far apart to notice while reading, but they grate in aggregate.

When to use it, and when not

Reach for this tool late in a draft, when the words are mostly settled. It answers a narrow question: which exact strings do you lean on too hard? That is a copy-editing question, not a drafting one.

It is well suited to three jobs. First, catching copy-paste leftovers, such as a boilerplate paragraph pasted into four product descriptions. Second, finding crutch phrases like it is clear that or the fact that. Third, spotting accidental self-plagiarism, where you explain the same idea in the same words in two chapters.

Repetition is not always a fault. Refrains, legal formulas, and defined technical terms are supposed to recur. The tool reports frequency; you decide whether each phrase earns its repeats.

Do not use it to judge overall vocabulary richness or to compare two authors. For word-level habits and common bigrams, the Word Frequency & N-gram Analyzer is the better fit. For authorship questions, use the Stylometry & Authorship Comparator.

How the matching works

Before anything is counted, the text is normalized. Matching ignores case and punctuation, so At the end... and at the end, are treated as the same words. Phrases never cross a sentence boundary, so the last word of one sentence and the first word of the next cannot join into a phrase.

Inside those rules, the tool looks for repeated word sequences. The core idea is an n-gram: a contiguous run of n words. A sentence with w words contains this many n-grams of length n:

g(n) = w - n + 1

Here w is the word count of the sentence and n is the phrase length. A 9-word sentence holds 9 - 3 + 1 = 7 distinct trigrams, 9 - 4 + 1 = 6 four-grams, and so on down to a single 9-gram. The tool counts, across your whole text, how often each distinct n-gram occurs, for every length from 3 up to the longest repeat that exists.

Why only maximal repeats are listed

If you counted every n-gram length independently, one long repeat would flood the results with its own fragments. Take at the end of the day repeating 3 times. That five-word phrase contains shorter phrases that also repeat exactly 3 times:

Sub-phrases of a five-word repeat, all with the same count
LengthPhraseCount
5at the end of the day3
4at the end of the3
4the end of the day3
3at the end of3
3the end of the3
3end of the day3

The five-word phrase alone spawns 5 shorter phrases at count 3. Reporting all of them buries the finding you care about. A maximal repeat is a phrase that cannot be extended left or right without dropping its count. If the end of the day only ever appears inside at the end of the day, it is not maximal and is not listed with the same count of 3.

A shorter phrase does still appear on its own if it occurs somewhere the longer phrase does not. If the end of the day shows up 5 times total but at the end of the day only 3 of those, the shorter phrase is maximal at the extra 2 occurrences and will be reported.

A worked example you can reproduce

Counting the three sentences from the demo

Load the demo, which uses the three sentences from the first section. Work through the repeated trigrams by hand.

  1. Normalize: lowercase everything and strip commas and periods. Sentence 1 becomes at the end of the day the numbers have to add up (11 words).
  2. Sentence 2 becomes but at the end of the day readers want a story (10 words). Sentence 3 becomes and at the end of the day that is what sells (10 words). Total: 11 + 10 + 10 = 31 words.
  3. Slide a 5-word window through each sentence and tally. Only at the end of the day lands in all three sentences, so its count is 3.
  4. Check whether it can grow. To the left it is preceded by nothing, but, and and: three different words, so it cannot extend left at count 3. To the right it is followed by the, readers, and that: again three different words. It is maximal.
  5. No other 3-word or longer sequence appears more than once, so the result list has exactly one entry: at the end of the day, count 3, length 5.

That single phrase accounts for 5 \times 3 = 15 of the 31 words, or about 48% of the text. In a real manuscript the share is tiny, which is exactly why these phrases hide.

The top phrase holds a count of 3 at lengths 3, 4 and 5, then drops to 0 at length 6 because nothing longer recurs. The tool reports only the length-5 row, the maximal one.

Reading and interpreting the results

Two numbers drive every row: the count (how many times the phrase recurs) and the length (how many words it spans). Multiply them for a rough sense of wasted words. A 4-word phrase seen 6 times spends 4 \times 6 = 24 words on repetition; a 12-word sentence seen twice spends 12 \times 2 = 24 as well, but reads very differently.

Sort your attention by that product, then split it by length. Short high-count phrases are usually crutches to prune. Long repeats are usually structural: a duplicated definition, a template you forgot to edit, a quotation reused on purpose. Treat length as a signal for what kind of edit each finding needs.

A slider sets the minimum phrase length from 2 to 8 words. As you raise it, the list of reported phrases shrinks, because long exact repeats are rarer than short ones. At length 2, function-word pairs like "of the" dominate; by length 5 or more, only deliberate or accidental duplications survive.

Common mistakes

The first mistake is treating a high count as automatic guilt. In a 40,000-word book, a 3-word phrase repeated 8 times is 24 words, or 0.06% of the text. Whether that is a problem depends entirely on whether the phrase is distinctive. in the same way repeated 8 times reads worse than the United Nations repeated 8 times.

The second mistake is confusing this with vocabulary counting. This tool reports contiguous exact matches only. run quickly and quickly run are different phrases here, and run versus ran are different words. If you want lemma-level or single-word frequencies, that is a job for the frequency analyzer, not this one.

Because matching normalizes case and punctuation, phrases that differ only in an invisible character can still be reported as identical, and genuinely different Unicode look-alikes can slip past. If two "identical" strings behave oddly, inspect them with the Unicode Inspector before you trust the count.

The third mistake is running the tool too early. On a rough first draft, repetition is expected and mostly noise. Wait until the prose is close to final, or you will spend edits on sentences you were going to cut anyway.

Related tools

For plain totals of words, sentences and reading time, use the Word & Character Counter. If repeated crutch phrases are dragging your sentences long, the Readability Scorer shows the effect on Flesch and Gunning fog. To compare two versions of a passage after you cut the repeats, the Text Diff Checker highlights exactly what changed.

Frequently asked questions

Why is a phrase I can see twice not in the list?

The tool ignores phrases shorter than the minimum length (3 words by default), and it will not list a phrase if a longer phrase containing it has the same count, because only the maximal repeat is shown. Check whether your phrase is part of a longer repeat.

Does it count phrases that cross a sentence boundary?

No. Phrases stop at the end of a sentence. The last words of one sentence and the first words of the next never join into a single phrase, even if they would read as a run.

Are "the End" and "the end" counted together?

Yes. Matching lowercases everything and ignores punctuation, so The End, and the end are the same phrase for counting purposes.

Is my text sent anywhere?

No. The analysis runs entirely in your browser, so nothing you paste leaves your device.

What counts as one occurrence in overlapping text?

Occurrences are counted as non-overlapping where possible within a sentence. In practice this rarely matters, since natural repeated phrases sit in separate sentences and cannot overlap across a boundary anyway.