The GRIM and GRIMMER Tests, Explained
After reading this you can check whether a reported mean and standard deviation could have come from the sample size claimed, and spot the many published values that are arithmetically impossible.
What the GRIM test is
GRIM stands for Granularity-Related Inconsistency of Means. The idea, from Nick Brown and James Heathers in 2016, is simple arithmetic. If a study measures a whole-number score from 28 people and averages them, the total is an integer between (say) 0 and some maximum. Dividing an integer by 28 cannot produce just any decimal. It can only produce multiples of 1/28 = 0.03571.... So a reported mean of 5.19 is only believable if some integer total, divided by 28 and rounded to two places, actually gives 5.19.
Many reported means fail this check. When Brown and Heathers ran GRIM on 260 means from psychology papers, about half of the eligible papers contained at least one impossible value. Some were typos. Some were dropped participants never mentioned in the text. A few had no innocent explanation.
GRIMMER, added by Jordan Anaya in 2016, extends the same reasoning to the standard deviation. The sum of squared integer scores is also an integer, so the variance is similarly constrained. A reported SD that no integer data set could produce is inconsistent.
When to use it, and when not
GRIM works only when the raw data are integers or sums of integers. Good candidates: counts, Likert items summed or averaged, test scores out of a fixed number of questions, days, clicks. The test needs to know the granularity of the data, which is why the tool asks for items per participant. If each person answered 5 Likert items and you averaged them, the effective sample is N \times items.
Do not use GRIM on data that were never integers: reaction times in milliseconds, weights, concentrations, ratios, or any measurement on a continuous scale. Those means can land anywhere, so the test always passes and tells you nothing.
GRIM has power only when N \times items is small, below roughly 100 for a two-decimal mean. With 400 participants, every two-decimal value passes, because the spacing 1/400 = 0.0025 is finer than the rounding step. A pass is not a clean bill of health. It often just means the sample was too large for the test to say anything.
The math behind it
Let the reported mean be \bar{x}, rounded to d decimal places, from n = N \times items integer scores. The true mean is T/n for some integer total T. GRIM asks whether any integer T exists such that T/n rounds to the reported value.
Here \text{round}(\bar{x} \cdot n) is the nearest integer to the reported mean times the sample size, our best candidate total. We divide that integer back by n and check whether the result, once rounded to d decimals, matches what was printed. The tolerance \frac{1}{2} \cdot 10^{-d} is half a rounding step: half of 0.01 for two-decimal reporting, which is 0.005.
The intuition: reachable means sit on a grid with spacing 1/n. A reported value is consistent only if it is within half a rounding step of a grid point. When 1/n is larger than the rounding step, most decimals fall in the gaps and fail.
GRIMMER applies the same logic to the sum of squares. The variance times (n-1) plus n \bar{x}^2 must be an integer sum of squares, and that integer must have the right parity and range. Anaya's version rules out SD values that no integer data set can reproduce even when the mean passes.
A worked example with N = 28
Checking mean 5.19, SD 2.28, N = 28
These are the demo numbers. One participant, one score each, so n = 28 and d = 2.
- Multiply the mean by the sample size:
5.19 \times 28 = 145.32. - Round to the nearest integer total:
round(145.32) = 145. - Divide back:
145 / 28 = 5.17857.... - Round to two decimals:
5.18. - Compare with the reported
5.19. They differ. Check the neighbour totals too:146 / 28 = 5.2143, which rounds to5.21. Neither 145 nor 146 gives5.19.
No integer total between 0 and 28 produces a mean that rounds to 5.19. The value is GRIM inconsistent. With 28 whole-number scores, the reachable means near here are 5.18 (from 145) and 5.21 (from 146). The value 5.19 falls in the gap.
The gap is real, not a rounding quirk. The grid spacing is 1/28 = 0.0357, more than three times the rounding step of 0.005, so roughly two thirds of two-decimal values in this range are unreachable.
Reading the result
The tool returns one of a few states. Consistent means a valid integer total exists; the value could be genuine, and GRIM has nothing to add. GRIM inconsistent means no integer total produces the reported mean. GRIMMER inconsistent means the mean passes but the SD cannot arise from any integer data set with that mean and sample size.
Interpret a pass carefully. When N \times items is 100 or more, a pass is close to automatic and carries no information. The tool tells you the grid spacing so you can judge this yourself: if 1/n is smaller than half your rounding step, the test is silent by construction.
Reporting to more decimals sharpens the test. A mean given as 5.19 (two decimals) with N = 28 tests one grid; the same mean as 5.193 tests a finer one and is even harder to satisfy by chance. More precision in the paper means more scrutiny GRIM can apply.
Common mistakes
An inconsistent value is not proof of misconduct. The usual innocent causes, in rough order of frequency:
- Dropped participants
- The mean was computed on a subset (missing data, exclusions) different from the N printed in the table. Try nearby values of N.
- Non-integer data
- The scale allowed half points, or scores were weighted, so the data were never plain integers. GRIM does not apply.
- Wrong item count
- Averaging several items per person changes the granularity. Set items to the real number, not 1.
- Transcription typos
- A single wrong digit in the mean or the sample size flips the result. Check the source table against the abstract.
Because these are so common, treat a failure as a prompt to look closer, not a verdict. GRIM finds discrepancies; humans explain them.
Related tools
Once a data set passes GRIM, describe it with the Descriptive Statistics Calculator for the mean, variance and quartiles. If you suspect the impossibility comes from an extreme value pulling the mean, the Outlier Detector applies Grubbs' test and Tukey fences. For a different angle on data authenticity, the Benford's Law Checker tests whether leading digits follow the expected distribution, a complementary forensic check.
Frequently asked questions
Does GRIM prove a paper is fraudulent?
No. It proves a reported mean is arithmetically impossible for the stated N. The cause is often a dropped participant, a non-integer scale, or a typo. It flags where to look, nothing more.
Why does the test say nothing for my large sample?
With N \times items of 100 or more, the grid spacing 1/n falls below the two-decimal rounding step of 0.005, so every value is reachable. GRIM has no power there. It needs small samples or many reported decimals.
My mean came from averaging 5 questionnaire items. What N do I enter?
Set sample size to the number of participants and items to 5. The tool uses N \times items as the effective count. With 28 people and 5 items the grid spacing is 1/140 = 0.00714, still coarse enough to test two-decimal means.
Can GRIM check percentages or proportions?
Yes, if they come from integer counts. A percentage like 63.4% from 47 people is a mean of 0s and 1s scaled by 100, so the same grid logic applies. Enter the mean as the proportion and N as the count.
What does GRIMMER add over plain GRIM?
GRIMMER checks the standard deviation. A mean can pass GRIM while its reported SD is still impossible, because the sum of squared integers is also constrained. GRIMMER catches those cases, roughly doubling the reach of the check on typical tables.