Fitting a Distribution to Your Data, Explained
After reading this you will know how a curve or a pile of numbers gets matched to a named probability distribution, how the fit is scored, and how to tell a real fit from a flattering one.
What the tool does
You have a shape in mind, or a column of numbers on your clipboard. You want to know which named distribution describes it: is this normal, or is the right tail too heavy for that? Is it exponential, or does it really need a gamma with a rising then falling density?
The Distribution Sketcher answers that question two ways. If you draw a density curve, it rescales your drawing so the area under it equals 1, then fits each of 16 distributions to the curve by least squares. If you paste raw sample values, it fits each distribution by maximum likelihood and reports how well the fit tracks your data.
Here is the hook. Suppose you paste 200 waiting times with a mean of about 2 minutes and a long right tail. A normal fit will hug the middle and then predict negative waiting times, which is nonsense. An exponential fit will respect the floor at zero and follow the decay. The tool ranks the exponential above the normal, and the number that does the ranking is worth understanding.
When to use it, and when not
Reach for a distribution fit when you need a compact model of one variable: service times, error magnitudes, particle sizes, the spread of a measurement. A fitted distribution gives you a formula, a handful of parameters, and the ability to compute quantiles and tail probabilities you never observed directly.
Do not use it to prove that data came from a distribution. A fit can only fail to reject a family, never confirm it. With 30 points almost everything fits; with 30,000 points almost nothing fits perfectly, because tiny real deviations become statistically visible. Judge fits by whether the misfit matters for your decision, not by chasing a perfect score.
Fitting a single distribution assumes your values are independent draws from one fixed process. Time series with trend or autocorrelation, or data pooled from two machines, break that assumption. A two-normal mixture can absorb two subpopulations, but it cannot fix data that drifts over time.
The two fitting rules
Sketch mode and data mode optimize different objectives, so it helps to see both.
For a drawn curve, the tool samples your normalized sketch at many x values, giving heights y_i, and searches for parameters that make the candidate density f(x_i) match those heights. It minimizes squared error using Levenberg-Marquardt:
Here \theta is the parameter vector of the candidate (for a normal, \theta = (\mu, \sigma)), and the sum runs over the sampled x positions. The drawing is compared shape to shape.
For raw data, there is no drawn height to match. Instead the tool asks which parameters make your observed sample most probable, by maximizing the log-likelihood:
Each x_i is one of your n sample values. The optimizer (Nelder-Mead) walks the parameter space to push \ell as high as it will go. A larger log-likelihood means the model assigns more probability density to exactly the numbers you observed.
Scoring: closeness versus complexity
A five-parameter mixture can always match a curve at least as well as a one-parameter exponential, because it contains more freedom. If you ranked purely by fit, the most complicated family would always win and would happily memorize noise. The score therefore charges a fee for every parameter.
k is the number of fitted parameters, n is the sample size, and \ell(\hat{\theta}) is the maximized log-likelihood. Lower BIC is better. AIC uses the same shape with the penalty 2k instead of k \ln(n), so it forgives extra parameters more readily.
Work the trade for a concrete case. Suppose a normal (k = 2) reaches \ell = -305 on n = 200 points, and a two-normal mixture (k = 5) reaches \ell = -302. Under BIC the normal scores 2 \cdot \ln(200) - 2(-305) = 10.6 + 610 = 620.6. The mixture scores 5 \cdot 5.298 - 2(-302) = 26.5 + 604 = 630.5. The mixture fit is better by 3 in log-likelihood but loses by about 10 in BIC. Its three extra parameters were not worth it.
A worked example on the demo data
Load the “Reaction times (raw data)” sample and you get a right-skewed batch of positive values. Walk through what happens to it.
From numbers to a ranked list
- The values are binned into a histogram so you can see the shape. Say the sample has n = 200, mean \bar{x} = 2.0 and a long right tail reaching past 8.
- The exponential is the leanest serious candidate. The tool fits a shift x_0 alongside the rate (k = 2 parameters); with data starting near zero the shift lands at the sample floor and the rate estimate is simply the reciprocal mean above it: \hat{\lambda} \approx 1 / \bar{x} = 0.5. Its log-likelihood is \ell = n(\ln \hat{\lambda} - 1) = 200(\ln 0.5 - 1) = 200(-1.693) = -338.6.
- The gamma is fitted next: shape \alpha, scale and shift (k = 3). If the data curves up before it falls, the fit lands near \alpha = 1.6, and its extra shape parameter buys a higher \ell, say -330.0.
- BIC for the exponential: 2 \cdot \ln(200) - 2(-338.6) = 10.6 + 677.2 = 687.8. For the gamma: 3 \cdot 5.3 - 2(-330.0) = 15.9 + 660.0 = 675.9. The gamma wins by about 12, so here the extra shape parameter earned its keep.
- The normal is fitted for contrast. It scores worse and, more tellingly, places mass below zero where no data can live. The KS statistic flags that gap.
Reading the results
Every candidate reports its parameters, a BIC or AIC score, a KS statistic, and a QQ plot. Learn to read the last two.
- KS statistic
- The largest vertical gap between your empirical cumulative distribution and the fitted one. A value of
0.04means the two curves never differ by more than 4 percentage points of cumulative probability. Smaller is better. As a rough gauge, gaps above roughly 1.36 / \sqrt{n} are worth suspicion; for n = 200 that threshold is about0.096. - QQ plot
- Sample quantiles on one axis, fitted quantiles on the other. Points on the diagonal mean the fit reproduces the quantiles. A tail that bends above the line means your data has a heavier tail than the model predicts.
- Overlay
- The fitted density drawn on top of your histogram or sketch. Use it as the sanity check that a good number is not hiding a bad shape.
Let the ranking narrow the field, then let the QQ plot and overlay make the call. A model can win on BIC and still miss the tail that matters for your risk estimate.
Common mistakes
The errors below are the ones that quietly ruin fits.
- Reading BIC across incomparable data. A BIC of 670 means nothing on its own. It is only comparable between models fitted to the same dataset. Never compare a BIC from 200 points to one from 500 points.
- Trusting a hard edge that sits on your data. The uniform and scaled beta pin their support to the observed range and the triangular fits its own edges, so nothing lands outside them by construction. But a fitted edge resting exactly on your largest observation is a warning, not a feature: the very next data point may fall past it, where the model says probability zero.
- Trusting a sketch tail you did not draw carefully. In sketch mode the tool matches the curve you drew. If you let your stroke slam to zero at the right edge, no heavy-tailed family can win, because you told it the tail is thin.
- Confusing a good center with a good tail. The normal and the Student-t agree closely near the mean and disagree only far out. If your decision depends on the 99th percentile, the center fit is irrelevant; read the QQ plot at the extremes.
Related tools
The two-normal mixture here is the one-dimensional cousin of a full mixture model. To see the multivariate version with soft cluster assignments, open the Gaussian Mixture & EM tool. When your goal is grouping points rather than fitting a density, try k-Means Clustering for compact round clusters, DBSCAN Clustering for density-based clusters that ignore noise, or Hierarchical Clustering & Dendrogram when you want a merge tree you can cut at any level. If the tail is what you care about, the Anomaly Detection Playground scores outliers three ways.
Frequently asked questions
Why does a heavier model sometimes lose even though it fits better?
Because BIC and AIC subtract a penalty for every parameter. A five-parameter mixture must improve the log-likelihood by more than about 1.5 \ln(n) over a two-parameter normal just to break even on BIC. With n = 200 that is roughly 8 units of log-likelihood.
What sample size do I need for a trustworthy fit?
Rough shape distinctions (skewed versus symmetric) show up by 50 points. Distinguishing tail behavior, say Student-t from normal, often needs several hundred, because the difference lives in rare extreme values that a small sample may never produce.
Should I use BIC or AIC?
Use BIC when you want the simplest model that explains the data and you have a real sample size, since its \ln(n) penalty grows with data. Use AIC when your goal is prediction and you would rather keep a slightly richer model. They often agree; when they disagree, BIC picks the simpler candidate.
What does a low KS statistic but a bad QQ plot mean?
KS measures the single largest cumulative gap, which usually sits near the center where data is dense. The QQ plot spreads attention across all quantiles including the sparse tails. A small KS with a QQ plot that peels away at the ends means the center fits and the tails do not.
Can I fit data with negative values?
Yes. Families defined on the whole line — normal, Laplace, logistic, Cauchy, Student-t, Gumbel and the mixture — handle them natively, and in this tool the lognormal, exponential, gamma, Weibull and Rayleigh carry a fitted location shift, so they can slide left of zero and fit negative values too. Only the Pareto insists on strictly positive data and is skipped otherwise.