Auto Analyst: Instant Statistical Report illustration

Auto Analyst: Instant Statistical Report

Drop a spreadsheet and the analyst reads it like a statistician would. It works out what every column holds (numbers, money, percentages, dates, categories, IDs or free text), audits the data for problems such as missing values that follow a pattern, duplicates, outliers, placeholder codes like 999 and inconsistent spellings, and describes each distribution. Then it tests every pair of columns with the test that fits their types, controls false discoveries across all of them, and ranks the real relationships by how strong they are, in plain sentences with a chart for each. Pick any column to see what drives it, check trends over time, and get warned when a pattern reverses inside groups (Simpson’s paradox). Your data never leaves the browser.

Runs 100% in your browser — your files never leave your device.

Notes

  • Every pair of columns gets the test that fits their types: Spearman’s rank correlation for two numbers, Mann–Whitney or Kruskal–Wallis for a number across groups, and a chi-square test (Fisher’s exact test on small 2×2 tables) for two categories. The p-values are then adjusted together with the Benjamini–Hochberg procedure, so that on average no more than one in twenty reported relationships is a false alarm. The statistics agree with SciPy to about thirteen significant digits.
  • Strength uses one scale for every test: ρ, the rank-biserial r, √η² or Cramér’s V below 0.1 is negligible, below 0.3 weak, below 0.5 moderate, and 0.5 or more strong. Findings are ranked by strength, not by p-value, because with enough rows even trivial effects become “significant”.
  • The paradox scan compares each association across all rows with the same association inside every grouping column of two to eight groups: a pooled within-group correlation, a Mantel–Haenszel odds ratio and risk difference, or a weighted within-group difference of means. It reports patterns that reverse and patterns that vanish.
  • Key drivers come from a random forest grown in the browser. Its score and the importance of each column are measured on out-of-bag rows, the ones each tree never saw, and the curves are partial dependence: the average prediction as one column changes. Drivers are predictive, not causal.
  • Data checks: values missing more for some groups than others, columns missing together, exact duplicate rows and IDs, placeholder codes such as 999 or -1 (treated as missing), spelling variants of one label (merged), values beyond three interquartile ranges, ambiguous dates such as 03/04, and columns that repeat each other.
  • Large files: the pairwise tests run on a random sample of 20,000 rows and the forest on 6,000. At that size a correlation is pinned down to about ±0.01, so the sample costs almost nothing in precision.
  • Sample data: Palmer penguins from the palmerpenguins package by Horst, Hill and Gorman (CC0; collected by Gorman, Williams and Fraser at Palmer Station, Antarctica); Berkeley 1973 graduate admissions from Bickel, Hammel and O’Connell, Science 1975; the customer churn table is invented.
  • Runs 100% in your browser — your files and data never leave your device.