A method that claims to be right 95% of the time is making a statement anyone can count. Almost nobody counts it.

The interval taught first in every introductory course covers 87.6% of the time when it says 95%, and the failure is worst exactly where proportions are usually reported — a rare event in a small sample. That is not an estimate. For a proportion the sample space is finite, so the coverage is a sum over every outcome that could have occurred, and the number is exact. The same discipline applies to everything here: a test's p-values are checked for flatness, a simulation carries its seed, and every claim that a picture makes has been run across many seeds rather than one.

Coverage of four nominal 95% intervals, n = 30Computed exactly by summing over all 31 possible counts, not simulated. The Wald interval drops to 26.0% and is jagged everywhere; Clopper–Pearson never falls below 95% and pays for it in width.0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted
Fig. 1 Four intervals that all claim 95%, with their actual coverage computed by summing over every possible sample rather than simulated. The jagged line near the bottom is the one in the textbooks. The essay carries the same figure with the sample size on a slider, and the failure does not go away as the sample grows.

Start anywhere

19 essays

0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted Intervals, counted

What the 95% refers to

An interval that claims 95% is making a checkable statement about a procedure, not about the interval in front of you. Build every possible sample and count, and the interval taught first turns out to cover 87.6% of the time.

6 figures
00.1000.2000.3000.400-2024standardised sumdensityskew 0.7472/√n = 0.70740,000 sums, one seed eachthe rate is predicted, not just the shape The distribution itself

Sums of almost anything

The theorem says sums converge on one shape whatever they are sums of, which is remarkable and true. Watching it happen from a one-sided skewed source, with the rate of convergence predicted in advance, is more convincing than watching the shape appear.

6 figures
05001e+30.2000.4000.6000.800p-valuecountflatKS distance 0.0165a p-value that is not uniform is not a p-value Tests, and the second number

A p-value that is not flat is not a p-value

Under a true null, p-values are uniform. That is stronger than saying the test rejects 5% of the time, it constrains the whole distribution rather than one point of it, and it catches implementation errors that a rejection rate sails past.

6 figures
fraction of group A given the treatmentgroup B0.00.51.032% reverseshaded: the overall comparison reversesthe rates are identical everywhere here Reversals that are not errors

Simpson's reversal is a region, not a table

The treatment wins in both groups and loses overall. That is normally shown with one famous table, which cannot answer the two questions a reader has — how often, and how large. Swept, it turns out to occupy 31% of the allocation space.

6 figures
20 independent samples, 40 points eachall twenty are normalthis is what the noise looks like What makes it checkable

The seed is part of the figure

Every other site in this fleet draws from a deterministic rule, so a figure either is or is not what it claims. Here the figures are samples, and a sample can be right by luck. That changes what a figure has to carry.

5 figures
0.8000.8500.9000.950255075100sample sizecoverage at a true proportion of 0.15n = 19n = 20WilsonWaldexact coverage at every n from 10 to 120more data is not automatically better here Intervals, counted

More data is not monotonically better

Coverage of an interval for a proportion does not improve smoothly as the sample grows. It oscillates, and there are larger samples that cover materially worse than smaller ones — a sample of twenty covers twelve points worse than a sample of nineteen.

6 figures
sample sizerelative error10204080160320640128010%1%at the mediantwo sigma outthree sigma outexact binomial against its normal approximationthe tail converges last The distribution itself

The tail converges last

The central limit theorem is usually shown as a shape arriving. What the demonstration leaves out is the rate — and the rate is wildly different in the middle and in the tail, which is where every approximation in the subject is actually read.

6 figures
00.2500.5000.750100.5001true effect, in standard deviationsprobability of rejecting80%, the usual target56% at d = 0.54,000 simulated studies per pointthe second number a p-value needs Tests, and the second number

What a p-value does not say

The same p of 0.04 corresponds to a large effect in ten observations and a negligible one in two thousand. A p-value alone cannot be interpreted, and the number that makes it interpretable is almost never printed beside it.

6 figures
00.2500.5000.7501how common the condition ischance a positive result is true1 in 10,0001 in 1,0001 in 1001 in 101 in 11.8%66.7%one test, every prevalencethe base rate outweighs the test Reversals that are not errors

What a positive test is worth

A test that is 90% sensitive and 95% specific sounds accurate. For a condition affecting one person in a thousand, 98% of its positive results are wrong, and a worse test on a commoner condition beats a better test on a rare one.

6 figures
0.4000.6000.80010.2000.4000.6000.800the true proportionactual coverage of a nominal 95% intervalnominal 95%Wald — the textbook oneWilsonAgresti–CoullClopper–Pearsonsummed over all 31 outcomesno interval is 95% until it is counted What makes it checkable

Two routes to every number

A site about probability that only simulates has one route to each answer and no way to tell a right one from a plausible one. Every important number here is computed twice, by arithmetic that shares nothing, and the two are required to agree.

6 figures
20 independent samples, 40 points eachall twenty are normalthis is what the noise looks like The distribution itself

What normal actually looks like

A single quantile plot of forty normal points wanders enough to look suspicious. Twenty of them, all genuinely normal, show what the noise looks like — and any single panel a reader would have rejected is in there.

6 figures
the truth, 0.350.00.20.40.60.80 of 20 missedthe 95% belongs to the procedure, not to one interval Intervals, counted

Twenty intervals and one expected miss

The 95% belongs to the procedure, not to the interval in front of you. Twenty intervals from twenty samples make that visible in a way no definition does, and the one that misses is not a mistake.

6 figures
02505007501e+3-0.50000.5001effect the study reportsstudiesthe truth, 0.3what gets published, 0.61shaded: reached p below 0.05power 20%, inflation 2.04× Tests, and the second number

The winner's curse

Filter honest studies down to the ones that reached significance and the effects they report are systematically too large. At low power the inflation is a factor of two, nobody has done anything wrong, and the selection did all of it.

6 figures
first measurementsecond+0.55-0.59predicted 0.64 from the correlation alonenobody was treated Reversals that are not errors

Regression to the mean

Select the worst performers, measure them again, and they improve. Select the best and they decline. No intervention is required for either, the size of the apparent effect is predictable from the correlation alone, and it is the reason so many things appear to work.

6 figures
00.1000.2000.3000.400-4-202standard deviations from the meandensity68.27%95.45%99.73%the bands are integrated, not recalledsigma = 1.00 The distribution itself

The shape, and where its mass is

68, 95, 99.7 is recited more often than any other set of numbers in the subject. They are integrals of a specific curve, they are worth computing rather than remembering, and the third one is the one people misuse.

6 figures
Wald0.241covers 94.2%Wilson0.247covers 96.5%Agresti–Coull0.259covers 96.5%Clopper–Pearson0.273covers 98.3%expected width, and what it buysorange: fails its nominal levelthe shortest interval is the one that misses Intervals, counted

The shortest interval is the one that misses

Four intervals for the same data, with their widths and their coverage measured together. The narrowest is the one that fails its stated level, which is exactly why it looks the most appealing.

6 figures
00.2000.4000.6000.8005101520analyses available to the researcherchance of finding something significantthe nominal 5%if the analyses were independentcorrelated, as real ones areno effect present anywhereevery individual analysis is correct Tests, and the second number

Twenty analyses of nothing

Twenty honest, correct analyses of data with no effect in it find something significant 58% of the time. Nobody p-hacked, every individual p-value is right, and the reported one is the smallest of twenty.

6 figures
the sample mean93.6%nominal 95%the sample maximum0.0%nominal 95%3,000 samples of 40the same procedure, two statisticsresampling cannot see past the data Intervals, counted

Where the bootstrap lies

Resampling is the most generally useful trick in the subject and it has a failure mode that is easy to state: it cannot see past the data. For a statistic that lives at the edge of the sample, coverage collapses from 95% to almost nothing.

6 figures
00.1000.2000.3000.400-4-2024standard errors from the meandensityt 2.57z 1.96solid: t · dashed: normal31% wider at 5 df Intervals, counted

The correction for not knowing the spread

The t distribution exists because the standard deviation is estimated rather than known. At eight observations, using the normal instead makes every interval 12% too short — and the coverage that follows can be measured rather than argued about.

6 figures

Threads running through

themes, not chapters

Count it, do not claim it

An interval that says 95% is making a statement about a procedure, and a procedure can be run ten thousand times. Every coverage figure on this site is a count — for a proportion it is an exact sum over the whole sample space, so the answer carries no simulation noise at all.

11 essays

One run is an anecdote

A figure showing a single simulation is showing one draw from a distribution of figures. Where the claim is about behaviour rather than about one dataset, the figure runs across many seeds and reports what held for all of them — and where it shows one run, it shows twenty of them side by side.

5 essays

Two routes to a number

Every simulation here has a closed form beside it and every closed form has a simulation. A distribution function and its quantile must invert each other to ten digits; an exact coverage sum and a Monte Carlo count must agree. Neither route can confirm itself, and they share no arithmetic.

9 essays

The tail is where it is read

A p-value, a control limit and a risk figure are all tail statements, and the tail is exactly where every approximation in this subject is worst. The central limit theorem converges in the middle long before it converges where anyone looks.

2 essays

The second number

A p-value cannot be interpreted alone. The same 0.04 corresponds to a large effect at ten observations and a negligible one at two thousand, and it means nothing at all until you know how many analyses were available to produce it.

4 essays

Reversals that are nobody's mistake

Simpson's paradox, regression to the mean and the winner's curse all arise from correct arithmetic applied honestly. Each is shown as a region of a parameter space rather than one famous table, so the questions of how often and how large have answers.

7 essays