What makes a gap statistically significant?
Change the gap, the spread, or the sample size. Then draw a new sample.
Callback mode uses a large-sample t approximation. With few callbacks, its p-values and confidence intervals can be unreliable.
t = observed gap ÷ standard error of the gap
Testing H₀: μB − μA = 0. Here ȳ is a sample mean, s is a sample standard deviation, and n is the number in each group.
Set the true gap to zero: how often will chance alone produce p < 0.05?
Sliders reuse the random draws so you can isolate each change. “Draw another sample” generates fresh observations.
How the simulation works
Continuous mode draws independent normal observations: group A has mean 0; group B has the chosen true gap. Both have the chosen standard deviation, whose square is the population variance. Every displayed statistic is calculated from the generated sample, including the estimated standard error.
Callback mode draws independent Bernoulli outcomes. It starts from the observed rates in bm.dta: 235/2,435 for White-sounding names and 157/2,435 for Black-sounding names. These rates are used as simulation probabilities, not treated as known population truths. Adjusting the gap changes group B's probability. A binary outcome's variance is p(1−p), so its spread cannot be set independently.
The two-sided pooled t-test uses df = 2n−2 and SE = √[(s²A + s²B)/n] for equal group sizes. The sample variances use the n−1 denominator. This simplified formula is the pooled-test formula because both groups have the same n; unequal group sizes would require a different pooled standard error. It matches conventional OLS with an intercept for reg outcome group, with A = 0 and B = 1. Stata's default ttest instead reports A − B: the sign reverses, but the two-sided p-value is unchanged. The t-test is exact under the independent, normal, equal-variance continuous model; the callback version is a large-sample approximation, not an exact test for binary outcomes. If neither group has variation, no t-test is reported. This is an independent-sampling illustration, not a replication of the original within-advertisement randomization.
The p-value is the probability, under the zero-gap null model, of a t-statistic at least as extreme in absolute value as the one observed. It is not the probability that the null hypothesis is true. Statistical significance does not measure the size or practical importance of an effect.
Repeated experiments use fresh observations with the current population settings. Under the continuous zero-gap model, about 5% are significant in the long run; any batch of 200 will fluctuate. When the true gap is nonzero, the fraction detected illustrates statistical power.