Practical deep-dives into ab testing, experiment design, and statistics related. The core topics every data analyst needs to master.
Why peeking at results mid-experiment inflates false positives — and how sequential testing fixes it
Every team wants their experiment to win, so everyone keeps refreshing the dashboard until it looks good — and until the guardrail metric doesn't look obviously bad. Here's why that behavior quietly inflates your false positive rate.
A fixed check-in day fixes optional stopping on the primary metric — but guardrail metrics like revenue still only get one noisy snapshot, with no way to catch harm early.
Sign in required. What Group Sequential Testing actually is, when it's worth the complexity, and how it solves both the peeking problem and the fixed-window problem at once.
Premium. A concrete day 1 / day 3 / day 7 threshold schedule, the simplest way to sanity-check it, and why that simple check is only an approximation once you account for correlated looks.
Premium. Precomputed, Monte Carlo-verified boundary tables for the four most common interim schedules across three spending function styles — twelve cases in one reference.
Every experiment has metrics you're trying to move and metrics you're trying to protect. Testing both with the same superiority test is why guardrail failures slip through unnoticed.
"Not significant, so we're good" is the most common misreading in experimentation. A walkthrough of how an underpowered guardrail lets a real regression ship, quietly, again and again.
Sign in required. Flipping the null hypothesis to ask the question a guardrail actually needs answered — what the non-inferiority margin is, and how to read a one-sided confidence bound.
Ruling out harm down to a specific margin takes more data than detecting any difference at all — and the gap grows quadratically as the margin shrinks.
Watching several guardrails at once doesn't inflate false positives — it quietly erodes power instead. Why the fix looks like Bonferroni, applied to β instead of α.
Premium. Deriving a real revenue guardrail's margin from business impact, discovering it's statistically unaffordable, and the two honest levers that fix it.
Premium. Precomputed sample sizes across metric types, margin levels, and guardrail counts, plus a ship-decision matrix for turning goal and guardrail results into a launch call.
Two questions that sounded like the same instinct — narrow the analysis down to the users who matter. One breaks randomization outright. The other is encouragement design, and the fix keeps every user in the analysis instead of filtering any of them out.
Cross-checking the simple ITT-to-LATE division against Bloom (1984), Angrist & Pischke, an HHS policy brief, and the randomization-inference literature — what carries over exactly, and where the standard error needs 2SLS instead.
Premium. A GIF-thumbnail removal experiment, an ML engineer's question about dilution, and the encouragement-design math that turns a diluted ITT estimate into the effect that actually matters.
Premium. Why an activation subgroup defined on 'saw the treatment-relevant item' produced a 67:33 sample ratio mismatch, and how that mismatch manufactured a guardrail lift that was never real.
Premium. How the estimated treatment effect changes as the exposure rate shifts, why a lower exposure rate implies a larger true effect, and how this connects to the instrumental-variables literature.
Premium. When new, cold-start products and established products compete for the same feed slots, a simple two-arm test can't tell a direct effect from a reallocation effect. A factorial design that separates them, plus a proportionality check for spotting a hidden quality problem.
Premium. Arm D doesn't test whether a quality effect transfers to new products — that's already confounded into arm B. D exists for exactly one reason: measuring whether stronger established-product ranking quietly claws back the exposure a cold-start fix was built to win.