The last two articles explained two bad options. One is to peek freely and inflate your false-positive rate. The other is to lock yourself into a single check at a fixed horizon and lose the ability to react early entirely — whether that means shipping a good result sooner or rolling back to what you had before. Group Sequential Testing (GST), the framework introduced today, offers a third option: look several times, on a plan you commit to in advance, while keeping your error probability controlled within 5%.
What Group Sequential Testing Actually Is
Group Sequential Testing refers to any experiment design where you pre-specify a small number of interim checks — "looks" — at planned points during data collection, assign each one a statistical threshold, and calibrate those thresholds so the combined false-positive rate across every look still equals your target (typically 0.05). You're not testing just once — you're deliberately testing several times, designed so that the statistical error doesn't create either of the two problems described above.
This single idea solves both problems from the previous two articles at once:
- It solves uncontrolled peeking, because the number of looks and their timing are fixed before the experiment starts. There's no "check whenever you feel like it" — each look's threshold is already tied to the plan for every other look on the schedule.
- It solves the fixed-window problem, because you get more than one statistically legitimate checkpoint. If a guardrail metric is clearly collapsing at the first interim look, you can trigger an early stop for harm — instead of "wait until day 7 no matter what."
The Mechanics, at a Conceptual Level
Think of your total error budget — say, — as money you're allowed to spend across the whole experiment. Group sequential testing decides, in advance, how much of that budget gets spent at each look:
- Early looks are assigned only a small slice of the budget. Because so little alpha is being spent this early, the bar for declaring significance is very high — stopping on day one or two requires strong, unambiguous evidence.
- Later looks are assigned progressively more budget. As you approach the planned end of the experiment, the threshold relaxes toward the familiar value (1.96 for a two-sided 5% test).
- The looks add up to exactly your target. No matter how many times you check, as long as you stick to the pre-registered schedule and thresholds, the overall false positive rate stays at 5% — not 18%, not 64%.
The resulting sequence of thresholds — one per look — is called a boundary. Different rules for how aggressively the boundary shrinks over time produce differently named designs (O'Brien–Fleming, Pocock, and others), each making a different trade-off between "how easy it is to stop early" and "preserving power for the final look." How to construct these boundaries — including a worked numeric example — is covered in the next article, A Worked Example: Spending Your Alpha Across Three Interim Looks.
When Group Sequential Testing Is Worth It
Not every experiment needs this machinery. Group Sequential Testing earns its complexity when at least one of the following is true:
- There's real operational pressure to decide early — safety concerns, expensive traffic, or any other situation where waiting for the full fixed sample is impractical.
- Stakeholders are going to check the dashboard anyway. If informal peeking is happening regardless, a pre-registered schedule at least makes that error explicit and controlled.
How It Compares to the Alternatives
- Fixed-horizon test (one look): simplest and perfectly valid if you truly never peek before the predetermined end date (day 7, say). But as covered in the previous article, it has zero flexibility to catch a problem early.
- Uncontrolled peeking: what most teams do by default. As covered in the last two articles, it inflates the false-positive rate.
- Bonferroni correction across k looks: a valid but blunt fix. It treats every look as equally important, which not only ignores that consecutive looks share overlapping data, but also splits alpha too evenly across looks — leaving far less alpha available at the final check than a properly shaped boundary would.
- Fully continuous ("always-valid") sequential testing: instead of restricting looks to a plan, this lets you check literally at any moment. It's more flexible than GST, but it spends that much more alpha to buy the flexibility. There's also an operational cost to checking often — the query cost of every refresh.
What Comes Next
Everything above is theory. The next article, A Worked Example: Spending Your Alpha Across Three Interim Looks, puts it into practice with an actual numeric example — a real interim schedule with real thresholds you could use on tomorrow's experiment. We'll also look at the simplest way to sanity-check your alpha budget.