A Worked Example: Spending Your Alpha Across Three Interim Looks

In the last article we defined Group Sequential Testing conceptually: spend your alpha budget across a small number of pre-planned looks instead of all at once. This article makes that concrete with actual numbers — a three-look schedule you could use on a real experiment tomorrow, plus the single most common mistake teams make when they try to justify it.

The Schedule

Suppose you plan three interim looks on a one-week experiment: day 1, day 3, and day 7 (the final, planned look). You assign each one a nominal, two-sided p-value threshold that must be cleared to stop the experiment at that point:

  • Day 1: stop only if p<0.001p < 0.001
  • Day 3: stop only if p<0.01p < 0.01
  • Day 7 (final): stop if p<0.04p < 0.04

Notice the shape: the threshold is extremely strict on day 1, loosens by day 3, and is close to (but slightly under) the usual 0.05 by the final look. That shape is deliberate — it's the same idea from the previous article: spend very little of your error budget early, and let the threshold relax as you approach the end.

The Simplest Possible Justification

The fastest way to sanity-check a schedule like this is to ask: if all three looks were independent tests, what's the chance that at least one of them shows a false positive by chance, given no real effect?

Treating the three thresholds as independent, the probability of not crossing any of them is:

(1−0.001)×(1−0.01)×(1−0.04)=0.999×0.99×0.96=0.94945(1 - 0.001) \times (1 - 0.01) \times (1 - 0.04) = 0.999 \times 0.99 \times 0.96 = 0.94945

So the probability of crossing at least one threshold by chance — a false positive somewhere in the schedule — is:

1−0.94945=0.050551 - 0.94945 = 0.05055

That's slightly above the standard α=0.05\alpha = 0.05 target — not the reassuring near-exact match it's tempting to read into a number this close. Taken at face value, the independence math says checking three times instead of once has nudged your overall false positive rate a bit past where you wanted it, as long as the independence assumption holds. It doesn't hold, and — as the next section shows — that turns out to cut in your favor, not against it.

The Catch: The Looks Aren't Actually Independent

They aren't. The day 3 look uses all the data collected by day 1, plus two more days. The day 7 look uses everything from day 3, plus four more days. Every look shares data with every look before it. That means the test statistics at day 1, day 3, and day 7 are positively correlated — not independent — and multiplying survival probabilities as if they were independent is only an approximation.

In practice, this approximation tends to be reasonably close for a small number of looks with thresholds shaped this way (very strict early, relaxing later), which is exactly why it's a popular back-of-envelope check. But "reasonably close" is not the same as exact, and the size of the error depends on how much the looks overlap and how the thresholds are spread across the schedule.

Doing It Properly: Accounting for the Overlap

Above, we assumed the three looks — day 1, 3, and 7 — share data with each other, but the simple calculation ignored that overlap and just multiplied the three numbers together as if the looks were independent. Doing this properly means accounting for exactly how much the looks overlap, and that isn't a matter of a few multiplications anymore — it takes a substantially more involved computation.

There are tools built specifically to do this for you. They're called alpha spending functions, and names like O'Brien–Fleming, Pocock, and Kim–DeMets refer to exactly this kind of tool. Give one only "when you plan to look" and "what your total error budget is," and it computes the exact threshold at every look, overlap and all. The Kim–DeMets version is especially convenient in practice: a single parameter, ρ\rho, controls how strict the early looks are, while the underlying formula stays simple. Raise ρ\rho and the early looks get stricter; lower it and the schedule flattens out.

Properly solving our day 1 / day 3 / day 7 schedule under Kim–DeMets with ρ=2\rho = 2 gives:

Hand-picked thresholds vs. Kim–DeMets (ρ=2), day 1 / 3 / 7 schedule
LookHand-picked thresholdTool-computed threshold (ρ=2)
Day 10.001≈ 0.001
Day 30.01≈ 0.0084
Day 7 (final)0.04≈ 0.045

The interesting part is how close these two columns are. The shape we picked by gut instinct — extremely strict early, loosening up later — turns out to nearly match what a tool that properly accounts for the overlap produces on its own. That's a coincidence specific to this particular schedule, though, not a general rule: with more looks, or looks spaced less evenly than day 1 / 3 / 7, a hand-picked guess and a properly computed boundary can diverge a lot more than they do here.

Revisiting the Earlier Approximation

Recall the simple multiplication earlier gave an overall error rate of 5.055%5.055\% — a bit over 5%, which looked mildly concerning. Once you properly account for the overlap, the actual value comes out to 4.6%4.6\% — slightly under 5%, not over it.

Why does it come out lower? Because the looks share the same underlying data and move together, they're less likely to disagree with each other than truly independent looks would be. That same tendency to move together also lowers the chance of landing in the "happens to look like a win by chance" error somewhere along the schedule.

So the simple, hand-computed schedule turns out to be slightly conservative relative to the properly-computed version — not the other way around. Rather than concluding the approximation is inaccurate and abandoning it, the more useful takeaway is knowing which direction the error runs: it's a good tool for making a fast case to people, as long as you know it tends to lean conservative, not liberal.

What Comes Next

Deriving a correlation-aware boundary by hand for every possible schedule isn't practical for most teams. The next article, the GST Alpha Spending Playbook, provides a reference: boundary tables for the interim schedules teams actually use, precomputed and verified, across three common spending function styles.