GST Alpha Spending Playbook: A 12-Case Boundary Reference for Practical Interim Schedules

The previous article showed why a hand-picked threshold schedule is only an approximation, and that doing it properly means solving a correlation-aware boundary. Most teams don't want to derive that boundary from scratch every time they plan an experiment. This article is the reference table: precomputed, verified boundaries for the interim schedules teams actually use in practice.

What's in This Playbook

Four common interim schedules — 1-3-5, 1-3-7, 1-3-7-14, and 1-7-14 days — crossed with three spending function styles, for twelve total cases. Every boundary below was computed with a recursive numerical integration procedure (the standard Lan–DeMets group sequential method) and independently verified by Monte Carlo simulation to control the overall two-sided Type I error at almost exactly 5% for every case in this table.

Assumption: information fraction is approximated as calendar-day-elapsed divided by the final day (e.g. day 3 of a 7-day schedule is treated as t=3/7t = 3/7). This is a practical default for teams that don't track exact event-count accrual. If your experiment's traffic is highly uneven across days (e.g. a big weekend spike), event-count-based information fraction will be more accurate than calendar time — the boundary construction method is identical either way, only the tt values change.

The Three Spending Styles

  • OBF-type (O'Brien–Fleming): α∗(t)=2(1−Φ(zα/2/t))\alpha^*(t) = 2\left(1 - \Phi\left(z_{\alpha/2}/\sqrt{t}\right)\right). Spends almost nothing early, preserves nearly full power for the final look.
  • Pocock-type: α∗(t)=α⋅ln⁡(1+(e−1)t)\alpha^*(t) = \alpha \cdot \ln(1 + (e-1)t). Spends alpha much more evenly, giving a real chance to stop early at the cost of a stricter final-look threshold.
  • Moderate (power family, ρ = 2): α∗(t)=α⋅t2\alpha^*(t) = \alpha \cdot t^{2}. A middle ground between the two — noticeably more permissive early than OBF, noticeably more conservative early than Pocock.

Schedule: Day 1 – 3 – 5

1-3-5 day schedule (final look = day 5)
DayInfo fraction tOBF-type z (nominal p)Pocock-type z (nominal p)Moderate (ρ=2) z (nominal p)
10.2004.383 (p ≈ 1.2×10⁻⁵)2.438 (p<0.0148)3.090 (p<0.0020)
30.6002.530 (p<0.0114)2.260 (p<0.0238)2.396 (p<0.0166)
51.0001.999 (p<0.0456)2.264 (p<0.0236)2.042 (p<0.0411)

A short schedule like this leaves little room for OBF's early conservatism to matter much — day 1 is already 20% of total information, so the OBF boundary (z = 4.383) is high but not absurd. Pocock is nearly flat across all three looks, as designed. The moderate style sits between the two, but not symmetrically — closer to Pocock at day 1 (z = 3.090 vs. 2.438) and closer to OBF by the final look (z = 2.042 vs. 1.999).

Schedule: Day 1 – 3 – 7

1-3-7 day schedule (final look = day 7)
DayInfo fraction tOBF-type z (nominal p)Pocock-type z (nominal p)Moderate (ρ=2) z (nominal p)
10.1435.186 (p ≈ 2.2×10⁻⁷)2.543 (p<0.0110)3.285 (p<0.0010)
30.4292.994 (p<0.0028)2.349 (p<0.0188)2.635 (p<0.0084)
71.0001.970 (p<0.0489)2.192 (p<0.0284)2.007 (p<0.0448)

This is the schedule used as the worked example in the previous article. Compare the OBF row here to the hand-picked 0.001 / 0.01 / 0.04 schedule from that article: the properly-computed OBF boundary is far stricter at day 1 (z = 5.186, p ≈ 2.2×10⁻⁷) and only slightly looser at day 7 (p < 0.049 vs. 0.04). The hand-picked version wasn't unreasonable — it just underestimated how extreme the earliest look needs to be once you account for how much day 1 overlaps with the final analysis.

Schedule: Day 1 – 3 – 7 – 14

1-3-7-14 day schedule (final look = day 14)
DayInfo fraction tOBF-type z (nominal p)Pocock-type z (nominal p)Moderate (ρ=2) z (nominal p)
10.0717.334 (p ≈ 2.3×10⁻¹³)2.760 (p<0.0058)3.657 (p<0.0003)
30.2144.234 (p ≈ 2.3×10⁻⁵)2.546 (p<0.0109)3.078 (p<0.0021)
70.5002.772 (p<0.0056)2.355 (p<0.0185)2.546 (p<0.0109)
141.0001.979 (p<0.0478)2.229 (p<0.0258)2.022 (p<0.0432)

With four looks over two weeks, day 1 is only 7% of total information — and under OBF, that produces a boundary of z = 7.334, which in practice is uncrossable at any plausible effect size. That's not a bug in the calculation; it's the intended behavior of OBF-type spending. The day 1 and day 3 checkpoints here function as monitoring points, not realistic stopping opportunities — the schedule is really "look meaningfully at day 7 and day 14, glance at day 1 and day 3." If your team actually wants a shot at stopping early in week one, Pocock or the moderate style is the more honest choice for this schedule.

Schedule: Day 1 – 7 – 14

1-7-14 day schedule (final look = day 14)
DayInfo fraction tOBF-type z (nominal p)Pocock-type z (nominal p)Moderate (ρ=2) z (nominal p)
10.0717.334 (p ≈ 2.3×10⁻¹³)2.760 (p<0.0058)3.657 (p<0.0003)
70.5002.772 (p<0.0056)2.227 (p<0.0259)2.504 (p<0.0123)
141.0001.979 (p<0.0478)2.214 (p<0.0269)2.019 (p<0.0435)

Dropping the day 3 look changes very little for OBF and the moderate style at day 1 and day 14 (they're barely sensitive to whether day 3 exists, since so little alpha is spent that early anyway) — but it does shift Pocock's day 7 threshold slightly looser (z = 2.227 vs. 2.355 in the four-look schedule). The cumulative alpha spent by day 7 is the same α∗(0.5)\alpha^*(0.5) either way; what changes is how much of that budget is already gone by the time day 7 arrives. With day 3 in the schedule, α∗(3/14)\alpha^*(3/14) has already been spent going in. Without it, only α∗(1/14)\alpha^*(1/14) has — a much smaller prior spend under Pocock's near-linear schedule — which leaves a larger increment to spend at day 7 itself, and a looser boundary as a result.

Why Day 1 Boundaries Are So Sensitive

Every OBF-type boundary above follows z1=zα/2/t1z_1 = z_{\alpha/2} / \sqrt{t_1} at the first look (this holds exactly, since there's no earlier look to correlate with yet). When t1t_1 is small — 0.071 for the two schedules that check on day 1 of a 14-day test — the boundary scales with 1/t11/\sqrt{t_1}, which means a small error in your estimate of t1t_1 gets amplified disproportionately. If your actual information accrual on day 1 turns out to be half of what you assumed (say, a slow traffic ramp-up), the true t1t_1 is smaller than planned and the correct boundary is even higher than the table shows. This is the practical reason early-look boundaries in short information windows should be treated as directional, not load-bearing — lean on the later, more stable checkpoints for real decisions.

What If Your Schedule Isn't Here?

Twelve cases cover the schedules we see most often, but no static table covers every combination of look count, spacing, and spending style — and event-count-based information fractions will shift every number in this article once your actual traffic pattern is uneven. The Group Sequential Boundary Calculator computes a verified boundary for your exact schedule and spending function on demand, the same way the tables above were generated for this article — included free with Tier 1 membership.