The previous article worked through setting one margin for one guardrail by hand — including the point where the business's first number turned out to need a traffic level ten times the entire user base. Most launches carry several guardrails at once, each needing this same treatment. This article is the reference: precomputed sample sizes across metric types, margin levels, and guardrail counts, plus the decision matrix for turning all of it into an actual ship call.
What's in This Playbook
Two metric archetypes crossed with three margin levels and from 1 to 5 guardrails, for thirty total lookups. Every cell already has the power correction from the multiplicity article baked in, targeting an 80% family-wise power across the goal metric plus all guardrails, at one-sided .
Assumptions: both tables express margins the same way — as a percentage relative to the metric's own baseline — so they're directly comparable. The binomial table assumes a baseline rate of ; the percentage-point equivalent of each relative margin is shown alongside it, since that's the unit most teams actually think in for a rate metric. The continuous table assumes a coefficient of variation of 2.25 (matching the winsorized revenue example from the previous article, e.g. , ). Guardrails are treated as independent — correlated guardrails need a smaller effective in practice, which the workflow below addresses.
Table 1: Binomial Guardrails (e.g. conversion rate, p ≈ 10%)
| Guardrails (G) | Corrected power | n per arm (smallest δ) | n per arm (mid δ) | n per arm (largest δ) |
|---|---|---|---|---|
| 1 | 90.0% | 6,166,000 (δ=0.5% rel = 0.05pp) | 1,541,000 (δ=1% rel = 0.10pp) | 385,000 (δ=2% rel = 0.20pp) |
| 2 | 93.3% | 7,126,000 | 1,781,000 | 445,000 |
| 3 | 95.0% | 7,792,000 | 1,948,000 | 487,000 |
| 4 | 96.0% | 8,301,000 | 2,075,000 | 519,000 |
| 5 | 96.7% | 8,713,000 | 2,178,000 | 545,000 |
A relative percentage and a percentage point are not the same margin, and mixing them up is an easy way to badly underestimate how much data a rate-based guardrail actually needs. "1 percentage point" on a 10% baseline is a flat move to 9% or 11% — a 10% relative change. "1% relative" on that same 10% baseline is a move to 9.9% or 10.1% — a tenth the size in absolute terms, at just 0.1 percentage points. Since required sample size scales with , that tenth-the-size margin needs roughly a hundred times the sample, not ten times — which is exactly the gap between a percentage-point-denominated margin and a genuinely matched relative one.
Once Table 1's margins are expressed the same way as Table 2's, binomial guardrails turn out not to be the cheap case — at every margin level and every , this table's requirements run larger than Table 2's. That's not a coincidence: at , the binomial relative variance is larger than the continuous table's , and required sample size scales directly with that ratio (about 1.78×, which is exactly the gap between the two tables here). This is the same lesson the previous article landed on comparing a conversion guardrail against a revenue guardrail: whichever metric type looks cheaper is usually a function of how loose its conventional margin is (a percentage point or two, versus a tight relative percentage), not some inherent advantage in binomial variance. Increasing still matters far less than the margin does, though — going from one guardrail to five inflates the requirement by roughly 40% at any fixed margin in both tables, almost entirely from the power correction, not the margin itself.
Table 2: Continuous Guardrails (e.g. revenue per user, CV ≈ 2.25)
| Guardrails (G) | Corrected power | n per arm (smallest δ) | n per arm (mid δ) | n per arm (largest δ) |
|---|---|---|---|---|
| 1 | 90.0% | 3,470,000 (δ=0.5%) | 867,000 (δ=1%) | 217,000 (δ=2%) |
| 2 | 93.3% | 4,010,000 | 1,001,000 | 250,000 |
| 3 | 95.0% | 4,380,000 | 1,096,000 | 274,000 |
| 4 | 96.0% | 4,670,000 | 1,168,000 | 292,000 |
| 5 | 96.7% | 4,910,000 | 1,228,000 | 307,000 |
Continuous, heavy-tailed guardrails are expensive at a tight relative margin, same as Table 1 — and here too, the row (1 vs. 5) matters far less than the margin column. Moving from a 2% margin to a 0.5% margin costs roughly 16× regardless of how many guardrails you're running, in both tables; adding guardrails on top of that only adds another 40% or so, also in both tables. If you're tight on budget for any guardrail, continuous or binomial, loosen the margin before you consider dropping a guardrail from the suite.
The Workflow
- 1. List every guardrail the launch decision actually depends on — not everything on the dashboard, just the metrics that would genuinely block or delay a ship if they looked bad.
- 2. Classify each by metric type — binomial/rate-based, or continuous — and estimate its variance (or CV) from historical logging data, the same way the previous article pulled from a quarter of revenue history.
- 3. Set each margin via business impact inversion, exactly as in the worked example — convert an acceptable dollar or operational cost into a relative or absolute .
- 4. Count G and look up the corrected power for your guardrail count from either table above (or the formula directly, if your G exceeds 5).
- 5. Look up (or compute) required n for each guardrail at its own margin and corrected power.
- 6. Take the maximum across every guardrail's required and the goal metric's own sample size calculation. That maximum is the real sample size the experiment needs — not the goal metric's number alone, and not any single guardrail's number alone.
The Ship Decision Matrix
Once the experiment has run at the sample size the workflow above requires, the launch decision reduces to two binary questions: did the goal metric show a significant improvement, and did every guardrail clear its own non-inferiority test?
| All guardrails non-inferior | ≥1 guardrail fails non-inferiority | |
|---|---|---|
| Goal metric significant | Ship. Real improvement, no guardrail cost beyond the pre-agreed margin. | Hold. A real win exists, but at least one guardrail couldn't be confirmed safe — check whether its interval crosses meaningfully past −δ or the test was simply underpowered before deciding. |
| Goal metric not significant | Do not ship (as is). No demonstrated upside — confirm the goal metric itself had adequate power before shelving the idea entirely. | Do not ship. No proven benefit and no proof of safety — the clearest of the four outcomes. |
One nuance the matrix can't show: a guardrail whose one-sided interval clears by a wide margin is a different situation from one that clears it by a hair. For results sitting close to the boundary, a staged rollout with continued monitoring is often more honest than a binary ship/hold call — the test passed, but the margin for error in that pass is thin enough to keep watching.
What If Your Case Isn't Here?
Two archetypes and five guardrail counts cover the common shapes, but real metrics don't always fit cleanly into "binomial" or "CV ≈ 2.25 continuous" — ratio metrics, correlated guardrail sets, and unusual baseline rates all shift these numbers in ways a static table can't capture. The Guardrail Sample Size Calculator takes your metric type, variance, margin, and guardrail count directly and returns a verified sample size on demand, the same way the tables above were generated for this article — and the next article walks through exactly that case by hand first.