The playbook's two reference tables both did the same simplifying thing: G copies of one metric, same baseline, same variance, repeated. Real launches don't watch G copies of the same guardrail — they watch a handful of genuinely different ones, each with its own baseline and its own variance. This article works that case with real numbers: a payment rate, a payment amount, and an ad revenue metric, on their own and combined.
Three Real Guardrails
Payment rate at a 30% baseline — a binomial metric, so its variance isn't a separate input at all; it's fully determined by the baseline itself:
Payment amount and ad revenue are both continuous, revenue-shaped metrics, and their variance has to come from actual logging history, the same way the previous worked example pulled a CV from a quarter of revenue data. Say that comes out to a coefficient of variation of 3 for payment amount, and 2 for ad revenue. All three margins in what follows are set at a 1% relative change — "don't tolerate more than a 1% relative drop" — the same relative-margin convention the playbook's tables used.
Scenario 1: One Guardrail — Payment Rate
With a single guardrail, the correction is the familiar case: 90% per-test power, one-sided . Plugging the payment rate's implied variance into the same relative-margin formula the playbook used:
Roughly 400,000 per arm to watch payment rate alone at a 1% relative margin.
Scenario 2: One Guardrail — Payment Amount
Same margin, same correction, but now the CV is 3 instead of 1.53 — more than three times the relative variance:
Almost four times the payment-rate requirement, from the same 1% margin. This is the same lesson the margin-setting article landed on: what makes a guardrail expensive is its variance, not whether it happens to be "a rate" or "an amount." Payment rate and payment amount are two views of the exact same checkout event, and they still need wildly different sample sizes to protect at the same relative margin.
Scenario 3: Two Guardrails Together — Payment Amount + Ad Revenue
Watching payment amount and ad revenue at the same time means — the family now has three members once the goal metric is counted ( ), so each individual test needs power to hold the 80% family-wise target. Recomputing both metrics at that corrected power:
| Metric | CV | n per arm |
|---|---|---|
| Payment amount | 3 | ≈ 1,781,000 |
| Ad revenue | 2 | ≈ 792,000 |
Per the workflow from the previous article — take the maximum across every guardrail's required — the experiment needs ≈1,781,000 per arm, set entirely by payment amount. Ad revenue's own 792,000 requirement never binds; it's already covered by the sample the more variance-hungry guardrail was going to require anyway.
Adding a cheaper guardrail is never free, even when it doesn't bind: payment amount's own requirement still grew from 1,541,000 (alone) to 1,781,000 (paired with ad revenue) — a 15.6% increase purely from the power correction, before ad revenue's own variance enters the picture at all. A second guardrail either raises the bar directly (if it's the more expensive one) or raises it indirectly (via the correction), even in the cases where its own number never becomes the binding one.
The Full Picture
| Metric | Baseline / CV | Solo (G=1, 90.0% power) | Paired (G=2, 93.3% power) |
|---|---|---|---|
| Payment rate | p = 30% (CV ≈ 1.53) | ≈ 399,600 | — |
| Payment amount | CV = 3 | ≈ 1,541,000 | ≈ 1,781,000 |
| Ad revenue | CV = 2 | ≈ 685,000 | ≈ 792,000 |
Read down any column and the pattern from the playbook holds: variance, not metric type, sets the price. Read across any row and the second pattern shows up too: pairing a guardrail with another one costs roughly 15% more, everywhere, regardless of which metric it's paired with — that's the power correction, not the specific guardrails involved. The set's actual requirement is whichever single number in the "paired" column is largest, not a sum and not an average.
What Comes Next
Working this out by hand for three metrics is manageable. Real guardrail suites are often four or five metrics deep, each with its own baseline, its own logged variance, and sometimes its own margin — and doing this arithmetic by hand for every combination doesn't scale. The Guardrail Sample Size Calculator does exactly this: enter each guardrail's own baseline or CV, margin, and how many you're running together, and it returns the corrected power and required for the whole set directly — the same calculation this article just walked through by hand, done for whatever mix of guardrails your launch actually has.