Every article so far in this series has treated the non-inferiority margin as a given number. It never is one in practice. Picking is the single hardest, least standardized step in the whole exercise — and it's a business decision dressed up as a statistical one. This article works through picking a real margin for a real revenue guardrail, start to finish, including the point where the business's first answer turns out to be statistically unaffordable.
The Setup
An icon redesign is being tested on one of the app's bottom-tab icons — the one that navigates to the Content tab — on a product with 2,000,000 monthly active users and an average revenue per user (ARPU) of $8.00/month — $16,000,000 in monthly revenue exposed to the change once it's fully launched. The goal metric is the tab's click-through rate: the share of sessions where a user taps that icon and lands on the Content page. Revenue per user is the guardrail — nothing about a navigation icon should move revenue directly, but a redesign that pulls attention toward Content and away from monetized surfaces could plausibly cost the business money as a side effect, which is exactly the kind of drift a guardrail exists to catch. The question on the table: what margin defines "not meaningfully worse" for revenue?
Method 1: Business Impact Inversion
The cleanest way to set is to work backward from a dollar figure the business actually cares about, rather than picking a percentage out of the air. In a planning meeting, finance states their real threshold plainly: a launch that costs the company more than $80,000 in monthly revenue is something they'd want flagged and reconsidered, not shipped quietly.
Converting that into a per-user, relative margin is direct arithmetic:
That's the business-derived answer: don't tolerate a launch that drops ARPU by more than 0.5% ($0.04/user/month). It sounds precise and defensible. Whether it's actually usable depends entirely on the next step — nobody in that planning meeting asked how much traffic it would take to prove it.
Checking Whether the Margin Is Affordable
Revenue per user is a heavy-tailed metric — a small share of high-spending users drives most of the variance. Pulling the last quarter's logging data gives a per-user standard deviation of , against a mean of only $8 — a coefficient of variation over 5, typical of monetary metrics with a long right tail. Plugging into the non-inferiority sample size formula from earlier in this series — which assumes the true underlying difference is exactly zero, the standard conservative planning baseline — at 90% power and one-sided (a common convention in tech A/B testing; clinical trials conventionally use the stricter one-sided 0.025) (, ):
That's not a typo: roughly 21.7 million users per arm — more than ten times the product's entire monthly active user base. The business's first answer is a perfectly reasonable number to want. It is not a number this experiment, or any realistic experiment on this product, can ever measure. This is the moment most teams either give up on guardrail rigor entirely, or quietly ignore the mismatch and ship an underpowered test anyway — exactly the "flat means fine" trap from earlier in this series, just reached by a different road.
Two Levers, Not One
There are two honest ways to close this gap, and the right answer usually uses both rather than picking one.
Lever 1: Shrink the variance, not the ambition
Revenue's enormous variance is driven almost entirely by a handful of extreme spenders. Winsorizing the metric — capping any single user's monthly revenue at, say, $150 before computing the guardrail statistic — trades a small, known amount of bias for a large reduction in variance. On this dataset, capping at $150 brings down from $45 to roughly $18. Recomputing:
Still far too large, but a 6× improvement from a single change — and it's a well-established practice on real experimentation platforms specifically for monetary guardrails. Winsorizing changes what you're technically testing (capped revenue, not raw revenue), which is a real limitation worth documenting, not a free lunch. Concretely: losses among the users spending above the $150 cap become partly invisible to the guardrail. If the icon redesign happens to specifically suppress spending among heavy spenders, the raw revenue loss could exceed $80,000 even while the capped metric stays comfortably inside — the mapping from Method 1 was derived on raw ARPU, and it no longer applies exactly once the metric being tested is capped (capped ARPU's mean is a little below $8 too, which nudges the relative-margin conversion slightly). Worth stating plainly in the write-up, not just implying.
Lever 2: Renegotiate the margin, with the cost curve on the table
The other lever is going back to finance with a curve instead of a single number — showing what margin is actually affordable at a traffic level the team can commit to. Using the winsorized :
| δ (relative) | δ (absolute, $/user) | Required n per arm |
|---|---|---|
| 0.5% | $0.04 | ≈ 3,470,000 |
| 1.0% | $0.08 | ≈ 868,000 |
| 1.5% | $0.12 | ≈ 385,000 |
| 2.0% | $0.16 | ≈ 217,000 |
Presented with this table, finance revisits their original number. The $80,000 threshold was a reasonable-sounding round figure, not a carefully derived limit — and once it's clear that protecting it tightly would require dedicating the entire user base to a single test for months, the team agrees that a launch costing up to $240,000/month (1.5% of exposed revenue, matching ) is the real line worth drawing, given the traffic realistically available. That agreement — made before data collection, with the tradeoff explicit — is the actual deliverable of this whole exercise. The number itself matters less than the fact that it was chosen deliberately instead of defaulted into.
The Same Process for Other Metric Types
Revenue is the hard case because of its variance. Other metric types go through the same business-impact-inversion logic, but the statistics behind them are far more forgiving.
- Conversion rate (binomial): variance is , and margins are usually expressed in percentage points rather than relative terms. For a conversion-type guardrail with a baseline of and a margin of percentage point (won't tolerate conversion falling below 11%):That's roughly 190 times cheaper than the winsorized revenue guardrail's 3.47 million — but this isn't a same-margin comparison, and the framing matters. A 1-percentage-point margin on a 12% base is a relative margin of about 8.3%, far looser than revenue's 0.5%. Tighten conversion to a genuinely matched 0.5% relative margin (0.06 percentage points) and the requirement jumps to roughly 5.03 million per arm — actually slightly worse than the winsorized revenue case, since conversion's relative variance at () isn't actually smaller than revenue's winsorized . The real reason conversion guardrails are cheap in practice isn't that binomial variance is inherently smaller — it's that the conventional margin (a percentage point or two) is far looser than the tight relative margins monetary guardrails often need. Compare margins on equal footing before concluding one metric type is simply "easier."
- Ratio metrics (e.g. revenue per session): when the denominator itself varies across users (sessions per user isn't constant), the variance of the ratio needs the delta method rather than a plain per-user variance — it depends on the variance of the numerator, the variance of the denominator, and their covariance simultaneously. The business-impact-inversion step is identical; only the variance formula feeding into the sample size calculation changes. The reference tables in the next article handle this case directly so you don't have to re-derive it per metric.
The Operating Principle
One rule makes all of the above valid, and breaking it invalidates all of it: has to be fixed before the experiment's data is looked at, and documented alongside the pre-registered hypothesis, sample size, and analysis plan — the same discipline this series has argued for since the very first article on peeking. Choosing using historical data, business input, and a cost curve like the one above is legitimate. Adjusting after seeing the experiment's own point estimate — "the observed drop was only 1.8%, so let's call our margin 2%" — is not a margin anymore. It's a p-value dressed up as a business decision, and it reintroduces exactly the false-safety failure mode this entire series exists to prevent.
What Comes Next
This example worked through one metric, one margin, one guardrail at a time. Real launches usually carry several guardrails at once, each with its own margin, its own variance, and — per the multiplicity article earlier in this series — its own corrected power target. The next article is the reference: precomputed sample sizes across common metric types, margin levels, and guardrail counts, plus a decision matrix for turning "goal metric significant, all guardrails non-inferior" into an actual ship/no-ship call.