The last two articles established the problem: a superiority test can't tell you a guardrail metric is safe, and "not significant" on an underpowered guardrail is silent, not reassuring. Non-inferiority testing is the framework built specifically to answer the question a guardrail actually asks — not "is there a difference," but "is the treatment no meaningfully worse than the control."
Flipping the Null Hypothesis
A standard superiority test starts from a null hypothesis of "no difference" and looks for evidence against it. Non-inferiority testing starts from the opposite assumption entirely: it assumes the treatment is meaningfully worse, and requires the data to actively rule that out before you're allowed to conclude otherwise. Formally, for a guardrail metric where higher is better:
The quantity is the non-inferiority margin — the largest degradation you've decided, in advance, is acceptable. Rejecting doesn't mean "no difference was detected." It means you have positive statistical evidence that the treatment is not worse than the control by more than . That's a categorically different claim from "the p-value on a superiority test wasn't significant," and it's the claim a guardrail decision actually needs.
Why the Margin Has to Be Nonzero
It's tempting to ask why not just set and require proof of literally no degradation at all. Two reasons rule this out. First, no finite sample can ever prove an effect is exactly zero — you can only ever bound how far from zero it could plausibly be, and a margin of exactly zero would require infinite data to satisfy. Second, "exactly zero change" usually isn't even the right bar in practice: a guardrail metric that moves by an amount too small to matter to the business shouldn't block a launch, and is precisely the number that encodes "too small to matter" into the test itself. Choosing well is a business decision, not a statistical one — covered in depth later in this series — but the test's logic requires some nonzero value to even be well-posed.
Reading It as a Confidence Interval
In practice, nobody derives a non-inferiority conclusion by hand-computing a rejection region every time. The operational shortcut is a one-sided confidence bound: compute a one-sided confidence interval for , and check whether it stays entirely above :
This is a useful mirror of how superiority tests are usually read in practice: for superiority, you check whether a two-sided interval excludes zero. For non-inferiority, you check whether a one-sided interval clears instead of zero. Same mechanical habit — read the interval, not just the p-value — pointed at a different, pre-committed threshold.
Superiority and Non-Inferiority, Side by Side
- Question asked: superiority asks "is there a difference?" Non-inferiority asks "is the treatment no worse than a pre-agreed threshold?"
- Null hypothesis: superiority assumes equality by default. Non-inferiority assumes meaningful harm by default — the burden of proof is reversed.
- What a "pass" means: a superiority pass means you found evidence of a real difference. A non-inferiority pass means you found evidence ruling out a meaningful regression — it says nothing about whether the treatment is better, worse, or identical within that margin.
- What a "fail" means: a superiority fail means "inconclusive" — it's compatible with no effect and with a real effect you didn't have power to see. A non-inferiority fail means the data couldn't rule out meaningful harm — closer to "unsafe until proven otherwise" than "probably fine." And non-inferiority tests aren't immune to the underpowering problem either — a non-inferiority test can fail simply because there wasn't enough data to clear the margin, not because real harm was confirmed.
A Common Misreading, Addressed Early
Non-inferiority testing does not mean "we've decided we don't care if this metric gets worse." It means the opposite: the team has explicitly decided, before looking at any data, exactly how much worse is still acceptable — and the test only passes if the evidence actively supports staying within that pre-agreed bound. A vague, undefined tolerance for harm isn't non-inferiority testing. It's precisely the "flat means fine" failure mode from the previous article, just relabeled. The margin is what makes it rigorous instead of hand-wavy — which is exactly why choosing it correctly matters so much, and why it's worth its own dedicated article later in this series.
What Comes Next
Flipping the null hypothesis fixes the logic problem, but it introduces a new practical one: ruling out a range of harm down to a specific margin is a harder statistical task than just detecting any difference at all, and it typically demands more data than teams are used to collecting for the same metric. The next article works through why — and gives an approximation for exactly how much more.