The last two articles established the problem: a superiority test can't tell you a guardrail metric is safe, and "not significant" on an underpowered guardrail is silent, not reassuring. Non-inferiority testing is the framework built specifically to answer the question a guardrail actually asks — not "is there a difference," but "is the treatment no meaningfully worse than the control."
Flipping the Null Hypothesis
A standard superiority test starts from a null hypothesis of "no difference" and looks for evidence against it. Non-inferiority testing starts from the opposite assumption entirely...