Bayesian or frequentist for the actual launch decision? I went back and forth on this longer than I'd like to admit. What settled it wasn't a theoretical argument — it was noticing how simple the Bayesian version could be, and then finding a genuinely surprising wrinkle once a second metric enters the picture.
Why Bayesian, in Practice
Doing this properly with a non-inferiority test is entirely possible, but it adds real complexity — a margin to agree on, a one-sided correction, a second hypothesis test running alongside the primary one. If the experimentation platform you're already running on defaults to a Bayesian statistics engine, bolting a frequentist non-inferiority framework on top of it is extra machinery for a benefit that's easy to get for free another way.
There's also a plainer, less statistical reason: most people making the actual launch call aren't fluent in what a p-value or a significance level really means, and no amount of explaining fully fixes that in the room. "There's an 82% chance this is actually better" is a sentence everyone in the room already understands. That's a real advantage, not just a communication nicety.
The Simple Rule: Chance to Win ≥ 50%
The rule this settled on: if the treatment arm's probability of beating control — its "chance to win" — clears 50%, launch it. That's it. No margin, no separate guardrail test with its own machinery, just the posterior probability of being better.
Compare that to how a standard frequentist setup actually behaves in practice: a result only counts as a "win" when it's clearly, significantly positive, and a guardrail only triggers a rollback when it's clearly, significantly negative. Both bars — the one for success and the one for failure — are set high on purpose. Add ordinary power problems on top of that, and a genuinely clear-cut win becomes rare almost by design. Most real results just sit in the ambiguous middle, neither confidently good nor confidently bad, and get treated as "no result" by a framework that only recognizes the extremes.
Which is the same mechanism covered last time: requiring certainty before acting is exactly what makes success rare, and going with the plain probability of winning instead — no certainty requirement at all — ends up raising the realized rate of good launches, not lowering it.
One Metric: 50% Is (Almost) Exactly Optimal
The part that ties this back to the posterior-expected-value rule from a few weeks back: for a single metric, "launch whenever chance to win is at least 50%" and "launch whenever posterior expected value is positive" land on the same decision. That's because, for a symmetric posterior — the usual case for a well-behaved effect-size estimate — the point where the probability of a positive effect crosses 50% is the same point where the expected value of the effect crosses zero. Chance to win isn't a rough proxy for expected value here; under that symmetry assumption, it's the same rule wearing a more intuitive name.
Two Metrics: Why the AND Condition Breaks 50%
The genuinely interesting part showed up once a second metric got involved. Require the chance to win on both metrics to individually clear 50%, and the combined rule stops being jointly optimal — even though 50% was exactly right for each metric on its own.
The reason is the AND, not either metric individually. Stacking two independent 50%-or-better requirements is a stricter joint bar than either threshold alone implies — some experiments where both metrics are genuinely, modestly positive still get held back because one of them dips just under its individual 50% mark. That's forgone expected value on both metrics at once, for the same underlying reason a family of guardrails needs its own correction once you're requiring all of them to clear a bar together. Working through the math, the balance point that actually maximizes the combined expected value across both metrics lands closer to a 40% threshold on each — not 50%.
That number isn't a fixed constant — exactly how far below 50% the optimal per-metric threshold falls depends on how the two metrics' uncertainty and correlation relate to each other, so 40% is a useful ballpark from this case, not a universal rule to copy elsewhere without checking. But the direction is the lesson: an AND condition across metrics needs each individual threshold loosened to stay jointly optimal, in exactly the same way a family of guardrails needs its power target raised to stay jointly protective. It's the same correction, just running in the opposite direction.
The Practical Habit This Left Me With
"Chance to win ≥ 50%" is the right rule for a single metric, for a well-behaved symmetric posterior — treat it as exactly that, not as a threshold that's automatically safe to reuse once a second metric, or a third, joins the launch decision. Every AND condition you add on top changes what the individually-optimal threshold actually is, and the fix runs the opposite direction from where intuition points: you loosen each piece, you don't tighten it.