Maximizing Expected Return Isn't the Same Policy as Proving "Not Bad"

I keep coming back to the same question about the posterior-expected-value rollout rule from a couple of weeks ago: if the goal is genuinely to maximize expected return across every experiment you run, how does that rule actually compare to the two testing regimes everyone already uses? I can state the mechanism now. I still haven't found the clean, one-line way to say it out loud.

Three Policies, Side by Side

If the objective is purely to maximize expected return, the best policy is: for every experiment, ship it whenever its posterior expected return is positive. Where that posterior comes from is the mechanism covered last time — when the sample isn't large enough yet, the estimate leans more on the prior; once the sample is large enough, it leans on the new data instead.

Compare that policy against the two testing regimes already in common use:

  • The standard significance test — the default, commonly used A/B testing method — only ships when it's confident there's a genuine difference at all. It requires being sure the "no effect" null is wrong before it lets anything through.
  • The non-inferiority test flips that logic: it ships unless it can't rule out meaningful harm. It requires positively proving "not bad" before it lets anything through.

The posterior-positive rule maximizes expected return by construction: it ships exactly the set of experiments whose expected value is positive, and holds back exactly the set whose expected value is negative. There's no other threshold you could draw that produces a higher average payoff across the long run of experiments, because the rule's entire definition already is "ship whenever expected value says to."

Why "Prove Not Bad" Leaves Money on the Table

The real question isn't why the posterior-positive rule wins — that's true by definition. The interesting question is why a policy built around proving "not bad" ends up with lower expected return than that.

The mechanism is this: a policy that requires confidently proving no meaningful harm will decline to ship some experiments that are actually modestly good — they just don't clear the "definitely not bad" bar with enough certainty. Every one of those declined-but- actually-fine experiments is forgone profit. The posterior-positive rule would have shipped them, because their expected value was positive even though the evidence wasn't strong enough to satisfy a confidence requirement. The non-inferiority test's insistence on proof, not just a favorable expectation, is exactly what creates that gap.

The standard significance test loses expected value for the mirror-image reason: it demands proof of a real, positive effect before shipping at all, which is an even higher bar than "not bad." Both testing regimes share the same underlying shape — they require statistical certainty in some direction before they'll act — while the posterior-positive rule acts on the sign of the expectation alone, certainty or not. That's the entire source of the gap between them: certainty is exactly what the classical tests are buying, and exactly what they're paying for in unrealized expected return.

Still Looking for the Simple Version of This

I can walk through the mechanism above and it holds up. What I still don't have is a short, intuitive way to say it to someone who doesn't want the full derivation — something like the guardrail-testing correction-direction argument from a couple of retrospectives back, which fit in a sentence once I'd found the right framing. This one hasn't compressed that well yet. For now, the honest version is: every testing policy that requires certainty before acting is quietly trading expected return for that certainty, and the size of the trade is exactly the population of "actually fine, just not provably so" experiments it turns away.