Going Bayesian for A/B Testing: Two Easy Fits and One That Isn't

Once you've talked yourself into a chance-to-win launch rule over a frequentist one, the next question is less philosophical and more operational: does going Bayesian actually work cleanly across everything an experimentation program needs, or just for the one decision rule you started with? Going through it piece by piece, two of the three places it needs to fit turned out to be straightforward. One really isn't.

The Engine Is Already There

The first question was whether "going Bayesian" meant building new statistical machinery, or just using what was already sitting underneath. It's the latter. GrowthBook, the experimentation platform already in use, ships a native Bayesian testing mode: you supply a prior distribution, and once new data comes in, it computes a posterior by combining the two — exactly the same weighting logic covered in the posterior-expected-value piece from a few weeks back: the more data you accumulate, the more the posterior leans on that data; the less you have, the more it falls back on the prior. Going Bayesian here isn't inventing anything — it's treating what the engine already computes as the actual input to the rollout decision, instead of deriving a separate calculation on the side.

Multiple Hypotheses: Looks Solvable, Needs More Reading

Multiplicity was the item on this list I expected to be the sticking point. It might actually be the easier of the two open questions. There's a real Bayesian literature here — hierarchical, shrinkage-based models that handle multiple comparisons by pooling information across hypotheses, rather than the frequentist move of penalizing each individual threshold after the fact (Bonferroni, FDR). I haven't worked all the way through it yet, but it looks like a genuinely different mechanism for the same problem, not a Bayesian relabeling of the frequentist correction.

The angle I actually want to chase next is less about the correction itself and more about whether it connects to modeling users directly: can the relationship between two metrics — not just the multiplicity penalty for testing both at once — be represented inside the same hierarchical structure? That's a genuinely open question I want to read more into, not one I've resolved.

Sequential Testing: Where the Fit Breaks Down

Sequential testing is where this stopped being a straightforward translation. The entire reason sequential methods exist in the frequentist world is to stop optional stopping from inflating the false-positive rate — checking the data repeatedly and quietly borrowing extra chances at a false "significant" result beyond the α you signed up for. Bayesian inference, as far as I can tell, isn't built around that same failure mode in the first place, which sounds like it should make the comparison simpler. It doesn't — it makes it more confusing, not less.

The confusion shows up immediately in the language both frameworks use. A Bayesian A/B testing tool reports something like "chance to win: 95%," which reads as a direct stand-in for p<0.05p < 0.05 — the two numbers rhyme, so people treat them as the same guarantee. They aren't built on the same machinery at all. Underneath, the Bayesian version isn't controlling an error rate; it's an expected-loss or risk framework, asking something closer to "what do I stand to lose by shipping, or not shipping, given this posterior" — not "how often would this test wrongly cross a threshold if I kept looking at the data over and over." Whether, and exactly how, that risk-based framing relates to the frequentist guarantee it's implicitly being compared to — especially once you're peeking at the data repeatedly instead of checking once — is the part I still need to work through properly.

The Practical Habit This Left Me With

Two of these three questions resolved faster than expected: the engine already computes the right thing, and the multiplicity literature at least exists to go read. Sequential testing is the one where switching to Bayesian doesn't retire the original problem — it just changes what the problem looks like. Treating a 95% chance-to-win as a drop-in replacement for a 5% significance threshold, just because the two numbers happen to rhyme, is exactly the kind of shortcut likely to bite later, until the actual relationship between the two frameworks under repeated looks is worked out properly.