What Does Arm D Actually Measure? Interaction Effects, Not Quality Transfer

The previous article set up a four-arm factorial design — attributes on new products, attributes on established products, both, or neither — and left one decision unresolved: if only one new arm can run next, is it C or D? Answering that well requires being precise about what arm D actually estimates, because the obvious-sounding answer turns out to be wrong.

The Tempting, Wrong Framing

It's natural to reach for D as a way to check whether established products' quality improvement also shows up on the new-product side — the reasoning goes: "C tests whether attribute matching improves established-product quality at all; if we're not sure that same quality effect would also occur for new products, add D to test that too."

That framing doesn't hold up. Whether attribute matching improves new-product retrieval quality is already answered — in a confounded way — by the original new-products-only experiment, arm B versus A. B−AB - A is the effect of attribute matching on new products; it's just tangled up with two things at once: more new-product candidates surfacing at all (a retrieval, or volume, effect) and better matching quality among the new products that do surface (a quality effect). Adding arm D does not untangle this. D still has new-product attribute matching switched on, so it inherits exactly the same candidate-pool expansion that B has. D was never going to isolate "does the quality effect show up for new products" — that question was never cleanly separated to begin with, and it's baked into every arm where the new-product attribute is on, B and D alike.

What D Actually Uniquely Measures

D's only genuinely new information is the interaction — the reallocation effect between new products and established products. Compare (D−C)(D - C) against (B−A)(B - A), or equivalently (D−B)(D - B) against (C−A)(C - A): does turning both changes on at once produce less — or more — than the simple sum of running each one alone? That comparison, and only that comparison, is what a fourth arm buys you that a three-arm design can't.

This reframes the original question precisely. "Does D shrink the new-product exposure gain?" is exactly this interaction term, correctly named — not a check on whether quality effects transfer between new products and established products, which was never D's job.

The Mechanism Behind "Does D Shrink New-Product Exposure?"

In arm D, new-product retrieval itself is identical to arm B — the new-product candidate pool is exactly the same size, because nothing about new-product attribute logic changed between B and D. Retrieval isn't where anything could shrink. What can shrink is the next stage: blending and slot competition in the feed. If adding attributes to established products improves their ranking quality, established products can start winning more of the shared, limited feed slots — displacing the new-product impressions that arm B specifically worked to increase, even though nothing about how new products are retrieved has changed at all.

A distinction worth getting exactly right here: "established products already have a large candidate pool, so adding attributes there won't increase their exposure" is a statement about retrieval — candidate pool size. Slot win-rate is a separate mechanism entirely. Even with the number of established-product candidates completely unchanged, if ranking quality improves, established products can win more of a fixed set of feed slots simply by out-ranking new products more often at the margin. These are two different channels, and only one of them — slot win-rate — actually matters for the reallocation risk that motivated arm D in the first place.

The practical consequence: checking the assumption behind all of this — whether adding attributes to established products actually increases their exposure — has to look at final exposure share (how often established products win a shared slot against new products), not just retrieval or candidate-pool counts. Candidate-pool counts can look completely unchanged while slot share still shifts underneath them, quietly eating back the cold-start gains the first experiment was built to produce.

Choosing Between A/B/C and A/B/D

The right decision criterion here isn't statistical power or traffic cost — it's whether the new-product-attribute change is going to ship regardless of what an established-products experiment finds.

  • If the new-product side has already won and will be deployed, the world going forward is the "new-product-attribute-on" world by default. In that world, the decision-relevant comparison is D−BD - B — established products added on top of a new-product change that's happening either way — not C−AC - A, which answers a question about a counterfactual world (new-product attribute off) that won't actually exist once it ships. Run A/B/D.
  • If the new-product change is still undecided, or the team specifically wants a methodologically clean read on established products' effect in isolation — a pure quality effect, uncontaminated by any interaction with cold-start retrieval — run A/B/C instead.

Half the Reallocation Story Is Free, Even With Three Arms

A full four-arm test is the only way to measure the pure non-additivity itself — the gap between what actually happens when both changes are on and what you'd predict by simply adding the two individually-measured effects together. But a well-instrumented three-arm test already recovers half of the reallocation story for free:

  • In arm C (against A), monitor new-product impressions and clicks as guardrails. This shows how much established products' improvement encroaches on new-product exposure — exactly the "did we quietly undo the cold-start fix" question.
  • In arm B (against A), monitor established-product metrics. This shows the mirror direction — whether the new-product change encroaches on established products.

The only thing genuinely locked behind a fourth arm is the non-additivity itself — whether turning both on together produces something different from what the two separate measurements alone would predict. Everything else about the reallocation story is visible from a three-arm test with the right guardrails turned on.

Where This Leaves the Decision

Answer the "does the new-product change ship regardless" question first — it resolves the C-versus-D choice more often than a purely statistical framing of the decision ever will. And whichever arm runs next, treat "does the quality effect transfer to new products" as a question the original B-versus-A experiment already answers, in confounded form — not a job to hand to D. D exists for exactly one reason: measuring the interaction between two changes that compete for the same shared resource. Asking it to do anything else — including "did we accidentally undo our own cold-start fix," the question that actually motivated this whole design — is asking it to answer something only the interaction term, not D alone, was ever built to isolate.