A recommendations engineer's first experiment tested whether enriching how new, cold-start products get retrieved would help them surface at all. It shipped a result, and the very next planning conversation surfaced two genuinely interesting problems at once: how to design a follow-up that doesn't get confused by new products and established products competing for the same feed slots, and how to explain a metric decline that didn't behave the way the rest of the funnel did.
The Original Experiment
Retrieval for the recommendations feed leans heavily on engagement history — past clicks, saves, purchases, ratings — to decide which products are worth surfacing for a given user. That's a fine signal for established products: items that have been live long enough to accumulate months of behavioral data, giving the ranking model a large, confident pool to draw on. It's a broken signal for new products: a listing that went live yesterday has almost no clicks, purchases, or ratings to be judged on, so the model has nothing to go on and mostly leaves it out of the candidate pool entirely — the pool of plausible new-item candidates stays small precisely because the data needed to grow it doesn't exist yet. This is the classic cold-start problem, and it compounds on itself: a new product that's rarely retrieved never gets the engagement data that would eventually let it be retrieved on its own merits.
The experiment added structured attribute text — category, brand, material, style tags extracted from the listing itself — into new-product retrieval, as a substitute signal for the engagement history these items haven't had time to accumulate. The theory: even with zero clicks to learn from, a new product's own description already contains enough signal to judge whether it's a plausible match for a given query or feed slot, if the retrieval logic is willing to use it.
Planning the Next Iteration Raised a Design Problem
The obvious next question: if richer attribute matching helps new products get retrieved, would the same enrichment help established products too? A third arm — attributes added to established-product retrieval only — was proposed to find out. But new products and established products don't operate independently in the feed. They compete for the same limited set of slots. If attribute-rich established products become more relevant and start winning more of that limited exposure, they could easily do it partly at the expense of the new-product impressions the first experiment was specifically trying to increase — quietly undoing the cold-start fix in the process.
That competition means a simple "established-products-only" arm compared against a plain control would answer a blended question, not a clean one. Any effect it turns up would combine two very different things: the direct value of matching established products better, and the indirect effect of established products and new products reshuffling how much exposure each one gets as a side effect of that better matching. Reporting the blended number as "the effect of better established-product matching" would overstate or understate the real effect depending on which direction the reshuffling went.
The Factorial Fix
The way to untangle this is a small factorial design — four arms instead of two, crossing "attributes on new products" with "attributes on established products":
| Established products: attributes off | Established products: attributes on | |
|---|---|---|
| New products: attributes off | A — control | C — established products only (planned) |
| New products: attributes on | B — new products only (already run) | D — both (under consideration) |
With all four arms in hand, the comparisons available go well beyond a single "established products vs. control" difference. Comparing D to B isolates what adding attributes to established products changes, while new-product retrieval already has its own attribute matching held constant — which removes the "new-product exposure is also shifting for unrelated reasons" confound that a plain C-vs-A comparison can't rule out. Comparing D to C does the mirror version for the new-product side. The interaction between the two — whether combining both changes differs from what you'd expect by simply adding the two individual effects together — is precisely the exposure-reallocation dynamic the team was worried about, made measurable instead of merely suspected.
Checking the Assumption Before Building the Whole Design
A four-arm factorial design is more expensive to run than a two-arm test — more traffic split more ways, more time to reach a clean read on each cell. Before committing to it, it's worth checking the assumption the whole design problem rests on: does adding attributes to established-product retrieval actually increase established-product exposure at all? Established products already have a large candidate pool built from months of engagement history — if richer attribute matching doesn't meaningfully change how much exposure they win on top of that, the reallocation concern motivating arm D evaporates, and the simpler three-arm version (A, B, C) is enough. Validating that a complication is real before paying for the design that handles it is worth doing every time — it's the same instinct as checking whether an assumption in a formula actually holds before building the more elaborate machinery to work around it.
A Puzzle From the First Experiment's Results
Separately, the completed new-products-only experiment left one number that didn't fit the rest of the story. Impressions and clicks on new-product placements moved together — clicks rose in roughly the same proportion as impressions rose, meaning click-through rate itself barely budged. That's a clean, boring, easy-to-explain result: more new products were being retrieved and shown, and the ones that were shown converted to clicks at basically the same rate as before. A volume story, not a quality story.
Wishlist saves from those same impressions didn't follow that pattern. They rose by noticeably less than the impression or click increase would predict. If the wishlist gain were purely a downstream consequence of more people reaching that step, it should have grown by roughly the same proportion as clicks did. It didn't — which means something beyond simple volume gain is happening specifically at the wishlist step, not earlier in the funnel.
Why Checking Proportionality Is a Useful Habit
This is a generally useful diagnostic, worth running on any multi-step funnel: compare how much a downstream metric moved to how much the upstream metrics feeding it moved. If a downstream metric moves in roughly the same proportion as its inputs, the honest story is "more volume, same quality" — nothing new is happening, there's just more of the same thing happening. If a downstream metric moves by less than that, something is under-delivering specifically at that step, for people who did make it that far.
In this case, a plausible mechanism fits the pattern: a click can be earned by an attractive thumbnail or price even when the underlying attribute match to the user's actual taste is a little off — clicking is a low-commitment, low-information action. A wishlist save happens after a user has actually evaluated whether the item matches what they were looking for. It's a later, higher-intent step, and far more sensitive to whether the retrieved product genuinely matches the user's underlying taste rather than just looking superficially appealing. If attribute matching quality is weaker than engagement-history matching for at least some new products, the wishlist step is exactly where you'd expect that gap to show up first and most clearly.
Wanting a Metric That Points at Which Attribute Is the Problem
Since the retrieval logic now matches on several attribute fields simultaneously — category, brand, material, style, and others — the natural next question is which specific field is carrying real signal and which is mostly noise for new products. A per-field match-accuracy metric, scoring how well each attribute field predicts eventual engagement once it does accumulate, would let the team compare fields directly and find exactly which one is underperforming, instead of guessing across the whole set at once.
That metric doesn't exist yet, and building it isn't trivial — it needs some definition of "predictive accuracy" between a structured attribute and eventual user behavior, which usually means either a held-out cohort of products that have since accumulated real engagement data, or a labeled relevance judgment, neither of which gets built overnight. It's worth prioritizing anyway, precisely because the proportionality check above already narrowed the search: there's a specific, real quality gap hiding somewhere in the attribute set, not just diffuse noise. A related analysis on measuring how well thumbnail images match a query — done for a different part of the same recommendations system — is a useful precedent here: matching quality between a retrieval signal and a user's actual intent is a recurring measurement problem across a recommendations system, not a one-off need specific to this experiment.
Which Third Arm to Run First
A four-arm test is the complete answer, but traffic budgets rarely are. If only one new arm can run next, the choice between C and D is really a choice between two different questions. C answers "what does better established-product matching do in a world where new-product retrieval stays as it is" — clean, but a counterfactual world if the new-product change is shipping anyway. D answers "what does better established-product matching do on top of the new-product change" — the world users will actually live in. And D quietly carries a subtler question inside it: whether the new-product experiment's gains came partly from exposure that stronger ranking on the established-product (warm) side will now claw back. What D actually estimates — and what it doesn't — turns out to be less obvious than the 2×2 table suggests. That's the subject of the next article.
Two Techniques Worth Keeping
Neither of these ideas is specific to cold-start retrieval. A factorial design is the right tool whenever two things you're changing compete for the same limited resource — feed slots, budget, attention — and you need to separate the direct effect of each change from the effect of them reallocating that shared resource between each other. And checking whether a downstream funnel metric's movement is proportional to its upstream metrics is a fast, cheap first pass for telling a volume story from a quality story, well before committing to build the more expensive metric that would diagnose the quality gap in detail.