The experiment itself closed quietly, with no significant difference at the full-population level. What's worth writing up isn't the headline result — it's a secondary subgroup cut of the same data that looked like a clear win, and turned out to be a textbook case of sample ratio mismatch manufacturing a guardrail lift out of nothing. This is a companion piece to the encouragement-design article on the same experiment, focused specifically on how that subgroup went wrong.
The Experiment
A category browse page shows product cards, some of which carry an animated GIF preview instead of a static thumbnail. Control kept GIF previews on as normal. Treatment disabled GIF previews on that page entirely, replacing every card with a static image, over a four-day test window.
| Control (A) | Treatment (B) | |
|---|---|---|
| Condition | GIF previews shown on category page | GIF previews removed on category page |
| Units (full population) | 151,000 | 149,000 |
The Full-Population Result
Analyzed at the full population — the only analysis that actually preserves the original randomization — goal, secondary, and guardrail metrics were mostly flat, with two guardrails worth flagging:
- Orders per active user: −2.0% (p = 0.02) — a weak but real negative signal.
- Payment amount per active user: −1.3% (p = 0.06) — borderline, not conventionally significant.
Nothing here supports "removing GIFs helped." If anything, the honest read of the full-population analysis leans slightly negative. The experiment was closed without shipping the change.
The Subgroup That Raised a Flag
A follow-up cut filtered both arms down to "users who saw at least one GIF-preview card" during the test — the intent being to isolate people who were actually exposed to the thing being tested. The resulting sample split looked like this:
| Control (A) | Treatment (B) | |
|---|---|---|
| Units | 100,000 | 50,000 |
| Share | 66.7% | 33.3% |
A 50/50 randomized assignment should not produce a 67:33 split after any legitimate filter. A chi-square test on this split rejects equal allocation decisively (p < 0.001) — a clear sample ratio mismatch (SRM). Whenever an SRM this large shows up after a filter, the filter itself is the first suspect, not the treatment.
Why the SRM Happened
The activation condition — "saw a GIF card" — directly conflicts with the treatment itself, rather than sitting neutrally alongside it:
- Control: GIF cards appear on the category page as normal, so satisfying "saw a GIF card" is easy and passive — it happens to a broad, unremarkable cross-section of users just from normal scrolling.
- Treatment: the category page never shows a GIF card at all. Satisfying the same condition requires encountering a GIF elsewhere in the app — a product detail page, a home carousel, search results — surfaces the experiment never touched. Only users who happened to browse well beyond the category page could qualify.
The filter, in other words, means something categorically easier in one arm than the other. It isn't a neutral lens applied evenly to both groups — it's a different bar in each one, and the treatment arm's bar is only clearable by unusually active users.
What Survivorship Bias Means Here
Survivorship bias: the error of drawing conclusions from only the subset of subjects that passed through some selection process — the "survivors" — in a way that no longer represents the original population.
The 50,000 treatment users who cleared the filter aren't a random 33% of the treatment arm. They're users biased toward heavy, exploratory app usage — the only kind of user who could satisfy a condition the treatment page itself made impossible to meet locally:
| Biased characteristic | Why |
|---|---|
| Longer session times | Had to browse multiple surfaces to have any chance of hitting a GIF card |
| Higher purchase intent | Kept exploring even though the category page gave them nothing to click on |
| Naturally higher conversion | A general property of heavy, highly-engaged app users |
Why the Guardrails Looked Positive
Inside this subgroup, treatment's guardrail metrics looked meaningfully better than control's:
- Share of users with a purchase: +7.0% (p < 0.001)
- Orders per active user: +6.5% (p < 0.001)
- Payment amount per active user: +6.2% (p < 0.001)
None of this means removing GIFs increased purchasing. It means the users who "survived" the filter in the treatment arm were already strong purchasers before the experiment ever touched them. The lift is a difference in who's being compared, not a difference in what happened to them — a sample-composition artifact wearing the shape of a treatment effect.
Conclusion
- Trustworthy: the full-population analysis (151,000 vs. 149,000), which preserves the original randomization end to end.
- Not trustworthy: the "saw a GIF card at least once" subgroup (67:33 SRM, p < 0.001).
The general rule this leaves behind: an activation condition has to be set on behavior from before group assignment, or on behavior that's unrelated to the treatment itself.
| Condition | |
|---|---|
| Good | "Entered the category page" — a pre-treatment action |
| Good | "Made a purchase in the N days before the experiment started" — unrelated to the treatment |
| Bad | "Saw a GIF card" — the treatment arm's own logic makes this condition harder to satisfy |
The undiluted, full-population effect is real but small, and the question of what the effect looks like specifically among users who would have actually seen a GIF card is a legitimate one — it's just answered by dividing out the exposure rate, not by filtering on it after the fact. That derivation is worked through in the companion piece on this same experiment, with the full sensitivity of the estimate to the exposure rate covered in the follow-up reference article.