Encouragement Design in Practice: An ML Engineer's Question About a Diluted Experiment

An ML engineer on the team flagged something odd about a subgroup readout from a category-page experiment, and the question that followed turned into one of the cleaner real-world illustrations of encouragement design I've worked through — the framework for experiments where you can't force exposure to a treatment, only make it more likely. This is that walkthrough, numbers changed, names removed.

The Setup

A category browse page shows product cards, and a subset of products have an animated GIF preview instead of a static thumbnail — a short auto-playing loop showing the item in use. The experiment removed GIF previews from the category page entirely for one arm, replacing every card with a static image, to see whether the animation was helping or just adding noise and load time. Control (A) kept GIF previews as normal. Treatment (B) disabled them on the category page.

The First Red Flag: A Subgroup That Looked Too Good

The first readout that raised eyebrows was a subgroup cut on "users who saw at least one GIF-preview card." Purchase-related metrics in that subgroup looked 6–7% higher in the treatment arm than in control — which would be a strange result if GIFs were actually helping. Before treating that as a finding, it's worth asking what "saw a GIF-preview card" actually means differently in each arm.

In control, seeing a GIF card just means scrolling the category page normally — a low bar, met passively by a broad cross-section of users. In treatment, the category page never shows a GIF card at all, so the only way to land in that subgroup is to have encountered a GIF elsewhere in the app — a product detail page, a home carousel, search results — surfaces the experiment didn't touch. Qualifying for the subgroup in treatment requires extra, self-selected exploration. Qualifying for it in control doesn't. The two "activated" populations aren't comparable, and sure enough, the actual counts confirmed it: roughly 100,000 users in the control subgroup against 50,000 in the treatment subgroup — a 67:33 split where roughly even should be expected. That's a sample ratio mismatch, and it means the treatment arm's "activated" subgroup is quietly biased toward heavier, more exploratory users, which is more than enough on its own to produce a 6–7% gap in purchase behavior with no causal story attached to GIFs at all.

Analyzed at the full population instead — about 151,000 users in control against 149,000 in treatment, a properly balanced split — no significant effect showed up in either direction. The subgroup result wasn't a hidden win being masked by noise. It was a selection artifact from a subgroup filter that meant something different in each arm.

Getting the Activation Metric Right

The question that came in: "Shouldn't an activation condition be based on behavior from before the group split, or behavior that's unrelated to the treatment itself? The cleanest cut point should be right before treatment and control actually diverge. If we activate on 'entered the category page' — which happens before the GIF logic kicks in — note that even in control, some users never encounter a single GIF card during their session. Doesn't defining the subgroup that far from where the arms actually diverge risk diluting the result?"

Randomization in this experiment happens at category-page entry — that's the actual split point. So "entered the category page" is a legitimate activation metric: it's a pre-treatment condition, identical in meaning for both arms, and excluding users who never opened the category page at all removes a large pool of people the experiment could never have affected either way, without touching anything downstream of the split. That's different in kind from the GIF-card subgroup above, which conditions on something that only happens after the arms diverge and means something different in each one.

Before settling on that, though, it's worth being explicit about what hypothesis is actually being tested — in this case, whether GIF exposure specifically on the category page moves the metrics, as opposed to GIF exposure anywhere in the app. This experiment only disables GIF previews on the category page; other surfaces are untouched in both arms, which is what makes "entered the category page" the right scope for the activation condition.

General experimentation platform guidance backs this up directly: leaving in users who were never actually exposed to the surface where treatment and control differ adds noise without adding signal, which makes real effects harder to detect — exactly the caution behind excluding users who never opened the category page at all. The distinction worth holding onto: that exclusion is a pre-treatment condition, identical across arms, and doesn't touch randomization. The "saw a GIF card" subgroup from the previous section conditions on something that only happens after the split and means something different in each arm — which is what breaks randomization rather than just adding noise.

The Dilution Question That Remained

The follow-up question: "Activation is already gated on entering the category page, so that part seems settled. But control includes two kinds of users — (1) those who actually saw a GIF card, and (2) those who didn't happen to see one that session — while treatment only ever has type (2), since GIFs are off for everyone. Doesn't having type (2) mixed into control dilute the result? Is there an analysis method that accounts for this?"

This is the right question, and it's a different problem from the SRM issue above — this one doesn't break randomization, it just means the headline effect is measuring something softer than "the effect of actually seeing a GIF." It's measuring the effect of being assigned to a category page that's allowed to show GIFs, averaged across everyone in that arm whether or not a GIF card actually crossed their screen. That's an intent-to-treat (ITT) effect, and it's real — but it isn't the number an ML engineer asking about GIF exposure specifically is actually after.

From ITT to LATE: What "Encouragement Design" Means Here

This situation — an experiment that can only nudge the probability of exposure rather than force it — has a name: an encouragement design. Nobody can make a user actually notice a GIF card; the experiment can only make GIF cards available to be noticed. Compliance is imperfect on the control side by construction, the same way a study that mails a coupon can't force anyone to redeem it. The standard move in an encouragement design is to convert the diluted ITT estimate into a Local Average Treatment Effect (LATE) — the effect among the users who were actually exposed — using the exposure rate as a scaling factor, the same logic behind an instrumental-variables (Wald) estimator.

It's easier to reason about with control relabeled as the "encouraged" arm: overall, treatment (no GIFs) trails control (GIFs on) by some percentage on a guardrail metric — equivalently, control leads treatment by that same percentage. If only a fraction rr of control users were actually exposed to a GIF card at all, the observed effect is that true per-exposure effect, diluted by rr:

ITT effect=r×X\text{ITT effect} = r \times X

where rr is the exposure rate (measurable directly — the share of control users who saw at least one GIF card) and XX is the per-exposure effect you actually want. Solving for it:

X=ITT effectrX = \frac{\text{ITT effect}}{r}

As a toy check before plugging in real numbers: if r=0.45r = 0.45 and the observed ITT effect were 1.35%1.35\%, the per-exposure effect would be 1.35%/0.45=3.0%1.35\% / 0.45 = 3.0\% — three times larger than the diluted number, because fewer than half of control users ever actually saw what was being tested.

Applying the Real Numbers

On the actual guardrail metric — orders per active user — the experiment's observed effect was −2.0%-2.0\% (treatment relative to control, p = 0.02), and the measured exposure rate came back at r=33%r = 33\%: about a third of control users saw a GIF card on the category page during the experiment window. Applying the same formula:

X=−2%0.33=−6.06%X = \frac{-2\%}{0.33} = -6.06\%

The diluted, headline number said "removing GIFs cost about 2% of orders per user." The undiluted number — the effect among users who would have actually encountered a GIF card — is closer to 6%. The intent-to-treat effect wasn't wrong; it was just averaged across a lot of users the treatment could never have affected in the first place, which is exactly the dilution the follow-up question was pointing at.

"Does That Mean We Just Read It As Is?"

The last question: "Given r comes out to about 33%, does that mean we can just take the treatment arm's number at face value?"

Worth pinning down first what "at face value" would mean here — the platform's reported effect and p-value describe the diluted ITT comparison, which is a perfectly valid, real result on its own. Dividing by rr doesn't replace that number; it answers a different, narrower question (the effect on exposed users specifically) that the ITT number was never designed to answer. Strictly, rescaling the point estimate this way also rescales its variance — a proper Wald-estimator standard error isn't simply the original standard error divided by rr, so a fully rigorous significance test on XX needs its own derivation. As a rough guide, though, the sign and rough significance of the diluted effect carries over: if the platform found the−2.0%-2.0\% ITT effect significant, the −6.06%-6.06\% exposed-user effect can reasonably be treated as at least that significant, since dividing by a fixed, precisely estimated rr doesn't introduce a reason for the direction of the result to reverse.

Why This Generalizes

Category pages and GIF cards are incidental. The pattern underneath is not: any experiment where the treatment arm can only make something possible to encounter rather than guaranteed — a notification that might not get opened, a feature that might not get discovered, a redesign that only some fraction of sessions actually render — is an encouragement design, whether or not anyone on the team calls it that. The ITT effect from the platform is always real and always worth reporting. But whenever exposure is partial, it's worth asking the same question that started this whole thread: what fraction of the treatment arm was actually exposed, and what does the effect look like once that dilution is divided back out?