Field notes from real experiment reviews and postmortems — the ambiguous rollout calls, the metric traps, and the reasoning that actually resolved them.
A redesigned placement's own click-through rate dropped, but every platform-level guardrail went up. Ship it or roll it back? A look back at how that call actually got made.
When nobody commits to a single goal metric before launch, every hypothesis becomes another roll of the dice. Why that habit is mathematically identical to p-hacking, and how to close it off.
Replacing a 'see more' page with an inline carousel cut a navigation step — and made browsing worse. Why fewer clicks isn't automatically better UX, and why click and like aren't the same signal.
Planning docs, Slack threads, and results all live in different tools with no canonical write-up. An idea for an agent that drafts the one-pager the day an experiment starts, and closes it out when it ends.
A standard significance test and a non-inferiority test both get called 'the guardrail test,' but they compound differently as guardrails multiply — one needs Bonferroni, the other needs a power correction.