Retrospective

Field notes from real experiment reviews and postmortems — the ambiguous rollout calls, the metric traps, and the reasoning that actually resolved them.

Worse on the Surface, Better for the Platform: A Rollout Call Worth Revisiting

A redesigned placement's own click-through rate dropped, but every platform-level guardrail went up. Ship it or roll it back? A look back at how that call actually got made.

RetrospectiveGuardrail MetricsProduct Analytics
Read article →

Hunting for a Goal Metric After the Fact Is Just Multiple Testing in Disguise

When nobody commits to a single goal metric before launch, every hypothesis becomes another roll of the dice. Why that habit is mathematically identical to p-hacking, and how to close it off.

RetrospectiveMultiple TestingGoal Metrics
Read article →

Not Every CTR Drop Means "Exploring Elsewhere" — A Case Where It Didn't

Replacing a 'see more' page with an inline carousel cut a navigation step — and made browsing worse. Why fewer clicks isn't automatically better UX, and why click and like aren't the same signal.

RetrospectiveProduct AnalyticsUX Measurement
Read article →

Every Experiment Doc Lives in Three Different Tools. What If an Agent Just Wrote the One-Pager?

Planning docs, Slack threads, and results all live in different tools with no canonical write-up. An idea for an agent that drafts the one-pager the day an experiment starts, and closes it out when it ends.

RetrospectiveExperiment OpsTooling
Read article →

Two Kinds of Guardrail Tests, Two Different Corrections (and Where I Mixed Them Up)

A standard significance test and a non-inferiority test both get called 'the guardrail test,' but they compound differently as guardrails multiply — one needs Bonferroni, the other needs a power correction.

RetrospectiveGuardrail MetricsStatistics
Read article →