2023–2026 · Product manager and part-owner; author; instructor · Research

Experimentation at scale

A/B testing as it is actually done, inside a small e-commerce retailer, written up as a textbook chapter and taught as a graduate course.

Experimentation at scale

Outcome

A Springer chapter (forthcoming), a CSCW 2023 workshop paper, and Analytics and User Experience at Waterloo, 87 graduate students.

Path

  1. 2023 · Question

    Why do so many A/B tests produce quick wins that never show up in revenue?

  2. 2023 · Workshop

    Considerations for experimental design within funnels. CSCW 2023 workshop.

  3. 2023 · Course

    MSCI 543, Analytics and User Experience: experimentation taught from live funnels, failure modes included.

  4. 2024–25 · Field

    A year of live experiments inside a healthcare e-commerce retailer, run and documented from the inside.

  5. 2026 · Chapter

    A/B Testing IRL. Textbook chapter, Springer, forthcoming.

The setting

  • A mid-sized healthcare e-commerce retailer, about 10,000 transactions a year, with an analytics team of a product manager, a UX designer, and a data analyst.
  • Shopify, Amplitude, and LaunchDarkly for funnels, cohorts, and feature flags. The tools every small retailer has.
  • I was the product manager and a part-owner, so the account is auto-ethnographic: the experiments as they were actually run, warts and all.

The experiments

  • Five families: cohort-based retention, funnel friction reduction, device-specific optimisation, personalised engagement, and content and trust building.
  • Localised wins were real. Checkout abandonment fell 8–12% in the best cases.
  • Most of them didn’t move revenue or repeat purchase. A lot of A/B testing produced quick wins that never turned into long-term behaviour.

Five tensions

  • Friction vs. engagement: removing steps speeds people up, and some of those steps are what make people trust you.
  • Short-term wins vs. long-term impact: a better checkout number isn’t a better customer.
  • Personalisation vs. experiment integrity: the recommender keeps learning while you’re trying to hold it still.
  • Funnel-stage optimisation vs. holistic gains: fixing one stage moves the drop-off to the next.
  • Experiment volume vs. data reliability: with modest traffic, every extra concurrent test costs you statistical power.

What it argues

  • Customer paths aren’t linear, so funnel diagrams mislead. Interventions have to be pathway-specific.
  • Structured experiment tracking, holdout groups, and adaptive funnel strategies make small-traffic experimentation trustworthy.
  • Experimentation is organisational learning, and should be taught that way.

The course

  • MSCI 543, Analytics and User Experience, University of Waterloo, Spring 2023. Graduate. 87 students.
  • Experimentation taught from the funnels and instrumentation I used in industry, with the failure modes included.

Papers from this work

  • A/B Testing IRL. Textbook chapter, Springer, forthcoming 2026.
  • Considerations for experimental design within funnels. CSCW 2023 workshop.