Filter results

How Clover experiments when billions of dollars flow through daily

Why JobLeads says one test won't move you, but 100 will

Inside Aspen Dental's 100-test-a-year experimentation program

Synthetic audiences meet real A/B tests at Principal Financial Group

Why US Bank considers missing even 1% of customers unacceptable

Why Farfetch manages by learning rate, not win rate

How Cogniteer Built an Experimentation Engine From Scratch

How Fin does 1,000,000 A/B Tests in 24 Hours

How Kargo turns losing experiments into competitive edges

The 'wine effect' and other surprises that reshaped how Box runs e-commerce experiments

Diligent explains why moving on from an experiment might cost you

The metric Stitch Fix says every experimenter should chase

What the Expedia Group cannot measure, it cannot ship
Top takeaways from our favorite conversations

Purge “anti-knowledge” by standardizing design, instituting cross-functional reviews, and only codifying learnings supported by repeatable data.

DoorDash's price experiment proved price by itself doesn't predict orders. Different customers want different things at different times, which pushed the team toward personalization.

Unblock teams: create a center of excellence for data science and enable rapid variants with AI-powered tooling.
.svg.avif)
A failed test can hold the real winner; contextual onboarding matched to user intent roughly doubled activation and became the default variant after the bundling experiment was rolled back.

Top-down buy-in shifts the conversation from "why test?" to "how do we test?": When leadership treats data as the tiebreaker, teams stop defending opinions and start building better experiments.

Twitch used geo-fenced experiments with matched markets and causal inference to measure true price elasticity, turning a feared pricing decision into a measured, accretive one.

Win rate matters less than learnings per test — DoorDash ships company-wide experiment summaries (win or lose) that the CEO actively reads and responds to, creating cultural accountability around testing rigor.

Simplification has a limit. Removing too much can strip away the cues and context buyers actually need to decide.

Massey's first test removed navigation from UPS's shipping checkout flow and delivered $35 million in incremental revenue—proving e-commerce best practices apply even when customers think "this is just a tool, not e-commerce."

Persistence pays: four months and three to four rounds of trial-model testing at Codecademy produced a 35% conversion increase.

When you struggle to land a result, lead with the story of what the customer did, then bring the numbers.

False negatives are more dangerous than false positives — they get institutionalized as "we tried that, it didn't work" and quietly kill good ideas for years.

Test metrics before you test features — usage time could signal engagement or just mean your product takes too long to do its job.

Scale experimentation with AI: use Cursor desktop/cloud agents for parallel builds and visual QA; orchestrate docs/analysis via Claude; automate cleanups and reporting.

A feature that fails early in a flow can succeed later; placement and timing often matter more than the idea itself.

Build composite metrics (e.g., CPQI) to align finance, engineering, and data science around shared outcomes.

Share losses as openly as wins. Wins build credibility, and losses build the psychological safety a testing culture runs on.




