Skip to content

Debates

What makes an A/B test result valuable rather than a wasted effort?

8 recorded positions from 6 people, first said Oct 12, 2022. They do not agree — the readings below are what each one actually argued.

Enter with strong intuition about expected outcome to know what to watch

Harry Stebbings · Nov 20, 2023

Teams should list their top three priorities and start with one, and define what success looks like before testing anything

Too many people run tests and then can't tell if the result matched expectations because they never set the target, so they can't decide whether to continue

28:53 20VC: The Ultimate Hiring Playbook: Five Questions to Ask Every New Hire | What Makes Truly Great Leaders and How They Give Feedback | Do VCs Really Add Value; Lessons from Hard Fundraises with Matteo Franceschetti, Co-Founder @ Eight Sleep

Raman Malik · Nov 15, 2024

You should enter an A/B test with strong intuition about what will happen, including which secondary metric might go down; without that intuition you're flying blind.

Intuition about the expected outcome is what tells you what to watch out for and how to interpret the result.

13:13 20Growth: Inside Perplexity's Growth Machine: What Worked, What Did Not Work | Why Paid Acquisition is a Drug and Brand Marketing is BS | The Good, Bad and Ugly of A/B Tests and Why Micro-Optimisations are Under-Rated with Raman Malik

Testing velocity and approach must scale with audience volume and funnel length

Mike Krieger · Feb 22, 2023

Process and technology choices should be rightsized to the company's stage: practices appropriate for a multi-office 2,000-engineer team are not right for a seven-person startup.

They ran a test on the Artifact beta group, did the math, and found the required effect size was implausible; only rare launches like personalized Explore or Stories delivered 2x deltas

Scope: changed once they went public and had enough users to get real results

37:59 20VC: Instagram Founders Kevin Systrom and Mike Krieger on Why Social Networks Should Be Less Social & The Next Wave of Social | Why San Francisco Will Return with a Vengeance and The Future For Remote Work | Let's Get Personal: Relationships to Money, Be

Adam Grenier · Mar 15, 2023

How fast you can read a message's performance depends on volume and funnel length, so the testing approach must differ radically between high-volume and long-funnel businesses

At Uber's billion-dollar-a-year spend he could get ad-layer learnings within minutes and run 50 messages at once through a creative system, whereas Lambda School's year-long funnel and hundred-person classes meant testing one thing for three weeks and leaning on qualitative student research for signal

32:46 20Growth: The Inside Story to Uber's Hypergrowth Scaling; What Worked, What Did Not? | Spending a $1BN Budget at Uber and Why China was the Wild West for Uber | Why You Do Not Need a Growth Team with Adam Grenier

Also on the record

Raman Malik · Nov 15, 2024

A well-scoped A/B test's real value is directional — it teaches you what to work on next and reveals what is actually incremental, whereas a poorly scoped test with a low minimum detectable effect just wastes a week

You need to know how every metric moves and whether you're accidentally harming metrics; badly scoped tests yield no information

11:55 Well scoped tests teach directionally what to work on next poorly scoped tests waste time

Raman Malik · Nov 15, 2024

A/B tests with mixed results — one metric up, another down — are the trickiest to navigate; rather than running the test for sixty days to see if retention gains offset the loss, you have to make the call right away or find a way to protect the downside.

Long-term retention gains might make up for a short-term query volume loss, but you can't run a test for sixty days before deciding; you have to decide in the moment or protect the downside.

13:53 Mixed signal tests require an immediate call or downside protection not extended waiting

Marty Cagan · Dec 7, 2022

You can usually determine in hours — not weeks — that an idea isn't right, and should iterate as soon as the team is convinced rather than waiting a month for statistically significant results

Putting prototypes in front of real users to test whether they could use it and whether they would use it allows a dozen iterations in a day

20:08 Qualitative prototype testing with real users yields decisive signal within hours

Gustav Söderström · Oct 12, 2022

Test changes separately on new versus existing users to distinguish trained habit from genuine quality

Existing users have a habit so metrics can fall, while new users who were never taught the old paradigm reveal whether the new design is actually better in a global sense — which then lets you decide whether to take the pain of retraining the base

18:31 Test changes on new users not just existing ones to separate trained habit from genuine quality

Your assistant can query this graph directly — 8 positions here, 19,646 across the corpus. Add 996.fm over MCP.