What makes an A/B test result valuable rather than a wasted effort?
8 recorded positions from 6 people, first said Oct 12, 2022. They do not agree — the readings below are what each one actually argued.
Enter with strong intuition about expected outcome to know what to watch
Harry Stebbings · Nov 20, 2023
Teams should list their top three priorities and start with one, and define what success looks like before testing anything
Too many people run tests and then can't tell if the result matched expectations because they never set the target, so they can't decide whether to continue
28:53 20VC: The Ultimate Hiring Playbook: Five Questions to Ask Every New Hire | What Makes Truly Great Leaders and How They Give Feedback | Do VCs Really Add Value; Lessons from Hard Fundraises with Matteo Franceschetti, Co-Founder @ Eight Sleep
Raman Malik · Nov 15, 2024
You should enter an A/B test with strong intuition about what will happen, including which secondary metric might go down; without that intuition you're flying blind.
Intuition about the expected outcome is what tells you what to watch out for and how to interpret the result.
13:13 20Growth: Inside Perplexity's Growth Machine: What Worked, What Did Not Work | Why Paid Acquisition is a Drug and Brand Marketing is BS | The Good, Bad and Ugly of A/B Tests and Why Micro-Optimisations are Under-Rated with Raman Malik
Testing velocity and approach must scale with audience volume and funnel length
Mike Krieger · Feb 22, 2023
Process and technology choices should be rightsized to the company's stage: practices appropriate for a multi-office 2,000-engineer team are not right for a seven-person startup.
They ran a test on the Artifact beta group, did the math, and found the required effect size was implausible; only rare launches like personalized Explore or Stories delivered 2x deltas
Scope: changed once they went public and had enough users to get real results
37:59 20VC: Instagram Founders Kevin Systrom and Mike Krieger on Why Social Networks Should Be Less Social & The Next Wave of Social | Why San Francisco Will Return with a Vengeance and The Future For Remote Work | Let's Get Personal: Relationships to Money, Be
Adam Grenier · Mar 15, 2023
How fast you can read a message's performance depends on volume and funnel length, so the testing approach must differ radically between high-volume and long-funnel businesses
At Uber's billion-dollar-a-year spend he could get ad-layer learnings within minutes and run 50 messages at once through a creative system, whereas Lambda School's year-long funnel and hundred-person classes meant testing one thing for three weeks and leaning on qualitative student research for signal
32:46 20Growth: The Inside Story to Uber's Hypergrowth Scaling; What Worked, What Did Not? | Spending a $1BN Budget at Uber and Why China was the Wild West for Uber | Why You Do Not Need a Growth Team with Adam Grenier
Also on the record
Raman Malik · Nov 15, 2024
A well-scoped A/B test's real value is directional — it teaches you what to work on next and reveals what is actually incremental, whereas a poorly scoped test with a low minimum detectable effect just wastes a week
You need to know how every metric moves and whether you're accidentally harming metrics; badly scoped tests yield no information
11:55 Well scoped tests teach directionally what to work on next poorly scoped tests waste time
Raman Malik · Nov 15, 2024
A/B tests with mixed results — one metric up, another down — are the trickiest to navigate; rather than running the test for sixty days to see if retention gains offset the loss, you have to make the call right away or find a way to protect the downside.
Long-term retention gains might make up for a short-term query volume loss, but you can't run a test for sixty days before deciding; you have to decide in the moment or protect the downside.
13:53 Mixed signal tests require an immediate call or downside protection not extended waiting
Marty Cagan · Dec 7, 2022
You can usually determine in hours — not weeks — that an idea isn't right, and should iterate as soon as the team is convinced rather than waiting a month for statistically significant results
Putting prototypes in front of real users to test whether they could use it and whether they would use it allows a dozen iterations in a day
20:08 Qualitative prototype testing with real users yields decisive signal within hours
Gustav Söderström · Oct 12, 2022
Test changes separately on new versus existing users to distinguish trained habit from genuine quality
Existing users have a habit so metrics can fall, while new users who were never taught the old paradigm reveal whether the new design is actually better in a global sense — which then lets you decide whether to take the pain of retraining the base
18:31 Test changes on new users not just existing ones to separate trained habit from genuine quality
Your assistant can query this graph directly — 8 positions here, 19,646 across the corpus. Add 996.fm over MCP.