Skip to content

Debates

Is inference speed a decisive competitive differentiator for AI products?

7 recorded positions from 4 people, first said Mar 24, 2025. They do not agree — the readings below are what each one actually argued.

Users wont tolerate slow background prompts speed matters viscerally

Varun Mohan · Jun 2, 2025

Users care about latency a great deal — in an autocomplete/tab product every 10 milliseconds moves acceptance rate by percentage points

Scope: evidence drawn from their tab/autocomplete product

39:20 20VC: Windsurf Founder on Will Model Companies Own the App Layer | Why Moats Do Not Exist in a World of AI | Why the Notion of Single Person $BN Companies is BS | Lovable vs Bolt & Cursor vs Windsurf: How Does it All End with Varun Mohan

Jonathan Ross · Sep 29, 2025

The view that users will be content with slow, long-running background prompts is completely wrong

People are consistently bad at predicting what drives engagement and outcomes; the same skepticism greeted fast-loading web pages, and lessons from the early internet show the visceral importance of speed

13:59 20VC: OpenAI and Anthropic Will Build Their Own Chips | NVIDIA Will Be Worth $10TRN | How to Solve the Energy Required for AI... Nuclear | Why China is Behind the US in the Race for AGI with Jonathan Ross, Groq Founder

Speed has no upper bound for hard problems

Alexander Embiricos · Feb 21, 2026 · hedged

Inference speed matters a lot, and OpenAI is attacking it from all angles — hardware, how inference is served (40% faster in the API, 25% faster in Codex), and model efficiency

GPT-5.3 Codex is significantly more efficient and users report it feels competitively fast; serving changes delivered 40% faster API and 25% faster Codex responses

16:56 20VC: Codex vs Claude Code vs Cursor: Who Wins, Who Loses | Will All Coding Be Automated - Do We Need PMs | The Real Bottleneck to AGI | The Three Phases of Agents and What You Need to Know with Alex Embiricos, Head of Codex at OpenAI

Andrew Feldman · May 26, 2026

For hard problems there is no upper bound on the value of inference speed; speed will be decisive across coding, agentic flows and every part of AI

If you solve in three minutes what a competitor takes twenty minutes to solve, you get vastly more problems solved per day and they get smoked; and historically there is no market for slow technology — nobody would take $1,000 a month to have slow internet

Scope: specifically for hard problems

20:37 20VC: Cerebras CEO on the Future of Data Centres, Token Costs and Memory | We are Not in an Infra Bubble & Dario Got a Bad Deal with Elon for Compute | Should US Companies Sell to China & Why Most Layoffs are AI Washed with Andrew Feldman

Also on the record

Jonathan Ross · Sep 29, 2025

Customers say they want speed but the real value proposition is compute capacity

Every customer arrives asking about speed, but none keep asking once they hit the supply constraint; two weeks ago a customer requested five times Groq's total capacity that no hyperscaler could supply either

22:00 Real customer value is compute capacity not raw speed

Andrew Feldman · Mar 24, 2025

There are many viable architectural approaches to AI compute, but Cerebras's has been the fastest across a broad set of models every day since it launched inference in August 2024

Independent benchmarking by Artificial Analysis and others shows this

11:27 Benchmarked speed leadership sustained across many models proves differentiation

Andrew Feldman · Mar 24, 2025

What customers prioritize — speed, cost, or accuracy — varies entirely by workload: accuracy dominates in medical diagnosis, cost dominates in batch jobs like synthetic data generation, and speed dominates in interactive use

A week's wait and higher price is worth one percentage point of accuracy in a cancer diagnosis, while generating data to tune a smaller model has no urgency at all

15:23 Workload type determines whether speed cost or accuracy matters most

Your assistant can query this graph directly — 7 positions here, 19,646 across the corpus. Add 996.fm over MCP.