What is the binding constraint on frontier AI research progress?
27 recorded positions from 16 people, first said Aug 28, 2023. They do not agree — the readings below are what each one actually argued.
Data quality is the primary bottleneck ahead of compute and algorithms
Alexandr Wang · Jun 12, 2024
AI progress requires compute, data, and algorithms to advance in tandem; the industry has scaled compute dramatically without building up the other two pillars
The history of AI is that progress comes from all three pillars being built together — compute plus algorithmic advances like the transformer or RLHF plus supporting data
5:32 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's
Alexandr Wang · Jun 12, 2024
AI progress is now fundamentally data-bottlenecked rather than compute-bottlenecked
If data production scaled in lockstep with NVIDIA's chip output, models would be astronomically more capable than they are
16:15 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's
Eiso Kant · Oct 7, 2024
Algorithmic and compute-efficiency work is table stakes rather than a differentiator; the real differentiation between two models is the data
Every lab — OpenAI, Anthropic, Google — is constantly improving compute efficiency, so it only lets you keep up; data is what varies
9:49 20VC: Raising $500M To Compete in the Race for AGI | Will Scaling Laws Continue: Is Access to Compute Everything | Will Nvidia Continue To Dominate | The Biggest Bottlenecks in the Race for AGI with Eiso Kant, CTO @ Poolside
Eiso Kant · Oct 7, 2024
The scale of data matters more than he previously thought — compute alone is not the determining factor
Scope: describes it as a continued realization over the last twelve months
57:42 20VC: Raising $500M To Compete in the Race for AGI | Will Scaling Laws Continue: Is Access to Compute Everything | Will Nvidia Continue To Dominate | The Biggest Bottlenecks in the Race for AGI with Eiso Kant, CTO @ Poolside
Edwin Chen · Jul 21, 2025
Data quality is the most pressing bottleneck to AI progress, followed by compute, with algorithms last
42:41 20VC: Scaling to $1BN+ in Revenue with No Funding: Surge AI | The Most Insane Scaling Story in Tech |
Edwin Chen · Jul 21, 2025
Data quality problems have already caused major setbacks at frontier labs, with teams discovering after six to twelve months that their training and eval data was bad and their apparent progress was illusory or negative
Teams repeatedly report to Surge that they trained and evaluated on data sourced elsewhere, watched metrics rise, then realized the data was garbage and the models had made no progress or regressed
Scope: based on what customer teams report to Surge before working with them
42:58 20VC: Scaling to $1BN+ in Revenue with No Funding: Surge AI | The Most Insane Scaling Story in Tech |
Compute because every idea must be tested at scale
Cristóbal Valenzuela · Aug 28, 2023
Speed is Runway's biggest rate-limiting factor, with industry-wide compute scarcity as a recurring constraint
Deploying and training models efficiently is constrained by available compute and by use case deployment
Scope: also sometimes constrained by use cases and model efficiency
20:51 20VC: Why AI Models are not a Moat, Where Does the Value in AI Accrue; Startups or Incumbents, What the World Has Got Wrong About AI, Why AI Needs a New Story and Who is the Right People to Tell it with Cris Valenzuela, Co-Founder & CEO @ Runway
Reid Hoffman · Jun 10, 2024
AI's progress comes not from discovering the right algorithm but from applying scale compute and large data with scale teams, which makes compute central to both training and inference
Each new level of compute scale has so far brought serious new capabilities to the table
Scope: all j-curves eventually become s-curves; true 'at the moment'
11:37 20VC: Reid Hoffman on Foundation Models: Who Wins & How Do Incumbents Respond | The Inflection AI Deal: How it Went Down | Why Trump is a Threat to Democracy | The Future of TikTok | Lessons from Sam Altman, Brian Chesky and the OpenAI Board
Eiso Kant · Oct 7, 2024
The biggest misconception about AI over the next ten years is that progress will halt; the only thing that would actually halt it is a global conflict disrupting the chip supply chain
58:52 20VC: Raising $500M To Compete in the Race for AGI | Will Scaling Laws Continue: Is Access to Compute Everything | Will Nvidia Continue To Dominate | The Biggest Bottlenecks in the Race for AGI with Eiso Kant, CTO @ Poolside
Jonathan Ross · Sep 29, 2025
Compute is the easiest and most predictable knob to turn in AI — algorithms rarely improve and more data is hard to get, whereas writing a check and waiting reliably yields more compute
Compute keeps improving every year and is purchasable, while synthetic data generation isn't yet good enough to convert compute directly into data
Scope: synthetic data generation is improving and getting there
40:40 20VC: OpenAI and Anthropic Will Build Their Own Chips | NVIDIA Will Be Worth $10TRN | How to Solve the Energy Required for AI... Nuclear | Why China is Behind the US in the Race for AGI with Jonathan Ross, Groq Founder
Demis Hassabis · Apr 7, 2026
Compute is the biggest bottleneck in AI progress today, not only for scaling systems up but because experiments must be run at meaningful scale to be valid
The cloud is the researcher's workbench: a new algorithmic idea has to be tested at reasonable scale or the result won't hold when put into the main system, so many researchers with many ideas require a lot of compute
6:11 20VC: DeepMind's Demis Hassabis on Why AGI is Bigger than the Industrial Revolution | Why LLMs Will Not Commoditise & We Have Not Hit Scaling Laws | Bottlenecks in AI & The Energy Crisis Caused By AI | Whether AI Will Do More to Harm or Help Inequality
Internet lacks reasoning trace data not reasoning itself is the bottleneck
Alexandr Wang · Jun 12, 2024
Current models are better than any human at emulating the internet, but AGI-level systems that do tasks and solve hard problems will not be reached from internet data
Pre-training optimizes for emulating the internet, while what we want from powerful AI — doing tasks, solving difficult problems, collaborating with humans — is a different capability, and the internet data is already exhausted
6:54 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's
Alexandr Wang · Jun 12, 2024
None of the reasoning and thinking that powers the economy today gets written down on the internet, so models trained only on the internet cannot learn it
A fraud analyst deducing whether a transaction is fraudulent never writes down their step-by-step analysis anywhere crawlable
8:02 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's
Aidan Gomez · Aug 19, 2024
The bottleneck on model reasoning is not that reasoning is intrinsically hard but that the internet contains almost no training data demonstrating the reasoning process itself
The web is the output of reasoning — people present conclusions without showing their work — so the demonstrations of thinking simply aren't freely available and have to be built
12:50 20VC: Chips, Models or Applications; Where is the Value in AI | Is Compute the Answer to All Model Performance Questions | Why Open AI Shelved AGI & Is There Any Value in Models with OpenAI Price Dumping with Aidan, Gomez, Co-Founder @ Cohere
Total compute not model size or data size alone is the binding constraint
Noam Shazeer · Aug 31, 2023
The total amount of computation spent training a model matters more than model size or data size individually; model size is the harder of those two constraints, but compute operations are the real binding constraint.
You want both a bigger model and longer training, and both are gated by how many compute operations you can afford.
0:00 20VC: Spending $2M to Train a Single AI Model: What Matters More; Model Size or Data Size | Hallucinations: Feature or Bug | Will Everyone Have an AI Friend in the Future & Raising $150M from a16z with Noam Shazeer, Co-Founder & CEO @ Character.ai
Noam Shazeer · Aug 31, 2023 · hedged
Model size is a bigger challenge than data size, but the real constraining factor is the amount of computation required to train the model
Data is relatively easy to get; making the model bigger and training it longer both multiply into training time, so compute operations are the binding constraint
Scope: both model size and data size matter
16:38 20VC: Spending $2M to Train a Single AI Model: What Matters More; Model Size or Data Size | Hallucinations: Feature or Bug | Will Everyone Have an AI Friend in the Future & Raising $150M from a16z with Noam Shazeer, Co-Founder & CEO @ Character.ai
Noam Shazeer · Aug 31, 2023
Computation is the biggest constraint on Character.ai's models, and better hardware plus longer training will straightforwardly yield smarter models
The currently served model was trained with about $2M of compute; more and better hardware plus longer training produces something smarter, just as 2016-era compute only bought translation-level ability
17:14 20VC: Spending $2M to Train a Single AI Model: What Matters More; Model Size or Data Size | Hallucinations: Feature or Bug | Will Everyone Have an AI Friend in the Future & Raising $150M from a16z with Noam Shazeer, Co-Founder & CEO @ Character.ai
Compute alone is insufficient data is an equally necessary bottleneck
Arthur Mensch · Apr 29, 2024
Scale matters but is not the only ingredient: compressing models requires training compute, and results also depend on data quality and training technique ('compute multipliers')
Without proper data you hit a data-quality limit; compute is expensive so efficiency gains that don't cost compute are the thing to harvest
9:38 20VC: Mistral's Arthur Mensch: Are Foundation Models Commoditising | How Do We Solve the Problem of Compute | Is There Value in the Application Layer | Open vs Closed: Who Wins and Mistral's Position
Alexandr Wang · Jun 12, 2024
The biggest misconception about AI today is that compute is all that stands between us and AGI; data is also required to get there
53:00 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's
Also on the record
Joelle Pineau · Nov 3, 2025
Algorithmic innovation is the hardest and most creative of the three AI scaling inputs to make progress on
The space of possible ideas is enormous and you cannot tell whether a direction was the right one until you get there — much like reinforcement learning
13:51 Algorithmic creativity is the hardest input to advance
Joelle Pineau · Nov 3, 2025
Labs need an equilibrium between talent and compute — excess talent without compute is wasted — and the importance of data is often underestimated
If you have too much talent and not enough compute you're wasting your time; data is getting more and more expensive and deserves a large share of spend
31:32 Equilibrium across talent compute and data not any single input
Sam Altman · Apr 15, 2024
The biggest risks to OpenAI's velocity are losing its researchers or research culture, and not having enough compute.
OpenAI exists to do useful things for people, so even the best research is blocked if there isn't enough compute to serve everyone who wants to use it as models improve.
12:57 Losing researchers or compute shortage are the two biggest velocity risks
Sam Altman · Apr 15, 2024
OpenAI's near-term challenge over the next twelve months is doing the best research and best productization/innovation, and a further challenge is sufficient supply chain and compute.
42:16 Compute and supply chain scarcity is the binding near term constraint
Bret Taylor · Oct 2, 2024 · hedged
Progress in AI has three inputs — data, compute, and algorithms/methodology — and each still has room for major gains, so a plateau or diminishing returns is not inevitable.
Algorithmic work continues (post-transformer research, instruction tuning precedent), compute clusters keep growing, and data constraints are being addressed via simulation, synthetic data and multimodality.
22:40 Data compute and algorithms each still have headroom so no single bottleneck
Anjney Midha · Apr 14, 2026
Algorithmic innovation is no longer a bottleneck; it is downstream of culture and solves itself if you get culture right
The right culture attracts the best researchers, who are mission-driven and architecture-agnostic rather than committed to transformers over diffusion models, so algorithmic progress falls out of having a flexible frontier team
5:30 Culture and researcher attraction not algorithms is the bottleneck
Ethan Mollick · Jul 31, 2024 · hedged
AI will hit a succession of 'reverse salients' — whichever input is lagging (data, then something else) becomes the focus of all effort and then gets solved and forgotten
History of science pattern: generators then transmission in early electricity, solar then batteries today; once all of science concentrates on one bottleneck we tend to find ways forward
12:39 Successive reverse salients each bottleneck becomes the focus then gets solved
Aravind Srinivas · Jun 5, 2024
There is still real gain left in making models bigger and training on more tokens, but only for labs that get the fine details of data curation and architecture right
Many labs have trained very large models on lots of data and ended up with nothing; the returns depend on the data mix — English vs other languages, code, math, chain-of-thought — plus scaling-law choices and mixture-of-experts efficiency
7:29 Data curation and architecture mastery not brute force scale determines remaining scaling gains
Your assistant can query this graph directly — 27 positions here, 19,646 across the corpus. Add 996.fm over MCP.