Skip to content

Debates

How much does training-data difference actually differentiate frontier AI models?

9 recorded positions from 8 people, first said Jun 30, 2023. They do not agree — the readings below are what each one actually argued.

Marginal data is noise proprietary domain data is what matters

Christian Lanng · Sep 27, 2023

A 10% difference in training data is a rounding error at trillion-parameter scale; what matters is years of proprietary domain data such as enzyme manufacturing or biology

At that scale the marginal data barely changes outputs, whereas deep domain data is unavailable elsewhere

60:20 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng

Anish Acharya · Feb 9, 2026

Data network effects were mostly a fake moat historically, but proprietary AND live data is now a genuinely powerful moat

'Data network effect' used to be what you said when you couldn't name a real moat; today a commodity model sitting on proprietary live data outperforms the most cutting-edge model without access to it

Scope: 'live' means continuously changing data such as health data or telemetry from a running product; open question how proprietary such data can remain

25:52 20VC: Is SaaS Dead in a World of AI | Do Margins Matter Anymore | Is Triple, Triple, Double, Double Dead Today? | Who Wins the Dev Market: Cursor or Claude Code | Why We Are Not in an AI Bubble with Anish Acharya @ a16z

Data secrecy implies real differences between labs

Douwe Kiela · Jun 30, 2023 · hedged

Part of the secret sauce making GPT-4 so good is that OpenAI went to enormous lengths to acquire special data nobody else has, such as using Whisper to transcribe the world's podcasts for high quality language data

Having high quality training data that no one else possesses puts you in a position of advantage

Scope: allegedly, on the Whisper transcription project

13:53 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI

Harry Stebbings · Sep 27, 2023 · hedged

The major labs may not actually be trained on the same data, since engineers who happily discuss their models refuse to discuss underlying data

Their secrecy specifically about training data suggests meaningful differences

60:10 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng

Data not algorithms is the scarce input so providers capture huge value

Nick Frosst · Sep 1, 2025

Algorithms are not the bottleneck to making models more useful; data quality is

Algorithms have changed only a little (base training plus RLHF and other RL techniques), whereas the constraint is getting good quality real data and generating good synthetic data from it

9:35 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst

Anastasios Angelopoulos · Aug 3, 2026

Leading data providers will easily be worth hundreds of billions of dollars, and possibly more

Data is the hardest part of model training — it must be sourced, it is dirty, and nobody wants to do it — while algorithms have become somewhat commoditised now that everyone knows how to use the transformer

43:02 20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena

Also on the record

Jeff Seibert · Nov 22, 2023

Acquiring high-quality clean training data is extremely challenging and genuinely valuable as a defense

Reddit, Twitter and others are shutting off APIs and tightening rate limits — a reversal of twenty years of opening up — showing there's been a clear realization data is valuable

36:03 Clean proprietary training data is scarce and valuable as a competitive defense

Mike Krieger · Mar 3, 2025

For labs, the value lies in team quality and in whether the model can reliably perform real-world actions in deployment, not primarily in proprietary data

Usefulness in actual use cases is the highest order bit; evals are useful for hill climbing and internal research but don't tell you whether a model will be excellent at what it's deployed for, or only in narrow situations

31:46 Team quality and deployment reliability not data is the real differentiator

Nick Frosst · Sep 1, 2025

Data is still a bottleneck despite synthetic data

You need real-world high-quality data to seed the synthetic data process; Cohere still uses in-house annotators making real data

9:05 Synthetic data still needs real world seed data so data remains scarce

Your assistant can query this graph directly — 9 positions here, 19,646 across the corpus. Add 996.fm over MCP.