How much does training-data difference actually differentiate frontier AI models?
9 recorded positions from 8 people, first said Jun 30, 2023. They do not agree — the readings below are what each one actually argued.
Marginal data is noise proprietary domain data is what matters
Christian Lanng · Sep 27, 2023
A 10% difference in training data is a rounding error at trillion-parameter scale; what matters is years of proprietary domain data such as enzyme manufacturing or biology
At that scale the marginal data barely changes outputs, whereas deep domain data is unavailable elsewhere
60:20 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng
Anish Acharya · Feb 9, 2026
Data network effects were mostly a fake moat historically, but proprietary AND live data is now a genuinely powerful moat
'Data network effect' used to be what you said when you couldn't name a real moat; today a commodity model sitting on proprietary live data outperforms the most cutting-edge model without access to it
Scope: 'live' means continuously changing data such as health data or telemetry from a running product; open question how proprietary such data can remain
25:52 20VC: Is SaaS Dead in a World of AI | Do Margins Matter Anymore | Is Triple, Triple, Double, Double Dead Today? | Who Wins the Dev Market: Cursor or Claude Code | Why We Are Not in an AI Bubble with Anish Acharya @ a16z
Data secrecy implies real differences between labs
Douwe Kiela · Jun 30, 2023 · hedged
Part of the secret sauce making GPT-4 so good is that OpenAI went to enormous lengths to acquire special data nobody else has, such as using Whisper to transcribe the world's podcasts for high quality language data
Having high quality training data that no one else possesses puts you in a position of advantage
Scope: allegedly, on the Whisper transcription project
13:53 20VC: Why Data Size Matters More Than Model Size, Why The Google Employee Was Wrong; OpenAI and Google Have the Advantage & Why Open Source is Not Going to Win with Douwe Kiela, Co-Founder @ Contextual AI
Harry Stebbings · Sep 27, 2023 · hedged
The major labs may not actually be trained on the same data, since engineers who happily discuss their models refuse to discuss underlying data
Their secrecy specifically about training data suggests meaningful differences
60:10 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng
Data not algorithms is the scarce input so providers capture huge value
Nick Frosst · Sep 1, 2025
Algorithms are not the bottleneck to making models more useful; data quality is
Algorithms have changed only a little (base training plus RLHF and other RL techniques), whereas the constraint is getting good quality real data and generating good synthetic data from it
9:35 20VC: Cohere Founder on How Cohere Compete with OpenAI and Anthropic $BNs | Why Counties Should Fund Their Own Models & the Need for Model Sovereignty | How Sam Altman Has Done a Disservice to AI with Nick Frosst
Anastasios Angelopoulos · Aug 3, 2026
Leading data providers will easily be worth hundreds of billions of dollars, and possibly more
Data is the hardest part of model training — it must be sourced, it is dirty, and nobody wants to do it — while algorithms have become somewhat commoditised now that everyone knows how to use the transformer
43:02 20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena
Also on the record
Jeff Seibert · Nov 22, 2023
Acquiring high-quality clean training data is extremely challenging and genuinely valuable as a defense
Reddit, Twitter and others are shutting off APIs and tightening rate limits — a reversal of twenty years of opening up — showing there's been a clear realization data is valuable
36:03 Clean proprietary training data is scarce and valuable as a competitive defense
Mike Krieger · Mar 3, 2025
For labs, the value lies in team quality and in whether the model can reliably perform real-world actions in deployment, not primarily in proprietary data
Usefulness in actual use cases is the highest order bit; evals are useful for hill climbing and internal research but don't tell you whether a model will be excellent at what it's deployed for, or only in narrow situations
31:46 Team quality and deployment reliability not data is the real differentiator
Nick Frosst · Sep 1, 2025
Data is still a bottleneck despite synthetic data
You need real-world high-quality data to seed the synthetic data process; Cohere still uses in-house annotators making real data
9:05 Synthetic data still needs real world seed data so data remains scarce
Your assistant can query this graph directly — 9 positions here, 19,646 across the corpus. Add 996.fm over MCP.