Skip to content

Debates

Who is best positioned to capture the value of proprietary enterprise data: startups, private data-rich companies, or large incumbents?

12 recorded positions from 11 people, first said May 17, 2023. They do not agree — the readings below are what each one actually argued.

Private data rich companies without compliance drag win

Harry Stebbings · Sep 27, 2023

The best position in AI today is the data-ownership 'no man's land': private companies with years of proprietary data but without incumbent compliance drag and public-company slowness

Startups have no data; incumbents have data but huge compliance burden and slowness

60:50 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng

Christian Lanng · Sep 27, 2023

Access to proprietary enterprise data requires paying a tax of compliance, regulatory and security capability, which giants like Google cannot pay because their lawyers will say no

He spent years convincing the largest companies to put supply chain data in the cloud; that capability is the entry price, and very large companies have too many lawyers to clear it

61:21 20VC: "How Being a Founder Almost Killed Me"; We Have Lied to a Generation of Founders | The Hardest Truths About Being a Founder Revealed | Why AI Co-Pilot is BS, Seat Pricing is Over & User Interfaces are Stupid with Christian Lanng

Harry Stebbings · Oct 27, 2023

The companies best positioned to win in AI are high-growth pre-IPO companies like Canva, which have the data without the regulatory constraints of being public

They hold proprietary data but avoid the incumbent downsides that come with public-company status

28:27 20VC: The Three Types of Seed Round Today, Why Seed Has Never Been More Competitive, Why Pricing Has Never Been Higher, Why Boards at Pre-Seed Can Be Helpful & How Too Much Cash Too Soon Can Harm Companies with Ed Sim, Founder @ Boldstart

Incumbents must convert their data corpus into self improving products

Ryan Petersen · Nov 13, 2023

Flexport is unusually well positioned to implement AI in global trade because it combines incumbent-scale structured data with a modern technology stack and team

It is the third largest international forwarding provider in the US, so it has incumbent-level data, but unlike incumbents it has the tech stack and team to actually implement

51:41 20VC: Flexport's Ryan Petersen: Reflections on Leadership from 13 Years Leading Flexport, Why Velocity not Speed is Most Important in Company Building, How Money Creates Inefficiencies in Scaling, The Future of Trade with China & Why Remote Work is so Cha

Anastasios Angelopoulos · Aug 3, 2026

Companies building self-improving products on their own proprietary data is inevitable, because in the age of AI only network effects and data moats remain as moats.

Software itself will no longer be a moat within about five years because it can be produced instantaneously, so a firm like Coca-Cola or Cisco must convert its data corpus into a self-improving product to stave off competitors.

Scope: projecting roughly five years out for software ceasing to be a moat

9:23 20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena

Enterprise frontier data stays proprietary and must be self mined

Emad Mostaque · May 17, 2023

The private data held inside institutions is far more valuable than the data sent to proprietary models, so bringing open, interpretable models to that private data is a durable business rather than a race to the bottom

Models can be taken to the data via partners and system integrators, standardizing complexity into building blocks that don't require continuous innovation, only data and distribution

22:39 20VC: Why the AI Bubble Will Be Bigger Than The Dot Com Bubble, Why AI Will Have a Bigger Impact Than COVID, Why No Models Used Today Will Be Used in a Year, Why All Models are Biased and How AI Kills Traditional Media with Emad Mostaque, Founder & CEO @

Alexandr Wang · Jun 12, 2024

Enterprise frontier data will not be open-sourced; each enterprise must mine and refine its own proprietary data to solve its own problems

10:26 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's

First deployer of fine tuned workflow models starts the data flywheel

Jake Saper · Mar 10, 2025

Domain-specific models trained on outcome data from a network of users will produce insights that even OpenAI cannot have.

Emergence's 2017 'coaching networks' thesis: the AI observes what action was taken and what outcome followed, then improves recommendations for everyone else in that network — data a general model has no access to.

Scope: 'coaching networks' was poorly branded; 'copilot' became the term that took off

46:00 20VC: Lessons from Investing $2BN and Returning $8BN in Cash | Why Most Venture Partnerships are Broken | We Sold Salesforce Early and Lost Out on Billions | Are The Best Deals Always Expensive and Competitive with Jake Saper @ Emergence Capital

Jonathan Siddharth · Dec 1, 2025

The enterprise data feedback loop is still wide open — models have touched reality in consumer but not in enterprise, and whoever deploys custom fine-tuned models for specific workflows first starts the flywheel

Deploying first reveals where the models fail, and that failure data can be used to generate additional training data to plug the gap; deployment is the only route to improvement

28:21 20VC: Scale, Surge, Turing, Mercor: Who Wins & Who Loses in Data Labelling | Is Revenue in Data Labelling Real or GMV? | Why 99% of Knowledge Work Will Go and What Happens Then? | Why SaaS is Dead in a World of AI with Jonathan Siddharth @ Turing

Also on the record

Jerry Murdock · Aug 22, 2026

Enterprise data is valuable through highly unique, company-specific use cases rather than as a static commodity, and only if the system lets the data and its utilization keep evolving

Data isn't static — it gives context and memory, and two companies in the same industry will utilize the same kind of data differently

50:42 Value is in evolving company specific utilization

Christian Kleinerman · Sep 22, 2023

Value in the next decade of AI will accrue mostly to incumbents that own the data rather than to startups

Data is what powers outcomes, so companies like Google with user data or Snowflake with private enterprise data are best positioned

30:30 Incumbents that own proprietary data capture most of the next decades ai value

Eiso Kant · Oct 7, 2024

GitHub's dataset confers no inherent capabilities advantage in the AI race, because private code cannot be trained on by anyone and all players therefore have access to the same public code

Private repos are off-limits to Poolside, OpenAI and everyone else, and public code is merely output data that is equally available

24:13 Github ownership confers no advantage since private code is untrainable and public code is equally available

Your assistant can query this graph directly — 12 positions here, 19,646 across the corpus. Add 996.fm over MCP.