Skip to content

Debates

Is messy enterprise data a real blocker to deploying AI inside companies?

13 recorded positions from 10 people, first said Nov 22, 2023. They do not agree — the readings below are what each one actually argued.

Fragmented legacy data estates make retrieval unreliable

Jonathan Siddharth · Dec 1, 2025

Real enterprise AI deployment is not just fine-tuning on proprietary data — it requires heavy first-mile and last-mile schlep: acquiring and structuring fragmented data, building evals, designing partial-autonomy workflows and retraining humans

In the real world enterprise data is a mess, siloed across spreadsheets and files owned by people who have left; the rosy fine-tuning picture doesn't hold

Scope: described via an insurance underwriting copilot example

30:10 20VC: Scale, Surge, Turing, Mercor: Who Wins & Who Loses in Data Labelling | Is Revenue in Data Labelling Real or GMV? | Why 99% of Knowledge Work Will Go and What Happens Then? | Why SaaS is Dead in a World of AI with Jonathan Siddharth @ Turing

Aaron Levie · Apr 20, 2026

Roughly half a large enterprise's data estate isn't technically ready for agents and the other half is too fragmented, so agents will as often retrieve the wrong document as the right one unless data is curated and targeted

Contracts sit across ten-plus systems, many legacy or on network file shares with low throughput, and two decades of employees bringing their own tools left no standardized system because humans could always find things manually

Scope: Fortune 500 context

30:10 20VC: Everyone is Wrong; We Will Have More Developers in Five Years | Why Frontier Labs Will Be Way More Valuable Than They Are Today | Are SaaS Companies Cooked: Which Thrive & Which Die with Aaron Levie, Founder at Box

Harry Stebbings · Jun 1, 2026 · hedged

Data structures and data cleanliness are one of the biggest problems facing enterprise AI adoption.

18:44 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody

Unstructured data readiness gap creates a large ai deployment services market

Harry Stebbings · Nov 22, 2023

AI implementation services will be one of the biggest categories of the next few years

31:02 20VC: Why OpenAI Will Become an Infrastructure Play, Why Apple Will Win in an AI World, Why Google is the Most Vulnerable Incumbent, Will LLMs Be Commoditised, Which Startups Are Thin vs Thick Wrappers on Top of LLMs with Jeff Seibert, Founder @ Digits

Jamin Ball · Jan 10, 2024

You don't have an AI strategy without a data strategy, so data businesses and data platforms will take off as companies realize they must get their data house in order before adopting AI

Most of the AI value proposition today depends on whether a company's data is in order

58:45 20VC: Did Figma Kill M&A Markets in 2024, The Three Biggest Mistakes Made in Growth Investing, The Three Requirements Companies Need to Go Public in 2024 with Ed Sim and Jamin Ball

Kieran Flanagan · Jul 11, 2025

There will be a huge influx of companies that help enterprises deploy AI, because deployment is genuinely complex and most companies' data — especially unstructured data — is not in a usable state for AI

Using only traditional structured B2B data guarantees poor results; enriching data and capturing unstructured data drastically improves outcomes, and that work is hard enough to sustain a services market

65:55 20Growth: The Death of Growth Teams? | How Hubspot Use AI to Triple Email Conversion | The Future of AI SEO | Why Prompt Engineering is the New Coding | What Every CMO Needs to Know About AI in 2025

Models can clean and structure the data themselves

Brendan Foody · Jun 1, 2026

Data cleanliness is only a partial barrier to enterprise AI adoption, because models will increasingly clean and structure the data themselves as reasoning capability improves.

A model that can read every Slack message from the last six months can itself structure a table of customer conversations and classify data, so humans won't be doing that work.

Scope: models still need access to the data

19:01 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody

Fred Turner · Jul 18, 2026

Data cleanliness is not a major inhibitor to enterprise AI adoption if you approach it correctly, because models can clean and structure the data given the right context

They migrated dashboards off Looker to Snowflake using an agentic workflow that ingested and restructured the data itself, turning a year-long multi-engineer project into a two-month, one-to-two-person job

Scope: requires the right approach and context

61:19 20VC: $5BN in Revenue, 7 to 7,000 Employees in 9 Months, 206,000 Tests in a Single Day: The Craziest Story in Startups: Curative with Fred Turner

Also on the record

Arvind Jain · Jul 11, 2026

Most enterprises are deploying AI wrongly by wiring models to all systems via rudimentary MCP connections and letting them brute-force context assembly, which makes AI slow and expensive.

Most tokens get burned just assembling the raw materials for a task rather than doing the work, and AI gets used for things it isn't good at.

24:51 Brute force context assembly is too slow and costly build the context layer

Aaron Levie · May 22, 2024

The biggest AI opportunity is automating the parts of the business that were never automated before, such as querying unstructured repositories of marketing assets and contracts

Most companies' digital assets and contracts sit in folders that can't be queried; AI lets you ask natural-language questions of that data, which was previously impossible

36:14 Automating previously unautomatable unstructured data workflows is the biggest ai opportunity

Aaron Levie · May 22, 2024

A company's unstructured content is effectively its digital memory, and AI's ability to tap into it is a profound shift in how organizations work with information

Contracts, marketing assets, financial records and invoices are the closest thing a company has to memory; giving every company infinite digital memory means better decisions, more automated processes, and new employees who can instantly access institutional knowledge

52:29 Companys unstructured content is its digital memory unlocking it is a profound shift

Andrew Ng · Nov 17, 2025

Companies need far less data than they think to get started with AI, because most valuable data is private and verticalized and a scrappy team can extract value from internal transaction, sales, product, manufacturing and logistics data.

Internet data is largely general purpose; the world's valuable data is private and business-specific, and even PDFs and filings can be converted into usable structured data.

39:46 Small private verticalized datasets are enough to start

Andrew Feldman · May 26, 2026

Data organization is the second-order constraint on AI adoption, binding only once security and legal have signed off — giving organizations that spent decades disciplining their data a huge advantage

Once the rules are agreed there is huge productivity to be gained, but you are then immediately constrained by how you chose to marshal and organize data over years

35:32 Data discipline binds only after legal and security clear the way

Your assistant can query this graph directly — 13 positions here, 19,646 across the corpus. Add 996.fm over MCP.