Is messy enterprise data a real blocker to deploying AI inside companies?
13 recorded positions from 10 people, first said Nov 22, 2023. They do not agree — the readings below are what each one actually argued.
Fragmented legacy data estates make retrieval unreliable
Jonathan Siddharth · Dec 1, 2025
Real enterprise AI deployment is not just fine-tuning on proprietary data — it requires heavy first-mile and last-mile schlep: acquiring and structuring fragmented data, building evals, designing partial-autonomy workflows and retraining humans
In the real world enterprise data is a mess, siloed across spreadsheets and files owned by people who have left; the rosy fine-tuning picture doesn't hold
Scope: described via an insurance underwriting copilot example
30:10 20VC: Scale, Surge, Turing, Mercor: Who Wins & Who Loses in Data Labelling | Is Revenue in Data Labelling Real or GMV? | Why 99% of Knowledge Work Will Go and What Happens Then? | Why SaaS is Dead in a World of AI with Jonathan Siddharth @ Turing
Aaron Levie · Apr 20, 2026
Roughly half a large enterprise's data estate isn't technically ready for agents and the other half is too fragmented, so agents will as often retrieve the wrong document as the right one unless data is curated and targeted
Contracts sit across ten-plus systems, many legacy or on network file shares with low throughput, and two decades of employees bringing their own tools left no standardized system because humans could always find things manually
Scope: Fortune 500 context
30:10 20VC: Everyone is Wrong; We Will Have More Developers in Five Years | Why Frontier Labs Will Be Way More Valuable Than They Are Today | Are SaaS Companies Cooked: Which Thrive & Which Die with Aaron Levie, Founder at Box
Harry Stebbings · Jun 1, 2026 · hedged
Data structures and data cleanliness are one of the biggest problems facing enterprise AI adoption.
18:44 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody
Unstructured data readiness gap creates a large ai deployment services market
Harry Stebbings · Nov 22, 2023
AI implementation services will be one of the biggest categories of the next few years
31:02 20VC: Why OpenAI Will Become an Infrastructure Play, Why Apple Will Win in an AI World, Why Google is the Most Vulnerable Incumbent, Will LLMs Be Commoditised, Which Startups Are Thin vs Thick Wrappers on Top of LLMs with Jeff Seibert, Founder @ Digits
Jamin Ball · Jan 10, 2024
You don't have an AI strategy without a data strategy, so data businesses and data platforms will take off as companies realize they must get their data house in order before adopting AI
Most of the AI value proposition today depends on whether a company's data is in order
58:45 20VC: Did Figma Kill M&A Markets in 2024, The Three Biggest Mistakes Made in Growth Investing, The Three Requirements Companies Need to Go Public in 2024 with Ed Sim and Jamin Ball
Kieran Flanagan · Jul 11, 2025
There will be a huge influx of companies that help enterprises deploy AI, because deployment is genuinely complex and most companies' data — especially unstructured data — is not in a usable state for AI
Using only traditional structured B2B data guarantees poor results; enriching data and capturing unstructured data drastically improves outcomes, and that work is hard enough to sustain a services market
65:55 20Growth: The Death of Growth Teams? | How Hubspot Use AI to Triple Email Conversion | The Future of AI SEO | Why Prompt Engineering is the New Coding | What Every CMO Needs to Know About AI in 2025
Models can clean and structure the data themselves
Brendan Foody · Jun 1, 2026
Data cleanliness is only a partial barrier to enterprise AI adoption, because models will increasingly clean and structure the data themselves as reasoning capability improves.
A model that can read every Slack message from the last six months can itself structure a table of customer conversations and classify data, so humans won't be doing that work.
Scope: models still need access to the data
19:01 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody
Fred Turner · Jul 18, 2026
Data cleanliness is not a major inhibitor to enterprise AI adoption if you approach it correctly, because models can clean and structure the data given the right context
They migrated dashboards off Looker to Snowflake using an agentic workflow that ingested and restructured the data itself, turning a year-long multi-engineer project into a two-month, one-to-two-person job
Scope: requires the right approach and context
61:19 20VC: $5BN in Revenue, 7 to 7,000 Employees in 9 Months, 206,000 Tests in a Single Day: The Craziest Story in Startups: Curative with Fred Turner
Also on the record
Arvind Jain · Jul 11, 2026
Most enterprises are deploying AI wrongly by wiring models to all systems via rudimentary MCP connections and letting them brute-force context assembly, which makes AI slow and expensive.
Most tokens get burned just assembling the raw materials for a task rather than doing the work, and AI gets used for things it isn't good at.
24:51 Brute force context assembly is too slow and costly build the context layer
Aaron Levie · May 22, 2024
The biggest AI opportunity is automating the parts of the business that were never automated before, such as querying unstructured repositories of marketing assets and contracts
Most companies' digital assets and contracts sit in folders that can't be queried; AI lets you ask natural-language questions of that data, which was previously impossible
36:14 Automating previously unautomatable unstructured data workflows is the biggest ai opportunity
Aaron Levie · May 22, 2024
A company's unstructured content is effectively its digital memory, and AI's ability to tap into it is a profound shift in how organizations work with information
Contracts, marketing assets, financial records and invoices are the closest thing a company has to memory; giving every company infinite digital memory means better decisions, more automated processes, and new employees who can instantly access institutional knowledge
52:29 Companys unstructured content is its digital memory unlocking it is a profound shift
Andrew Ng · Nov 17, 2025
Companies need far less data than they think to get started with AI, because most valuable data is private and verticalized and a scrappy team can extract value from internal transaction, sales, product, manufacturing and logistics data.
Internet data is largely general purpose; the world's valuable data is private and business-specific, and even PDFs and filings can be converted into usable structured data.
39:46 Small private verticalized datasets are enough to start
Andrew Feldman · May 26, 2026
Data organization is the second-order constraint on AI adoption, binding only once security and legal have signed off — giving organizations that spent decades disciplining their data a huge advantage
Once the rules are agreed there is huge productivity to be gained, but you are then immediately constrained by how you chose to marshal and organize data over years
35:32 Data discipline binds only after legal and security clear the way
Your assistant can query this graph directly — 13 positions here, 19,646 across the corpus. Add 996.fm over MCP.