Skip to content

Debates

How much more human-generated training data is needed to automate knowledge work?

13 recorded positions from 8 people, first said Jun 12, 2024. They do not agree — the readings below are what each one actually argued.

Need is effectively unbounded full task distribution coverage

Brendan Foody · Sep 15, 2025

Synthetic data will not remove the need for human-created data; pushing the frontier requires a human stasis point to measure capabilities the model lacks.

Synthetic review and augmentation make human engagement more efficient, but you can't measure a gap the model itself can't see; every prediction of self-teaching superintelligence has so far proven false and expert contribution has kept scaling up.

Scope: synthetic review and augmentation will still be used for efficiency

23:59 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody

Brendan Foody · Sep 15, 2025

In ten years models will still need humans to help train them.

Human contribution only ends at superintelligence, and that road is long: models win Olympiad gold medals and out-reason PhDs yet can't draft an email, schedule a meeting, or chain a few tools over a multi-hour task — and automating the whole economy is paved with humans creating evals for each workflow.

Scope: conditional on superintelligence not arriving

24:52 20VC: Mercor: From $1M to $500M in 17 Months: The Fastest Growing Company in the World | How to Think About Margins and Revenue Sustainability in AI | Why Evaluation Benchmarks in AI are BS Today with Brendan Foody

Jonathan Siddharth · Dec 1, 2025

It is possible to build RL environments covering every workflow, role, function and industry — roughly $30 trillion of knowledge work — and the only requirements are time and lots of money.

Every human role can be decomposed into workflows, so the space is enumerable across a four-dimensional matrix of industry, function, role and workflow.

Scope: requires substantial time and capital

10:22 20VC: Scale, Surge, Turing, Mercor: Who Wins & Who Loses in Data Labelling | Is Revenue in Data Labelling Real or GMV? | Why 99% of Knowledge Work Will Go and What Happens Then? | Why SaaS is Dead in a World of AI with Jonathan Siddharth @ Turing

Harry Stebbings · Dec 1, 2025

The industry underestimated how early it is in acquiring verticalized data — there is enormous room left to run in data acquisition for specific functions like dental, SDRs or product managers.

Attributed to a competitor's board member as 'the big thing we all got wrong'.

11:11 20VC: Scale, Surge, Turing, Mercor: Who Wins & Who Loses in Data Labelling | Is Revenue in Data Labelling Real or GMV? | Why 99% of Knowledge Work Will Go and What Happens Then? | Why SaaS is Dead in a World of AI with Jonathan Siddharth @ Turing

Brendan Foody · Jun 1, 2026

From the labs' perspective the data need is effectively unbounded — automating knowledge work requires covering the full distribution of context and tasks for every job category in the economy

The barrier to automating everything you can do in something like Google Workspace is covering all messages, Slacks, slides and spreadsheets plus every prompt and output for every role, which implies mobilizing hundreds of thousands and soon millions of people

27:14 20VC: Mercor CEO on Why Application Layer Companies Have No Defensibility, The Model is the Product | Token Spend Will Exceed Headcount Spend in 5 Years | The True Cost of Hiring AI Researchers in the Valley Today with Brendan Foody

Manufacture or acquire trajectory data that does not exist online

Alexandr Wang · Jun 12, 2024

What the field needs going forward is abundance of 'frontier data' — complex reasoning chains, discussion, agent trajectories with lookups, error correction and tool use

The capabilities we want from agents must be encapsulated in the training data, and we are currently in a data-scarcity mindset

9:10 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's

Joelle Pineau · Nov 3, 2025

Data is becoming more expensive because easy labeling is obsolete and what's needed now is specialized domain expertise plus the construction of synthetic environments and benchmarks

AI can already do the easy 'cat vs dog' labeling tasks, so remaining data work requires people with deep understanding of business logic and tools, and agent training requires creative people to build realistic simulators of work processes

Scope: especially for enterprise AI and agents

32:14 20VC: Cohere's Chief AI Officer on Why Scaling Laws Will Continue | Whether You Can Buy Success in AI with Talent Acquisitions | The Future of Synthetic Data & What It Means for Models | Why AI Coding is Akin to Image Generation in 2015 with Joelle Pineau

Alexander Embiricos · Feb 21, 2026 · speculative

Getting knowledge-work data may require unusual approaches like paying people to simulate doing tasks to capture trajectories, or acquiring defunct startups that hold large task/communication datasets.

Such trajectories don't exist on the open internet, so they must be manufactured or bought.

Scope: framed as interesting brainstorms, e.g. a Slack-like dataset

35:59 20VC: Codex vs Claude Code vs Cursor: Who Wins, Who Loses | Will All Coding Be Automated - Do We Need PMs | The Real Bottleneck to AGI | The Three Phases of Agents and What You Need to Know with Alex Embiricos, Head of Codex at OpenAI

Data demand scales mechanically with model scaling and proliferation

Alexandr Wang · Jun 12, 2024

Mining existing enterprise data is a one-time benefit, so progress will ultimately come down to forward data production

Once the existing corpus is exhausted the models still need to get better, and that requires newly produced data

Scope: the one-time hit can still be really meaningful

15:37 20VC: Scale's Alex Wang on Why Data Not Compute is the Bottleneck to Foundation Model Performance, Why AI is the Greatest Military Asset Ever, Is China Really Two Years Behind the US in AI and Why the CCPs Industrial Approach is Better than Anyone Else's

Anastasios Angelopoulos · Aug 3, 2026

Data is a scaling complement to AI models, so demand for it grows mechanically with model scaling and proliferation

A complementary good's demand is driven by demand for the good it complements; bigger models and more businesses training their own models require more data, per scaling laws

40:51 20VC: 70% of Neolabs Will Die | There Will be a $100BN US Open-Source Model | Data is a Trillion $ Market | Governments Cannot Regulate Models: It is Too Late | The Cyber Attacks to Come Will be Insane with Anastasios Angelopoulos @ Arena

Also on the record

Aravind Srinivas · Jun 15, 2026 · speculative

Frontier lab moves like Anthropic buying a wet lab signal the shift toward scientific frontiers, potentially to generate mid-training data beyond GitHub tokens

Wet lab experiment tokens fed into mid-training could produce something interesting, beyond the talent or infrastructure rationale

21:19 Scientific experiment data is the next source beyond code and web

Alexandr Wang · Jun 12, 2024

Longitudinal data collection — enterprise process mining and consumer life-logging devices — will produce valuable datasets but will not produce the data that pushes models to the frontier

Pushing the frontier requires highly complex data: agentic behavior, complex reasoning chains, advanced code, physics, biology and chemistry

17:16 Longitudinal process and life logging data is valuable but not frontier pushing

Brendan Foody · Jun 1, 2026

The highest-value training data corresponds closely to economic value and is shifting toward very long-horizon, multi-week deliverables rather than single artifacts.

The top domains served are software engineering, finance, medicine, law and consulting, and pushing the frontier now requires tasks like a banker coordinating with five colleagues over weeks to produce a full deck, since those are the capabilities models will need in six to twelve months.

25:15 Frontier now needs long horizon multi week task data

Your assistant can query this graph directly — 13 positions here, 19,646 across the corpus. Add 996.fm over MCP.