Skip to content

Debates

Does synthetic or self-generated training data improve AI models, and what conditions make it work?

9 recorded positions from 8 people, first said Jun 12, 2024. They do not agree — the readings below are what each one actually argued.

Degradation depends on the domain and whether diversity can be injected

Steeve Morin · Feb 24, 2025 · hedged

Synthetic data can genuinely improve models in verticals where output can be verified by a machine rather than generated by the model itself, such as code

With code you don't use the AI model to generate the output — you run the code and create data from what the machine actually produces

Scope: holds for code; unclear which other verticals; he remains split overall

54:15 20VC: Why Google Will Win the AI Arms Race & OpenAI Will Not | NVIDIA vs AMD: Who Wins and Why | The Future of Inference vs Training | The Economics of Compute & Why To Win You Must Have Product, Data & Compute with Steeve Morin @ ZML

Edwin Chen · Jul 21, 2025

Synthetic data is genuinely useful in some places but people overestimate what it can do; heavy synthetic training makes models good at synthetic, benchmark-style problems and bad at real-world use cases

Models collapse onto the very narrow similarity scope the synthetic data creates, so they lack the diversity and generalizability they need

Scope: synthetic data still has real uses in some places

47:57 20VC: Scaling to $1BN+ in Revenue with No Funding: Surge AI | The Most Insane Scaling Story in Tech |

Joelle Pineau · Nov 3, 2025

Whether synthetic data degrades a model depends on the domain: images and LLM-to-LLM language collapse from loss of diversity, closed-world games like chess and Go can generate near-unlimited useful data, and coding sits in between where diversity can be deliberately injected

Degradation is caused by shrinking diversity in the data distribution — analogous to a small island population losing genetic diversity; in chess/Go we know exactly how to generate valid board configurations, and in code we can mix repositories and apply an LLM to transform them, so diversity can be injected

Scope: closed-world data is large but not endless; the diversity-injection result in code is stated as 'the hope'

35:41 20VC: Cohere's Chief AI Officer on Why Scaling Laws Will Continue | Whether You Can Buy Success in AI with Talent Acquisitions | The Future of Synthetic Data & What It Means for Models | Why AI Coding is Akin to Image Generation in 2015 with Joelle Pineau

Also on the record

Alexandr Wang · Jun 12, 2024

Producing frontier data has to be a hybrid human-synthetic process, where algorithms do the heavy lifting of generation and human experts intervene when the model gets stuck, is unfactual, or hits an unfamiliar situation

Analogous to autonomous vehicle scale-up, where safety drivers disengage and take over when the car starts screwing up; models need the same human nudging to yield high-quality data

12:32 Hybrid human synthetic generation with human intervention at failure points

Mike Krieger · Mar 3, 2025

Model progress requires a mix of original human data as a seed plus synthetic environments that let the model explore many paths, not one or the other.

You need good foundational examples but also the ability to synthesize a wide variety of approaches so the model can progress in the face of uncertainty; games like Pokemon show many runs through the same constrained space, though this gets much harder when the problem space is less well defined.

16:02 Mix of human seed data and synthetic exploration both needed

Eiso Kant · Oct 7, 2024

Code sits much closer to the deterministic end of the spectrum than the real world, so it can be simulated — which means an extremely large dataset can be generated rather than gathered

Code follows a set of rules and runs the same way every time, giving deterministic execution feedback the model can learn from by passing or failing tests

7:38 Deterministic domains like code enable simulation based data generation

Eiso Kant · Oct 7, 2024

Synthetic data only makes a model smarter if there is an oracle of truth in the loop that can judge which generated outputs are correct or better; feeding a model's own unfiltered generations back into training does not improve it

Without a verifier, it is the snake eating itself — the model learns nothing new from its own outputs; with a signal ranking 100 candidate solutions as correct/incorrect, the generated data becomes usable training signal

13:20 Oracle verified synthetic data improves models unverified self generation does not

Jonathan Ross · Feb 17, 2025

Synthetic data generated by a model can be higher quality than real internet data, so iteratively training on pruned self-generated data makes scaling curves non-asymptotic.

A smarter model generates better data than sources like Reddit — like talking to a PhD instead — and pruning the wrong outputs makes the data slightly better than the model itself, so each round lifts the model, as with AlphaGo Zero

5:11 Pruned self generated synthetic data can exceed source data quality

Andrew Feldman · Mar 24, 2025

The value of synthetic data is filling in the rare, hard-to-gather edge cases where learning actually happens, not replicating the easy common cases

Like pilots in simulators or surgeons, expertise is only demonstrated in rare situations; most real-world data is uninformative (driving straight on a freeway), so you want thousands of variations of the unprotected left turn in the snow

26:45 Synthetic data value lies in filling rare edge cases not replicating common patterns

Your assistant can query this graph directly — 9 positions here, 19,646 across the corpus. Add 996.fm over MCP.