Skip to content

Debates

Is reinforcement learning a viable path to general intelligence, or is its inefficiency a fundamental blocker?

8 recorded positions from 5 people, first said Jun 24, 2024. They do not agree — the readings below are what each one actually argued.

Rl plus compute reliably surpasses any benchmarkable task

David Luan · Jun 24, 2024 · hedged

Reasoning will be solved by training a base model, giving it access to a wide range of environments in which to attempt hard problems, and combining its self-generated attempts with human feedback on whether it did well.

Scope: framed as his and the field's strong suspicion, not a demonstrated result

21:05 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept

Victor Lazarte · Apr 14, 2025

Pure reinforcement learning is the most interesting thing happening in AI right now: if you can build a benchmark for a task, RL plus lots of compute can produce a model that surpasses it

Experts write questions models can't answer plus grading rubrics, the model generates many answers, and the graded pairs train a better model

30:01 20VC: Benchmark's Victor Lazarte on Why Portfolio Construction is BS| Why SaaS Spreadsheet Investing is Dead | Why China is a Stabilising Force for the US | Three Traits All the Best Founders Have & The Lie All Big Tech Companies Have Been Telling

Also on the record

Harry Stebbings · Nov 3, 2025

Karpathy holds that reinforcement learning is terrible, just less terrible than it was twenty years ago.

4:32 Rl is marginally improved but still fundamentally bad

Joelle Pineau · Nov 3, 2025

Expecting RL out of the box to deliver AGI is getting ahead of reality; RL is terribly inefficient and the learning-efficiency problem must be solved first.

The amount of signal needed to shape a model's behavior is far beyond what current methods can deliver.

4:48 Learning efficiency must be solved before rl can deliver agi

Joelle Pineau · Nov 3, 2025

RL is inefficient for two structural reasons: errors compound across sequential decisions, and learning requires taking actions rather than static data, which demands expensive simulators and synthetic data.

Sequential decision making means every wrong branch compounds over the action sequence, making the right solution a needle in a haystack; and you cannot get the right policy from static data, so you need environments to test in, which are scarce and costly.

5:42 Compounding errors and environment requirements make rl structurally inefficient

Joelle Pineau · Nov 3, 2025

RL costs are coming down and progress is fast in domains where the reward function can be written down precisely, but shaping model behavior in social contexts remains unsolved because no one knows how to express that mathematically.

AlphaGo showed RL can beat world champions once the goal is clearly specified, which is why math, well-defined reasoning tasks and games are progressing; social behavior has no writable reward function.

7:04 Rl works where the reward function is writable not in social domains

David Luan · Jun 24, 2024

Pure model scaling will not deliver reasoning; reasoning requires training models in environments where they can try composing what they know, and that still needs new research.

Reasoning means composing existing thoughts to discover a new thought, which is not instilled by asking a model to regurgitate the internet's worth of data; a human mathematician works by trying compositions of known truths in a proving environment.

18:53 Reasoning requires environment based composition training not pure scaling

Nick Frosst · Sep 1, 2025

Reinforcement learning from human feedback is far more data-efficient than he believed — a small dataset of human feedback can make a model meaningfully better

He argued the opposite around 2020 and it proved to be a technological misstep

62:50 Rlhf is far more data efficient than previously believed

Your assistant can query this graph directly — 8 positions here, 19,646 across the corpus. Add 996.fm over MCP.