Is reinforcement learning a viable path to general intelligence, or is its inefficiency a fundamental blocker?
8 recorded positions from 5 people, first said Jun 24, 2024. They do not agree — the readings below are what each one actually argued.
Rl plus compute reliably surpasses any benchmarkable task
David Luan · Jun 24, 2024 · hedged
Reasoning will be solved by training a base model, giving it access to a wide range of environments in which to attempt hard problems, and combining its self-generated attempts with human feedback on whether it did well.
Scope: framed as his and the field's strong suspicion, not a demonstrated result
21:05 20VC: Why Foundation Model Performance is Not Diminishing But Models Are Commoditising, Why Nvidia Will Enter the Model Space and Models Will Enter the Chip Space & The Right Business Model for AI Software with David Luan, Co-Founder @ Adept
Victor Lazarte · Apr 14, 2025
Pure reinforcement learning is the most interesting thing happening in AI right now: if you can build a benchmark for a task, RL plus lots of compute can produce a model that surpasses it
Experts write questions models can't answer plus grading rubrics, the model generates many answers, and the graded pairs train a better model
30:01 20VC: Benchmark's Victor Lazarte on Why Portfolio Construction is BS| Why SaaS Spreadsheet Investing is Dead | Why China is a Stabilising Force for the US | Three Traits All the Best Founders Have & The Lie All Big Tech Companies Have Been Telling
Also on the record
Harry Stebbings · Nov 3, 2025
Karpathy holds that reinforcement learning is terrible, just less terrible than it was twenty years ago.
4:32 Rl is marginally improved but still fundamentally bad
Joelle Pineau · Nov 3, 2025
Expecting RL out of the box to deliver AGI is getting ahead of reality; RL is terribly inefficient and the learning-efficiency problem must be solved first.
The amount of signal needed to shape a model's behavior is far beyond what current methods can deliver.
4:48 Learning efficiency must be solved before rl can deliver agi
Joelle Pineau · Nov 3, 2025
RL is inefficient for two structural reasons: errors compound across sequential decisions, and learning requires taking actions rather than static data, which demands expensive simulators and synthetic data.
Sequential decision making means every wrong branch compounds over the action sequence, making the right solution a needle in a haystack; and you cannot get the right policy from static data, so you need environments to test in, which are scarce and costly.
5:42 Compounding errors and environment requirements make rl structurally inefficient
Joelle Pineau · Nov 3, 2025
RL costs are coming down and progress is fast in domains where the reward function can be written down precisely, but shaping model behavior in social contexts remains unsolved because no one knows how to express that mathematically.
AlphaGo showed RL can beat world champions once the goal is clearly specified, which is why math, well-defined reasoning tasks and games are progressing; social behavior has no writable reward function.
7:04 Rl works where the reward function is writable not in social domains
David Luan · Jun 24, 2024
Pure model scaling will not deliver reasoning; reasoning requires training models in environments where they can try composing what they know, and that still needs new research.
Reasoning means composing existing thoughts to discover a new thought, which is not instilled by asking a model to regurgitate the internet's worth of data; a human mathematician works by trying compositions of known truths in a proving environment.
18:53 Reasoning requires environment based composition training not pure scaling
Nick Frosst · Sep 1, 2025
Reinforcement learning from human feedback is far more data-efficient than he believed — a small dataset of human feedback can make a model meaningfully better
He argued the opposite around 2020 and it proved to be a technological misstep
62:50 Rlhf is far more data efficient than previously believed
Your assistant can query this graph directly — 8 positions here, 19,646 across the corpus. Add 996.fm over MCP.