Dwarkesh Patel Podcast
Episode overview

The next big breakthrough will be AIs learning on the job

Dwarkesh Patel Podcast · 5 Egleze moments
The next big breakthrough will be AIs learning on the job
Episode summary

In this monologue episode, AI researcher and podcast host Dwarkesh Patel examines the fundamental strategic bet major AI labs are making: that training models on millions of verifiable tasks across thousands of reinforcement learning environments will create artificial general intelligence. Patel reveals that current models are one-millionth as sample efficient as humans during training, though labs argue this inefficiency is a one-time cost amortized across billions of deployment sessions. He identifies an underrated bottleneck in AI progress: computer use capabilities lag because training requires replayable simulators, and companies like Amazon block bot training on real websites, forcing labs to build labor-intensive application clones. Patel argues that critical real-world skills like building businesses, winning elections, or succeeding in markets cannot be trained through current RL methods because they require months of real-world interaction that cannot be simulated in data centers. He cites a revealing quote from Anthropic CEO Dario Amodei suggesting short-horizon RL training may not generalize to long-horizon performance, potentially undermining the core AGI scaling hypothesis. The episode explores why continual learning and sample efficiency are deeply connected problems, discussing architectural innovations and alternative training methods like on-policy self-distillation and speculative "dreaming" approaches where AIs build and train against self-generated simulations. Patel concludes with a 2027-2028 scenario where deployed AIs learn primarily from real-world interactions across users rather than pre-deployment training, fundamentally changing how AI capabilities improve.

Key points
Watch original episode More from Dwarkesh Patel Podcast

5 moments from this episode

Source-linked · editorially selected
01
AI & Tech

Computer Use Progress Slower Because AI Cannot Grind Against Real Websites

The speaker reveals that progress on AI computer use capabilities lags behind other domains because training requires replayable simulators, and companies like Amazon will block bot training on real websites. This forces labs to build labor-intensive clones of applications, highlighting an underrated bottleneck in AI development that won't be solved until AIs can build high-fidelity application clones themselves.

Read this moment →
02
AI & Tech

AI Cannot Learn Real World Skills Like Building Businesses Without Sample Efficiency

The speaker argues that critical real-world skills like entrepreneurship, litigation, trading, and political strategy cannot be trained through current reinforcement learning methods because they require months or years of real-world interaction that cannot be simulated in parallel rollouts. This represents a fundamental limitation in the path to AGI that scaling compute alone cannot solve.

Read this moment →
03
AI & Tech

AI Labs Bet Millions of Tasks Across RL Environments Will Create AGI

Major AI research labs are making a central bet that scaling reinforcement learning across massive numbers of verifiable tasks will produce artificial general intelligence. This represents a fundamental strategic direction for companies investing billions in AI development, with the belief that current limitations in data efficiency and continual learning can be overcome through sheer compute scaling.

Read this moment →
04
AI & Tech

Current AI Models Are One Millionth as Sample Efficient as Humans

AI models require a million times more training samples than humans to learn the same tasks, according to the speaker's analysis. While AI labs argue this inefficiency is a one-time cost amortized across deployment, it reveals fundamental limitations in how current models learn and suggests barriers to achieving human-like learning capabilities.

Read this moment →
05
AI & Tech

Dario Amodei Quote Hints Short Horizon RL Does Not Generalize to Long Horizons

Anthropic CEO Dario Amodei's comment during a podcast suggests that reinforcement learning training at short time horizons may not generalize to long-horizon performance, potentially undermining the core bet that scaling RL environments will produce AGI. This raises questions about whether AIs trained on containerized tasks can develop the abilities of historical entrepreneurs and leaders.

Read this moment →