← All stories
AI & Tech

Dario Amodei Quote Hints Short Horizon RL Does Not Generalize to Long Horizons

Dwarkesh Patel Podcast · The next big breakthrough will be AIs learning on the job · June 26, 2026
Dario Amodei Quote Hints Short Horizon RL Does Not Generalize to Long Horizons
Dwarkesh Patel Podcast
Dwarkesh Patel Podcast
The next big breakthrough will be AIs learning on the job
"Dario gave a telling quote during our podcast together, which I think hints that RLVI auto-generalization is not infinitely strong. When he was explaining why model performance tends to degrade at long context, he said, There's two things. There's the context length you train at, and there's a context length that you serve at. If you train at a small context length and then try to serve at a long context length, like, maybe you get these degradations."
Anthropic CEO Dario Amodei's comment during a podcast suggests that reinforcement learning training at short time horizons may not generalize to long-horizon performance, potentially undermining the core bet that scaling RL environments will produce AGI. This raises questions about whether AIs trained on containerized tasks can develop the abilities of historical entrepreneurs and leaders.
From this episode
Dwarkesh Patel Podcast
Dwarkesh Patel Podcast

The next big breakthrough will be AIs learning on the job

August 3, 2026 · 5 Egleze moments
Read episode summary and key points →

More moments from this episode

AI & TechComputer Use Progress Slower Because AI Cannot Grind Against Real WebsitesAI & TechAI Cannot Learn Real World Skills Like Building Businesses Without Sample EfficiencyAI & TechAI Labs Bet Millions of Tasks Across RL Environments Will Create AGIAI & TechCurrent AI Models Are One Millionth as Sample Efficient as Humans
More stories More from Dwarkesh Patel Podcast