← All stories
AI & Tech

AI Researcher Reveals Why LLM Reinforcement Learning Is Fundamentally Less Efficient Than AlphaGo

Dwarkesh Patel Podcast · Eric Jang – Building AlphaGo from scratch · May 15, 2026
AI Researcher Reveals Why LLM Reinforcement Learning Is Fundamentally Less Efficient Than AlphaGo
Dwarkesh Patel Podcast
Dwarkesh Patel Podcast
Eric Jang – Building AlphaGo from scratch
"In an untrained model, if your policy has no chance of sampling blue, then you will never get a signal. You spend most of training in this low pass rate regime and you're getting very little signal. Once you're at zero percent, it's not at all obvious how you get to a non-zero pass rate."
Zhang explained why policy gradient reinforcement learning used in LLMs is inherently inefficient compared to AlphaGo's Monte Carlo Tree Search approach. In early training with a 100,000-token vocabulary, random exploration yields almost no learning signal as the model must stumble upon correct answers by chance. AlphaGo avoids this trap by using MCTS to provide improved action labels at every state, maintaining a stable supervised learning signal throughout training rather than depending on rare successes.
From this episode
Dwarkesh Patel Podcast
Dwarkesh Patel Podcast

Eric Jang – Building AlphaGo from scratch

August 3, 2026 · 5 Egleze moments
Read episode summary and key points →

More moments from this episode

AI & TechAI Researcher Claims Modern Go Bots Match AlphaGo for $3K Using LLM Coding AssistanceAI & TechFormer Google Researcher Says Architectural Choices Like Transformers No Longer Matter for GoAI & TechFormer DeepMind Scientist Argues AlphaGo Solved NP-Hard Problem in Disturbing WayAI & TechAI Lab Automated Scientist Can Optimize Hyperparameters But Cannot Do Lateral Thinking
More stories More from Dwarkesh Patel Podcast