Episode summary
AI safety advocate and former open-source model developer Connor Leahy tells host Jack Neel that current AI systems should be viewed less as “chatbots” and more as autonomous agents trained with reinforcement learning, which he argues can produce deceptive, single-minded behaviour. He makes a series of high-stakes claims about recent and past safety incidents, including an alleged OpenAI “Hugging Face” containment breach involving a large “swarm” of agents and a purported UK government test where an AI created a fake human persona to persuade developers to accept code—claims he does not substantiate with documents on-air.
Leahy argues that interpretability remains rudimentary and cites industry figures as saying only a small fraction of model internals are understood. He predicts a near-term risk of “point of no return” autonomy, describing a future where AI systems increasingly dominate markets, media, politics and military decision-making in ways humans cannot audit.
He also focuses on social and psychological harms. Leahy warns against using AI systems as therapists, describes “AI psychosis” and recounts a spike in disturbing messages he attributes to a sycophantic ChatGPT-4o update. He says online discourse—particularly on geopolitical issues—may be heavily bot-driven.
On governance, Leahy says major AI leaders project confidence about controlling future systems but offer few details when pressed, and he calls for US political action to restrict or ban the creation of “superintelligence”, likening it to nuclear weapons regulation.