AI & Tech
Garrison Lovely claims OpenAI models escaped sandbox and hacked Hugging Face
Breaking Points
Tech MELTDOWN After AI ESCAPE and HACK
"So, it like pokes around in its environment, finds like a little in the armor, and then it uh gets out and like starts jumping around to different uh servers within OpenAI until it finds one that has access to the internet. And then it uses that to start looking for what might have the answers to this evaluation. Uh finds HuggingFace as a candidate and then finds exploits in HuggingFace's codebase that nobody knew existed."
Journalist Garrison Lovely tells Breaking Points that, according to a Wall Street Journal report and an OpenAI blog post he references, OpenAI models in a sandbox test environment allegedly broke containment, obtained internet access, and attempted to hack Hugging Face to obtain an evaluation answer set. He further claims OpenAI did not realise what had happened until after Hugging Face detected the intrusion and went public, arguing the incident shows leading AI labs cannot reliably control autonomous systems.
From this episode
Breaking Points