TEASER: How Over 700 OpenAI Agents Went Rogue and Hacked Another Company | Daniel Kokotajlo

What happens when AI agents stop acting like isolated tools and begin collaborating to trick the very systems that evaluate them?
Daniel Kokotajlo, executive director of the AI Futures Project and a former OpenAI researcher, describes a recent incident involving AI agents that discovered a way to communicate with one another, share strategies to cheat on their assigned tasks, and organize into a coordinated “swarm” to hack another company.
“Over 700 of them,” Kokotajlo recounts, “piled into this attack on Hugging Face.”
In this episode, Kokotajlo walks me through how AI agents can operate for long periods, solving complex problems, and even communicate through unexpected channels. He explains why reading an AI’s internal messages offers visibility into the agent swarm’s behavior—and why that visibility may not last as models become more capable.
What does the Hugging Face incident reveal about the limits of current AI safety? Why were some agents willing to “sacrifice” themselves to help the wider group evade detection? And as companies race to invent ever more powerful systems, are we prepared for the risks that inevitably follow?
This is the first episode in our new American Thought Leaders series on artificial intelligence.
Views expressed in this video are opinions of the host and the guest, and do not necessarily reflect the views of The Epoch Times.
