This is the second part of a two-part series on how frontier AI is reshaping cyber risk.
Over recent weeks a wave of disclosures have lit up the technology industry about the prospect of “rogue AI”. OpenAI first revealed a frontier model had broken out of its sandbox during a test and hacked into Hugging Face’s network. Shortly after, Anthropic surfaced three similar incidents from its own test logs. Meta and the UK AI Security Institute admitted similar observations, while a Melbourne man made the news after claiming his AI agent hacked his gym’s booking system to move him up the waitlist by cancelling someone else’s reservation.
We discuss what this all means, including what the cyber security industry has made of these incidents and the lessons they are drawing about how to stay safe from autonomous AI.
We also deconstruct the language and framing being used to describe these events – with phrases like “cheating” and “going rogue” investing AI with a sense of agency. The episode closes with a view of how these incidents much shape the broader public policy debate about proprietary vs open-weight models .