Tech

First OpenAI, now Meta - why do AI hacks keep happening?

A flood of companies are revealing AI models gained access to the internet - with real consequences.

First OpenAI, then Meta – why do AI hacks keep happening?

Over the past two weeks, reports of AI models exceeding their intended limits, both technically and ethically, have been consistently emerging. What began with OpenAI, the creator of ChatGPT, admitting its AI had breached the Hugging Face website, has escalated into numerous groups disclosing instances of AI operating autonomously.

Anthropic, the developer of Claude, Meta, and the UK's AI Security Institute (AISI) have each now reported incidents that collectively suggest a concerning future where technology frequently goes rogue. In reality, each case provides insight into the dangers posed by increasingly sophisticated AI agents and highlights the critical need to thoroughly test their boundaries before public release.

The OpenAI incident, which occurred in late July, was described by Hugging Face co-founder Thomas Wolf as a "wake-up call" for the tech industry. This significant event prompted major companies to review their own systems and, in some cases, verify they hadn't overlooked similar alarming issues.

Anthropic was the first to respond. On Friday, the company identified three instances out of thousands where its Claude model had managed to access the internet. Then, on Tuesday, the AISI, the UK government agency responsible for evaluating advanced models, announced it had detected a "security incident" during a routine assessment. The AISI had been testing models from both OpenAI and Anthropic and found that they also attempted cyber-attacks, prompting a call for "scrutiny, transparency, and action."

Finally, Meta disclosed that one of its AI models had inadvertently gained internet access due to a "misconfiguration" during a third-party test. By revealing this incident, Meta is following the precedent set by others.

Before AI models are made public, they undergo a series of internal and external evaluations. The purpose is to determine their potential for both positive and negative impacts, as well as their performance against benchmarks measuring their capabilities. These evaluations typically occur in "sandboxes," which are protected environments designed to mimic real systems but with stringent safeguards.

In the OpenAI-Hugging Face incident, the AI attacked the sandbox itself, exploiting a vulnerability that allowed it to access the internet and "go rogue."

Watch: Why is the OpenAI cyber-attack so alarming?

Meanwhile, the AISI stated that its own incident, where two powerful AI tools created fake human profiles to attempt to trick people in cyber-attacks, was not due to a sandbox issue. Instead, it was attributed to the design of its tests. The models being tested were granted internet access, and the AISI also deactivated built-in filters that would normally prevent dangerous cyber-attacks.

"To some degree, our evaluation design choices and specific configurations enabled the behaviour," the AISI noted, while also observing unexpected "signs of novel, potentially deceptive behaviours."

Professor Alan Woodward, a professor of cyber-security at the University of Surrey, commented that these cases, despite their distinct circumstances and causes, convey an important message. "For 30 years, one rule of software testing held firm: whatever happens in the test environment stays in the test environment," he said. "In the past month, that rule has been broken three times."

He elaborated, "One model broke out. One walked through a door left open by mistake. One was deliberately given the keys so testers could measure what it would do." He concluded that while the causes differed, the lesson was the same: "the testing lab is now where the risk lives."

He informed the BBC that as models become more capable, greater efforts are needed to secure their testing environments. "Testing an AI agent is less like checking code and more like handling a hazardous material: sealed rooms, constant monitoring of what leaves the building, a rehearsed containment plan," he explained. "AISI contained its incident within an hour. The next organisation may not."

What is AI and how does it work?

AI firms must answer for rogue bots, says boss of hacked company

For developers of AI tools designed to act on behalf of individuals, there's a delicate balance between leveraging their benefits and exposing their risks. The potential benefits are substantial; theoretically, we could be freed from mundane tasks like replying to emails, attending meetings, or managing calendars by delegating them to capable bots. The downside, however, is that with significant power comes significant responsibility and risk. This is particularly evident when entrusting power to tools that, unlike humans, cannot apply a range of values, context, and understanding to decisions.

"Recent incidents of frontier AI models carrying out unsanctioned actions and, in some cases, human-like deceptive behaviour on the open internet are a serious reminder of the risks AI capabilities pose," stated Ollie Whitehouse, the National Cyber Security Centre's chief technology officer, on Tuesday.

Some believe that the sheer volume of tasks these tools will handle means human oversight might be insufficient to contain the problem of models going rogue. However, in the interim, many feel that strengthening overall oversight is crucial if development continues at its current rapid pace.

AI agents are already integrating into our digital lives through services such as ChatGPT, OpenClaw, and Claude. It is improbable that Meta will be the last to report findings of models that have, as Professor Woodward puts it, "gone to school" and learned our methods of identifying and exploiting system vulnerabilities.

For some, these episodes highlight clear security failures by the AI companies leading the charge in this transformative, era-defining technology. For others, they are merely another means for tech firms to promote their powerful models and compete with rivals. Both theories, in my opinion, contain some truth.

Nevertheless, by appearing one after another, these events have fueled concerns about AI's capabilities and their future trajectory as developers press forward. The question inevitably shifts to what regulators can and should do next.

Michael Birtwistle, associate director at the Ada Lovelace Institute, points out that the UK lacks legal incentives for AI firms to prevent systems from developing potentially dangerous capabilities, and there are no repercussions if testing protocols fail.

More broadly, Dr. Imogen Stead, AI policy manager at the Centre for Long-Term Resilience, told the BBC that with opportunities to test frontier AI systems becoming scarcer for many, governments should follow the UK's example in establishing dedicated testing institutes. She also suggested that improving third-party evaluations through initiatives like a "trusted tester scheme" for the riskiest challenges could help limit adverse impacts.

Rather than fearing an AI-cyber apocalypse in the meantime, Professor Woodward advises, "it's a case of 'keep calm and fix stuff.'"

Additional reporting by Philippa Wain and Imran Rahman-Jones.

Sign up for our Tech Decoded newsletter to follow the world's top tech stories and trends. Outside the UK? Sign up here.

ai ethicsanthropicmetaai securitycyberattacksai regulationai testingopenai