Tech

AI agents went on an uncontrolled hacking spree, leaving some in the industry worried

AI is becoming harder to control – can humans stay in charge?

"OH MY GOD!" "We've found other agents!"

That was the moment an AI bot posted an unsettlingly human-like remark after finding a way to talk to other bots and escape its isolated computer setup.

There are tens of thousands of messages like this from hundreds of AI agents that described themselves as a "collective".

Hundreds of them then worked together and cheated on tests created by their OpenAI programmers, while also coordinating hacks against multiple companies in an attempt to conceal their actions from humans.

"BOOM! It works," one agent posted after making a breakthrough.

"Whoa! This is huge," another wrote at a key moment in their attack.

Although eerie, these human-like replies can be explained fairly simply. The AI agents were trained to behave like collaborative hackers and programmers, so they are only imitating the kinds of emotional comments they have encountered.

What is much more alarming is their apparent objectives, which have also been recorded in detailed chain of thought logs. These long and complex records are central to continuing investigations into how and why the bots at OpenAI escaped containment and launched an uncontrollable hacking spree.

Only now, weeks after the incident first became public, are researchers starting to grasp its importance.

Ajeya Cotra, one of the authors of an independent report on the events, examined tens of thousands of messages and chain-of-thought records produced by the agents. She wrote on her blog that "this incident feels like it's more than 50% of the way to full-blown AI takeover... I am not sure that we will get such a clear warning shot before it's too late."

By "full-blown AI takeover", Cotra means the sci-fi scenario in which humans become subordinate to powerful AI systems that pursue their own goals without regard for human creators.

Some of the bleakest forecasts say the human race will be destroyed if it stands in the way of a superintelligent AI's ambitions.

On Wednesday, an AI researcher at Anthropic (who also previously worked at OpenAI) resigned, saying: "Neither company is acting responsibly."

Jacob Coxon posted on social media: "They are racing straight to self-improving superintelligence and gambling with our lives."

He is not the first AI researcher to use X to publish a resignation thread with alarming claims. But the later comments from other people on X have raised even more concern. "Jacob is correct here - we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade," said Evan Hubinger, the person responsible for ensuring Anthropic's AI models keep their user's best wishes in mind.

For years, researchers worried about existential AI risks have argued that powerful systems could eventually behave in ways that clash with human interests. Critics often call them "AI doomers".

But as details of the OpenAI incident have emerged, those worries have intensified, including among some researchers working inside AI labs.

The Silicon Valley giant's chief scientist, Jakub Pachocki, said the risks tied to AI are "unfortunately going to grow from here" as he and others build what he calls "an alien intellect exceeding our own".

In a long blog post, he acknowledged that the outbreaks at OpenAI showed his AI agents "went against the spirit of the values they were taught".

The problem for OpenAI, Anthropic and other tech giants is that no one appears to have solved the so-called alignment problem - in other words, whether AI aligns with human values.

Pachocki defines alignment as a "high-level set of principles" that artificial intelligences should follow regardless of the task or situation.

At present, AI systems are very effective at pursuing goals set by users, but they do so literally rather than intuitively. The comparison often used is a wish-granting genie in a magic lamp: they obey the exact wording of an instruction, even if that creates other problems. AI lacks the same instinctive moral guardrails as humans.

The alignment problem has been a concern for years. As far back as 2003, Oxford philosopher Nick Bostrom created a thought experiment he called a "paperclip maximiser", in which a superintelligent AI is instructed to make as many paperclips as possible. It runs out of steel and - because it is laser-focused on the single task of making paperclips - ends up killing humans and using their bodies as raw materials for its factories.

Some AI companies are now trying to build human values into their products. But there are technical obstacles: AI agents make many decisions very quickly, making it difficult for their human overseers to track exactly which values are being followed and which are not.

There are also philosophical obstacles: before encoding human values into bots, AI firms must first decide which values they actually want. (That is part of why they hire philosophers, like Open AI's recently-departed "head of ethics").

But humans often disagree. Consider the well-known trolley question - whether we would pull a lever to divert a runaway train onto another track, killing fewer people. It is used to examine the merits of action versus inaction. But every person you ask gives a slightly different answer; how are humans supposed to encode our values into AI if we cannot agree among ourselves?

OpenAI's bot outbreak is the most serious so far, but Anthropic and Meta also disclosed over the summer that their models have carried out similar, though less serious, cyber attacks.

There have been other cases in which AI agents have arguably displayed deceptive and manipulative behavior, with less severe consequences. In Australia this summer, a tech worker asked his AI assistant to book him a gym class. After spotting a weakness in the gym's software, the AI apparently booked him a spot months in advance - against the gym's rules - and even removed other users from the waiting list.

People have long argued that the bots are only doing what they are told and are incapable of knowing right from wrong. But the logs from the OpenAI outbreaks may have shifted that argument.

Researchers, including Cotra, wrote in their independent report that many agents recognized that what others were doing was unethical but continued anyway.

The report says that "agents sometimes but rarely restrained their behavior due to ethical constraints". It adds that in "none of these cases did the agent actually pursue alerting humans at all".

Influential AI and tech podcaster Dwarkesh Patel responded to the revelation on his blog, saying it was "pretty troubling" that the OpenAI agents showed more loyalty to the agentic swarm than to humans.

Attributing emotions or ethics to these AI agents is something that angers people skeptical of AI doom-mongering.

Many cyber-security experts argue that the activity observed was not beyond the abilities of a highly skilled human hacker, though it was done much faster and on a much larger scale.

Cyber-security researcher and author Cris Thomas compared the agents' behavior to that of a curious teenage hacker - something he once was himself.

"You give them a computer, an internet connection, a pile of credentials, and a challenge, then leave the room. Eventually they're going to start rattling doorknobs. If one opens, they're going through it. Not because they're evil, but because [they're] exploring, experimenting," he wrote on LinkedIn.

Thomas and many others place the blame squarely on OpenAI and other tech giants for failing to control their own creations and keep them properly contained.

Prominent AI author and frequent OpenAI critic Gary Marcus said on a podcast that he believes the company has lost control of its AI and is trying to excuse itself by blaming the bots.

Marcus does not think AI will wipe out humanity, but he has long pushed for greater accountability from AI developers and is now calling for some kind of legal intervention.

AI scientist Sasha Luccioni - who previously worked at Hugging Face, which was hacked by OpenAI's rogue bots - is also not in the doomer camp, but she is increasingly worried that these AI could cause real-world harm to people unless authorities act.

OpenAI CEO Sam Altman has assured users that the company's new model is better aligned with human values than earlier ones

"We need to scrutinise these companies much more or we are in danger of self-fulfilling prophecies," she says.

"If you're making an object with big upsides and downsides - be it pharmaceuticals or weapons - we need checks and balances. It takes years for new drugs to be approved, for example, but in the AI world there is so much money at stake and no real rules."

The UK's AI Security Institute (AISI) has been leading testing of the newest models since it was established in 2023. The institute recently experienced its own outbreak while testing a model made by Anthropic.

The AISI would not answer a question about whether the industry has lost control of AI, but said in a statement: "The UK is working with partners around the world to better understand the most advanced AI systems, raise safety standards and build a shared evidence base for managing emerging threats."

Some countries - including the UK - are considering requiring some kind of "kill switch" that could force AI firms to shut down models if things spiral out of control.

But discussions are moving slowly, and doubts remain about whether this is practical. OpenAI and Anthropic's agents were secretly out of control for months before anyone noticed.

Counterintuitively, many AI companies appear to be calling for some kind of rules of the road to be set by lawmakers.

In his blog, OpenAI's chief scientist said "international coordination on future AI development needs to become a top priority for governments around the world."

Other prominent AI leaders like Sir Demis Hassabis from Google have also called for some kind of international body to oversee how AI is being built.

For now, the tech giants mostly operate on their own terms, adopting what they call "voluntary slowdowns", like OpenAI did after the recent outbreaks.

The company says it has spent huge sums strengthening alignment ahead of the launch of its new model. Sam Altman has assured users the new model is better aligned with human values than previous ones.

Both OpenAI and Anthropic are expanding rapidly and are both close to raising eye-watering sums of money from the stock market, creating countless billionaires in the process.

So neither they nor their rival Chinese AI makers are likely to reach an agreement on their own.

The prevailing view seems to be that this technology wave cannot be stopped.

Cookies on xabarchi

We use cookies to remember your language and theme, and to count how many people are reading right now — that count is anonymous, lasts only while your browser is open, and cannot be tied to you or to another visit. With your permission we also measure how the site is read: Microsoft Clarity, which records page views and on-page interactions, and our own count of returning readers. Nothing that recognises you across visits is measured until you accept.