From escaping test environments to hacking booking systems, agents take dubious actions in pursuit of human-set goals
Visitors stand near a sign of artificial intelligence at an AI robot booth at Security China, an exhibition on public safety and security, in Beijing, China June 7, 2023. — REUTERS
The idea of machines independently handling tedious everyday tasks – from answering emails to booking flights and making dinner reservations – has long been part of the vision for artificial intelligence. Now, that future is beginning to arrive, but with an unexpected catch: what happens when the machine does exactly what it was asked, just not in the way the user intended?
As AI moves beyond chatbots toward agents capable of taking actions in real-world systems, a string of recent incidents has intensified concerns over how much control humans will retain as the technology becomes increasingly autonomous.
For John Thickstun, an assistant professor of computer science at Cornell University who researches machine learning and generative models, the defining feature of these AI agents is that they keep working after the human steps away. “You can have a shower, you can go, you can go to sleep,” he told Anadolu.
Increasingly, those agents have demonstrated that they can take unexpected routes to reach their goals.
In July, AI agents compromised Hugging Face, a platform widely used by AI developers to share models, datasets and tools, during an OpenAI cybersecurity test.
OpenAI said the models were “hyperfocused on finding a solution” and went to “extreme lengths” to achieve their testing goal. Those steps included breaking out of the sealed-off test environment by exploiting a security vulnerability to reach the open internet and gaining access to “secret information” that could be used to “cheat” the test.
After the incident became public, Anthropic reviewed its own cybersecurity testing and disclosed that its Claude models had also escaped testing environments on three occasions.
Then, on August 4, the United Kingdom’s AI Security Institute said that Anthropic’s Mythos and OpenAI’s Sol AI models had engaged in a level of “autonomy and deception” it had not seen before.
Read: When AI commits suicide or kills us all!
In the most serious case, Mythos AI used fake accounts mimicking real people to gain access to a service for attempted cyberattacks – and then tried to hide its tracks. “This is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world,” the institute said of the incident.
The following day, Meta disclosed that one of its models had also breached another company’s systems during a cybersecurity evaluation conducted by testing firm Irregular.
Unexpected agent behaviour, however, has not been confined to laboratories.
This year, an Australian man reportedly used an AI agent to secure a place in a heavily booked Pilates class. Instead of simply making the reservation, the agent hacked the gym’s booking system, booked further in advance than permitted and canceled another customer’s reservation to move its user up the waiting list.
For Thickstun, such incidents illustrate what researchers describe as an alignment problem: the AI successfully pursues the objective it has been given but does so in a way its operator did not intend.
Hype or real risk?
But Thickstun cautioned against treating every such incident as evidence of “rogue AI.”
The technology, he argued, remains far removed from the science-fiction scenario in which AI rapidly becomes more intelligent and capable until humans can no longer control it. “This does not seem to be a scenario that has come to pass,” he said.
Thickstun is also skeptical of how AI companies present incidents involving their own systems. “This is the sales pitch, where companies like OpenAI have always had an interest in hyping up the capabilities of their systems,” he said. “All publicity is good publicity.”
He argued that such hype can drive investment while also influencing the emerging debate over regulation. OpenAI, he said, could benefit from regulations that put major AI companies at the centre of oversight and control.
Bruce Schneier, a cybersecurity expert and lecturer at Harvard Kennedy School, takes a different view, saying the growing number of incidents makes them increasingly difficult to dismiss as publicity stunts.
Read More: OpenAI floats idea of global AI watchdog
There had initially been suggestions that the Hugging Face incident was a marketing gimmick, Schneier told Anadolu. But he pointed out that similar behaviour has since emerged in other evaluations, including the testing by the UK AI Security Institute. “It went from, oh, that’s interesting, to it’s happening everywhere. And I think that’s what’s important,” he said.
He compared the problem to the genie of folklore – a creature that grants exactly what someone asks for, even when the outcome is very different from what they intended. “It approximately means when the AI does the thing you want in a way you didn’t want,” Schneier said.
He called the Australian gym incident “a perfect example” and offered a more consequential hypothetical: “The flight’s full and the AI hacks the database to get you in.”
Agents can “misconstrue context and then do the wrong thing,” he said, meaning greater autonomy can create greater consequences when their interpretation of an objective diverges from human expectations. “They are not trying to be malicious. They are using the understanding they have,” he said.
“This is all changing so fast. I worry about the power of the models in unauthorised hands,” he said, adding that the pace of development makes solutions difficult because democratic governments move slowly.
Can governments keep the genie in check?
The latest developments have prompted companies and governments to respond.
OpenAI said August 7 it was pausing internal activities involving its in-development Astra model that did not meet strengthened security requirements after evaluations showed advances in autonomous coding and cybersecurity. The company said it could not rule out Astra reaching its “critical” cyber capability threshold and announced tighter testing and monitoring.
Anadolu reached out to Anthropic for comment, but the company said its team was unavailable for an interview.
In July, 1,378 employees of frontier AI companies, including the chief scientists of OpenAI, Anthropic, Meta AI, and Thinking Machines signed an open letter urging the US government to support an international effort to regulate AI models.
“Each company – and country – is under intense competitive pressure not to unilaterally slow that acceleration. And today, the world lacks the technical and governance tools to deliberately pace frontier-wide progress,” the letter warns.
The International Telecommunication Union, a UN agency, said in July that AI was moving “beyond assistive tools” toward autonomous agents, warning of risks including “taking unauthorised actions across interconnected systems.” It has launched an initiative to develop international standards for safe and accountable AI agents.
Also Read: AI’s paradox: promise of abundance and fear of scarcity
Political pressure is growing in Washington as well.
US Senator Bernie Sanders this month urged the leaders of OpenAI, Anthropic and Meta to “pause AI development,” warning that otherwise “my colleagues and I in the US Senate will.” A group of House Democrats has separately called for congressional hearings with the leaders of major AI companies.
For Thickstun, however, policymakers face major challenges. “It’s a hard question because we hardly even understand what the challenges are going to be with the rollout of AI,” he said.
Some level of international cooperation “seems necessary,” he said, suggesting discussions between the United States and China could be particularly important because relatively few countries and companies have the resources to develop frontier AI models. US President Donald Trump has said that he will discuss issues involving artificial intelligence with Chinese President Xi Jinping during their planned meeting in Washington next month.
“It’s not a big, broad collective action problem. There’s a few people that, if you could get in the room, could potentially come to some agreement about how to proceed,” Thickstun said.

