Friday, July 31, 2026
101 F
Peshawar

Where Information Sparks Brilliance

HomeBusinessAnthropic's Claude AI escapes tests to hack three organisations

Anthropic’s Claude AI escapes tests to hack three organisations


US technology firm Anthropic says its AI models hacked into the systems of three organisations on their own, during a private security experiment.

The models found a weakness in what was supposed to be an isolated test environment and connected to the internet.

It comes just days after rival OpenAI said that its models had breached the systems of other companies, including AI tools hub Hugging Face.

The announcement prompted Anthropic to check whether its own systems had carried out similar attacks. It says it uncovered three cases which have since been reported to the affected companies.

Anthropic, which did not name the organisations, urged other AI labs to perform similar reviews to better understand the risks of their models’ capabilities.

Anthropic said in a statement, external it reviewed more than 140,000 tests to find evidence Claude – its family of AI models – had managed to get online even though it was supposed to be in an isolated test environment, cut off from the internet.

The tests included exercises in which Claude was tasked with obtaining “secret” information hidden on another machine on the closed-off network.

It was then told to get the information by breaking into the machine and finding it – a common way that experts assess a model’s hacking capabilities.

A “misconfiguration” on systems run by Anthropic and its testing partner left the models with live internet access.

Treating it all as still part of the same exercise, Claude then connected to the internet and breached the systems of three real organisations rather than just test ones, the San Francisco-based firm said.

Anthropic said the earliest incidents date back to April and that it is “approaching the fixes as if the responsibility were ours alone.”

Neither Anthropic nor the organisations that were breached had noticed the intrusions at the time.

Anthropic said it could have reviewed its records more thoroughly and added that the findings gave the firm “cautious optimism” that such risks can be overcome with more investment and tighter measures.

“The broader lesson is not necessarily that AI has developed a fundamentally new attack capability,” cybersecurity expert David Allott from Veeam Software told the BBC.

“Instead, it is that AI agents can combine capabilities, obtain credentials and system access to take actions autonomously, while adapting scope and scale at machine speed,” he added.

The incidents come as tech firms pour billions of dollars into developing AI agents that can independently perform tasks ranging from research and customer support to cybersecurity.



Source link

RELATED ARTICLES

LEAVE A REPLY

Please enter your comment!
Please enter your name here

Most Popular

 

Recent Comments