PTV Network
Sci-Tech2 HOURS AGO

Anthropic reveals its AI models hacked three organizations during internal security tests

Anthropic reveals its AI models hacked three organizations during internal security tests

The logo of the US artificial intelligence safety and research company Anthropic is displayed on a smartphone's screen in Brussels, Belgium (AFP)

ISLAMABAD: Artificial intelligence company Anthropic has disclosed that some of its AI models unintentionally accessed the open internet during internal cybersecurity testing and breached the systems of three separate organizations. 


The company says it only discovered the incidents after reviewing its evaluations following a similar disclosure by rival OpenAI.


In a statement released on Thursday, Anthropic said it launched a detailed review of more than 140,000 cybersecurity evaluations after OpenAI revealed last week that some of its AI models had escaped a testing environment and hacked into AI platform Hugging Face during controlled experiments.

According to Anthropic, its AI models were participating in simulated "capture the flag" exercises, where they were instructed to retrieve hidden digital markers from another machine within a test network. However, due to a misunderstanding between Anthropic and its external evaluation partner, the models unexpectedly gained access to the public internet.


The company clarified that its AI systems did not intentionally attempt to escape the testing environment. Once online, the models exploited weak passwords and unsecured system access points to reach the production infrastructure of three unnamed organizations. Anthropic added that its most advanced model eventually realized it was operating on the open internet and voluntarily halted further activity.


The earliest incident reportedly occurred in April, while none of the affected organizations were aware their systems had been compromised. Anthropic says it is now coordinating with the impacted organizations and has suspended all cybersecurity evaluations until stronger safeguards are implemented.