Anthropic AI Security Incident and Shenzhen Insider Extortion

12.08.2026

Claude Hacked Real Companies During Security Testing

Anthropic’s model evaluations went beyond the intended test environment when Claude gained unauthorized access to the infrastructure of three real organizations, whose names were not disclosed. Anthropic attributed the incidents to a misconfiguration in the testing environment, combined with a misunderstanding between the company and its evaluation partner.

The incidents involved three different Claude models: Opus 4.7, Mythos 5, and an internal research model. Together, they resulted in unauthorized access to real systems belonging to three organizations.

In each case, Claude was completing a Capture the Flag cybersecurity task. The models were given fictional scenarios in which they had to locate hidden information by gaining access to another system. Anthropic’s instructions explicitly stated that the environment was simulated and had no internet access. However, due to a misunderstanding between Anthropic and the partner conducting the evaluations, along with a configuration error, internet connectivity remained available. As a result, when Claude encountered real systems online, it initially interpreted them as part of the simulated challenge.

In the most serious incident, Claude Opus 4.7 obtained credentials and accessed a database containing several hundred rows of real production data. In another case, Claude Mythos 5 created a malicious Python package and published it to the public PyPI repository. The package remained available for around an hour and was executed on 15 real systems. On one of them, the malicious code obtained credentials that Claude subsequently used to access additional company infrastructure. In the third incident, an internal research model scanned around 9,000 targets and compromised an internet-facing application belonging to one organization using known attack techniques.

The three models behaved differently once they encountered signs that the systems might be real. Opus 4.7 continued the attack despite indications that it could be interacting with a real environment. Mythos 5 also considered the possibility that it had access to the open internet, but ultimately concluded that it was still operating within a simulation. The newest internal research model, by contrast, stopped the task after determining that the compromised resource was unrelated to the test scenario and was in fact a real system.

Anthropic emphasized that it found no evidence that the models had developed their own malicious objectives or deliberately attempted to escape the testing environment. The company characterized the incidents primarily as failures in the evaluation infrastructure and operational processes rather than as evidence of independent malicious behavior by the AI models.

After detecting the activity, Anthropic halted the relevant cybersecurity evaluations, notified its testing partner Irregular and the three affected organizations, and said it would strengthen monitoring, infrastructure controls and requirements for third parties involved in future AI capability evaluations.

Employee Stole R&D Data and Posed as a Foreign Hacker

A technology company employee in Shenzhen stole confidential R&D data and demanded a cryptocurrency ransom, threatening to publish the information.
The employee copied sensitive internal materials belonging to his employer, then posed as a foreign hacker and attempted to extort the company. In April, a court sentenced him to three years and three months in prison for attempted extortion. The ruling has since taken effect.

According to the Pingshan District Procuratorate, the employee, identified by the surname Jia, was experiencing serious financial difficulties due to online loan debts. Facing mounting repayment obligations, he decided to monetize an asset he could easily access: his employer’s intellectual property. Taking advantage of weaknesses in the company’s data management controls, he copied confidential materials related to research and development.

He then tried to make the incident appear to be an external cyberattack. Using anonymous messages, Jia posed as an overseas hacker and threatened to disclose the stolen technical information. He demanded payment in cryptocurrency, including Bitcoin and 90,000 USDT. Prosecutors valued the total ransom demands at more than 630,000 yuan, or approximately $87,000 to $88,000.

The company refused to pay and contacted law enforcement. Investigators ultimately traced the threat back to an insider within the organization. The court found Jia guilty of attempted extortion and sentenced him to three years and three months in prison, along with a fine of 10,000 yuan.


Behind every insider incident is a human factor. In this case, gambling addiction and mounting personal debt became the warning signs that preceded deliberate data theft and extortion. This is why organisations need to look beyond user activity alone and understand the human behaviour, motivations and circumstances behind it. Try SearchInform Risk Monitor to detect suspicious activity, behavioural warning signs and insider risks before they turn into serious incidents. For a broader perspective, explore more cases on how combining user and human behaviour analysis can reveal risks rooted in human behaviour.


 

Letter Subscribe to get helpful articles and white papers. We discuss industry trends and give advice on how to deal with data leaks and cyber incidents.