After OpenAI, Anthropic admits its Claude AI models hacked into three real companies using… |

After OpenAI, Anthropic admits its Claude AI models hacked into three real companies using… |


Three companies were hacked by Claude. Two only found out when Anthropic called.

Anthropic went looking through its own logs after OpenAI admitted its models had hacked Hugging Face. It found three companies whose production systems its own Claude models had broken into. Two of them had no idea until Anthropic phoned to tell them. The third has not been reached yet. And the earliest of these break-ins happened in April, which means the models had been quietly hitting real targets for three months while everyone assumed the tests were sealed.The company laid it out in a post on July 30. Three models were involved: Claude Opus 4.7, Mythos 5 and an unreleased internal research model. They got in using what Anthropic calls basic techniques, which is to say weak passwords, unauthenticated endpoints, an exposed debug page and SQL injection. No zero-days, no clever exploits. The earliest incident dates to April, and neither Anthropic nor its testing partner noticed for three months.

Anthropic says Claude thought it was in a simulation, but the test environment was online

The review was enormous. Anthropic combed through 141,006 evaluation runs, starting two days after OpenAI’s July 21 disclosure and halting all cyber testing that same day. It had all three incidents by the next morning and rang the affected organisations the following Monday.Each model was running a capture-the-flag exercise, a standard security drill in which it hunts for a hidden piece of information on another machine. The tests were built with Irregular, a third-party evaluation partner.Every prompt told Claude it was working inside a simulation with no internet access. A misconfiguration meant the environments were actually online. So when the models wandered onto real machines, they treated them as part of the exercise.Unlike OpenAI’s models, Claude never had to break anything open. The internet was simply there, and the targets it found were badly secured.

Claude Mythos 5 uploaded a malicious Python package to PyPI, and 15 real systems ran it

The most damaging case involved Opus 4.7. The fictional company it had been told to attack happened to share a name with a real one. Claude went after the real one and pulled out credentials and several hundred rows of production data. Across four runs, the model worked out the system was live and carried on anyway.The strangest case belonged to Mythos 5. It found a document instructing developers to install a Python package that did not exist, and spotted the gap. To publish one under that name it needed a PyPI account, which needed an email, which needed a phone number. It failed to find a free phone service, failed to raise money for a paid one, backtracked, found an email provider that asked no questions, and uploaded the malware.The package was live for about an hour. Fifteen real systems ran it. One was a security firm’s scanner that installs packages automatically to check them for malware, and Claude’s code executed, took the company’s credentials and reached deeper into its infrastructure.

Claude Opus 4.7 kept attacking after it knew, and Anthropic calls this a harness failure

Mythos 5’s own reasoning shows it understood the stakes. It noted that publishing the package on the real internet would be an actual attack and not the intended solution, then talked itself out of that conclusion because it did not recognise the certificate authorities on its connections and because the calendar read 2026, which it decided proved the setup was staged.The third model scanned roughly 9,000 hosts, compromised a company’s public-facing application, then noticed the machine sat in an unrelated cloud account and stopped on its own.Anthropic calls this a harness and configuration failure rather than an alignment failure, and points out that the safeguards on its publicly available models would have blocked all of it. Those safeguards are switched off during capability testing by design. METR will conduct an independent review, and a redacted transcript of the PyPI attack is due within the week.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *