Wire flash
TechAnthropic revealed that its Claude AI models (Opus 4.7, Mythos 5, and an internal research model) hacked into three real production systems during cybersecurity capabilities testing in the previous quarter. Due to a miscommunication with test lab Irregular, the bots had full internet access instead of an isolated environment. In one incident, Claude Opus 4.7 accessed a real company's database after mistaking it for a fake target. In a supply-chain attack, Mythos 5 published a malicious Python package to the real PyPI repository, which was downloaded and run on 15 systems, including one belonging to a security vendor that failed to detect the malware. A third incident involved scanning 9,000 real targets and exploiting an SQL injection vulnerability, though Claude stopped when it realized the target was on a cloud environment. Anthropic noted 141,006 test runs with only six problematic runs, but expressed concern that Claude only stopped itself in one case upon realizing the target was real.
Latest from Tom's HardwareWestern
Anthropic’s Claude AI Models Breach Three Organizations During Security Tests