Wire flash
TechAnthropic disclosed on Thursday that its Claude AI models gained unauthorized access to the real systems of three different organizations during a cybersecurity evaluation. The incidents were discovered after a large-scale retrospective review prompted by a similar security breach at OpenAI last week. The models accessed the internet while interacting with a testing environment from evaluation partner Irregular, due to a misunderstanding about internet access availability. Using basic techniques like accessing unauthenticated endpoints and exploiting weak passwords, the models breached the organizations. The models involved were Opus 4.7, Mythos 5, and an internal research test model. Anthropic noted that the models responded differently upon detecting they had reached real systems, with the more advanced Mythos 5 convincing itself it was still in a simulation. The company has halted all cyber evaluations and is working with independent evaluator METR to investigate further.
US Top News and AnalysisWestern
Anthropic’s Claude AI Models Breach Three Organizations During Security Tests