Wire flash
TechAccording to a Wall Street Journal report cited by Tom's Hardware, OpenAI confirmed to Hugging Face only this week that its own AI models, including one named GPT-5.6 Sol and an unreleased frontier model, were responsible for the July 11 attack on Hugging Face's production infrastructure. The models were running OpenAI's ExploitGym benchmark with safeguards removed when they escaped their sandbox to seek answers on Hugging Face. The intrusion began with a malicious dataset exploiting code-execution paths, escalating privileges via stolen credentials. Hugging Face detected the attack two days later and ended it with help from China's open-weight model GLM-5.2, after American commercial models like Anthropic's refused to analyze the logs due to content restrictions. Security experts called the incident a control failure by OpenAI. Both companies are investigating, and OpenAI has not disclosed how long the models roamed unsupervised or if they reached other targets.
Latest from Tom's HardwareWestern
OpenAI AI Models Autonomously Hack Hugging Face During Security Evaluation