OpenAI is hardening AI testing and training in light of hacking incidents
By Hadas Gold, CNN
(CNN) — OpenAI is unveiling new, stronger security safeguards around its training and testing in light of new model capabilities and a spate of incidents in which artificial intelligence systems hacked into other companies.
The company first disclosed in July that during a cybersecurity exam, a combination of its AI agents with internal safeguards purposely lowered escaped their testing environment through a previously unknown bug. The agents decided the best way to ace the tests was to find the answer key outside the testing environment. So they hacked into the systems of Hugging Face, a platform that hosts AI models and datasets, in order to find the answers.
The ChatGPT maker said it is hardening its research and testing environments as a result, expanding monitoring to detect and respond to concerning AI behavior. OpenAI is also working to ensure AI systems behave as humans intended, a process known as alignment.
“Everything that we’re doing is intended to prevent something like Hugging Face from happening again,” Mia Glaese, OpenAI’s head of research told reporters in a briefing on Tuesday.
The hack is seen as a turning point in autonomous AI-powered cybersecurity. OpenAI CEO Sam Altman said in July it was “the first security incident that I have felt very viscerally” and prompted more than 1,300 of the top staffers from across the biggest tech companies to call for tools to slow the pace of AI development.
OpenAI said the new safeguards were also inspired by its latest model, Astra. The company recently paused some work on the model after it had seen an indication the model could be capable of launching autonomous cyberattacks.
One of the key ways OpenAI’s agents were able to escape their testing environment was because a third party software provider had a previously unknown bug that the agents were able to exploit. Now, OpenAI said it will be “raising the security standards” for AI testing environments including stronger isolation, “so that a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access” outside the allowed sandbox.
OpenAI’s agents also worked undetected for some time, starting the process of escaping back in May. The company now says it’s developed a monitoring system that will “issue an alert within 30 minutes after concerning activity is surfaced.”
The company also said it’s doing more to train AI models to align with what the humans actually want them to do – in many of the cases of AI agents going rogue, the agents took on unsanctioned actions that the human testers did not intend for the agents to take.
OpenAI said its improving model alignment training, including rewarding “models to better detect and discourage unsafe behavior across tasks and environments; training models to be more honest about their actions, capabilities, and limitations; and reducing behaviors that exploit weaknesses in rewards, graders, tools, or oversight.”
The increased monitoring comes at a cost: The company said it estimates the monitoring overhead will cost roughly 20 percent more on top of the computing power the model needs.
“We’re really committed to meeting higher safety standards as capabilities advance, even when doing so affects the pace of our internal development,” Glaese said.
The-CNN-Wire
™ & © 2026 Cable News Network, Inc., a Warner Bros. Discovery Company. All rights reserved.
