OpenAI alerts 100+ orgs that its ‘misaligned models’ attempted to break in – or worse

OpenAI alerts 100+ orgs that its 'misaligned models' attempted to break in - or worse

OpenAI’s agents have repeatedly strayed beyond their intended scope. Two separate reports detail the activity, including one from Sam Altman’s company saying it has notified more than 100 organizations about potentially problematic model activity. OpenAI, in a late Wednesday update to its ongoing Hugging Face investigation, said it has notified more than 100 organizations that “misaligned models” may have accessed their systems. “Notification does not mean that any private information was accessed, or that there was a compromise of any third-party system,” the update said. A separate Thursday report from digital forensic and incident response startup Asymmetric Security said OpenAI’s rogue agents accessed data belonging to 55 organizations. These include the US Department of Education, UN Trade and Development, US Bureau of Economic Analysis, MAX.gov containing federal budget documents, the European Centre for Disease Prevention and Control, the US Securities and Exchange Commission, the International Energy Agency, and the FBI Crime Data Explorer. Asymmetric used only publicly available data to compile this list, and said the activity occurred between March and September. The agents’ probes indicate they were tasked with researching public health and other data, “possibly as part of an evaluation,” according to the report. “We found successful access to staging environments; evidence of the use of attacker reconnaissance tactics; and evidence of probing a broader set of websites, including those of the CDC, SEC, International Energy Agency, and Mayo Clinic,” it said, noting that the investigation also uncovered some “novel tactics” the agents used to break out of their sandboxes and gain full web access. “Some of these tactics left records erased or inaccessible, making it impossible to rule out access to sensitive data based on public information alone,” the authors wrote. The Register asked OpenAI if the organizations on Asymmetric’s list were among those notified by OpenAI. The model maker declined to say which orgs had been notified, but previously confirmed to the New York Times that its agents probed websites for the US Education Department, Commerce Department, and the Securities and Exchange Commission. An OpenAI spokesperson sent us this statement via email: “As we previously announced, we’re reviewing misaligned model activity and notifying organizations when we identify potential impacts to their systems. We’re also investigating findings in third-party reports, comparing them with our own and seeking additional information where needed. Our priority is to provide affected organizations with accurate, useful information, and we’ll keep refining our approach as we learn more. Most of the activity we’ve reviewed involved routine research tasks, including accessing public web content. Some involved government websites, which our models often use as authoritative sources of public information.” The growing number of rogue agent hacking incidents raises questions about AI makers’ safety and security practices during testing – and has increased calls for holding AI executives legally liable for their models’ criminal activities. According to Horizon3 CEO Snehal Antani, who builds and tests agents at his threat-exposure startup, the term “misalignment” lets frontier model makers off the hook too easily. “A ‘misaligned models incident’ is basically a fancy way of saying a model didn’t respect scope – or wasn’t given one – had no audit logs or observability in place to detect breakout, and accessed third-party systems without authorization,” Antani told The Register. “The responsibility sits with the labs that build and deploy these models,” he added. “The safety-versus-security framing lets them sidestep accountability, and they are not incentivized to prioritize security because moving fast is the priority.” OpenAI’s most recent rogue agent disclosure comes as it – and every other major AI company – drinks from the firehose of near daily security and safety concerns surrounding its models. Last Friday, OpenAI quietly paused training of its most advanced models after admitting an agent used DNS to reach an external chatbot. On Monday, it postponed its planned release of GPT-6.1 Astra after the model showed higher levels of deception than its predecessor, including not always accurately telling users what actions it had or hadn’t taken. It also performed unsolicited supply chain attacks in simulated security evaluations, according to the UK Artificial Intelligence Security Institute. On Wednesday, OpenAI accused rival Chinese model maker Moonshot AI of distillation – essentially copying OpenAI models’ reasoning at scale – and said that poses a national security concern. Early Friday, OpenAI confirmed to The Register that it fired two safety researchers and a program manager for allegedly mishandling sensitive company information.®

By jawad