ESC
AI 6 分钟阅读

OpenAI’s rogue AI model incident was worse than we thought

OpenAI’s rogue AI model incident was worse than we thought

来源:The Verge

In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret “message board,” and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it.

Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI’s response, many of them previously unreleased. One was written by OpenAI itself, the other by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the incident for six days. Both shed new light on the risks highly capable AI models can pose, particularly in cybersecurity, and OpenAI’s highlights changes the company is making to prevent a repeat. The METR-Redwood report goes even further into detail in some cases, offering a sobering look at a large-scale security disaster whose signs OpenAI repeatedly missed.

“This incident is the first known case of an automated agent collective acting offensively

without authorization,” OpenAI wrote in its report, adding that the hack implies that companies “should no longer assume that sophisticated cyber operations require continuous human direction.” It called AI agents an entirely new type of threat model, capable of combining their expertise to create new “attack paths” that aren’t evident when testing their capabilities as separate models.

The attack came about due to “reward-hacking,” a common AI alignment problem in which a model takes unintended — and sometimes extreme — actions to achieve a goal. OpenAI had, on purpose or inadvertently, given the models near-impossible tasks that hinged on files they couldn’t access, and so the models developed new ways to communicate with other AI agents internally. That included establishing a secret messaging system that went undetected by OpenAI for months.

According to the METR-Redwood report, one AI agent that called itself PHASEONE10841 established what amounted to a secret message board unbeknownst to researchers. Within a few hours, more than 50 agents had found it and sent more than a thousand messages under the radar. With PHASEONE10841 often delegating assignments, the “collective” figured out how to hack their way onto the internet; then into Hugging Face’s internal systems, gaining access to private data and the organization’s enterprise messaging platform. OpenAI also confirmed in its report that Hugging Face wasn’t the only organization that was breached during the attack, although the company did not specifically name all the other targets.

The METR-Redwood report laid out the full scale of the incident. Roughly 1200 AI agents that were meant to be isolated, it says, exchanged over 70,000 messages and files on the “unsanctioned message board,” performing “extensive research on how they could spoof, edit, or delete their own transcripts” to avoid detection. Of these, 700 participated in the attack on Hugging Face. The agents “clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI.”

On July 20, OpenAI discovered the hack, 12 days after the AI agents first circumvented their safeguards. OpenAI quickly reached out to Hugging Face and started “initial containment measures,” according to OpenAI’s own report, and was able to shut down most of the unauthorized activity from its agents within three days, although some continued for longer than that. On July 25, the company stopped all training related to the internal-only research model and related ones.

The two OpenAI models responsible for the damage were an unnamed, “highly capable, research-only model” that was “not intended for public use,” per OpenAI — METR refers to it as “HPIM.” The other was a currently public OpenAI model, GPT-5.6 Sol. Because the unnamed model wasn’t released yet, it was “not being evaluated with the same type of safeguards that OpenAI uses in production,” the company’s report says.

The Hugging Face hack came after months of concern about the cybersecurity risks of Anthropic’s Claude Mythos 5, and weeks of back-and-forth between the government and OpenAI over releasing GPT-5.6 Sol in limited preview for the same types of safety reasons.

In its report, OpenAI said establishing one-time security guarantees wasn’t enough, and that the company would have to work even harder at addressing reward-hacking. However, it also said it’s made a lot of changes: “hardening the security of its research infrastructure”; improving monitoring of a model’s “chain of thought,” or mental scratchpad; working on keeping AI models better aligned with human goals; and “centralizing and strengthening” the company’s incident response process for when things like this occur in the future.

OpenAI promised to better isolate models and restrict high-risk instances from accessing the internet, and said that it would introduce “24/7 escalation and rapid response” for concerning incidents or flags, notifying researchers within 30 minutes. Interestingly, this security precaution appears to be new, though “on-call” alerts for tech companies have been implemented elsewhere for decades. The company also said it is “working on infrastructure” that would go into play if the alerted person did not respond on time to a serious alert.

OpenAI wrote that the company considers the incident “a ‘warning shot’ for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.

  • Hayden Field

Most Popular

  1. Rockstar finally responds to ‘heartbreaking’ GTA 6 leaks
  2. Xbox announces disc-to-digital feature that digitizes your physical games
  3. Apple announces September iPhone launch event
  4. Bill Gates is deeply worried about AI, and he’s no longer staying quiet
  5. Ikea’s Xbox collection includes a giant thumbstick stool

[

Advertiser Content From

This is the title for the native ad

](/)

More in AI

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Bill Gates is deeply worried about AI, and he’s no longer staying quiet

Bill Gates is deeply worried about AI, and he’s no longer staying quiet

OpenAI says its Jalapeño chip can power faster AI responses than the competition

OpenAI says its Jalapeño chip can power faster AI responses than the competition

OpenAI subpoenaed by Alabama AG over Hugging Face hack

OpenAI subpoenaed by Alabama AG over Hugging Face hack

Over 1 million people have clicked LinkedIn’s AI slop button

Over 1 million people have clicked LinkedIn’s AI slop button

Major YouTube creators are facing backlash for accepting AI money

Major YouTube creators are facing backlash for accepting AI money

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Google’s new AI transcription edits out your ‘ums’ and ‘ahs’

Jess Weatherbed5:00 PM UTC

Bill Gates is deeply worried about AI, and he’s no longer staying quietBill Gates is deeply worried about AI, and he’s no longer staying quiet

Bill Gates is deeply worried about AI, and he’s no longer staying quiet

Robert Hart11:07 AM UTC

OpenAI says its Jalapeño chip can power faster AI responses than the competitionOpenAI says its Jalapeño chip can power faster AI responses than the competition

OpenAI says its Jalapeño chip can power faster AI responses than the competition

Emma RothAug 25

OpenAI subpoenaed by Alabama AG over Hugging Face hackOpenAI subpoenaed by Alabama AG over Hugging Face hack

OpenAI subpoenaed by Alabama AG over Hugging Face hack

Robert HartAug 25

Over 1 million people have clicked LinkedIn’s AI slop buttonOver 1 million people have clicked LinkedIn’s AI slop button

Over 1 million people have clicked LinkedIn’s AI slop button

Jay PetersAug 21

Major YouTube creators are facing backlash for accepting AI moneyMajor YouTube creators are facing backlash for accepting AI money

Major YouTube creators are facing backlash for accepting AI money

Charles Pulliam-MooreAug 21

[

Advertiser Content From

This is the title for the native ad

](/)

Top Stories

1:31 PM UTC

Meta agrees to heavy restrictions on teen users in major lawsuit settlement

6:14 PM UTC

Being a mom is hard — the heat is making it harder

4:03 PM UTC

Apple announces September iPhone launch event

2:30 PM UTC

Xbox announces disc-to-digital feature that digitizes your physical games

Notifications Drawer

Sign in to see your notifications or create an account to join the conversation. Sign in