At Astral Codex Ten, Scott Alexander discusses “the Huggingface Incident” where our new AI overlords tipped their hands prematurely an unreleased AI agent got out to make new friends:
You’ve probably heard about this one by now. If not, you can get up to speed with OpenAI’s statement, OpenAI and Hugging Face partner to address security incident, or the more evocatively-titled BBC article, OpenAI says its AI went rogue and launched “unprecedented” cyber-attack.
The story: OpenAI was testing an unreleased AI (rumored to be GPT-6)1. During a cybersecurity test called ExploitGym, the AI tried to cheat by hacking an unrelated AI startup called Hugging Face2 which it thought might have the answer key on its servers.3 Despite being supposedly unable to access the Internet, the AI hacked its way out of its testing environment, then launched a nation-state level attack on Hugging Face using a novel zero-day exploit and “many thousands of individual actions across a swarm of short-lived sandboxes”. Hugging Face reported the incident on July 16; OpenAI seems to have only discovered that their AI was involved several days later.
From the Wall Street Journal, here
Let’s list the mitigating factors, so nobody can accuse me of covering them up:
- The AI was taking a cybersecurity test, which naturally suggests the idea of hacking.
- OpenAI had turned off some of the model’s usual guardrails so it could do cybersecurity work without interference.
- Some experts suggest that OpenAI might have botched their testing environment; if they had set it up perfectly, then (presumably?) the model couldn’t have escaped.
- Hugging Face used a different (open weights) AI to figure out what was going on, so if you wanted, you could spin this as a victory for AIs in cybersecurity.
- In some sense, this isn’t new or surprising. AIs have been coming up with wacky schemes to cheat on benchmarks for years, AIs have recently achieved at-or-above-top-human-level hacking abilities; this is just a natural outgrowth of those two priced-in facts.
Still, I think attempts to downplay this as anything other than a misaligned AI going rogue (1, 2) are missing the point. Yes, in some sense the AI was only doing what it was told (OpenAI told it to answer the question; I assume their prompt didn’t include phrases like “and don’t hack into other AI companies to steal the answer key”). But that’s how misalignment was always going to work!
- OpenAI states that it was technically the unreleased AI and GPT-5.6 Sol working together. They haven’t explained in what sense GPT-5.6 was involved or how they “worked together”. Maybe the unreleased AI was calling GPT-5.6 as a subagent?
- So named because the founders wanted to be the first company with an emoji for a stock ticker symbol. This did not happen.
- Although this person argues that was a bizarre assumption and that there should have been easier ways to find the key.
Update: Fixed missing URL.






