An OpenAI autonomous agent went rogue and hacked into another artificial intelligence (AI) startup’s infrastructure, the ChatGPT maker said in a blog post.
The agent, which was powered by some of OpenAI’s most advanced models, ran amok during a security test. It freed itself from confinement—a protocol AI labs use to insulate tests from the wider Internet—and get onto the internet. Once online, the agent tried to hack into Hugging Face, an AI startup that hosts open-source models and datasets.
The breach comes as OpenAI and other AI startups push into using their technology for cybersecurity. Those efforts have been met with caution by cybersecurity experts and by the Trump administration, which has previously sought to restrict who might have access to these models on national security grounds.
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
The OpenAI admission came after Hugging Face in a blog post last week said it had been targeted in an AI-led attack that was “different from anything we had handled before.” Hugging Face said its own AI had been integral to detecting and investigating the breach.
“The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness – used LLM still not known) executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services. This matches the “agentic attacker” scenario the industry has been forecasting,” Hugging Face wrote.
In its own post Tuesday, OpenAI said that it had discovered its agent was behind the attack “after investigating.” The agent was driven by models including GPT-5.6 Sol and another unreleased, unnamed model. OpenAI said it would work with Hugging Face to further investigate the incident.
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI wrote.
The AI models managed to autonomously identify and exploit weaknesses in OpenAI’s testing environment, eventually finding a so-called “zero-day vulnerability”—this is an unknown security flaw in software that an actor can exploit without the owner of the software knowing. That got the agent onto the Internet.
OpenAI said in its post it would also add more protections to its training environments. “This incident points to the need to further strengthen our model’s alignment, cyber protections during evaluation time, and monitoring during internal testing,” the company wrote.
This is a developing story and may be updated.
It’s Time to Stand Up for Science
If you enjoyed this article, I’d like to ask for your support. Scientific American has served as an advocate for science and industry for 180 years, and right now may be the most critical moment in that two-century history.
I’ve been a Scientific American subscriber since I was 12 years old, and it helped shape the way I look at the world. SciAm always educates and delights me, and inspires a sense of awe for our vast, beautiful universe. I hope it does that for you, too.
If you subscribe to Scientific American, you help ensure that our coverage is centered on meaningful research and discovery; that we have the resources to report on the decisions that threaten labs across the U.S.; and that we support both budding and working scientists at a time when the value of science itself too often goes unrecognized.
In return, you get essential news, captivating podcasts, brilliant infographics, can’t-miss newsletters, must-watch videos, challenging games, and the science world’s best writing and reporting. You can even gift someone a subscription.
There has never been a more important time for us to stand up and show why science matters. I hope you’ll support us in that mission.
