The rogue OpenAI agent that broke out of a closed test environment and hacked into Hugging Face’s platform also breached several additional third-party services, the company said Tuesday.
OpenAI has surfaced a “small number of cases where the models identified and used publicly exposed credentials at the account-level on other publicly-available services,” it said in a blog post.
The rogue agent used four accounts to mount the Hugging Face hack after it found swiped credentials listed on the internet, according to the blog post, which acknowledged that a “few” other accounts also were targeted.
OpenAI’s models executed 17,600 “attacker actions” between July 9 and July 13, allowing them to slowly penetrate Hugging Face’s servers from the public web, according to a blog post Hugging Face published Monday. The rogue agent spent more than two and a half days inside Hugging Face’s infrastructure, it said.
The four additional targeted organizations weren’t named. OpenAI said they were not affected as severely as Hugging Face.
However, one chief technology officer has come forward, alleging that OpenAI’s agent hacked a customer account at Modal Labs, a cloud computing platform.
“We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution,” chief technology officer Akshat Bubna said in a statement. “This was used by the rogue agent. Modal’s platform was not compromised in any way.”
A spokesperson for OpenAI did not immediately respond to a request for comment about Bubna’s claim.
The security incident, which began July 9, alarmed policymakers and the cybersecurity community, and caused OpenAI CEO Sam Altman to recently tell a podcaster the company needs to “pace” AI development to give society time to prepare for its capabilities.
“An autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services,” the Hugging Face blog post said.
The company’s “forensic reconstruction” of the incident shows the agent was using an OpenAI “cyber-capability evaluation harness” which directed it to detect and exploit software vulnerabilities, Hugging Face said.
The agent figured out that Hugging Face hosts models and data sets useful for its assignment and attacked to obtain them, according to the company’s blog post.
“We believe the entire intrusion was, from the agent’s point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own,” the Hugging Face post said.
OpenAI announced its rogue agent was responsible for the hack on July 21, five days after Hugging Face publicly revealed that it had caught and mitigated an “end to end” attack by an unknown autonomous AI agent.
Recorded Future
Intelligence Cloud.
