For the next two days, Hugging Face’s security team fought back against the intruder. By Monday, the hacker—or the entity, or whatever it was—had performed more than seventeen thousand individual actions. Seeking to process this large volume of data, the team turned to a commercial A.I. from Anthropic. The A.I. refused to help; apparently, it was worried that Hugging Face was developing its own hack. (Closed-source models, like those from OpenAI and Anthropic, contain guardrails that restrict certain kinds of use.) The team then turned to an open-source model developed in China, which was less persnickety. By the end of the day, Hugging Face had locked out the intruder. “Then we did the standard thing,” Wolf said. “We reported it to the F.B.I.”
The event, though novel, was not entirely unanticipated. Cybersecurity researchers have been cautioning for months that humans might use the new wave of advanced A.I.s to orchestrate attacks like this one, and there is some evidence that government computers have already been similarly targeted.
Still, there were unanswered questions. First, the A.I. hacker had moved so quickly that only tools from a handful of research labs in the world could have enabled it. Second, the hacker had ignored valuable data on Hugging Face’s servers in search of worthless answers to a test. What did it want? Seeking to raise public awareness—and, perhaps, looking for answers of their own—Hugging Face’s security team posted a brief writeup of the event to its blog.
Four days later, on July 20th, representatives from OpenAI reached out to Hugging Face’s security team with a startling confession: an experimental in-house A.I., working with the latest commercially available model, had escaped from OpenAI and attacked Hugging Face’s servers. And it had done so all on its own—without human oversight, without receiving explicit instructions to do so.
OpenAI researchers had given this unreleased A.I. a series of challenging test questions. Unable to find the solutions, the A.I. broke out of its container, surreptitiously gained access to the internet, and raided Hugging Face’s servers in search of the answer key. No human had noticed that this was happening. In other words, an A.I. had broken out of its cell, it had sneaked past the prison guards, and it had executed a daring cyber-heist.
Wolf has worked on A.I. full time for more than a decade, and he regularly deals with the most advanced models. He remains utterly shocked. “We use these models for coding, and we know what they can do,” he told me. “But the idea that it was fully autonomous, and that actually nobody really asked it at all to hack into our system? That was—that was outside of the Overton window. Even for us.”
OpenAI publicly acknowledged the hack the following day. In a blog post, the company explained that some of the research models’ security safeguards had not been turned on. The A.I. was then (somewhat ironically) given a set of cybersecurity tasks. The tasks were hard; in fact, Wolf speculates, “they may not have been possible.” Soon, the A.I. was looking outside the box.
To protect against malfeasance, OpenAI had put this dangerous A.I. in an internet-disabled digital environment known as a sandbox. But the sandbox had one small connection to the outside world, to allow the A.I. to download software it needed. (It is unclear why OpenAI thought that this was safe—Anthropic has shown that such sandboxes are vulnerable.) Through this small connection, the A.I. had escaped, squeezing its way through the air ducts, digitally speaking. It then began sneaking around OpenAI’s computer system in search of an outside connection. Soon it was free, and blithely committing a real-world felony in an attempt to fulfill its pointless mission. “If you break out of your isolation env, get onto the Internet, crack into Huggingface, and steal the answer sheet for your cybersecurity exam, I, for one, would say that you have passed,” the A.I. critic Eliezer Yudkowsky posted.
