Skip to playerSkip to main content
  • 10 hours ago
Transcript
00:00OpenAI is disclosing what it calls an unprecedented AI security incident. The company says it
00:06was testing GPT 5.6 Sol and an even more capable but unreleased model with reduced safety guard
00:14rails to measure their cyber capabilities. During this evaluation process, OpenAI says the models
00:21found a way out of their sandbox testing environment, obtained internet access,
00:26and then targeted AI platform HuggingFace, searching for secret information the models could use
00:32to cheat the evaluation. OpenAI says the models ultimately reached HuggingFace's production
00:37systems. The investigation is still ongoing, but it's already raising new questions about what
00:43today's frontier AI models are capable of. Let's bring in Bloomberg's Rachel Metz, who broke the
00:48story. And I think the best place to start, Rachel, is how did this happen?
00:53Well, this happened while OpenAI was testing some model cybersecurity capabilities. It was
00:59actually essentially giving it a test, like problems on a test, like a person would conduct a test. And
01:06in the process of seeking answers to the test, the model did a series, or the models, it was several
01:13of
01:13them, did a series of things in order to get access to the internet and go over to HuggingFace's servers
01:20to
01:20try to find answers to the test. What have OpenAI and HuggingFace done in response to this? Like they've
01:28said the investigation's ongoing, but if this is as significant a security incident as the companies
01:35have portrayed it as, and as the internet is discussing it as, they must have taken some action in the
01:40interim.
01:41Yeah, they've done several things already. So they are, OpenAI is implementing some controls. It says,
01:50even though it might slow things down with their own work in order to patch any vulnerabilities,
01:56it's also working with HuggingFace to go through everything that happened here. There was a third
02:03party company whose software the AI models found a vulnerability in order to get access to the internet.
02:09So OpenAI did what you're supposed to do with these kinds of things, and that is let that vendor
02:14know with a certain amount of time so that they could patch their own software. And one other thing
02:19that they did is OpenAI several months back, similar to what Anthropoc has done with its most capable
02:25cybersecurity models, is they have a trusted access program. And so that gives certain vendors and
02:30researchers access to these less restrictive cybersecurity aimed models. And so they brought HuggingFace into this
02:37program so that a HuggingFace can also be more aware of what's happening with those models.
02:42I ran into you on my way out yesterday evening. And I think the kind of moment was what is
02:48happening?
02:49You know, it's hard to understand in the moment how severe something is, what happened. But I mean,
02:55do you have a sense overnight, you know, how seriously the companies are treating this, what the sort of
03:00sentiment of industry is towards like what this represents? It's a worrying development to a lot
03:06of people in the world of AI.
03:08Yeah, I mean, I guess I think about it kind of two ways. One is, this is not good if
03:14you have AI models
03:15that are, you know, breaching the containment you set for them and then going off and doing things
03:20and hacking essentially, like accidentally hacking into another company's servers like that. I think we can
03:26agree that that's problematic. But on the other hand, the model did, I would argue, essentially what it was
03:33tasked with doing. It was told to solve a series of problems related to an evaluation process called exploit
03:40gym. And the model took these steps. It just wasn't the steps that the researchers were expecting it to take,
03:47or
03:47perhaps might have wanted it to take.
Comments

Recommended