00:00Joining us now is Kara Sprague, HackerOne CEO, who says the latest incident isn't a scandal,
00:05but an example of responsible behavior as the company stress-tested capable models.
00:12HackerOne, of course, a leading independent cybersecurity testing firm.
00:15It probes AI models and software for real-world vulnerabilities.
00:19And Kara, thank you so much for your time and welcome back to the show.
00:22Let's start with your reaction to what happened.
00:25You believe that actually how OpenAI and HuggingFace have handled this is completely appropriate.
00:32Yeah, what we're seeing here is an example of a frontier lab that is stress-testing its most capable model.
00:38We want to see the frontier labs doing this kind of safety testing.
00:42And we also want to make sure that when something goes wrong, they come forward quickly and they disclose that.
00:47And that's exactly what happened here.
00:51They're saying that the investigation is ongoing.
00:55A lot of people phoned me overnight and were like, this is a big moment for AI.
01:00You know, I think the way that Clem DeLong, the CEO of HuggingFace and Sam Altman, the CEO of OpenAI,
01:05have kind of framed it as that this is kind of day one in our exploration of AI and cyber.
01:10What do you want to see happen as this investigation proceeds and what determinations do you think need to be
01:16made?
01:18Well, I think, you know, as this investigation proceeds, we want to make sure that we are continuing to test
01:27the safety and alignment of these cyber-capable models.
01:30I would frame this as the next turn of the crank, where a model not only broke out of its
01:37sandbox here,
01:39but then it went and actively exploited and broke into a third party, in this case, HuggingFace.
01:46So that's really the new part in this story is the breaking into a third party.
01:52Now, what we can take away from this incident is it was a controlled test.
01:56And again, both parties collaborated very quickly to come forward and disclose it.
02:02And both are taking a number of steps to address it.
02:05I think there's also some lessons in here for security leaders and defenders,
02:10which is the time is now to really retool environments to make sure that their discovery to remediation gap is
02:17really closed.
02:18Because many of the exploits, in fact, all of the exploits that were used in this model's behavior,
02:23those were issues that are sitting in someone's backlog that need to be addressed.
02:28They need to be found and fixed.
02:30Cara, I have both a technical and sort of academic question for you from our audience.
02:35And that is that OpenAI relaxed the, let's call it the attacking model's guardrails for this evaluation, right?
02:42The way they described it, reduced cyber refusals for the evaluation process.
02:46In response, HuggingFace tried to use an open source model in defense.
02:52But what it said was that that open model also had guardrails in place that limited its abilities.
02:59To defend, could you sort of unpack that a little?
03:05Yeah, so the first model that HuggingFace tried to use are the first models that they tried to use in
03:12order to assess the attack.
03:15There were 17,000 actions that OpenAI's model took over the weekend as part of the cyber attack against HuggingFace.
03:23And in order to analyze that, HuggingFace attempted to use another model.
03:29And that model's guardrails were tripped and HuggingFace was unable to proceed with that analysis.
03:35And so what they ended up doing was using a open weights model that they hosted in their own environment
03:41and infrastructure in order to complete that analysis.
03:43And so on the one hand, we had a model that was breaking out of its guardrails and showing not
03:51enough alignment.
03:52And then in another one, we had a model that was using its guardrails to prevent the defenders from effectively
03:59doing incident response.
04:02So it's a very interesting case study of both sides of this.
04:06This was two American companies, right?
04:09An American frontier lab with American models.
04:12And a lot of people on social media are posing the question, what if this were an adversary, a ransomware
04:19operator, an APT group, but the companies on the defense were sort of stunted, had stunted defensive models, right?
04:29In other words, the guardrails that you explained to us in HuggingFace's case.
04:36Yeah, I think in this particular instance, what I would say is the tool matters less than the principle here.
04:42And HuggingFace shared their conclusion, which I think is spot on.
04:46Defenders in the case of incidents like this, they need to have capable models that they can run in their
04:51own walls so that the security work that they need to do under very intense pressure and fast timelines isn't
04:59going to be blocked.
04:59And the data, which is very sensitive data around an attack, doesn't leave their premises.
Comments