Skip to playerSkip to main content
  • 7 weeks ago
Transcript
00:00Campbell, first of all, great to have you. And I want to draw on your journalistic background
00:03because this is near and dear to us. Anytime we do anything as journalists, accuracy is the most
00:08important thing. And as we start to use more of these tools, you know, there's always that
00:13asterisk as to can I trust everything we're getting out of it? Now, you went out there
00:17and you really analyzed this mainly through the perspective of the election. What did you actually
00:22find? So we looked at politics and elections in particular, Romaine, because it's one of
00:29the hardest areas that you can evaluate or test for because so many issues around politics in
00:37general are very nuanced, very complex, subjective. There's not always a yes or no answer. And across
00:45the board that we found, and I'll tell you, the categories we measured in are just factual
00:50accuracy, the quality of the sources that the different chatbots used, and neutrality, whether
00:56they were leaning one way or another on political bias. And on those three categories across the
01:02board, there was a 90% failure rate, meaning all of the chatbots failed in at least one of these
01:08categories. That's not great. And we have to fix it. There is a trust problem generally right now with
01:17AI. We're sort of that's creating, I think, a bottleneck in terms of adoption, not only at the
01:22consumer level, but also with businesses. And it is in order to address it is going to require AI
01:30companies and tech companies to open themselves up for scrutiny in a way that they really haven't
01:34before. Today, they're grading their own homework. There is no accountability. There's no evaluation
01:40system in place or an ecosystem of evaluation that tests these models and can give you a real sense of
01:48how they perform. And that's what we're advocating for. It's obviously what my company does, but there
01:53needs to be a lot more of it. Well, that's what I'm curious about. I mean, what is the solution?
01:57Because I mean, a lot of these companies say, look, I mean, we're just retrieving information
02:00that's out there. We put it together to some degree or another. There might be, you know, one mistake
02:05here or maybe some nuance that we didn't pick up. But, you know, there's only so much that the
02:09computers themselves can do without the human beings actually making a judgment call.
02:13Except that, Romain, they're selling these products to us as a tool that is essentially
02:20going to be the, you know, the infrastructure, which are all of our future work is built on top
02:26of. They're selling to businesses to use in all kinds of high stakes use cases. And that's mostly
02:33what we focus on at Forum AI is evaluating high stakes use cases. Geopolitics is one. Financial,
02:40advisory, anything in fintech and in financial services. In healthcare, people are using these
02:48chatbots to ask medical questions. And if you think a business is going to accept, oh, it's okay,
02:55we make mistakes every now and then, when they're using these tools to make crediting lending decisions,
03:00to talk to customers directly, to give financial advice, that's insane. They're not going to accept
03:06that. And I think we've sort of reached a point from the enterprises that we work with and that
03:11we've talked to where there's been this huge push to spend and adopt AI, deploy as quickly as you can
03:19so that you don't fall behind the competition. And we're starting to hear from compliance departments
03:24and CEOs and boards who are saying, wait a minute, can we be sure that this is a accurate, safe?
03:33And we
03:34can't be. I mean, you just saw what happened in the news over the last 24 hours. Open AI decided
03:38to hold
03:38back on the release of its model at the request of the government, because there is concern about
03:44whether it can be safe. What's the safety regimen or testing regimen rather around that, that,
03:49you know, opens, again, opens themselves up or the companies voluntarily opening themselves up to
03:56more testing so we can better understand what's going on. Well, Campbell, I'm glad you brought that
04:01up because I wanted your opinion exactly on that. Chat GBT, the GBT 5.6 model series, as you said,
04:08Open AI holding that back. That is at the request of the Trump administration. We know that we're
04:14waiting for the effects of that executive order that was signed by President Donald Trump to,
04:19you know, maybe see more instances of this. And I wonder, you know, what you make, first of all,
04:24of this push by the government and whether or not you think it's appropriate.
04:29So I think on issues related to security, it, of course, it's appropriate. What's hard for us to
04:37know is what's really going on. I think, you know, we're not getting that much information out of the
04:42administration or out of the AI companies as to what the challenges may be, who brought them to
04:50the attention of Open AI. Did Open AI bring them to the attention of the administration? That's not
04:55clear to us because there actually is no sort of standardized way that we approach this today.
05:01And that is around security, things that are deeply concerning. So, you know, forget about something
05:07like, you know, the quality of the information around elections or politics or anything like
05:11that. That's obviously going to be secondary. But until there is a, whether it's government
05:18regulation, and I think in the area of security, even though there's no formal regulation at this
05:24point, you will see government stepping in when they have legitimate security concerns.
05:29My hope is, is that we don't need regulation when it comes to things like that, that we're evaluating
05:35around is that the industry itself will take the step necessary to do it voluntarily. And there's
05:42certainly examples of this. You look at, you know, crash tests for cars. That's funded by the insurance
05:47institute and insurance companies. And, you know, because they want to get the information out there.
05:54Consumers want to know the information before they make a purchase.
05:56Consumers. This does not require government regulation to do. But I think if you want to
06:02establish trust in AI and ensure adoption beyond where we are today, you have to figure out how to
06:09make those products more trustworthy. Well, I'm curious, too. I mean, if the solution sort of lies
06:14more in the companies or I guess the institutions, if you will, pushing back a little bit harder on this
06:19and demanding that accuracy. I mean, do we have to wait till we get to the breaking point? And I
06:24don't want
06:25to harp too much on your previous employer, but there are a lot of people that would look at
06:28some of the benefits and the promises that we saw on social media in the early days and the slop
06:34that a lot of it has become today. And I think there is a worry that for all the promises
06:38we have
06:38with AI, that we might end up with a similar slop somewhere down the road if we don't find a
06:44way to
06:44arrest this. So, Romaine, this is actually a great point. And I it's one of the reasons I'm very
06:51optimistic. If you think about social media, and this was always one of our challenges at Meta, social media
06:56is optimizing for engagement and you can't optimize for engagement and get accuracy at the same time. You
07:02know, what do people engage with the most hyperbolic content out there? But if you are a company, you're a
07:08big
07:08bank, you are an insurance company, you are, you know, a hospital, you are going to demand that AI optimizes
07:16for
07:17accuracy. You're not going to say it's OK to optimize for engagement. And so there there is a different
07:22incentive here that I think pushes us potentially in the right direction longer term. Right now, these
07:29companies, their businesses are focused on enterprise. That is where they're making their money. And if
07:35those companies are going to adopt these tools, they are going to demand more accuracy and more clarity
07:41around around how the process, the transparency around the process and how they're being evaluated in order to
07:50continue deploying beyond where they are today or in any area where there's real risk. So far, what we've seen
07:57the
07:57courts do is demonstrate that companies who deploy AI bear the responsibility for what that AI does. And if you
08:05are
08:05taking on that risk, potential litigation, then you're going to demand the AI companies do something more than they're doing
08:11today. And so I do think because those incentives are different that we do have an opportunity here
08:16potentially to get closer to accuracy being being the goal than we did with social media.
Comments

Recommended