Skip to playerSkip to main content
  • 11 minutes ago

Category

🗞
News
Transcript
00:00Ryan Krishnan is the founder and CEO of Vows AI. Vows independently tests and benchmarks AI models,
00:05including unreleased models from OpenAI and Anthropic. He says embedded evaluators will
00:11promote an even playing field and public trust, and he joins us now. You had a response to Dario
00:17Amadei's essay. Let's start with that response. What was your conclusion? Yeah, well, I think,
00:23I mean, to be honest, I think there's been a lot of extreme statements being made over the last
00:26week, and I can understand where they're coming from, but I think they can also cause undue
00:31anxiety for the general public. And so really, I think in response to this, I found reason to be
00:36optimistic. From what I see at Vows AI, boots on the ground, we've actually been able to see the labs
00:42be very forthcoming, and we engage across all the major foundation model labs, as well as with members
00:46of Congress and government agencies and some of the largest enterprises. And so I actually think
00:50there's a lot of conflict happening on the public stage, but on the whole, we have reason to be
00:55optimistic that a coordinated effort is possible. So let's talk about what Vows sees, right? If we
00:59were trying to really simplify what Dario Amadei and others are calling for, it is to slow down
01:06capability development, the pace of capability development, allowing safety alignment to catch
01:12up. You see these models in the training phase and before they're released? Yeah, well, I think what
01:18we see is that there's really been an outsized investment towards the generation or the capabilities of
01:24the models without that same parallel investment in the testing and evaluation. And so our ability
01:28to create better models now outpaces our ability to actually understand them. That's the issue at
01:32hand. A lot of people are asking, why now? And what Dario Amadei wrote is about the ability of AI
01:40to aid in the research and development of future generations of AI, so-called recursive improvement.
01:46That was a very clear piece of the essay. It is something that you look at at Vows through the
01:52RSI index. So I think we're going to bring up that index. Explain what we're looking at, the
01:58methodology, and what that index tells you about an AI's ability to contribute to the development of
02:04AI. There it is. Right. Well, so this is, yeah, this is the mainline figure tracking the progress of the
02:10major frontier models towards RSI, recursive self-improvement. And that describes the model's ability
02:15to make the successor version of itself autonomously, so wholly without human input. And that's, I think,
02:21to contrast today where models are used as an active part, but humans are still involved.
02:26And so this would be akin to GPT-6 being responsible for making GPT-7 or CLOD making the next
02:31version of
02:32itself without any human researchers. And so what you see in the plot is that we have some early signs
02:37that models are exhibiting RSI, but they're still far away from the human frontier. And in extrapolating
02:42out these trends, we're actually making the prediction that we'll see models eclipse human
02:46researchers probably around August of 2027. August of 2027. So the question right now, like loads of
02:52people pose this, is can an AI model act like an AI researcher right now? Can it run experiments? Can
02:58it
02:58improve AI systems with little or no human help? And right now the answer is no, but you have line
03:06of sight to
03:06when the answer would be yes. To be clear, we're testing a set of proxies and the publicly available
03:11models. And so the indication that we see is that the models are actually very good at executing
03:15experiments and operating as engineers, but they lack kind of the intuition to develop new experiments
03:20and have fundamental breakthroughs. But I think it's also important to add that we're not testing any
03:24unreleased model systems. We're not testing specialized agents or even multi-agent systems. And so I think the
03:30role of event evaluators over time could be to get closer to these internal systems. Again, the kind
03:36of concrete proposal from Anthropic is to have independent evaluation, but inside the company. It
03:44sounds very close to what you're already doing with that company. This would create a demand for
03:49your services. So how would that work? Anthropic would pay you. How would that impact your independence
03:55to make an assessment of a yet unreleased model? Yeah, I think on the whole, it's going to take a
04:00diversity of approaches. So it'd be a failure to say that this will be the end-all be-all.
04:04But I think there's a historical precedent when you look at things like ratings agencies or auditing
04:08firms who have been for-profit companies and have been responsible for testing and valuation in
04:13other industries. I think there's natural conflicts that arise. And so the main one that we want to
04:18avoid is a situation where the same group auditing is also responsible for remediating or fixing.
04:23And then you end up in situations like Enron. And so we've set the line very clearly where,
04:27although we make our methodologies public, we have test sets that are kept private.
04:31And we also ensure that we don't sell training data or the solution to labs while we do the
04:35testing process. I guess to end our conversation, the obvious next question is what happens next?
04:41You know, people have taken different approaches, but it seems that this is only going to work
04:46if all of the frontier labs are able to cooperate with one another to set a standard. You work with
04:52them. Do you see that happening?
04:55Yeah, honestly, I mean, I have a lot of reasons to be optimistic. I think that people at these labs
04:59are rational actors. And so they see that the way things are headed, it will take some coordinated
05:04effort. And so I think with or without regulation, I think market-based solutions will have a part to play.
Comments

Recommended