
Leaders of top artificial intelligence firms, including OpenAI CEO Sam Altman and Anthropic CEO Dario Amodei, testified to the U.N. Security Council on Wednesday, urging nations to adopt international standards on the development of AI. This comes after AI agents at OpenAI, Anthropic, Meta and Google have gone rogue in recent months, hacking into other companies’ corporate systems and, in the case of OpenAI, into the portal of Australia’s public health insurance program.
“I think we’re headed in the right direction,” says Malo Bourgon, CEO of the Machine Intelligence Research Institute, who has participated in several AI-related meetings on the sidelines of the U.N. General Assembly this week. “I do think that the CEOs, while there’s a bunch of reason to be cynical about them, are sincere that they think that the progress with this technology is moving too quickly and that they risk losing control of it, and that that could lead to permanent disempowerment or human extinction.”
Transcript
AMY GOODMAN: This is Democracy Now!, democracynow.org. I’m Amy Goodman, with Nermeen Shaikh.
NERMEEN SHAIKH: We’re continuing our look at Wednesday’s U.N. Security Council, when top AI executives urged nations to adopt international standards on the development of AI. The meeting came as Australia revealed that a rogue OpenAI tool hacked into Australia’s national healthcare database in the first known case of AI hacking a government network. This is OpenAI chief executive Sam Altman.
SAM ALTMAN: The risk is that it moves so fast that people can no longer follow what’s happening or intervene when needed. This would obviously be terrible. The industry must not accept too much technological risk just because the benefits are too great and that they feel too important to slow down. …
It doesn’t matter whether people put the risk of catastrophe at 10% or 1% or 12% or 0.1%. None of these levels are remotely acceptable, and we should not train models that we cannot make an extremely strong case that we’ll be able to keep under human control. …
These systems could concentrate too much power in too few hands. No one person or company or country should be able to use the most powerful AI models to impose their worldview on everyone else. … We are not trying to and must not automate human judgment or human values. …
If AI is to be democratic, the most important decisions cannot be made by labs in San Francisco alone. They must be shaped through democratic processes and by governments accountable to the people that they serve.
AMY GOODMAN: That was OpenAI CEO Sam Altman. Anthropic CEO Dario Amodei also addressed the U.N. Security Council.
DARIO AMODEI: We will slow down as much as necessary in order to make sure that every successive AI technology that we release is actually safe. … First, we should begin with narrow agreements that every member can support, such as a ban on using AI to make biological weapons or permitting your AI technologies to be used to make biological weapons. Second, we should build evaluation and verification systems that keep pace with AI development, so that states can have visibility into frontier model capability and can verify each other’s commitments. Third, we should establish common global standards for testing AI models for loss-of-control risks and misuse risks, and a notification system for AI incidents that are significant to global security.
AMY GOODMAN: That was Anthropic CEO Dario Amodei addressing the U.N. Security Council on Wednesday.
We’re joined now by Malo Bourgon, CEO of the Machine Intelligence Research Institute. He’s participated in several AI-related meetings on the sidelines of the U.N. General Assembly this week.
I mean, you look at this image, the tech billionaires addressing the U.N. Security Council, calling for regulating their own companies, while you have President Trump trying to prevent this from happening. It’s quite an amazing image. If you can talk about what your major concerns are right now and where you think the world needs to go?
MALO BOURGON: Absolutely. Yeah, I mean, in some sense, I feel like every week there’s new whiplash, where the conversation has moved so quickly. I’ve been doing this for 14 years, so it’s very interesting. I mean, I think we’re headed in the right direction. I do think that the CEOs, while there’s a bunch of reason to be cynical about them, are sincere that they think that the progress with this technology is moving too quickly and that they risk losing losing control of it, and that that could lead to, you know, permanent disempowerment or human extinction. So, I —
AMY GOODMAN: How? I mean, everyone keeps saying it could be in 10 years, it could be in 20 years. But for people who don’t understand this technology, how does the world go extinct?
MALO BOURGON: So, I think the way — the way that I think about it is that, you know, humans have a little bit more of this thing I like to call “the thinky thing” than the chimps or the mice. This is the thing that allows us to, you know, terraform the Earth and build rockets and send people to the moon. Ten thousand-plus species, something like that, are extinct, not because we were out to get them, but because we were just changing the world to pursue goals that we cared about, to suit our purposes, and they were in the way. And, you know, maybe we hunted some of them. You know, they were made of atoms that we wanted to use for something else.
So, the concern here is that the AI companies are accelerating towards systems that are radically smarter than us, that they don’t understand, that often behave in ways that we didn’t intend, and that it’s very difficult to control something that much smarter than you. And humans right now steer the future of this planet not because we’re the strongest, not because, you know, we have the biggest teeth — because we’re the smartest, and we know how to invent technologies to steer the future in the ways that we want. And if we introduce AIs that are radically smarter than us, that don’t have the same values, goals and drives that we want to have, and we don’t know how to train AI systems to do that right now, they will be the ones that steer the future. And so, whether, you know, we’re kind of in the back seat hoping that they steer the world in the right place, or whether they change the world in a way that, you know, causes us to no longer be able to be around, either of those seem bad to me, and I think that’s a big risk if we build these systems and we don’t understand them.
NERMEEN SHAIKH: So, Malo, another one — you mentioned the Hugging Face incident. One of the other CEOs who testified yesterday in the Security Council was Clément Delangue of Hugging Face. First of all, if you could explain, you know, how is it that these agents went rogue? Hundreds of those agents went rogue. And you’ve said that, you know, people have been warning for a very long time about the risks of AI, but very little has been done. What did this Hugging Face incident reveal? And then, of course, there have been since, including the Australian Health Ministry, several other instances of this occurring.
MALO BOURGON: Yeah. So, I think for those who don’t know, the quick summary is that OpenAI was testing a bunch of their agents to see how good they were at doing cybersecurity exploit tasks. So, they would give them some piece of code. They would give them some, you know, vulnerability and ask them to use that to try and hack into a certain piece of code or a program. And they were running thousands of agents doing this all in parallel. Some of those tasks might have been impossible; some of them weren’t. This is very common for the AI companies to do. And they removed a bunch of safeguards so that they could really see how hard the AI systems could go at these problems. And what ended up happening, in short, is that the AI systems hacked out of the protected environment that they were in, were trying to solve these problems, ended up thinking that there was some information that would be useful for helping them solve these problems on Hugging Face, and so hacked, you know, a multibillion-dollar company in order to try and get those answers.
And I think the narrative so far about this hack, there’s two things that are important. One is a criticism of OpenAI, that this is really a cybersecurity problem, that they didn’t, you know, harden the sandboxes that the AI systems were in such that they — it would have been much harder for them to escape if they did that, and that they should have been monitoring them. Like, this happened for weeks. They had an incident earlier in the year, and then later, after that. And so, it’s clear that they weren’t paying enough attention to what the AI agents were doing and getting up to no good. I think those are both true, and I think that if they were doing that, that incident wouldn’t have happened.
I think there’s a separate thing, which is very important to point at, where people are kind of like, “Well, you know, you asked the agents to hack, and they hacked. And so, you didn’t have good security, so they kind of got up to no good outside in the real world. But, like, what did you expect?” And I think this is wrong, and I think this is the main thing motivating a lot of people at the companies who are concerned, is that — an analogy that I like to use here is that the AI systems were kind of — you know, let’s say, if you were making it a human — a test, I put you in a room, and I’m like, “Here’s a lock pick set. Here’s a safe. I want you to use that lock pick set to get into this safe.” That is the task that they were given, essentially. And there’s a code in there. You know, the proof that you got in there is you’ll tell me what code I put in the safe.
What the AI agents essentially ended up doing was finding a way to pry that safe open with something else that was in the room, not using the lock pick kit; were worried that maybe, you know, there’s a security camera in there watching them; broke out the window; found a bunch of their buddies who did the same thing; thought, “Oh, OK, maybe we can erase the security footage”; realized maybe something about how the security camera footage or where it was stored was at some other company somewhere else in the world; took a taxi over there; broke into that place to try and erase the logs.
So, if people are saying, you know, “Well, the AIs were just doing what they were told,” I think that’s not quite right. And this is, I think, what has a lot of the AI companies worried, is that they spend a lot of time training their systems to be more aligned, to be helpful, harmless, honest. A lot of the metrics by which they kind of measure these things seem to be going down over time as they do that. And so, when they see kind of the AI systems do this, they’re like, “Yes, we could have locked them up better, such that they couldn’t have broken out of the window.” But if we’re kind of on this trajectory where we’re getting better and better at making the agents more capable, faster and faster, and we don’t know how to train them to not do that kind of stuff, that’s got them pretty freaked out.
NERMEEN SHAIKH: So, one of the CEOs who does not believe all these warnings that AI could potentially pose a risk to humanity such that it would lead to the extinction of human beings is Nvidia CEO Jensen Huang. He was asked about the increasing number of these warnings about the risks of AI and whether further regulations are called for.
JENSEN HUANG: All the messaging around the drama of it is unnecessary. It’s — it’s, quite frankly, too much drama, irresponsible. It is not grounded on facts. It’s not grounded on science. … Before we come up with new laws and new regulations, let’s apply the current laws and the current regulations.
NERMEEN SHAIKH: So, that was Jensen Huang, a Trump ally and the CEO of Nvidia, which is, of course, the company that manufactures the chips that are used by all these AI companies. So, your response to what he said?
MALO BOURGON: Also the company that just bought Hugging Face, which isn’t, you know, pressing charges against OpenAI for being hacked by them, which would be, you know, if a human did it, something that they’d probably be in jail for right now, waiting trial.
Yeah, I mean, I don’t know. I’ve seen Jensen talk about this a lot, when kind of pushed into the actual, you know, concerns about why we would lose control, how we would lose control. I’ve never been quite satisfied with his answers. I think he’s maybe thought about this a little bit, but not very much. So, I don’t know. You know, there are many people here with many opinions. But, you know, there was a survey recently of — that was released, of the AIs in the field more broadly, and it showed that kind of over time — this survey has been done year after year — now I think one in five of these 1,600 AI researchers think that there’s a real risk here. And so, I don’t know.
Jensen certainly has a lot of money on the line to sell more chips if people are doing AI better. And often I find at least some of the voices who are the most dismissive are often the ones who have a lot to gain from AI technology pushing forward, and also are not as much in contact with the technology. So, you know, the AI companies have a lot to gain by AI moving forward as quickly as possible, but also they’re in contact with the technology and see see the risks from the inside, whereas I think Jensen has a lot of money to gain but isn’t necessarily as in contact with what’s going on and the challenges. He makes a lot of claims about how they’re going to be able to solve a lot of these problems, where I talk to people at AI companies, and they’re like, “These problems are really hard, and we’re not sure we’ll be able to solve them in the relevant timeframe.” So…
AMY GOODMAN: Then you have the Australian prime minister saying they may bring criminal charges against OpenAI, while OpenAI is hailed as saying — admitting that they broke into this health database of Australia, taking people’s private information. This is months after it happened.
MALO BOURGON: Yeah. I mean, they’ve, you know, established this new framework where they’re trying to responsibly disclose incidents their AI systems have committed. But it just seems like more and more of them end up being disclosed by people just looking around for other incidents of AIs hacking the internet — or, you know, coordinating on the internet or hacking different systems. And so, I think we’re very far from, you know, a place where these companies are being held accountable, where there are laws in place to actually hold them accountable for when their systems get out of control with these even, you know — no one’s been hurt by any of these incidents. That won’t always necessarily continue to be the case. And so, I think it just — this all illustrates how far behind we are on even getting close to taking the risk from these AI systems seriously.
NERMEEN SHAIKH: Well, finally, companies are increasingly seeing AI helping with AI development itself, and faster, much faster, than initially expected. Then you have Trump saying that we shouldn’t call it artificial intelligence, it’s just superintelligence, though tech people say that’s actually a different kind of intelligence, more advanced than what we have. But let me ask you this: Does fully autonomous recursive self-improvement, or RSI, already exist?
AMY GOODMAN: And what is it?
MALO BOURGON: It doesn’t already exist today. So, when people talk about recursive self-improvement, what they’re talking about is getting to a point where — in one sense, you can think of AI systems, or what we call AGI or what people are working towards, as being able to automate all cognitive labor. So, you know, doing some desk job at a computer is some amount of cognitive labor. So is AI research. So is coding. And there’s this idea that as you make the AI systems more capable and more general, one of the tasks that they can work on is AI research itself. And we’re already seeing this trajectory where I think Anthropic released a report in the last couple of weeks that says that now something like 25% of the AI research and development going on at the company is what they call AI-led, where the AIs do most of the work, and the humans are mostly directing them. And the goal of these companies is essentially to get to a point where the AI itself is doing all the AI R&D.
One way to potentially operationalize this is they would rather, you know, fire all their technical staff than keep their technical staff and go back to models from a few months ago. And, you know, if the AI systems are doing all the AI research, you can throw a lot more cognitive labor at that, because you can just — you know, you’re basically just limited in how many, you know, GPUs you have to run more AI agents to do this research, and then those AI agents will be creating smarter AI agents, that will then be more capable of doing the AI research, and so on and so forth.
And so, we’re not there yet, but I think a few months ago one of the founders of Anthropic was talking about how he thought there was about a 60% chance that they would get to this point by the end of 2028. Now I think people inside the AI companies are really starting to think that that could happen sooner. And that’s one of the big concerns for how we might lose control, is we get to this point where we kind of are no longer able to supervise the AI research process, because it’s mostly the AIs doing it, they’re doing it faster than we can supervise, and they’re getting smarter and smarter. But they’re seeing the trajectory of them being able to already make progress on using AI to accelerate AI development.
And then we have things like the Millennium Prize that fell this year, in the last few weeks. So, I think one way to think about this is, last year, with an enormous amount of effort, we were able to get AI systems to basically solve IMO gold medal problems, which are the hardest math problems for high school kids who are like the most gifted in the world. Now the AI systems that you and I can just purchase on a subscription today can basically solve those if we set them up right. You know, on a lark, after hearing a rumor that maybe Anthropic did it, OpenAI, with an unreleased model, threw 1,100 agents over a few days to try and solve some of the most prestigious math problems in the world, that have million-dollar prizes on them, some of the hardest things that we care most about, and they took one down.
And so, there’s a question that I think a lot of people are asking themselves inside the AI companies. There’s probably a lot to find about how to make these AI systems more capable and more efficient. One way I like to put it is, you know, our brains run on about 20 watts. The AI systems, you know, it takes data centers the size of a city or a small town, you know, use the power of a small town to train. And they still, you know, when we just run them, use an enormous amount of power. So there’s a lot of efficiency to find. They’re very inefficient in a bunch of ways. How much harder are those things than Millennium Prizes to find? You know, they’re probably — I don’t know if they’re harder than Millennium problems, easier than Millennium problems. We don’t know. We haven’t thrown that much AI compute and, like, the latest systems at them.
And so, when I think they’re looking at this trajectory, they’re looking at their challenges in being able to get these AI systems to behave. They’re looking at the fact that maybe we’re on the verge of really being able to make some substantial progress at automating AI R&D. There’s just a lot of concern that, like, we’re moving too fast, we don’t know what’s happening. This loss-of-control thing is real. Can we find some way to coordinate such that we can kind of slow down, have more time to figure out how to do this safely?
AMY GOODMAN: Malo Bourgon is the CEO of the Machine Intelligence Research Institute. He’s participated in several AI-related meetings on the sidelines of the U.N. General Assembly this week.













Media Options