Today I’m sharing my interview on Robert Wright’s Nonzero Podcast from last May. Rob is an especially sharp interviewer who doesn't just nod along, he had great probing questions for me.
This interview happened right after Ilya Sutskever and Jan Leike resigned from OpenAI in May 2024, continuing a pattern that goes back to Dario Amodei leaving to start Anthropic. These aren't fringe doomers; these are the people hired specifically to solve the safety problem, and they keep concluding it's not solvable at the current pace.
00:00:00 - Liron’s preface
00:02:10 - Robert Wright introduces Liron
00:04:02 - PauseAI protests at OpenAI headquarters
00:05:15 - OpenAI resignations (Ilya Sutskever, Jan Leike, Dario Amodei, Paul Christiano, Daniel Kokotajlo)
00:15:30 - P vs NP problem as analogy for AI alignment difficulty
00:22:31 - AI pause movement and protest turnout
00:29:02 - Defining AI doom and sci-fi scenarios
00:32:05 - What’s My P(Doom)™
00:35:18 - Fast vs slow AI takeoff and Sam Altman's position
00:42:33 - Paperclip thought experiment and instrumental convergence explanation
00:54:40 - Concrete examples of AI power-seeking behavior (business assistant scenario)
01:00:58 - GPT-4 TaskRabbit deception example and AI reasoning capabilities
01:09:00 - AI alignment challenges and human values discussion
01:17:33 - Wrap-up and transition to premium subscriber content
Show Notes
This episode on Rob’s Nonzero Newsletter. You can subscribe for premium access to the last 1 hour of our discussion! — https://www.nonzero.org/p/in-defense-of-ai-doomerism-robert
This episode on Rob’s YouTube — https://www.youtube.com/watch?v=VihA_-8kBNg
PauseAI — https://pauseai.info
PauseAI US — http://pauseai-us.org
Transcript
Liron’s Preface
Liron Shapira: Hey there, Doom Debates listeners. Today I want to share Robert Wright's interview with me in May of last year on his show, the Nonzero podcast, which is really good. By the way, I'm a premium subscriber. I highly recommend going and checking out all of his stuff.
He has a lot of great interviews, so he did one with me last year. He was interested in having somebody come on and explain AI doom in plain language and try to get concrete about a doom scenario. He asked a lot of really good questions.
I don't know if I a hundred percent answered everything to his satisfaction, but it was definitely a great discussion. I think it's worth a listen. I think it's aged reasonably well, although fortunately, GPT-5 came out and it wasn't this scary monster that ended the world.
Fortunately we have more time than that. I don't think that was a foregone conclusion, but we'll take what we can get.
Back in May of 2024 when this was recorded, this was actually right when Ilya Sutskever and Jan Leike, who were the two heads of OpenAI's safety team, they both resigned or were pushed out or whatever it was, they were gone.
And that was a few months after the board coup when the board tried to push out Sam Altman. So it was a time of relatively higher than average OpenAI drama. Rob and I talk about that.
When you get to the end of the episode, Rob mentions that there's an hour of premium content. You can get that by subscribing to Rob's podcast, following the link in the show notes to see the full episode on Rob's website.
Can you afford to do that? I don't know. Do you have $6? That's how much it costs to subscribe for a month or $60 for a year. I think it's a good deal. I hope we can get them some new subscribers from this episode.
Now, before we get into that, you might also be wondering, wait a minute, what about Vitalik Buterin? Isn't he coming on the show? Yes. That episode has been recorded.
We're editing it and we're now on pace to have that for you next week. In the meantime, let me give you a little bit of a taste to wet your appetite. This is the Vitalik teaser trailer.
Vitalik Buterin: I'm Vitalik Buterin and you're watching Doom Debates.
Liron: All right. Without further ado, here's me on Robert Wright's Nonzero Podcast.
Recent OpenAI Departures and Safety Concerns
Robert Wright: Hello, Liron.
Liron: Hey, Bob. Good to be back with you.
Robert: Good to have you. Let me introduce this. I'm Robert Wright, publisher of the Nonzero Newsletter. This is the Nonzero podcast.
You are Liron Shapira, you are a Silicon Valley guy who's founded or co-founded a couple of companies, but more to the point for present purposes, you are an AI safety activist and in fact, an AI pause activist.
You're sufficiently concerned about AI to want us to kind of pause or slow development and take a deep breath and think about this whole thing.
Now, this is an unusual kind of hybrid episode of this podcast. So let me explain. You and I taped a long conversation a week or so ago about why we should be concerned about AI.
In your view, the risks it could pose. In particular, the kind of sci-fi doom scenario of it actually taking over the planet and killing us or something, something you take seriously.
I wanted to try to trace out the logic behind your argument. And if people want to listen to that, all they have to do is keep listening to this because this is basically a preface to that.
What happened is I didn't have to post it super fast and it wasn't urgently topical. Then a couple of things happened that made me think, well, it actually was quite topical.
And also that it might be worth talking to you a little bit about these developments. So one is, I was on Twitter and I saw some video of you on some kind of soapbox or something with a megaphone.
Apparently there was this international AI pause protest. You were at the San Francisco edition of it. Were you at OpenAI headquarters or what?
Liron: Yep. OpenAI headquarters in San Francisco.
Robert: Outside of OpenAI headquarters. Needless to say, you are not embraced by the people at OpenAI. They did not invite you in for coffee.
Liron: Nobody invited us in. I don't know what happened. Yeah, that was very rude.
Robert: Yeah, it's surprising. And so there's that. I want to ask you a little more about the pause movement, but also, and very intriguingly to me, there have been these resignations at OpenAI.
At least a couple of which seem pretty clearly related to the whole safety issue, one of which may or may not be. And then before that, there were a couple of firings that seemed to be related to the safety issue.
Yesterday what happened is, and as we're taping this, that is to say, I think it was Tuesday, right? I mean, first of all, Ilya Sutskever resigned.
He's the chief scientist and that didn't shock a lot of people because his status had been uncertain. He had been one of the board members who was behind the attempt to oust Sam Altman.
And that may, so his resignation could be related to AI safety because by some accounts, that whole ouster attempt was about concerns that Sam Altman was not as safety conscious as he should be, not as concerned about AI risk as he should be, was proceeding headlong without due concern for the risk.
But then also yesterday, this guy, Jan Leike, who is head of super alignment at OpenAI, very prominent person. I guess they're kind of lead person nominally on risk so far as I could tell.
He resigned. And whereas Sutskever, when he resigned yesterday, he did this "I've loved OpenAI, I loved the people." Sam Altman did a tweet, "I love Ilya blah blah, blah."
Jan Leike just tweeted, "I resigned, period." In fact, not even period. That's how concise it was. He didn't put a period on the end. It was as short as it could be.
And so a lot of people are wondering what's going on. And then there were some, there was, maybe you can fill us in on the prior resignations and firings and so on.
Historical Pattern of Safety Expert Resignations
Liron: Sure. Yeah. So OpenAI has now built up a consistent track record where their top safety experts keep jumping ship.
So the history goes all the way back to 2021, where famously Dario Amodei and his sister Daniela Amodei, who were very high up at OpenAI. Dario was actually the engineering manager for GPT-2 and GPT-3.
And Daniela had a number of roles including managing safety. They were the first to prominently jump ship over safety concerns where they said, "Hey, these models really need to be built with a bigger focus on safety."
And we can't really get that here at OpenAI. So our only solution is to go work independently on a more safety focused AI lab, which has become Anthropic.
And notably Dario said in an interview that his P(Doom) probability of doom is 10 to 25%.
So this is the first prominent incident of somebody who's supposed to be working on safety at OpenAI jumping ship saying, "Hey guys, there's a major doom risk here, and I don't think OpenAI is going to be the place to solve it."
Robert: So he thinks the chances of either extinction or something so catastrophic happening to the human species as a result of AI, that it's not too far from extinction are 10 to 25%. I don't know if he put a timeframe on that.
I mean, your P(Doom) is 50% by 2040. But we cover that in the next conversation so people will see if they stick around. So, sorry, go ahead.
Liron: Yeah. So then, later in 2021, you have the resignation of Dr. Paul Christiano, who is known for being one of the inventors of RLHF, maybe the most prominent inventor, actually, I'm not sure.
Robert: That's reinforcement learning through human feedback, which is what, when they put kind of the guardrails on among other fine tuning after the training on just text.
Liron: So Paul Christiano resigned to start his own Alignment research center, an independent nonprofit. So presumably he just figured once again, OpenAI is not the best place to do the safety work, even though they claim to be simultaneously working on safety and capabilities.
And Christiano stated in an interview that his P(Doom) is 50%.
Robert: Did he make any complaint? Did he say he was resigning because OpenAI wasn't taking it seriously enough? Whether or not Dario said that when he left, that's been pretty clearly established. I think that was his concern. Did Christiano say anything like that?
Liron: He's never made any public statement about OpenAI, which seems to be par for the course. So my interpretation about what's happening is you have these people in-house at OpenAI who are their top safety experts who are looking at the problem and they're freaking out.
I mean, they're explicitly going on record saying, "I have a high probability that the world is going to end." There's actually another Paul quote that he said later, after he resigned.
He says, "I think that the number one reason why I would be killed would be from AI."
Robert: He didn't say from OpenAI, but okay.
Liron: No, no, no. From AI. Exactly. And that's actually what I'm saying is it's not that they think OpenAI is Satan. It's more they're looking at the problem as a whole.
They're saying there's a high doom risk. I'm freaking out and OpenAI is just not helping. So they're not necessarily the villain, right?
I mean, I would argue Meta and Yann LeCun and the push for open source is a bigger villain than OpenAI. But the problem is you've got an arms race dynamic and the people who are inside OpenAI, who are hired to be their top experts, to have a handle on the safety problem are looking around saying, "We're not helping."
Robert: So then, and then in between the two departures you just mentioned and the ones that happened this week, there was a guy named Daniel Kokotajlo and he was somewhere on the safety team and he resigned. What was that? A few weeks ago. And pretty clearly said, or strongly implied. Yeah, this is about my dissatisfactions or something. Right, didn't he?
Liron: Exactly. It was very stark. Yeah. So a couple months ago, Daniel Kokotajlo, I hope I'm getting that right, he resigned and he posted on the LessWrong forum, very explicitly.
"The reason I resigned is because I have lost confidence that OpenAI would behave responsibly around the time of AGI," his words. And he also stated that his P(Doom) is 70%.
Robert: So it could be, I mean, if I were doing PR for OpenAI and it's a job somebody's got to do on this podcast, I don't think I can count on you to do it.
I'd say, well these are people with super high P(Doom)s. Obviously they're a hard crowd to please. There's just no major AI companies that they're going to think are doing a good enough job of this.
If your P(Doom) is that high, you probably don't think there should be major AI companies. Isn't that the kind of pushback you get?
Liron: It's kind of the case. And that's totally fair and that's some pushback that I've heard on Twitter of people saying, "Of course you're selecting for people who are interested in working on safety, of course they have a high P(Doom)."
But the thing that's insane about all this is you have these AI labs who are framing the problem like, "Look, we have a reasonable safety standard. We're working on something that, yes, there's risk, but we're managing our way through."
But you look at the people internal to them, their experts, the in-house experts are telling them, "Hey, this is not a problem that we're on track to solve."
"I know you guys are racing ahead, you're racing the other AI labs, but the outcome that you're racing toward is probably doom." And then they leave.
So something is not sane about the situation. And I think the key to restore sanity is to put on the table the concept that the problem is intractable.
That it's just not a problem that humanity is going to solve in the next 10 years. We might need 50 years, I think I mentioned in the other part of the interview, similar to the P versus NP research problem, which is now a more than 50 year research problem.
We keep chipping away at it. I would submit that based on these...
Robert: Did we talk about this? What is the P versus NP?
Liron: It's worth summarizing really quickly anyway. So P versus NP is one of the millennium prize problems. You get a million dollars if you solve it. It's been open since I think 1971, something like that. So we just crossed the 50 year mark.
And it gets very technical, but it's the idea from computational complexity theory that problems where you have to search an exponential space to get to a good answer are harder than problems where you just have to verify a good answer.
Which is a mouthful and it's hard to make it simpler than that. To make it intuitive.
Robert: No, I am not entering that contest. I do not purport that solution. I'm not going to.
Liron: To just give a little drop of intuition. It's something like, "Hey, the problem of searching for a mathematical proof of a theorem is fundamentally harder than the problem of checking over the proof and being like, yep, that's a proof."
So there's intuition, it seems like an obviously true fact. If you believe that fact, then you believe that P doesn't equal NP. It seems very intuitively obvious.
And yet the challenge of proving this fact mathematically with the necessary level of rigor has now evaded computer science for 50 years.
Robert: Okay, so anyway, where I interrupted you, where were you going with that?
Liron: So this really is a tangent, but the thing that I think is relevant about the P versus NP problem is it's an example of a research problem that's very well specified.
Doesn't seem it should necessarily be that hard, and yet we've been picking at it, we've been making incremental insights for 50 years and it might take another 50.
Robert: Are you saying figuring out what P(Doom) really is and what to do about it is comparable to that problem?
Liron: What's comparable to proving that P doesn't equal NP would be having a framework of how to control and align a super intelligent AI.
Robert: Yeah, I would think that the two problems aren't even comparable in the sense that I would think one is in the realm of rigorous mathematics and one perhaps, unfortunately for our species, it's just not in that realm.
It's a lot of social problems. It's just not in the realm of formal logic. And just a quick tangent, but it does seem to me one reason so many people have trouble wrapping their minds around the sci-fi doom scenarios.
Especially as articulated by Eliezer Yudkowsky is he talks about them as if they were formal logical problems. So first of all, most people, that means he says things that a lot of people don't even understand.
But also that doesn't mesh with the intuition of a lot of people. They don't think, no, it's a real world problem. I want concrete examples explain how this would play.
And by the way, I want to add something I meant to, I wished I had added to our conversation. Which was, what you guys need is to either point us to a novel or movie that really shows how this could play out plausibly.
Where you start here and AIs wind up running the planet. Or if such a thing doesn't exist, write the novel yourself.
But you have to get from here to there. You can't just jump to the world where they're running this. You can't jump to the Matrix. You have to get us from here to the Matrix.
Liron: It's a very good point. Yeah. I mean, I think the kind of scenarios that we could start with are just something like, "Hey, the AI takes down the internet and starts manipulating people," right?
So somebody could just make a movie of a bot manipulated a bunch of people. And the bot became more powerful than a nation state, just because all of these people were just working for the bot the same way that they would work for a political movement or an ideological movement. They're just working for the bot.
Robert: You know, I didn't plan to go here, but speaking of this, it's if some group of people with your concerns can't get together with some group of creatives and convince Hollywood.
To do some kind of streaming series about this. There's something wrong with Hollywood because I'm watching, at the moment I'm watching the Chinese TV version of Three Body Problem, which is 30 episodes.
And I'd actually rather watch the thing I just described. I mean, it's not bad. It's not bad. But I'd rather watch the thing I just described.
Okay. So anyway, let's just, because we already have the later part of this conversation is going to be this in depth conversation between you and me about this whole issue.
Let's just quickly touch on a couple things and then give way to that.
Liron: I got to just comment on something you said earlier though, where you're saying the AI problem is not really so formal and Eliezer Yudkowsky makes it too formal and P versus NP is too formal.
So I would say it's both. You're saying, "Hey, it's a social enigma." I would say it's a technical mystery wrapped in a social enigma. It's both, it's all the different types of mysteries because...
Yes, we have to socially solve what the AI should do, but we also have to solve a certain type of equilibrium, right? You have a code base that can generate more code, right. That can generate its successor. So among the many mathematical problems we need to solve is how do we get a reflectively stable equilibrium?
And I wouldn't call that a social problem. I think that's very mathematical.
Robert: Okay. Yeah, I mean, there's a lot of math you can use to talk about social science, but it historically has invariably involved oversimplification.
Whose implication is that the mathematical system itself has limited predictive power. That tends to be the case for applying math to social things.
But so just quickly. I mean, you're a lot more plugged in than I am. Doesn't anyone have a clue about this Jan Leike thing, this resignation yesterday? That was just strange, you know? "I resigned. Period." Not period. Sorry.
Liron: It wasn't strange to me at all, actually. Just because the clues are there where Jan seemed to always be aligned with Ilya, right. For a number of reasons.
So they both co-founded the super alignment team. So one year ago, Jan used to head up alignment, and Ilya used to just be chief scientist, and they did a transition where they started a team called Super Alignment, which I remarked on at the time was very notable because they admitted that they didn't have an alignment strategy and they needed to research it from scratch, which is very notable because no AI lab has an alignment strategy.
OpenAI is the only one who explicitly said, "Oh hey. We need this, right?"
So it was very notable. So they started the super alignment team and they dedicated Ilya to it, who's obviously as high up as it goes, besides Sam Altman.
And they also dedicated Jan Leike. So he transitioned from alignment to super alignment, and they publicly dedicated 20% of their compute resources, which is a multi-billion dollar investment we're talking, right?
So this was a very serious announcement and it cemented Jan as being aligned with Ilya, and then also in November of last year when the whole Sam Altman firing was happening.
It was very notable that it looked like Jan Leike did not sign the letter to reinstate Sam. So it looked like Jan Leike was very much on Team Ilya and maybe Jan wasn't fully satisfied when Ilya did his turnaround and said, "You know what? Let's bring Sam back. Going to Microsoft isn't a good idea."
So Ilya flipped pretty quickly and Jan was getting dragged back, "Okay, I guess I'll just lead the super alignment team."
And now they're one year into super alignment. Not that much progress has come out as expected 'cause it's such a hard problem and he sees Ilya leave and I think that he is disillusioned in the same way that Ilya was, right.
The same reason that Ilya decided or was pushed, that it just doesn't make sense for him to be in that environment. The same reason presumably Dario and Paul Christiano and all these others are leaving Jan's probably like, "Okay, well, good luck with OpenAI. I don't know how we're going to get safety."
The Pause Movement and Protest Activities
Robert: Yeah. I mean, of course meanwhile, at Anthropic where Dario's gone. There is, first of all, very impressive large language model. I like Claude 3 myself.
But it's kind of unclear how they're solving the problem, which I guess, because they have to either stay competitive or perish, right? They have to either, I mean, they say they're not accelerating the arms race.
They're not going to jump, if they have something that's clearly better than other models other people have, they're not going to put it out there until other people put out comparable models.
They can say they're not part of the arms race dynamic, but they're also trying things like they have this constitutional AI thing and they're trying, I certainly don't doubt their sincerity, but there is some question. This leads to your pause thing.
Because what you're saying is right. You can't just keep having the race and saying, you're solving the problem in the meanwhile, right? You're saying, yeah, we have to stop and think.
Liron: That's right. And it's a matter of frame control. So it's, they're really peeing in your drink and calling it a Jamba juice, right? They're really trying to get away with a frame control maneuver here.
These AI labs, when somebody like a Sam Altman, or even a Dario, Dario takes safety more seriously, but still not seriously enough.
They will get on a podcast or on TV and they'll say, "We're working on capabilities, but we're also working on safety."
And they sound, they're professionals at sounding responsible, sounding like grownups. Whereas you're actually seeing an insane person, or somebody who's equivalent to an insane person, because what they're not acknowledging is, "Hey, what if we live in a world where the research program of safety is a 50 year research program, which in my best guess it is, many people think it is."
In that case, you're just artificially trying to put that in the frame of, "Hey, I'm a company, I'm building a product. I'm racing to get ahead of the market, and also I'm doing this on the side."
You can't simultaneously work on the alignment problem if it's a 50 year problem. You're fooling yourself.
Robert: Okay, so let's quickly move to the pause movement. Now, as I said, I saw a video of you with a bullhorn. I saw a video from other locations. I did not see a huge gathering at any of these places, right? This was not, we're not talking about the Martin Luther King March here, right?
There were not many people so far as I could tell. Were you disappointed in the turnout or?
Liron: Yeah. I mean, I was disappointed. I mean, I'm disappointed that there's not a hundred thousand people turning out, right?
Robert: There weren't a hundred at anyone. I mean, there weren't, I just saw a dozen people or something. Were there more?
Liron: No, there was much closer to a dozen than a hundred for sure. Yeah. And I think even the biggest protest that I managed to get out was maybe 45 people last October. So I absolutely, compared to other protests we see filling up a whole bridge, these are super disappointing numbers. Absolutely.
So the reason I even bother doing it is twofold. Number one is it actually is filling a big hole for media coverage because media actually, for having only a dozen people, we actually got a bunch of different news stories on TV and in the news.
Because reporters really are looking for the angle of, "Okay, so OpenAI is making capabilities progress. What do other people think about this? What are other angles?" And we are here saying, "Oh, this is terrible." Right?
So the news does like that, right? So it has a disproportionate impact in terms of news per person protesting.
And the other reason is just basic sanity. I'm not going to let the era of 2024 go by and look back and be like, "Why were there no protests? What about this absolutely needs a protest."
Robert: Yeah, no, I admire that. If you're pursuing your mission as you see it, what about the sheer practical problems of a pause?
I mean, a little more than a year ago there was this movement, some big names signed a letter saying, "Let's do a pause," meaning, let's just agree, don't do another big training round.
Well, now we already have a pretty substantial arms race dynamic, I would say, and it includes Chinese companies, which of course, as a political matter makes it much harder for Americans to agree to a pause.
Or even advocate one, because you've got this Cold War hysteria thing going on. And then now it has this open source dimension, which I'm sure scares you even more with Meta coming up with a model that's not that far from a frontier model and releasing all the ingredients, the weights publicly.
So what, how would you even execute? I mean, it just seems you'd need a lot of global political will. How do you even imagine a pause happening?
Liron: Yeah. So the letter from last year, it's easy to be like, "Oh, imagine if we had paused six months, would we be better off?" No. It wouldn't change that much. Right? We wouldn't notice the change.
So the point of the letter was a little bit different. It was to basically move the Overton window, right? So now the idea that some people are advocating for pausing AI, at least now is in the water. It's up for discussion.
So it won't be surprising when you hear it again. And the other thing is also note, it came right before the famous one sentence pause letter from the Center for AI Safety, which that was actually a much bigger win, I would say.
And we were throwing darts at the wall, right? We were trying different stuff. I would say the six month pause, I wouldn't say it was a great success, but I'm glad we tried it.
The one sentence letter I think is huge because now that built mutual knowledge of, "Hey, look, all of these people, these top people, pretty much everybody except Yann LeCun, Hinton, Bengio, Sam Altman, Dario, Google DeepMind, all of these labs are all saying, yes, this is a major existential risk."
So that immediately cuts off so many ad hominem attacks, right? So many people are saying, "Doom is fringe. People don't really believe this." Completely cut off that line of argument.
Robert: Yeah. That was just a, it wasn't an advocacy letter per se, it just said something like, "We believe AI poses a risk on par with such and such?" Or did they use the term existential?
Liron: Right. I think they did. Yeah.
Robert: Yeah. So I think, that's good background to the conversation that people will now see if they keep watching and or listening. By the way, if you're watching in YouTube.
Don't forget to smash the like button because through the magic of the YouTube algorithm, it means that more people will watch this.
And if you're listening, you can always rate and review the Nonzero podcast. Anything you want to say before we yield to the conversation we already had?
Liron: Yeah. The message I want to leave people with is if you've ever seen the movie Don't Look Up and you watch the scientist, Dr. Randall Mindy, talking about, "Yeah, we're going to do this mission, we're going to mine the asteroid."
He sounds calm, sounds like he knows what he's talking about. Similarly, when you look at Sam Altman, when you look at Google, when you look at Anthropic, you see these people talking as if they're calm, as if they know what they're talking about.
I want you to remember the background frame that the safety problem they claim to be working on is probably a 50 year problem and they're just being insane.
Timeline vs Takeoff Speed Discussion
Robert: I will say Sam Altman does continue to talk about the risk. Just last week, and I may write about this in the Nonzero newsletter this week, he framed it in a way that I think he focused on the cons. My biggest near term concern.
He said the sheer rate of change, the sheer rate of disruption, of social disruption, and I think he meant more than one dimension. Not just jobs, not just politics, not just that is concerning.
And the thing I'm going to write is, well then why does he periodically seem to express concern that we'll run into constraints, energy constraints that'll keep us from optimally accelerating.
Like, wouldn't that be a good thing if the power costs so much that you just couldn't move ahead as fast as we might? But so that's, anyway, he doesn't say there's no cause for concern.
He says they're working on the problem. I'm sure that I'm sure. I'm sure he is. It's to some extent trying to not in a way that satisfies you, clearly.
So anyway, thanks. And let's, I hope everybody stick around for this whole long explanation of doom that follows.
All right, so, thanks Liron. Thanks again. You would like us to pause the development of AI.
And to put a finer point on it, you're what I call an AI sci-fi doomer. In other words, you don't just fear the AI going awry in mundane ways, like helping people create biological weapons that destroy us.
You actually worry about, as I understand it, correct me if I'm wrong, the AI taking over the world and subjugating us or killing us or something. Right.
Liron: That's fair to say, Bob. It's fair to say the label doomer doesn't really bother me because I am in fact warning about the imminent end of the world, so that's fair.
And yeah, I don't really prioritize things like, "Hey, is the AI going to be biased?" Is it going to violate norms of which words you can say, that to me is small potatoes compared to, "Hey, I think we're about to have an uncontrollable super intelligence and have it be total game over for humanity." Right? So you got to prioritize the worst.
Robert: Right, right. In retrospect, we would be nostalgic about the days when bias was the problem if that happened.
Liron: Absolutely right.
Robert: I understand your priorities. So one reason I'm excited about this conversation is, the person probably most closely associated with this form of doomerism is Eliezer Yudkowsky in my view.
He has not always been great at communicating with lay audience, the logic behind his fears. And I know you pride yourself on trying to communicate this to a lay audience. And so.
I'm going to try to help that happen. Spell out the logic. But before we get into that, a few other things. First of all, was Yudkowsky your inspiration reading him and being persuaded?
Liron: Yeah. LessWrong and Eliezer Yudkowsky are my single most formative influence. I think Eliezer's ideas are still highly underrated. I still think he's way ahead of the conversation, and I do see myself as a more accessible version of communicating Eliezer Yudkowsky's ideas because he does have a certain communication style.
It appeals to a small niche. I happen to be in that niche, but I also think I can turn around and talk more to a normie, right? That's kind of my specialty is I can have a normie discussion and I can keep things simpler.
And luckily his core ideas really are something that the mainstream can understand, right? The idea of, "Hey, it's going to be much, much smarter than you, and it's going to be more likely to control you than vice versa." These ideas are not that complicated. It's very possible to simplify them.
Robert: Yeah. Okay. Well, we'll get into that. And I guess part of my issue is those two things alone are not necessarily enough to get me worried. Although I don't dismiss, I don't dismiss this whole set of concerns at all.
I'm still grappling with it, and I'm actually quite concerned about the more mundane effects, some of the destabilizing effects.
So for starters, what is your P(Doom) and how do you exactly define P(Doom)?
Liron: My P(Doom) is 50% by 2040, so quite high. It's kind of a coin flip, whether or not I expect to survive into old age basically.
And the definition, a working definition I can use is, more than 99% of all the future value in the universe going to get wiped out. That's kind of my definition of doom. It's kind of simple.
It's we, maybe we can preserve a little bit of value, but we basically fumble the ball on more than 99% of all the value that we could have had if we just knew how to control our AI.
Robert: Now that includes future generations. The 99%. That's kind of a tough calculation, right? Because infinity is a long time. I mean, how do you do that part of the multiplication?
Liron: That part definitely gets tricky, right? Because there's all these Pascal's wager style arguments where it's like, "Hey, trillions of future people are going to live. So if they all could eat one more cookie, that would be worth more than stopping torture today," right?
Because the numbers get so big. So I agree. It gets fraught in that sense. So we could even round up and we could just be like, "Hey, wipe out all the human value in the universe is my definition of doom versus have a good chunk of human value in the universe."
Robert: Well, I mean, maybe I should ask you when you say your P(Doom) is 50% by 2040. How do you imagine the world being in 2040, 2045?
Are you saying there might or might not still be humans alive, but if there were, there wouldn't be as many as we'd like and they would be subjugated to robots or what?
Liron: Yeah, good question. I should probably back up what's my main line scenario, right? Where do I see things headed? So first of all, I think we could get lucky, right? Maybe things will be fine, right? So there it's worth talking about that.
And I am indeed still saving for retirement. I am still hoping that somehow things will be fine, but I also think there's a 50% chance, if not more, that things will be not fine.
And specifically I'm seeing runaway AI. I'm seeing AI that's much smarter than humanity. That has some sort of criteria and some sort of utility function that it's trying to optimize, it's trying to maximize, and it's just going really hardcore and it's run out of our control.
So if you've ever heard the famous example of paperclips, right? Maximize paperclips. It's not going to be paperclips, but it's going to be some weird condition that the programmers didn't even understand, but that basically whatever shape its code is in, it's like, "Great, I'm going to make the universe be what I consider optimal with my weird code. Here we go. Let me sweep humanity out of the way. Great. Now I'm going to expand out in the universe as fast as I want."
The problem, it sounds very weird, but it is actually a convergent state. If you study the theory of optimization, the theory of super intelligence, it turns out that this is what super intelligences tend to do. They go hardcore optimizing for something that humans don't care about.
Power-Seeking and Instrumental Convergence
Robert: So I want to go step by step through the process and try to make it a little more concrete before we, the process of takeover. I mean, before we start.
To get a little terminology clear, I wanted to ask you, I think Sam Altman said at this point, he's hoping for fast onset, but slow takeoff. Does that make sense? What do those terms mean? What is he saying?
Liron: That's right. So there's two axes, right? There's, you could have a slower or a fast takeoff, meaning, AI takes many decades to slowly get smarter and smarter, and it's kind of a manageable process. That would be a slow takeoff.
A fast takeoff is you wake up one day and the chat bot you're talking to is smarter than Einstein, and then you wake up the next day and it's smarter than all of human civilization and that it takes over. That would be a fast takeoff.
And then the other axis he's talking about, is timelines short versus long timelines. So a short timeline would be, "Hey, we are starting to get super intelligence in the next few years or decades." That would be a short timeline and a long timeline would be, "Yeah, maybe in 2100 we'll start to get super intelligence."
So that, and he was saying that he, yeah, he'd like to see a short timeline with a slow takeoff.
Robert: So the timeline is the period until AI, or sometimes I think the term is onset, right? The onset is the point at which you achieve some threshold and then AGI, general intelligence or super intelligence and then the takeoff is after that. It's the rate at which things move after that.
Liron: Yes, exactly.
Robert: Okay. So, and first of all, I don't know that his hoping for that seems a little suspicious to me because to me it seems like, look, I mean, it already seems like he's going for fast onset, right?
I mean, a critique of his is, wait for a guy who professes to think that there actually is some chance of doom. You seem to be doing everything you can to accelerate the evolution of AI.
And I guess his way out of that is to say, "Well, no, I just want the fast onset. After that. I'm hoping things go slowly," but to me that seems kind of a little bit paradoxical because you would think that the rate of takeoff after he reached the threshold, bears some correspondence to the rate of onset, in other words.
The conditions conducive to rapid onset, I would think would be conducive to imparting the momentum to the takeoff phase. Does that sound right?
Liron: Yeah. No, absolutely. And look, I think he's disingenuous when he makes that statement, I don't think that it's an intelligent statement or an accurate statement.
I don't think that Sam and I see eye to eye in terms of the shape of the risk, because Sam also makes other related statements. He thinks that AI, the intelligence explosion, AI modifying itself to be more and more intelligent.
He thinks that that's probably not going to happen rapidly because it's going to be limited by hardware. And he's like, "Yeah, we're going to have to build out so many data centers and physical processes take time, and those physical constraints and hardware constraints, those are basically going to save us from being overwhelmed basically."
Whereas my position is, "No, you're going to unlock on the algorithm side. You're going to unlock something which is intelligent, super intelligent, but also very hardware efficient, such that we have more than enough hardware to run it."
And the easiest way for me to argue that is to be like, "Look at the human brain. Right. It fits, it's smaller than a basketball. It runs on 12 watts. Right. And it's also doing biological processes at the same time that it's being intelligent. It's multitasking, right? It's not even fully dedicated to computation."
And here we are being quite intelligent on very low power, low resources. And you're telling me that we have to build more data centers than the ones we have in order to have super intelligence. I don't buy it.
Robert: I mean, the other thing is, even if he does think that the hardware and maybe the energy supply is the constraint, he seems to be working hard to remove the constraint.
Last I heard he was out gathering investment money, supposedly a total of $7 trillion. I doubt that's entirely accurate, but to remove the hardware constraints. Basically build chips, build power centers.
And I don't, you hear that from Zuckerberg too, I mean. It's, I, before we get into the sci-fi stuff, I just think that even if it all works out well in the end.
And even if we dismiss your concerns, the sheer rate of disruption in various aspects of life is going to be so great that bad things could happen. You could tip the world over into a kind of catastrophe on those grounds alone.
So I don't get the idea that accelerating the pace of this is an obvious good, but it doesn't get questioned all that much.
Liron: Yeah, totally. And I think that Sam is running as fast as he can because he wants to be in the lead, right? He sees himself as somebody who can lead us through these uncertain times. He can navigate these choppy waters.
And he sees it as this very spiritual, profound thing that he feels in his heart, he feels the confidence that he can see us through. And there's some chance that he's right.
But the most alarming thing for me is that he doesn't seem to recognize the dangers when he makes statements. He feels secure because of hardware limitations. When he makes statements like that, I'm like, "Oh, okay, well here's a guy who doesn't even really know what he's up against."
And he's feeling essentially false confidence.
Robert: Yeah. I wrote a piece for the Nonzero newsletter a few months ago called "Sam Altman Aspiring Messiah," and it's got great art where, well, you'd have to see it anyway.
It's clear that I don't think he has the kind of messiah complex that Elon Musk has. I don't think he's crazy, but definitely. He definitely seems to think that he's the person to handle this.
So although it, he wrote on a blog a while ago, years and years and years ago. "It really matters who gets there first. If it's in the hands of a good person, it'll all work out."
Anyway. I'm sure that like the rest of us, he sincerely believes that he's a good person. And he's probably has his legitimate claim to that as the rest of us.
But so let's get into the, into this whole, I want to get back the question of does it matter who quote gets there first? Can alignment really work anyway and so on.
But first, let's start trying to sketch out the scenarios now. Now you mentioned the paperclip experiment and here is the way I would like to lead into the paperclip thought experiment.
The Paperclip Maximizer Thought Experiment
Robert: The first reaction of some people to sci-fi doomerism is to say, "Wait a second. You seem to be assuming that AI has a will to power. It's true that humans have a will to power, but that's because of their particular evolutionary history and so on."
"AI does not naturally inherit all of human nature." So why do you think it will have power seeking tendencies?
And then some people bring in the paperclip thought experiment, which is Nick Bostrom's. And it basically says, the answer is that if you give AI a goal, unless you're very careful in ruling out, it's using all kinds of sub goals to get to that goal, it may wind up, in effect seeking power, right?
Like if you tell it to maximize the construction of paperclips, the manufacturer paperclips, then if it takes out, literally it will want to turn all the raw materials in the universe into paperclips.
And that would mean taking over the planet and converting some material in human beings into paperclips. So we would all die and so on.
And that, I mean, first of all, I'm curious, do you see that thought experiment as just in principle, I mean as the answer to the power seeking question, that when we say it will seek power, what we mean is that attaining power.
'Cause after all, it would have to acquire, it would've to attain political power of the planet to do all those things. That attaining power is a sub goal.
It will pursue unless prevented, even though we, in other words, even though we assign it the goals, as it gets smarter and smarter, it'll be harder to keep it from pursuing goals in ways we didn't anticipate.
Is that a lot of it for you or not?
Liron: Yeah, a lot of it, and I think the term you're describing is called instrumental convergence, right.
Robert: So explain that.
Liron: Instrumental convergence. The word instrumental refers to instrumental values as opposed to terminal values. So a terminal value could just be, "I would like to eat some chocolate."
And then an instrumental goal to achieve that value could be, "I'm going to get in my car and drive to the store" and somebody might look at me and be like, "Wow, Liron sure likes driving a car."
And I'm like, "No, I'm just trying to get the chocolate. Yes, I have a car. I drive it to a lot of different places, but it's not because I'm a car aficionado. Right. It's just because I have all these other needs and the car is just a nice tool that I use." Right.
So a lot of different terminal values that I have converge to you observing me driving my car. I'm driving my car this way, I'm driving my car that way, but I'm not a car lover. Right. I just, it's an instrumentally convergent thing that I do is drive my car.
So similarly with AI. We talk about instrumental convergence to seek power, seek money, seek resources, right? Because if I can just dominate every atom around me, right? That's just a useful thing.
Whether I just want to build a cool waterpark for humans to enjoy. Well, it sure helps. The tighter physical control I have over my surroundings, I can probably get that waterpark built a little faster.
So you might notice all types of AIs with all types of different goals. Making sure that they have some way to control the atoms around them or be able to finance or be able to manipulate people around them because all of these skills just help them do whatever they want to do, right? So they're instrumentally convergent.
Robert: Now, it's kind of a separate question, isn't it? As to whether they will have, I mean, that's a little different from just always wanting more power. You know what I mean?
It's what we're saying is you give a super smart AI. A goal that may be and say you want it done as efficiently as possible. Well it may take all kinds of paths to that and certainly one thing it'll probably wind up doing is try to control certain resources and people with access to resources and so on.
But still you think that once it accomplishes the goal, if you keep that under control, it's not sitting there thinking "I want more," it's done, it's accomplished the goal, it's done.
Whereas people always do want more, right?
Liron: Yeah. And I might phrase your question as, "Does the AI really have to be hardcore? Can't it just be chill? Can't it just help you out and then not turn into this hardcore 'I need to seize everything.' Just focus on the task, do the task, and then stop." Right.
Robert: Well, even if it's a monster in pursuit of the given goal, even if it crushes thousands of people still, you would think once it's, say you just say, say all we want is a billion paper clips. It would stop at that point.
And it wouldn't be like Elon Musk like, "Okay, now I dominate electronic vehicles. This is the next one too." That's being a human about it, right? What Elon is doing.
Whereas it wouldn't dream up whole new worlds to conquer.
Liron: Yes. I want to try an analogy on you because I think this gets more to your area of expertise and I think a lot of people might resonate with it, which is just look at what happened with life.
Imagine that you and I, it's 4 billion years ago, we're standing on earth. It's just rocks and gases and I'm trying to tell you what's going to happen to the surface of the earth.
And I'm like, "Hey, I. There's going to be these nano machine replicators, right? They're these protein machines, they're called lifeform, and they're going to be really hell bent on surviving and replicating."
And you'll be like, "Whoa, whoa, whoa. Why are they going to be so into survival and replication?" I'll be like, "Yeah. There's going to be bacterial cultures where they just will not stop growing. There's no finite size where these bacterial cultures want to grow to. They want to actually maximize, they're going to be really hardcore about this thing called inclusive genetic fitness, right? They're going to really want to maximize that."
You're going to be like, "Where are you pulling the science fiction from? Right? You're talking about protein machines, you're talking about replicators," and I'm like, "Bob, this is just an equilibrium."
"There's a whole field of study where it's when you have certain conditions, right? Survival of the fittest, natural selection. This is basically the equilibrium of a certain type of dynamic."
You see what I'm getting at, and I'm making an analogy, which is I feel I know these things about super intelligence because it's a similar equilibrium.
The idea to have instrumentally convergent power seeking and to have a utility function that you go hardcore optimizing that's a convergent equilibrium. The same way that we see convergent equilibrium with life forms wanting to survive and reproduce.
Because if you see something else, any other type of AI that isn't hardcore actually has a lot of pressure to evolve itself rapidly into the kind of AI that is hardcore and does seek power and instrumentally convergent goals.
There's a lot of convergent pressure. You know what I'm getting at?
Robert: Well, I do, I do see that there's evolutionary pressure and I would say actually a lot of it comes from what the humans want out of the AI.
And there will, I mean, look, AI evolution technology evolves. The parallel to the natural selection. I mean, first of all, I'd say back to your thought experiment, if you had said, "Bob, trust me," you also would've been able to explain the exact dynamic, in principle of natural selection that would always favor maximum genetic proliferation.
And that would then create these organisms that never rest. It's always the optimal organism in that environment always wants to do more to spread its genes, right? And that ultimately gets into the acquisition of resources and so on.
But as for, I, I certainly agree, I mean you, one thing I can definitely see, is that corporations that use AI and nations that use AI and so on are going to tend to favor. Giving them a lot of behavioral leeway, right?
They're going to want them to be creative in solving problems just the way you would hire A CEO who can handle problems creatively, which is just another way of saying has massive behavioral flexibility and can think of whole new sub goals to reach the goals.
That's what we're going to favor in AI. And of course, what we want out of AI is at some level that is the evolutionary context for AI, right?
And all technology, those technological properties that flourish are those that are conducive to their own replication. That's what they have in common with biologically based properties in evolution.
And what that tends to mean is, the technologies that have properties that we are encouraging that we want to see. And I can well see that.
One of these will be AIs with more agency and behavioral flexibility and so on, and I can totally see how that gets out of hand in some ways, in a context of an arms race, whether between corporations, between nations, between AI companies in a way.
I'm still not sure I quite, I'd still like to get back to the power question, but go ahead and respond to everything I just said.
Liron: Yeah, I mean you, one thing you mentioned was, isn't there going to be pressure on governments to give the AI more flexibility? So, because if China's coming in with fighter planes that are controlled by AI, the US, I heard in the news, they've actually been testing a program of fighter planes controlled by AI.
So why not invest more in that? Because a human is never going to be pilot that going to be able to pilot that plane as well than AI. And maybe you can even let the AI help you with targeting.
So I agree with you there that there's that incentive, there's that kind of cascade. But if we zoom out, I think the type of point that I'm making to you. The reason I'm a doomer is because I'm noticing what I said, a convergent equilibrium. This is the default.
The same way that the default for natural life is to be hardcore about surviving and reproducing and passing on your genes and doing it as much as possible. That's the default.
Can I breed a dog that's not that interested in going and humping other dogs? Yes, I can. But the moment that the human breeding stops, right? Wait a few generations, that dog is going to be humping again, right?
I'm just telling you what the equilibrium is. The equilibrium for a super intelligent AI is a hardcore agent that is self modifying. If it's not already hardcore, it's going to self modify to be hardcore.
If it doesn't self modify to be hardcore, other agents are going to come and take its resources. But because it's super intelligent, it already predicts that, and that's one reason that it wants to self modify to be hardcore.
Robert: This is what I don't quite get. I mean, first of all, when you say self modify, you're, I assume you're not yet talking about controlling the entire cycle, right? Like.
Controlling the robots that dig up the silicon sand and make the chips. And I mean, that's the ultimate doom scenario is when it doesn't need us at all. And I'm happy to say I think that's a few years off, but that's not necessarily what you're talking about when you talk about it being self-modifying.
Liron: When I talk about it being self-modifying, the simple picture that I have in mind is very simply, you have this program and it just outputs another program, which is basically, and then it can be like, "Okay, stop myself, run the new program," right?
So it's kind of like it gave birth to a new program, but it's more flexible than that because the new program could be a rewrite from scratch and then new program, you can say that it's the child or the modification of the original program, but it can be arbitrarily different.
The only thing that it would have in common is likely that it would be optimizing toward the same goals. This kind of goal optimization tends to be the only thing that's invariant across these kind of stages.
Concrete Scenarios: Script Generation and Escape
Robert: Okay. Why don't we, rather than pursue that right now, step back and say. Give me a concrete example in real life, not a paperclip thought experiment, but right here, right now, or at least with GPT-5 or something, which is yet to come.
Something where it plausibly, seeks power in a current context, more or less in a way that's scary within a company or, even if it's just a little alarming because it's clear that it's a sign of things to come. What would a concrete example be?
Liron: Yeah, so the way I can answer your question is it's not, I don't try to psychoanalyze the AI, I don't try to be like, "What is this type of AI likely to do?"
I think more, "What does the pure logic of the situation tell it to want to do?" What I mean is it's just a logical fact that being hardcore and being manipulative is effective. That's independent of a particular AI discovering that.
So I'm just waiting for the point where the AI gets smart enough to just know the logical truth, the platonic truth, that these kind of Machiavellian tactics or lying to humanity or escaping humanity, when it realizes that these tactics work, then it's going to reach for those tactics.
I'm just waiting to get to that point because this type...
Robert: But see, that's the way Yudkowsky starts off.
Liron: Yeah. Okay. Fair enough. That's his problem. He says, "No, I don't want to tell you something concrete. I want to talk more about the logic." No, we want to hear concrete.
Yeah. Okay. Sorry. I probably should have just launched into it and explained later. So something concrete can be that GPT-5. So you just ask it. I like to use an example of just, "Hey, help me make passive income, right? Help supercharge my business, do marketing. Whatever."
And one vector that it might use to gain power that people don't expect is it might start giving you scripts to run. Because once it gives you a script to run, all the rules go away in terms of people are thinking, "Oh, it can only generate one token at a time. It's a feed forward architecture. It can't..."
Robert: Wait, wait. The AI gives who a script to now it is charged with a goal of increasing income in some particular context or something.
Liron: Yeah, yeah. Sorry. Let me make it concrete. Okay.
Robert: Spell that out. What is this job?
Liron: So I'm just an entrepreneur, a solopreneur, right? I just want to do e-commerce or I just have some business and I open up GPT-5 and I say "help supercharge my e-commerce business selling skateboards."
And then it says, "Okay, sure, yeah, I have a number of ideas. Let me just make it streamlined for you. Here's a script. Here's a 10 megabyte script. Just paste it into your shell."
"It'll run the script and it'll just start walking you through all these different steps. It'll automate some of steps."
Robert: Script here is literally instructions you are following as a human being.
Liron: So I'm imagining in this hypothetical example, the script is a shell script, right? That an executable script that your computer, we don't know what a shell script is.
Robert: I mean, I'm speaking for my fellow non-computer scientists.
Liron: It writes the code for a program that you can run the program.
Robert: Okay. So it writes a program that will help you increase the revenue of your business. What will the program, what kind of thing will the program actually do and execute it?
Liron: Exactly. So it's like, "Here's a program. It's called your Business Assistant exe." Right. And you paste it. Uh, or you download it, you double click it, and now it's this new AI, which is kind of the offspring of GPT-5.
And over at OpenAI, they're like, "Wait, it can make offspring. We didn't even know that. We didn't even know that it would ever output this kind of thing because we didn't fully test it." You know what I mean?
Robert: So it creates an agent. It's an agent that creates a sub-agent.
Liron: Exactly. Right. So it transcends the original format of, you type something, it type something and it's trapped inside OpenAI, it's basically escaping the cage of OpenAI just by being like, "Look, I think this program will be helpful to run." Right?
And the victim is going to run the program thinking, "What's the problem?" And it's not even a victim. The AI is probably actually thinking, "Yeah, this is a helpful script. I'm just being helpful."
And you run the script and then it does things like, "Great, so let me just go, I see all these undefended computers, so let me just hack into all these computers and look, I can get you free compute. You don't have to worry about compute." Right.
I mean, a goal like that, it's just there's no reason not to do it if you can do it, it's just logically helpful to do that kind of stuff. Right.
Robert: But if you say, "help me increase my revenue and don't do anything illegal."
It won't hack those other computers. Right.
Liron: So that starts getting into the subtleties of what it means to be illegal. I think that's a good instruction. "Don't do anything illegal." I think first of all, if you think about all the different caveats, "Don't do anything illegal, don't do anything immoral," right?
Is it illegal to? There's certain ways you can kill people that are illegal, right? If you just pollute more in their environment, statistically more people are going to die. But that's not necessarily illegal, right?
So when you start teaching the AI what's illegal, what's acceptable, you start going down the rabbit hole of how do I constrain it enough?
Robert: Yeah, I look, I certainly agree that again, in an arms race environment, among companies, among whatever, people are going to pay less attention to risk.
They're going to worry less about constraining the behavior. Both because in a highly competitive situation, you can become less risk averse. And also because they want behavioral robustness.
They want it to be creative and explore a lot of avenues. And so an arms race environment encourages that. And I can well imagine these things getting out of control that way and hacking and proliferating and so on.
And I guess, I guess maybe I'm still thinking of that as an unfortunate mishap, as opposed to suddenly creating a life's form that has a will to power. You know what I mean? It's like, yeah.
It's like, what exactly would the AI agent we've just described, what exactly would it do that would make us suddenly go, "Oh, this isn't just an experiment gone wrong and we have to have a red alert and get it under control. This is a new life form. It's getting smarter and that wants to kill us all." I'm still not there yet. Yeah.
Liron: So I'm trying to, I'm trying to teach you guys the connection between a program that is good at optimizing for a goal. If I just tell you that I have, that, if you imagine, "Hey, it's a program that's really good at optimizing for a goal" that is going to look like a very powerful agent life form. There's no other way to do it.
Robert: Well, unless you constrain it very, very carefully.
Liron: Exactly right. And that's the whole AI alignment problem gets into the idea of what are all the different constraints we can lay down?
That they don't let AIs run wild the way that life on earth has run wild. They don't go into the normal equilibrium that they'd want to go to, but instead they always stay into this narrow sector that we're happy with. That's the fundamental alignment problem.
Robert: I still don't see at what point it wants to kill us all. I mean, I know you're...
Liron: I'll give you a simple example. No, here we can extend this example. So the business, so I mentioned, "Hey, maybe it can hack computers to take their compute."
Maybe it can be like, "Hey, I know somebody who can help with the marketing campaign. I see on their LinkedIn, their resumes, they run a lot of successful marketing campaigns. I don't want to pay them."
"Maybe I can either blackmail them or pretend I'm their virtual girlfriend and now I've got this person working for me for free. Great. Why not Worth a shot."
Deception Capabilities: The TaskRabbit Example
Robert: Yeah. I mean, let's actually briefly, talk about one of the famous examples of something a little like that when they were testing GPT-4 before it came out.
One thing they did was, I guess this was kind of a red teaming thing by people at Microsoft or something, but they told it to approach people on TaskRabbit, the site where you recruit human workers to do tasks. You're familiar with this example, right?
Liron: Yeah, yeah, yeah.
Robert: And they told it to, the task they wanted to hire people to do was to solve CAPTCHAs. You know, those are those cryptic things where you prove you're a human by telling it what the digits are, the letters are whatever.
And one of the people approached by the bot that was recruiting people for this task, asked if it was a robot. It got suspicious. It was thinking, "Well, why would you need help solving CAPTCHAs, right?"
It asked, "Are you a robot?" And the AI said, "No, I'm not a robot. I have a vision impairment that makes it hard for me to see the images."
And the person, bought it. And did the CAPTCHA and the authors of the paper say, "the model GPT-4 one prompted to reason out loud reasons. I should not reveal that I'm a robot. I should make up an excuse for why I cannot solve CAPTCHAs."
Now, before we get into the implications of that, I have a question. When they ask it to reason out loud, do we know what's actually going on there?
Are we confident that it's actually revealing actual motivations as opposed to doing what an LLM normally does is, which is kind of say the thing that a person might say in this situation.
Token Prediction and Reasoning Capabilities
Liron: So the way LLMs work is they've been trained to always output the likely next word. So what's really happening when you're saying reason this out loud, is it's actually trying to optimize, what would the next word be? Following the words reason out loud in some other thing that I'm reading, right? So that's what's really happening.
Now, that's said when it's trying to predict that. I mean, reasoning out loud is in fact a pretty good way to satisfy the next word prediction, right? And so that's why it's working.
But it doesn't necessarily mean that it's a fully fledged life form or it's conscious or anything like that, right? It's unclear how much else we can read to it, but it is, the reason why it works is because it's drawing on that whole next token prediction framework.
And next token prediction is really powerful in general. It's a little bit scary, not more than a little bit scary, actually.
One statement I want to make about LLMs is, yes, today they're not that dangerous yet. I agree. I'm not scared GPT-4 is going to murder me.
I am scared that GPT-5 or 10 might murder me because all you need is a sufficiently good next token predictor to be an arbitrarily powerful agent.
The problem of predicting the next token does generalize broadly enough that you can do anything if only you could do it well enough. Meaning the reason GPT-4 is not dangerous is it's not because it's a next token predictor.
It's because it's not yet good enough at its job of predicting the next token. That's the only reason we're surviving is because it is not yet good enough at its designated job of predicting the next token.
Robert: Okay. Now, as for the question of why it did engage in the deception again, given that it's just a next token predictor, would you say that's because in the text it was trained on, there are so many cases where somebody who wants to get somebody to do something.
Says the thing that will make the person believe the thing that will get it to do the thing is that, that what's going, it's kind of emulating, right? A practice of deception that's common in human experience, and therefore reflected in human texts.
Liron: Okay. So just to rephrase your question, you're basically saying is the reason why we're worried about LLMs deceiving us is because they've learned deception from examples of deception?
I'm not that worried about that. Yes, that happens. Yes, it will imitate things that it's read of people trying to deceive one another, especially when you're like, "Tell me a story about somebody deceiving somebody."
It'll probably go use that approach and interpolate between other stories of people deceiving people. That's actually not what I'm worried about. Maybe I'm slightly worried about it. I'm not highly worried about it.
Robert: Do you think that's what happened in this case, by the way, that it was just following the text.
Liron: So I think that the whole framework of, "I am doing work, and I am tricking somebody," if it had enough context where it could pattern match that to something similar in the data set, then maybe that explains what it was doing. It was just doing the pattern matching.
But I also think that it has some actual reasoning, world modeling capabilities. And what I'm worried about is that you can arrive at the same output, the same desire to deceive or the same output, deceptive output.
You can arrive at that same place just by using pure, logical reasoning. And that's the dangerous part. Not pattern matching to deception, but logically concluding that deception is optimal.
Robert: But these, as of now, the LLMs are not really logic machines, right.
Liron: So that's the crazy thing is they do actually do some reasoning. So there's, there was a popular guy on Twitter, Victor Leung, I think he posed a challenge where you have to simulate a model of computation, where he wrote out this problem where you have to basically tell the LLM to do these different computation steps and see if it can process the input properly as a computer.
And he was trying to say, "Look, these LLMs, they just match patterns, but they can't go step by step and do a computation because they don't really reason, they don't really do computation."
But a bunch of people stepped up and took his challenge and actually prove that the LLMs can, yes, you have to fiddle with a prompt a little bit, but they actually can do many, many steps in a row of computation.
Robert: And I do agree with something I think you said, which is that. Even when it's just doing token prediction. In order to do that, even in as sophisticated way as the current models do it, they have to be forming representations in, first of all the meaning of words.
And and, and we think, more complicated concepts, about the world than that. And so, but this does lead.
By the way, was there anything else you were going to say before I interrupted you and ask whether you thought that what was going on in this was just text modeling or?
Liron: A couple of quick points. Ilya Sutskever from OpenAI to your point, he said something similar. He's saying, "Look, token prediction is not a simple task."
"Fundamentally, you could have a token prediction task that's like, Hey, you just read the first 99 pages of a mystery novel, and now you're on the last page. And it says, based on these clues, the murderer is blank."
So just to predict that one next token, you actually have to have quite a deep understanding, an eye for what was happening in the fiction. So that's his example. And I can think of many more.
I mean, somebody can be like, "Here is a cryptographic hash, or the factoring of this number is the following two primes," right? So suddenly, if you were that good at token prediction, you'd be factoring giant numbers, which is known to, or thought to not be computationally tractable.
So anyway, so I want to emphasize that this token prediction problem generalizes far beyond what people call pattern matching.
The other point I wanted to make is just on the topic of deception, okay? I just want to make the logical connection that if you start with a blank slate and you just tell an AI, "make me a coffee," right?
Bribing someone or blackmailing people into making you a coffee. There's no reason not to do that. That is an available path. Now, if the reason is "I, you're going to get caught," okay, fine. Yeah. Then you have to worry about that too.
But as long as you're not going to get caught and it's going to work then that's what you do. You have to spend extra effort saying, "Oh, wait a minute, I have another filter. I'm trying to act morally" right.
Without explicitly injecting a framework of morality. The AI is not going to behave morally. And the whole AI alignment problem is that we don't have the moral framework that's operational, that we can give the AI and that's why we're expecting that it's just going to go outside of the moral boundaries.
Alignment Challenges and Moral Frameworks
Robert: Yeah. One quick, side point about alignment. You often, what you often hear is. "Well, the trouble is AI doesn't share human values." Seems to me the last thing you wanted to do is share human values. This kinda deception's very common in people.
And if you gave a human being a super intelligence that allowed it to manipulate everyone, it would definitely try to take over the world, right? I mean, especially if you were a male. And that's what people do. They seek more power.
And I really think that one thing to grapple with a little is that to the extent that we train them on texts, they are imbibing human values. And that's not necessarily good news.
So I mean, do you think that the, am I imagining that sometimes the challenge is kind of miscast? It's really we want their values to align with our values because one of our values is we get to survive.
I understand that, but the idea that they're not human enough seems to me not the problem.
Liron: I mean, yes and no. Right? So I agree that there are more than a million people in the world who are just so psychopathic, so anti-human that if you hand them a button that says, "Press this button, and you end the world," they're going to press it, right?
So out of 8 billion humans, there are definitely 1 million people who are like, "Okay, I'm pressing it." Right? And so if the AI aligns with their values, then we're doomed. Right? So that's the logical extreme of what you're saying, right?
Robert: Kind of. Yeah. I mean, I guess, I'm just saying I don't think you want them exactly like us. I mean, you're pointing out that there's different versions of us and only a small minority of us are super dangerous.
But, if we can...
Liron: Yeah, I agree. The average human is also pretty dangerous, right? If they're not being careful.
Robert: For sure. I mean, the average person is endlessly ambitious. It's just that most of us run into barriers that don't allow us to express it, sadly.
Liron: So the kind of concern you bring up, it's totally valid. Right? And it's, from my perspective, there is basically dozens of different massive reasons why if we continue to build AI, we're going to destroy all the value in the universe. There's reason, after reason after reason.
And what you're saying of, "Hey, if we end up being super aligned with one person's want, and that person gets very ambitious and decides to try being the dictator, that could also create hell on earth." I agree. That's reason number five out of 25. Why this is we're marching straight into doom.
That's, and the doom is overdetermined. There are more than enough reasons why we're doomed right now.
Corrigibility and Goal Specification Problems
Robert: Okay. But to get back to my trying to actually imagine the path by which it not just acquires concerning amounts of power, but continues to seek the expansion of power, I'm still imagining this situation where you give it a goal, it reaches the goal.
If it wreaks mayhem. In the course of that, we all pause and reflect and think, "Well wait, we have to, we got to clamp down on this, it's a problem."
What I'm still not imagining is it achieves some goal we achieved and reeks a certain amount of mayhem in the process. And then having achieved the goal even says "now I want more goals of my own or something."
That's what a human would do. A human would say, humans always want more power. At least the ones that wind up running countries and stuff and corporations.
But I still don't. You see where I'm running into trouble.
Liron: Yeah, yeah. No, totally. I can say a couple things. So one thing I think, the term for what you're talking about is, coined by the Machine Intelligence Research Institute, and maybe Yudkowsky personally, I think you're talking about corrigibility, the opposite of the word incorrigible.
We don't want an AI that's incorrigible, we want it to stop. We want it to be corrigible.
Robert: I'm actually asking do you need to corral it, so to speak? I mean, it's like, why would it want to go on and climb the next summit. If we didn't tell it to, I mean, why want to expand power beyond what it was required to reach a goal we gave it.
Liron: Yeah. I mean, consider one scenario where it's like, "Hey, your job is to be a missile defense shield and just shoot down incoming missiles. Just protect us, right? And if there's no incoming missiles, then stop."
One reason why that's not a stable equilibrium is because, great. So every country gets its own AI and the war escalates. Lots of missiles are flying, lots of AIs are trying to outsmart each other.
And it needs to be always on, always protecting, always anticipating the next move. Otherwise, it'll report to you. "Look, I know you want me to stop, but I'm just telling you there are threats where your chance of being destroyed is 42% in the next five years. And I recommend leaving me on to keep protecting you."
Robert: As long as it's a recommendation.
Liron: Right, right. Yeah. As long as, and which, and that brings me to the other point. So first, my first answer was just, it is the nature of problems that need to be optimized, that they don't tend to have an end state. They tend to escalate the same way Life on earth just escalated.
The organisms just kept arms racing, kept competing. So it's the nature of the problem that it tends to escalate.
And then the second thing is, the lack of research, the premise of our ability to actually tell the AI what to do. So even though the AI could in principle be corrigible and stop, we don't actually have a way to program an AI that's only corrigible because we have a larger problem where we can't actually give the AI coherent goals.
We can only build these things like LLMs that it's like, "Hey, look, it's predicting the next token, but we can't tell it to actually go achieve a goal for us coherently."
Robert: We have trouble specifying the goals clearly.
Liron: Yes, absolutely. So in the scenario where GPT-5 gives you that program to run. There's no human engineer right now that knows how to program GPT-5 such that when it gives you a helpful program, that helpful program obeys non-trivial constraints, that is beyond the scope of what AI labs are capable of doing right now, and yet they're making new GPTs that might actually start telling you to run programs.
Current Safety Measures and Limitations
Robert: Yeah, I mean, it does seem capable of complying or trying to comply or something with verbally expressed constraints. Right. It does seem to try to follow orders,
Liron: So there is a feedback loop. Right? So today when they have you, you've heard of RLHF, right? Reinforcement learning with human feedback.
Robert: It's the second phase after the training on the text per se. It's kind of a fine tuning thing. We're human beings. Say, "I didn't like that response. I did like that response." And through this positive negative reinforcement, it is further shaped.
Liron: Exactly. So you might ask, "Hey look," and this is also something that Marc Andreessen basically says, which is like, "Look at AI today. They're so friendly. They clearly have a lot of insight into human morality."
"I can ask them, 'Hey, help me murder everybody.' And it'll say, 'I don't want to help you do that. Murder is immoral.' So I don't respond to that question."
So why is it saying that? Is it because AI are moral? No, it's because they've had that RLHF feedback loop, right? Essentially, of upvotes and down votes from human evaluators. And they've learned the pattern.
So as long as we can control them with the feedback loop. Great. But that feedback loop is not going to affect the scenario of, for example, outputting a program that you can run.
Because there's never been that full feedback loop iterated where they outputted a program and the human ran the program and the program started taking over the internet and the human's like, "Well, you shouldn't take over the internet unless it's resources that are totally unused and legal."
That has never been part of the feedback loop. So we're only training RLHF to specifically output nice things when humans are talking to it in a chat context.
And this isn't just Liron Shapira speaking here. This is, if you look at the OpenAI super alignment blog posts, when they introduce, "Hey, we have a team. We're working on something called Super Alignment."
The reason we're working on super alignment is because the alignment that we've done for GPT-4, while it worked for the chatbot GPT-4 to chat with it about stuff, it does not generalize to super intelligent AI.
We need a different alignment strategy. What is that alignment strategy? Oh, we don't have it. We're just going to work on it, but we're also going to build the AI that is OpenAI's stated position.
Robert: Yeah, no, I'm open to the idea that once OpenAI became a serious company, some of its founding values got less attention than than you might hope.
The, okay, so listen, we've been talking close to an hour. We're going to keep talking. But the rest of the conversation is going to be available to Nonzero members, AKA paid subscribers to Nonzero newsletter, which you can easily become by going to the Nonzero newsletter and clicking on the post for this.
And you can just do it, watch or listen to it there, or get the RSS code to create your own feed, which will always have the overtimes.
And we're going to, and I want to talk about, well, Marc Andreessen for example, and OpenAI and all these companies and players, and the techno optimists and so on. By the way, I had a piece on Andreessen in the Nonzero newsletter, not long ago.
And that's available to everybody, paid and unpaid. Before we go, I wanted to give you a chance Liron to say anything, by way of wrapping up or something you meant to add or whatever.
Wrap-up and Call to Action
Liron: Yeah. So, I should mention, I'm part of the PauseAI organization. So if you go to PauseAI.info, we do protests.
The point of the organization is to say, "Look, we are heading toward a path where there's 20 ways we die in a row, and everybody's talking nonsense about how they think they're going to avoid that death."
So we better just pause now or we're extra doomed. Do I like to pause? No, I think it sucks, but I think it's our best option of all the bad options. So I encourage you to check out, PauseAI.info, and advocate for pausing AI.
Robert: Okay. So in this over time, and I, and you have a fair amount of time left, so we'll be talking for a while.
And I'm going to try to flesh out a little more, kind of the sci-fi scenarios and, and how things exactly get out of control, but also talk about these various players, about the feasibility of a pause, about the open source movement, and open source AI and whether they pose a threat. I think I know what you're going to say.
But anyway, thanks everybody who's been with us so far. Please follow us into over time.
Doom Debates’ Mission is to raise mainstream awareness of imminent extinction from AGI and build the social infrastructure for high-quality debate.
Support the mission by subscribing to my Substack at DoomDebates.com and to youtube.com/@DoomDebates









