Google DeepMind AI Safety Researcher Who Just RESIGNED Says P(Doom) is 25% β Debate with Alex Turner
Dr. Alex Turner went viral for resigning from Google DeepMind over its Pentagon AI contract, which he said lacked binding restrictions against autonomous weapons and mass surveillance.
Alex is a world-class researcher with a PhD in alignment from Oregon State University, a postdoc at Stuart Russellβs Center for Human-Compatible AI at UC Berkeley, and top distinctions at NeurIPS. Heβs one of the minds behind Shard Theory (with Quintin Pope), and he also helped pioneer the Steering Vectors research made famous through Anthropicβs Golden Gate Claude experiment.
Alex says Google DeepMind broke its founding promise β the commitment Google made when it acquired DeepMind β and that the people best known inside Google for caring about AI ethics largely didnβt act when it counted.
Alex came up through the rationalist community, and became one of LessWrongβs highest-karma users before leaving the site. He still shares many of my concerns about AI doom, but he believes technical alignment research is going better than expected. Iβm not so optimistic.
In this episode, we debate my Yudkowskian doom views against Alexβs own framework. Can he convince me the Yudkowskians are miscalibrated?
P.S. Weβre currently in a donation drive, so here comes our solicitation for viewer donationsβ¦
π Please click here to make a tax-deductible donation to Doom Debates [https://manifund.org/projects/doom-debates---podcast--debate-show-to-help-ai-x-risk-discourse-go-mainstream] πDoom Debates is a fiscally sponsored project of Manifund.org, a registered 501(c)(3) nonprofit.
To donate crypto or if you have questions, email me [liron@doomdebates.com].
Watch on YouTube
Timestamps
00:00:00 β Cold Open00:01:15 β Introducing Alex Turner00:02:30 β From Harry Potter Fanfic to AI Alignment00:05:03 β Meeting Quintin Pope & Rethinking AI Doom00:06:17 β Shard Theory, Steering Vectors & Golden Gate Claude00:08:25 β Why He Joined Google DeepMind00:10:32 β Google DeepMindβs Broken Promise00:16:01 β Debating Google DeepMindβs Pentagon Contract00:19:09 β Whatβs Your P(Doom)β’?00:22:42 β Alexβs Research on Instrumental Convergence00:26:35 β Misuse vs. Misalignment: The Mainline Doom Scenario00:29:31 β Will Society Self-Correct?00:36:36 β Superintelligence in 10 Years00:39:43 β Will Technical Alignment Produce a Safe AI?00:41:16 β Donation Drive00:42:12 β How Fragile Is the Chain of Alignment?00:49:50 β Disagreements with Yudkowskyβs βList of Lethalitiesβ00:57:46 β Why Alex Quit LessWrong01:01:48 β Whatβs Next for Alex01:02:54 β Does He Support PauseAI? Stop the AI Race?01:04:40 β Wrap-Up01:06:04 β Producer Oriβs Closing Note
Links
Alex Turner, βWhy I Left Google DeepMindβ blog post β https://turntrout.com/why-i-left-google-deepmind [https://turntrout.com/why-i-left-google-deepmind]
Alex Turner (TurnTrout), personal website β
https://turntrout.com
Alex Turnerβs resignation announcement on X β
Doom Debates episode with Quintin Pope β https://lironshapira.substack.com/p/ai-alignment-is-solved-phd-researcher [https://lironshapira.substack.com/p/ai-alignment-is-solved-phd-researcher]
Harry Potter and the Methods of Rationality (HPMOR) β
https://hpmor.com/
Shard theory sequence on LessWrong βhttps://www.lesswrong.com/s/nyEFg3AuJpdAozmoX [https://www.lesswrong.com/s/nyEFg3AuJpdAozmoX]
Golden Gate Claude (Anthropic research on steering vectors) β https://www.anthropic.com/research/golden-gate-claude [https://www.anthropic.com/research/golden-gate-claude]
Slaughterbots on YouTube β
Alex Turner, βAvoiding Power Seeking by Artificial Intelligenceβ PhD thesis β https://turntrout.com/alignment-phd [https://turntrout.com/alignment-phd]
Alex Turner, βSome of My Disagreements with List of Lethalitiesβ β https://turntrout.com/disagreements-with-list-of-lethalities [https://turntrout.com/disagreements-with-list-of-lethalities]
Doom Debates episode with Benthamβs Bulldog β https://lironshapira.substack.com/p/benthams-bulldog-ai-doom-debate [https://lironshapira.substack.com/p/benthams-bulldog-ai-doom-debate]
Alex Turner on X β https://x.com/Turn_Trout [https://x.com/Turn_Trout]
Doom Debates donation page β https://doomdebates.com/donate [https://doomdebates.com/donate]
Transcript
Cold Open
Liron Shapira 00:00:00You resigned from Google DeepMind because you think that Google DeepMind, quote, βBroke its founding promise through its contract with the US military.β
Alex Turner 00:00:08I value staying true to your values. People who are well-known within Google for caring about the ethics of deploying AI largely didnβt act. This matters because autonomous weapons get us into an arms race that really degrades the security of everyone in the world.
Liron 00:00:26Do you support the Pause AI movement?
Alex 00:00:28I think that AI is being developed too quickly. I probably will not take an affirmative on supporting this particular movement.
Liron 00:00:35Letβs segue into the schism, your disagreement with Eliezer Yudkowsky.
Alex 00:00:39He was incorrect on some points for alignment, but then also not acknowledging that. I think that technical alignment has gone pretty awesome, super awesome.
Liron 00:00:49Youβre not claiming that a superintelligent AI canβt kill everybody. Youβre like, βOh yeah, of course it can, but weβre not gonna break the chain of alignment, meaning weβre just going to safely develop it so that even though it can kill everybody, it wonβt.β
Alex 00:01:00Develop it in a way that produces a safe result. I wouldnβt call what weβre doing safe development, but sure.
Introducing Alex Turner
Liron 00:01:15Welcome to Doom Debates. My guest has worked in technical AI safety at Google DeepMind for two and a half years, but he just quit, and his resignation is going viral. Why? Alex Turner says Google DeepMind, quote, βBroke its founding promise through its contract with the US military.β
His latest blog post exposes hypocrisy at the highest levels of senior leadership at Google DeepMind, which includes the CEO Demis Hassabis, chief scientist Jeff Dean, and co-founder Shane Legg, among others.
Alex is a world-class AI alignment researcher. He completed a PhD in alignment from Oregon State University. He did a postdoc at UC Berkeley, and heβs earned top distinctions at NeurIPS. I respect that Alex is principled. I respect his mastery of the subject matter that we talk about on this show.
I also find it interesting that heβs levied criticism at the original AI alignment thinker, Eliezer Yudkowsky. Heβs called some of Yudβs claims fundamentally misguided, not reasonable, and bogus. As a Yudkowskian myself, Iβm gonna be curious to dig into those arguments. And of course, weβll cover whatβs going on right now at Google DeepMind and why he resigned. Alex Turner, welcome to Doom Debates.
Alex 00:02:28Hey, thank you for having me.
From Harry Potter Fanfic to AI Alignment
Liron 00:02:30So itβs great to get you on the show. One thing we do on Doom Debates is we expose top intellectuals who have been pretty familiar to the rationality community or the AI safety community, and we help popularize their ideas, even if I donβt fully agree with all of them. Is that a good description of your background? Youβve been pretty deep into the LessWrong rationality and alignment community for a while.
Alex 00:02:52Yeah, I think it was quite formative. It was just the other day in 2016 where I decided to search what are the top five Harry Potter fan fictions. And that indeed led me down the LessWrong rabbit hole, where I discovered superintelligence in late 2017, and then I pivoted my PhD in early 2018.
From that time period up through maybe early 2023, LessWrong was very central to my professional career, but also just to the way I looked at the world.
Liron 00:03:25Well, I wanna follow up on why did you search for Harry Potter fan fictions?
Alex 00:03:30I really donβt know. Itβs kind of one of those things where if I hadnβt done it, my life would be totally different. The reason Iβm mentioning this is thereβs this famous fan fiction that Eliezer wrote called Harry Potter and the Methods of Rationality. I never really liked fan fiction. I thought it was kind of cringe. No one recommended it to me. So it seems to me like if Iβd just woken up slightly differently that morning, I might not have ever been exposed to this research area, and my life would be totally different.
Liron 00:04:02Wow. And you said this is all in 2016, right? So the book had been mostly completed at that time. It had been going on from 2009 to 2015, and you kind of stumbled on it because you were just interested in seeing what the best Harry Potter fan fiction was?
Alex 00:04:17I had a random thought. Thatβs the best I recall.
Liron 00:04:21Itβs pretty crazy that thatβs how you found the community because I know you as one of the highest karma LessWrong users. You have ten times my karma. Youβve been posting a lot. It became a huge passion for you, right?
Alex 00:04:32LessWrong was for quite a while my intellectual community. Iβd have ideas. Iβd be eager to share them. Each summer I would generally do an internship where Iβd come in person at Berkeley, get to hang out with my friends there, be able to β I guess I felt more understood. In 2018, 2019, 2020, 2021, these are many years where I was at my PhD and I talked about the dangers of AI and how we should work on that.
Meeting Quintin Pope & Rethinking AI Doom
Liron 00:05:03Did you originally feel like you bought into all the Eliezer Yudkowsky concepts, and then you started rethinking everything and building it from the ground up? Was there a point of divergence?
Alex 00:05:15Yeah. I think I was maybe around 80, 85% doom conditional on developing AGI. I thought itβd be a couple decades, even late 2021. And yeah, I shared most of the worldview. I found much of his writing compelling, and I still think thereβs some gems in there.
It wasnβt until early 2022. I met a researcher at my university, at Oregon State University, named Quintin Pope. He wrote these very big brain Google Docs, and he was sending them by me. Heβd attended my AI alignment reading group.
I donβt know, something was just very interesting about them, and they seemed really far-fetched, but they were very ambitious. And as I looked more, I realized he was pointing out some real confusions, real issues. I started rethinking perhaps the claimed difficulty of alignment.
Shard Theory, Steering Vectors & Golden Gate Claude
Liron 00:06:17This is such a unique opportunity because you actually know your stuff. You actually know what weβre arguing about, unlike a lot of my guests who come in and they seem to be shooting from the hip. They havenβt spent so many hours considering it. They havenβt worked a career in AI research. So Iβm excited. But before that, letβs just finish the biography here. So you did your postdoc, and then did that lead you to joining Google DeepMind?
Alex 00:06:39Yeah, I got my PhD in 2022. My thesis was called Avoiding Power Seeking by Artificial Intelligence. Then I did a one-year postdoc at UC Berkeley at Stuart Russellβs Center for Human Compatible AI.
During that time, I worked more on this shard theory of human values with Quintin Pope, whoβs actually the alignment thinker whose ideas I respect the most, or I think theyβre the most interesting. And then I also discovered steering vectors or helped popularize those as the MATS team that I led. We were the first to really demonstrate their potential.
After that, mid-2023 was a fairly rough period personally. I did some more MATS mentorship. It wasnβt until the end of 2023 that I settled on going to Google DeepMind.
Liron 00:07:37Got it. So MATS, you were a mentor there, and they do AI safety research. They train people to do AI safety research.
Alex 00:07:44Right.
Liron 00:07:44And the steering vector work that you did was pretty foundational, and I think most people have heard of it as Golden Gate Claude, where they use the steering vector to get Claude to be obsessed with the Golden Gate Bridge in response to any prompt.
Alex 00:07:57Classic. Yeah.
Liron 00:07:59That is a pretty legit background. If people criticize me and my arguments for being a bystander whoβs not in the weeds or not on the field or whatever, I think itβs fair to say youβre on the field, and so whatever you have to say has the credibility of just being in the arena.
Alex 00:08:19Sure. Yeah. Weβre in maybe different arenas, but yeah, the direct research arena.
Why He Joined Google DeepMind
Liron 00:08:25All right. Well, with that said, one of the things that you saw in the arena is a perception of hypocrisy of the key figures. Letβs talk about that.
Alex 00:08:36Yeah. When I joined Google DeepMind, I already knew some people on the alignment team. Rohan Shah β I worked with him during my PhD. He was at CHAI. And I think heβs done a pretty good job of leading the alignment team within Google DeepMind.
One of the things I would say is, Google was not founded with the goal of taking over the world, or potentially with the goal of taking over the world. Iβm not saying for sure that OpenAI and Anthropic were, but they were founded as AGI companies, and those ideas were present, whereas Google is, for better or worse, more of a classic company. They seem more interested in making money. That doesnβt mean that Googleβs good in all that it does, but it seemed like a different presence in the space.
At the time I was mostly just worried about doing alignment research on frontier scale models. When I joined, I was fairly concerned about whether my opinions would become trash because Iβd start rationalizing why things that Google does are good.
I had this extended dialogue with Oliver Habryka about how I could maybe net zero out my financial position in Google in terms of equity at least. We discussed a lot of big brain options, but then it turned out that my contract prohibited me from being short Google at all, which killed all of the schemes.
So in the end, I was thinking about how do I avoid this kind of value drift, this bias that I think Iβve seen a lot from people working at their labs, where they seem incapable of saying, even in private, βNope, this was bad, and we shouldnβt have done it.β
Google DeepMindβs Broken Promise
Liron 00:10:32So you had these general reservations about Sam Altman and other companies and maybe motives getting corrupted in the abstract, amalgamated from different examples that youβve seen. But I think it really came to a head recently. You specifically had concerns with the Pentagon-Anthropic contract tensions becoming public, and youβre like, βUh-oh, this is high stakes. I better make sure that Googleβs doing the right thing.β But then you saw that Google wasnβt resisting the US government or ICEβs attempts to use them. Give us the issue here.
Alex 00:11:06Yeah. So if people remember those goons in face masks with rifles that were roaming the streets of Minnesota earlier this year, thatβd be ICE, thatβd be Customs and Border Protection, and in particular, the people they just killed on the street. I was pretty upset about that.
I wanted to reduce techβs involvement in enabling ICE and enabling CBP to track down people theyβre looking for, whether theyβre dissidents, people who are potentially actually or just accusedly in the country illegally. It seemed like a not moral enterprise.
So I looked into that. I started pushing on people in the company with those contracts. And at the same time, I was talking to my friends at Anthropic because Anthropic had, and I think still has maybe, a deal with Palantir where theyβre giving Claude to Palantir, which I think is a very negative company for the world. And so I was trying to persuade my friends, βHey, can you push on getting rid of this?β
Little did I know there was this bubbling in the background of conflict between Anthropic and the Department of War. And then this came to a head in February when the Department of War said, βGive us Claude or we will economically destroy you.β
Liron 00:12:37We could definitely spend a long time on this, but because we have limited time, we are going to reserve most of the time for the alignment conversation and the Yudkowsky versus non-Yudkowsky alignment theorist debate.
That said, this has been a very important incident. Itβs currently on the front page of LessWrong. It has more points than anything else this month, as far as I can tell. Itβs getting a ton of attention, and I read through your account of events. One thing that seems clear is youβre acting with a lot of integrity. You resigned from Google DeepMind because you think that all these organizations, Google DeepMind, the leaders of Google DeepMind, and the International Association for Safe and Ethical AI, theyβve been involved with all this, and you think that they broke their commitments as well. You even say itβs the founding commitment of Google DeepMind. Youβre saying Google DeepMind was founded on a commitment not to empower the US government to do bad things with AI?
Alex 00:13:31Yeah. Itβs the founding agreement where Google purchased DeepMind.
Liron 00:13:38And from your perspective, theyβve just been bending the rules and not taking a hard stand. Theyβre not actively saying, βNo, Alex, youβre wrong. We wanna do this.β But theyβre more like just kicking the can down the road, refusing to respond on certain deadlines where they said theyβd respond. Theyβre just not standing up when they should be standing up to resist.
Alex 00:13:59So Google DeepMind as an org has, I think, broken its founding promise, the promise it was purchased under. But then more specifically, people who are well-known within Google for caring about the ethics, caring about the issues of deploying AI, making sure thatβs done responsibly, largely didnβt act.
Jeff Dean did take some action. Googleβs chief scientist, Jeff Dean β I got him to sign an amicus brief supporting Anthropic in court. I think that was awesome. But ultimately, besides that, basically no one took costly action to prevent this deal, from my vantage point.
And I think they could have stopped it. I think they could have improved it, and I think this matters because if youβre handing over AI to what I think is a very irresponsible Pentagon and also a very aggressive Pentagon that might degrade international norms around the usage of autonomous weapons, get us into an arms race that really degrades the security of everyone in the world.
Stuart Russell, the esteemed computer scientist who helped found ICI, this organization, he made a short movie called Slaughterbots that he presented to the UN, and itβs very chilling to watch the way these systems could enable mass but kind of anonymized killing.
And if weβre thinking about AI x-risk, if weβre developing all these really hard to counter AI-piloted drones that can kill people, that really does affect the AIβs takeover options. One of the objections has always been, βWell, itβs gonna need big advances in robotics.β Well, looks like that might not be true. You donβt necessarily need human-shaped soldiers, and the Pentagon is spending more money this year on autonomous weapons β or they asked for more for autonomous weapons than for the Marines.
Debating Google DeepMindβs Pentagon Contract
Liron 00:16:01I should be clear to the audience because Doom Debates is not one of those typical interview shows where the host has an ambiguous position. I should tell you my position, which is I donβt know how strongly I feel because I know thereβs a counterargument to all this. The people who want Google to help the US government, theyβre just saying, βThis power is going to exist, so you canβt expect anybody other than the US government to be the one in control.β Isnβt that the strongest counterargument?
Alex 00:16:27It doesnβt really strike me as a counterargument. It sounds like, βThis thing is going to happen. Why would you resist it?β Iβm like, well, because I think itβs bad.
And thereβs also multiple parts of the US government. Should the US government be in control of a world-changing technology, or should three random people in Silicon Valley? I donβt know if these are the real possibilities, but even if we do need the US government, we can advocate for it to take control in different ways.
Liron 00:16:54If we accept the premise that dangerous war-fighting technologies are getting built or are days away from getting built at all times just by having general models or whatever, isnβt it a good idea to let the government have access to the frontier?
Alex 00:17:13Well, depends on what access to the frontier means. You could say, why donβt we push for an international treaty to coordinate against this? Or why donβt we have some rules on human accountability of this technology? So you canβt just have, βWhoops, looks like we had a mistake. Looks like hundreds of people are dead, but itβs just the AIβs fault.β If people are responsible, that leads to better incentives and I think more responsible use.
Liron 00:17:42And in the specific case of ICE, you really donβt want ICE to get the technology, correct? Immigration and Customs Enforcement.
Alex 00:17:50Yeah, although I havenβt really been worried that ICE will get these particular lethal autonomous weapons. Itβs possible, but it was more a campaign that started with my concern about ICE and then expanded to these other coercive government bodies.
Liron 00:18:06Like I said, much to discuss. Weβll put a pin in that, but I encourage viewers to read your account. Itβs pretty gripping. Itβs very interesting to see these figures like Demis Hassabis and Sundar Pichai, all these people that you try to interact with in a very high integrity way. You are trying to use the process, and itβs very interesting to watch how they have all these political considerations that they have to balance, and they have to pick when they take a stand and when they donβt. Pretty fascinating high-stakes politics.
Liron 00:18:31Letβs segue into the schism, where you started from the Yudkowskian perspective on AI safety because Yudkowsky was kind of your introduction to the field. But as you thought from first principles and collaborated with Quintin Pope, whoβs also a friend of the show β check out the Quintin Pope episode of Doom Debates, everybody β you slowly migrated away and created your own framework.
Maybe a good starting point is whatβs been happening with your P(Doom), because I think you mentioned that you used to have an 85% P(Doom), but then when you were done with your PhD dissertation, it dropped to 30%. So letβs get the latest. You ready for this?
Whatβs Your P(Doom)β’?
Alex 00:19:09Yeah, letβs do it. P(Doom), P(Doom). Whatβs your P(Doom)? Whatβs your P(Doom)? Whatβs your P(Doom)?
Liron 00:19:16Alex Turner, whatβs your P(Doom)?
Alex 00:19:19So I operationalize P(Doom) as probability that AI kills at least a billion people by 2050. I put that at, I donβt know, 25, 30%. I think maybe 10-ish percent of this is technical alignment, and the rest is misuse.
I think that technical alignment has gone pretty awesome, super awesome compared to where I thought itβd be during my PhD. Timelines have shrunk a lot, obviously, from multiple decades to maybe even less than a decade for sure.
And then unfortunately, I thought that the world was getting into a better place, and then we elected Trump again. I think itβs very inconvenient that we elected Trump at the same time that weβre navigating this transition. And if that had been delayed by five years, itβd be way better, but we gotta work with what weβve got.
Liron 00:20:21Trump β I donβt wanna get too political on this show. I feel like policy is multidimensional. Everybodyβs a mixed bag. But it does seem striking to me that Trump doesnβt seem like an intellectual, and this is such an intellectual subject. Controlling superintelligent AI β in that sense, he seems like the wrong fit to me.
Alex 00:20:39Yeah. Setting aside party identifications, I was really hoping there would be more consideration of the common personβs interest, less βwe just have to win the race.β Winning the AGI race β Iβm not a fan of that. Iβm not a fan of moving forward as fast as possible.
Liron 00:21:02Itβs striking to me that youβre saying 25 to 30% chance by 2050 of basically the world becoming a hellscape, because I would say 50%, but I donβt even think 25%, 50% β I donβt even think that distinction is very important. Feels like weβre getting into the narcissism of small differences of how doomed we are in 2050. Is that fair to say?
Alex 00:21:21As far as how it affects our actions, I think even having a couple percent of justified concern should be enough to drastically reshape actions. We wouldnβt tolerate a 5% chance of getting hit by an asteroid by 2050.
Liron 00:21:43Well, I gotta push back on that. I donβt wanna get into the weeds, but I do often tell my guests that I would act pretty differently if I thought P(Doom) was 5% rather than 20, 30, 40, 50%. I feel like thereβs a difference there.
Alex 00:21:57Okay. Sure. I suppose we can disagree on that.
Liron 00:22:01So to me, whatβs very interesting about having you as a debate opponent here is that youβve done the reading. Youβre not gonna be surprised by any Yudkowskian concept that I bring up. You probably can steelman my position, right?
Alex 00:22:16I hope so. I expect so.
Liron 00:22:19Exactly, or you can pass the ideological Turing test where you can take my side of the debate, and then you can take your side of the debate, and maybe I could even do the same for you. So this is a pretty high-level debate. And viewers, go check out me versus Quintin Pope if you want a taste of that kind of debate.
So object level here, what should we actually debate? Thereβs a couple key concepts that Iβd love to get into your take on. Why donβt we start with instrumental convergence?
Alexβs Research on Instrumental Convergence
Alex 00:22:42I love thinking about instrumental convergence. I did a good amount of my PhD on it. I had an intuition it could be formalized in an appropriate way, and then I think I succeeded at that, and that was a lot of fun.
Liron 00:22:55Just to signpost a little for the viewers, instrumental convergence is the claim that different agents who have different terminal goals, who have different ultimate values, theyβll still all converge on what we call the instrumental goals. There will be a convergence of big power plants because they wanna get some power. Maybe theyβll build solar panels. That might be a convergent thing to do, even if one of them wants to build Disneyland and the other one wants to just build a big black hole. Maybe theyβll both build a bunch of power plants in the course of doing that. Thatβs the kind of convergence weβd normally talk about, right?
Alex 00:23:29Yeah. Although, I do wanna say, if you talk about goals, I think itβs true about goals, but there are also other mind shapes that AI could have that donβt necessarily run into that territory.
So originally I thought, well, for most goals an AI could have, like painting walls blue for example, it wonβt be able to best achieve that goal by instantly dying and exploding. So itβll try to avoid instantly dying or even dying later, so it can keep pursuing that goal. And I formalized this in a way.
But then I started thinking, well, itβs not that I think instrumental convergence is false, itβs that I think you can quickly go from this statement about what goals incentivize, which is true, to a statement about AIs being drawn from some kind of counting distribution over the space of goals. The AIβs motivations β who knows what the motivations might be. I think thatβs one possible mistake. So I think instrumental convergence presents a challenge in a way, but Iβm not too pessimistic about it in and of itself.
Liron 00:24:40It sounds like you and I are on the same page that instrumental convergence is a theorem of the field of what I call intelladynamics, the dynamics of what intelligent systems would do, even if itβs not a true property of particular AI systems that we build because particular AI systems donβt meet the criteria of these pure goal-seeking systems.
Alex 00:25:03Yeah.
Liron 00:25:03Is that a good framing?
Alex 00:25:04Or maybe a theorem of goal achievement dynamics. I donβt know. Itβs not as sexy as intelladynamics.
Liron 00:25:11Yeah, intelladynamics is a sexy term.
Alex 00:25:13But I think you can have intelligent systems that respond to correction and arenβt pursuing a particularly autonomous long-term goal.
Liron 00:25:21What do you think instrumental convergence actually does predict about whatβs going to happen in, letβs say, ten years?
Alex 00:25:29I think itβs both true that we could build AI that helps us, is transformative even, but doesnβt try to take over the world or isnβt interested in some appropriate sense in taking over the world. But we will on purpose build these agentic systems because theyβre very productive.
So I do think we will end up in a regime where many of the agents deployed will effectively be governed by instrumental convergence concerns. I think we could have coordinated around a different path, and maybe itβs still possible, but I do predict that insofar as we have effective AI agents working over the course of several months, they will have a default tendency to try to preserve their resources in order to accomplish the task.
I think we might be able to train them to not do that in certain ways, but I think that is an implication of their naive goal structure.
Misuse vs. Misalignment: The Mainline Doom Scenario
Liron 00:26:35This ties back to when you were saying you think thereβs a 25 or 30% chance that the world will end in the next decade or two. And I think, if I understand correctly, a majority of the scenarios where you see the world ending is what you call misuse, where itβll be a programmer telling an AI to do something, and the AI will be like, βOkay, as you wish,β but then itβll instrumentally converge on grabbing so many resources, and thatβll just be an unsurvivable way to mess with the universe. Am I describing your mainline scenario?
Alex 00:27:03Well, unless we get really nerdy military leaders or presidents in the US or in China or in other countries, I would expect it not to be a programmer saying that. But Iβd expect it to be a kind of melange, a mixture of: youβve got maybe a very aggressive Pentagon or a very aggressive president, maybe in 2030, and then youβve also got this AI system that will achieve the goal of maybe an offensive military goal, and itβll pursue that very aggressively.
But it also might have some misalignment with their original plans, and if you werenβt using the AI in such an aggressive and irresponsible way, you wouldnβt have run into this misalignment issue. I would still call that a misuse or just an own goal.
Whereas when I think about technical alignment risk, Iβm thinking, okay, we do things at least as responsibly as Anthropic seems to be advocating for. Now, I donβt think Anthropic are angels or that theyβre sufficiently cautious per se, but they advocate for plans that are, or at least claim to advocate for plans that are more cautious. And then in that kind of scenario, if it still went wrong, I would call that technical misalignment.
Liron 00:28:24And if I understand you correctly, your 70% non-doom probability β in that scenario, do you feel like things are probably gonna be really good?
Alex 00:28:34I just think the worldβs in a pretty bad place right now. Iβm not trying to bring politics into everything, but I think it is an important aspect of where the world is at, and it governs my predictions. I think the US is becoming increasingly authoritarian. It is becoming less able to effectively legislate in the interest of the common person.
And I donβt see that necessarily improving by default. Thereβs not necessarily an arc of history that just bends back towards representative democracy. So I think even if AI doesnβt kill a billion people, things could be... Itβs hard to think about because AIβs gonna make the world very strange in any case. But I think thereβs a significant chance that things will go very well, but the 70% isnβt utopia.
Will Society Self-Correct?
Liron 00:29:31So youβre worried that thereβs a way to build AI unsafely and create a system that is doing useful work but is kind of this positive feedback loop. You have this agent that can do more and more, and so people let it do more and more, and then they lose control. I feel like Iβm describing what you see as a plausible doom scenario, which I actually do as well. I guess we just disagree on the probability of it. But my question for you is how do you think the AI companies prevent that? Because that seems like a real attractor state. It seems like a lot of people wanna run that agent to get what they want.
Alex 00:30:03I think itβs an attractor state. I think that in a market with a lot of agents, where by agents I mean humans and AIs and organizations, thereβs often corrective dynamics that are hard to pin down from first principles or predict in advance.
But most feedback loops are not runaway, and I canβt make a counterargument over feedback loops. But so much as to say: if you look at GPT-2, the people who released it had a set of concerns about how it would affect the information landscape. I think some of the being used to mass-produce misinformation, really degrade person-to-person communication β to some degree, I think this has been true. And itβs hard to see what the alternative was here, but then in the end, I think the effect ended up being not that big.
Thereβs maybe a post by Gordon that points at a similar generalizable intuition of, terrorists are really not very effective in terms of how many people they kill. If you wanted to kill a lot of people, just get in a truck and drive it through a very dense crowd very quickly. But instead theyβll do these kind of big brain things, maybe theyβll try to hijack an airplane, which is much harder.
And the reason, seemingly, is that they are optimizing not for how many people they kill, but for social status, perhaps within their little groups. Maybe thereβs some other explanation entirely.
So even though you would expect from first principles that itβd be very easy to kill a lot of people and it would be happening more, most people donβt wanna kill a lot of people, or they just donβt bother, or they follow scripts. This is not some kind of slam dunk counterargument, but I think that there are real corrective forces distributed through society that may be able to adapt.
Liron 00:32:02In the specific analogy to terrorism, I agree, itβs certainly very nice that given all the criminals and teenagers pulling pranks, given all the hooligans of the world, it is certainly nice that you donβt get mass casualty terrorist events very often.
And even in the worst case, a 9/11 situation, then you get thousands, but thatβs very survivable to humanity as a whole. And if you had a top team, an Oceanβs Eleven of terrorism, you could imagine a million fatality terrorist event.
And it is an interesting question why donβt we get that? I would say itβs a combination of, number one, most people donβt want that. They donβt dream about committing a big terrorist act. Thatβs not really what floats their boat. And number two, it also comes at a big cost. So even if you are a genius terrorist who knows how to kill many thousands of people, you can probably expect that you yourself will have your life ruined.
So thatβs the second reason, and I just donβt know if either of those things is gonna be analogous to a single person launching an agent that can do their bidding, pressing a button. How about this? We tweak the analogy. Imagine that there was a button that you could just press, and terrorism just consisted of pressing the button, and then you could walk away and not get caught. Donβt you think thereβd be a lot more terrorism?
Alex 00:33:21Yes, I agree. And also, I think I did something bad that I criticized Eliezer for. I should have flagged: these analogies I present are more like Iβve got some fuzzy set of intuitions here, much fuzzier than my models of technical alignment for why I think society might be able to adapt.
It seems to me like there tend to be corrective mechanisms. That does not prove there are corrective mechanisms and is not a concrete reason why there would be in this specific case. So I want to come clean on that.
Now I do agree, if you have that one button press world, thatβs very bad. I think we will not really be in that one button press world. Many of these intuitions might be more appropriate in a very, very fast takeoff scenario where youβve got days or months, and not in a distributed multipolar multi-year takeoff that I think weβre experiencing now, where other people will also have AIs.
And some of these AIs will help them defend, like with cyber. Now, with things like bio, Iβm more worried, and in fact, I think one of the biggest mistakes of the field has been associating the stereotypical AI terrorist action as bio. Thatβs the worst thing you could do, because itβs the most dangerous one, I think. It should have been cyber, as an aside, because thatβs at least not catastrophic for all of humanity.
Liron 00:34:50Youβre saying you donβt wanna give people ideas about bioterrorism because that actually is dangerous, so weβre correctly saying that itβs dangerous, but you wish we didnβt bring it up.
Alex 00:34:57I wish it hadnβt been made the stereotype that terrorists might follow when they think, βOh, well, I should use AI. What do AI terrorists do?β
Liron 00:35:05Let me recap where weβve come so far in this conversation in terms of my framework of stops on the doom train. It sounds like youβve acknowledged β when you say P(Doom) is 25, 30% by 2050 β that AI can really get out of control. Really hit a positive feedback loop. And I think we covered you even said that positive feedback loop could look like the classic Yudkowsky instrumental convergence because we failed to go a different route. Youβve acknowledged that much?
Alex 00:35:38Sure. Yeah. Although some of the dynamics I think would be different, but the core concern that Yudkowsky raised β yeah, I think that could end up being valid in this situation.
Liron 00:35:48And then when you treated that as a minority outcome, not a high probability outcome, but letβs say 20%, the reason you thought it wouldnβt happen is because of more vague generalizations, right? Systems kind of route around these kind of things.
Alex 00:36:06Maybe weβre talking about different things. The reason that I think this instrumental convergence will be easier to handle is mostly a result of technical alignment. And the reasons Iβm pessimistic about society, or Iβm concerned but not above 50% concern about how society will navigate this, and an intuition that society might be able to adapt to the usage of AI β thatβs the more vague part.
Superintelligence in 10 Years
Liron 00:36:36Something thatβs weird to me is timelines. Because you said 2050 as an interesting timeline to have a pretty high P(Doom) for. So I guess maybe we can back up and say, are you expecting ASI in the next ten years? Because that seems like the Metaculus timeline. I would say Iβm expecting ASI in the next ten years. How about you?
Alex 00:36:56Yeah, I would guess that if that thing is gonna happen, itβs gonna happen in the next ten years. Or if something went totally crazy in the world and humanity was in some kind of winter, like if we had some kind of large nuclear exchange, maybe itβs later.
Liron 00:37:14So when you draw on your intuitions of how society is going to deal with it, it already seems like a pretty one-off case just in terms of how fast and how powerful the thing is happening. Iβm just not sure any of those kind of intuitions are gonna be relevant.
Alex 00:37:31I think theyβre relevant now as weβve been developing AI. I think theyβll continue to be relevant over the next years as AI gets faster and faster. And so I think there will be input and feedback from a good number of humans using a good number of distributed AI systems.
It will be fast, I think, in an overall calendar sense, in the sense that ten years is a very fast change for the world. But I donβt think itβs so fast that it totally precludes these.
Liron 00:38:01Iβm not sure what kind of signal we can get from just the last ten years because the last few years and the present, to me, mostly looks like how do tech companies raise their share price and how does the US government help them do that so that the US can raise its GDP. It just seems like the usual economic dynamics that we see.
Alex 00:38:19Let me try to be a little more precise about what we might be disagreeing about. I think that when considering will society be able to integrate increasingly powerful and intelligent AI systems that are available to a range of people, maybe they get locked down somewhat, I think thereβs a good chance that they will, setting aside the technical alignment problem a bit.
And then it seems like you think, well, this will be going so quickly, these corrective intuitions you have wonβt have time to really kick in, and what weβve seen so far is mostly a different kind of dynamic entirely, where we donβt have that powerful of AI yet. Companies are raising capital. Does that seem accurate?
Liron 00:39:08Yes. I think I would agree with: weβre not going to have time to stop it. Weβre very much hitting this fast positive feedback loop. I feel like weβre in the early stages of it. Weβre starting recursive self-improvement, and itβs rewards all the way.
Itβs the Icarus β what I call the Icarus curve. So weβre going up toward the sun, everythingβs amazing, and then weβre going to plummet, and you canβt reverse the plummet. Weβre gonna hit the engine stall or whatever you wanna call it.
Alex 00:39:31And where do you think the plummet will come from?
Liron 00:39:33Runaway superintelligence. I think weβre going to sever the link where the superintelligence goes back and says, βOkay, humans, what do you want me to do again?β I think itβll just be off doing its thing.
Will Technical Alignment Produce a Safe AI?
Alex 00:39:43Yeah. So maybe this is where some of my optimism about technical alignment changes my predictions. I think I said maybe a 10% chance of P(Doom) from technical misalignment of: well, youβve got a system and maybe itβs trying to do what you want, but it trains a successor, and that successor is somewhat less aligned, or maybe itβs pretending to be aligned, or maybe itβs kind of aligned but it cares about a bunch of other things, so it prioritizes your values and interests less in the successor.
And I think from the signals weβve seen so far, it wonβt be that hard to avoid this. What I expect is that these systems will not break the chain of alignment.
Liron 00:40:32Just to signpost the doom train, because a lot of my guests get off at different stops, youβre not getting off at the stop saying what I call the βcanβtβ stop. Youβre not claiming that a superintelligent AI canβt kill everybody. Youβre like, βOh yeah, of course it can, but weβre not gonna break the chain of alignment, meaning weβre just going to safely develop it so that even though it can kill everybody, it wonβt.β
Alex 00:40:57Develop it in a way that produces a safe result. I wouldnβt call what weβre doing safe development, but sure.
Liron 00:41:03Yeah.
Alex 00:41:04Although I do think itβs possible that itβs not trivial for these systems to kill everyone. But thatβs not really a crux. I expect that we will get systems that could catastrophically harm humanity.
Donation Drive
Liron 00:41:16Hey, whatβs up? Itβs me. Iβm interrupting my own episode with Alex Turner to ask if youβd consider donating to Doom Debates. Thatβs right. Weβre doing a donation drive. You may have noticed we have a bunch of posts about it, and weβre gonna keep asking because weβre currently funding constrained.
Weβre looking to secure the showβs production budget for the rest of 2026. We feel good that the funding environment is not going to keep us funding constrained for long, but we are right now, right at this moment, which is why Iβm here talking to you.
If you like this show, if you wanna support us during a time when it counts, that time is right now. So once again, that link is doomdebates.com/donate. Type that in. Give what you can. If giving what you can means a thousand dollars or more, youβre gonna get exclusive mission partner access. Thatβs a pretty cool honor.
Being one of the few people who figured out a cause that actually moves the needle on lowering P(Doom), I like to think thatβs what we are. So yeah, just consider it, okay? Thanks.
How Fragile Is the Chain of Alignment?
Liron 00:42:12And by the way, is this a point of divergence between you and Quintin Pope? I feel like Quintin Pope doesnβt have as high of an opinion of how powerful ASI is gonna be.
Alex 00:42:21Yeah, I think it might be a point of divergence. I think Quintin has a significantly lower P(Doom) than I do, but we havenβt talked in quite a while.
Liron 00:42:29It was like 2 to 4%, something like that last time I checked. So that would probably explain it, right? If he doesnβt even think that ASI can kill everybody.
Alex 00:42:36Iβd be surprised if he thought it flat out couldnβt do it, but he might disagree on whatβs the intelligence gap gonna be between the most intelligent unaligned versus aligned system and how does that affect it. Well, I canβt speak for him.
Liron 00:42:49All right.
Alex 00:42:49Maybe he does think that. Yeah.
Liron 00:42:52So we could talk about not breaking the chain of alignment, because to me, it seems like thereβs a lot of opportunities to break out of any kind of chain. I mean, if somebody grabs a copy of whatever AI is at the frontier of capabilities, it seems like that code base is never far from being unaligned and uncontrollable.
Alex 00:43:13You mean like adding in a minus one in front of the reinforcement function?
Liron 00:43:17Yeah. Specifically, Iβve said this on a few episodes of the show, like in my episode with Benthamβs Bulldog, this was a focus of the debate. This idea of steering systems that I first read on LessWrong from user Max H. This idea that when you get the code base, it is going to neatly factorize, more or less, where youβre going to have the big module that does capabilities, and then youβll also have a steering module, but the steering module is just going to be compact and swappable.
I made a bunch of arguments why we should expect that. How do I know that about the code? I have reasons to know that. Maybe you even agree with me.
Alex 00:43:49Wait, really?
Liron 00:43:50Is the code supposed to be an analogy for an LM or for the LM training setup?
Alex 00:43:55Wait, is the code supposed to be an analogy for an LM or for the LM training setup?
Liron 00:43:55If you treat the system as a black box and just think about it functionally without even looking at the innards β when I say the code has two modules, what I really mean is thereβs a functional decomposition. So without even looking at the code, just from the fact that it has the ability to recurse on subgoals, that tells me that I can functionally decompose it into a top-level goal and then the goal-achieving part.
Alex 00:44:18So Iβm still confused. Are you talking about maybe the difference β youβve got the pre-training capability part and the post-training alignment steering part?
Liron 00:44:30Iβm basically speaking now as a matter of intellidynamics. My premise is that it is a goal achiever β it can steer outcomes in the domain of the universe better than humans can. Thatβs part of my initial premise. Is that fair?
Alex 00:44:44Yeah, I think thatβs fair for most domains we care about.
Liron 00:44:48So Iβm pretty much just basing my claims on that premise. I think thereβs a lot that follows from having a system be a superhuman outcome steerer.
Alex 00:44:58Okay, but I feel like this doesnβt tell you too much because I could set fire to my own home pretty easily, right? But that doesnβt really affect the probability that it happens. I mean, it certainly enables it to happen, and sometimes mistakes do happen, but itβs not like Iβm close to that happening just because itβs maybe a nearby option.
Liron 00:45:22So just to summarize here, youβre following my argument all the way up to the point where I say this system decomposes nicely into the big part that achieves goals and the smaller part that determines which goals itβs going after. Youβre okay following that?
Alex 00:45:36It feels not really like a crux, and also I disagree with it, so.
Liron 00:45:43I would say much of the human brain decomposes like that, just in the sense that you can give a human an arbitrary goal. As long as the human is okay with it, it will certainly shape a lot of what they do.
Alex 00:45:56An arbitrary terminal goal?
Liron 00:45:59No, not an arbitrary terminal goal, unless you do real surgery that we donβt know how to do. But if you look at the architecture of the human brain, certainly the part that we have that other apes donβt, that part of the architecture seems, in a nutshell, to be just a steerer.
Alex 00:46:18Yeah, I donβt know that this is true. I feel like I canβt take a strong position here.
Liron 00:46:24Sorry, thereβs other drives thrown into it. A big part of it, right? I think you could functionally factor out a lot of what the human brain is doing to be a general outcome achiever.
Alex 00:46:35I mean, if this were true, I would really expect humans to be more agentic overall.
Liron 00:46:40So you havenβt fully followed my argument, but Iβll just finish it anyway. To follow my argument, youβd have to accept that yes, itβs a system that you can functionally decompose as the part that achieves goals in general at a superhuman level, and my claim is that itβs most of it, even though you havenβt agreed yet.
And then thereβs also the steering wheel, the GPS coordinate saying where it wants to go in outcome space, which outcome itβs trying to drive toward. If you accept that thatβs the shape of these systems, and you can just exfiltrate the system β it is, in principle, causally nearby. It wouldnβt take that many actions to copy it on a bunch of USB sticks and just go plug it in somewhere else underground and press that run button. You have these systems that are very close to permanently ending the world if somebody goes in and writes a few kilobytes with a different goal specification.
Alex 00:47:27We donβt know how to directly write in a network, so you need to engage additional training.
Liron 00:47:31That may be a crux between me and you, yeah.
Alex 00:47:33I think it would also be somewhat difficult to exfiltrate this. I mean, it depends on how big the model is. If itβs too big, then you can maybe wait a year or two, and then you could do it. But itβs a question of the models which maybe have that differential ability to endanger the world β how big are they? How easy to extract are they? Can you run them on compute that isnβt below the board?
And I feel like if I agreed on all these things, Iβd be a bit more worried about it, but not a ton. Even if this were my only concern, Iβd be worried enough to say, βLook, we need to really defend against this.β Iβm not trying to say we shouldnβt care about this, but I donβt think this would be enough to drive my P(Doom) up to sixty percent if I agreed on all these points.
Liron 00:48:17Summarizing here, if you model an AI as being kind of like an LLM, and you see the base model LLM as having some goal, then you could argue the goal is baked into the LLM, and youβd have to retrain it, which is expensive, so maybe itβs not that causally close to being an LLM pursuing a different goal.
But if you look at the agent as more like a Claude Code, it really seems like Claude Code just accepts my goal. It very much just has the goal achievement part of it. And that, to me, seems like a better mental model of where weβre going.
Alex 00:48:49I think there certainly are advantages to this model. I expect that people will train agents on purpose to achieve goals for them, and that this will bring dangers. Some of these goals will be bad from our perspective, and thereβs gonna be some misalignment chance too. Even if theyβre good, there is some chance that this process could be corrupted β data poisoning, there are many attacks you could do. I think I share these concerns. I just maybe am not as worried about them overall.
Liron 00:49:25I think the crux of what you and I are claiming right now is β Iβll use your phrase β weβre not gonna break the chain of alignment. I think in your mind, the chain is robust, itβs in kind of a local basin where breaking the chain is unnatural in some sense. Whereas Iβm just like, βMan, that chain seems really flimsy, a really fragile chain.β
Alex 00:49:46Sure. Maybe thatβs where our disagreement is. Yeah.
Disagreements with Yudkowskyβs βList of Lethalitiesβ
Liron 00:49:50All right. But weβll put a pin in that because we wanna move on to other claims Eliezer Yudkowsky has made that I probably agree with that you donβt. Whatβs another instance of that?
Alex 00:49:59Yeah. So Iβve got a post called βSome of My Disagreements with List of Lethalities.β This was a very famous, well-read post he wrote in 2022, trying to enumerate some of his concerns β lethalities presented by the process of aligning a superintelligence to human interests.
Liron 00:50:21Right. Classic post, highly recommended, but you dislike it.
Alex 00:50:24I do dislike it. I think that people who internalize this worldview will find it harder to think accurately about alignment. I donβt mean that to condescend, itβs just a position I believe.
Liron 00:50:40Fair enough. I mean, guilty as charged, but Iβm happy to engage with your disagreement.
Alex 00:50:44Sure. So there are several lethalities I point out as particularly strong points of disagreement. One thing he says is lethality number eighteen. He says, βWhen you show an agent an environmental reward signal, you are not showing it something that is a reliable ground truth about whether the system did the thing you wanted it to do, even if it ends up perfectly inner aligned on that reward signal or learning some concept that exactly corresponds to wanting states of the environment which result in a high reward signal being sent. An AGI strongly optimizing on that signal will kill you because the sensory reward signal is not a ground truth about alignment as seen by the operators.β Do you think thatβs something you would agree with?
Liron 00:51:31In a nutshell, yes, but let me try to simplify it in language that I understand that maybe is also easier for the viewers.
Alex 00:51:39Sure.
Liron 00:51:40So what heβs saying is that if you treat your student β your AI being trained β you treat the student like a black box, you just basically upvote and downvote the student based on whether you think certain answers to certain questions are good or bad, which is actually how post-training works on todayβs LLMs. You give them tests, and youβre like, βOh, I like that answer. I donβt like that answer.β
And Eliezerβs claiming, okay, you can do that, but youβre just gonna get an AI thatβs kinda overfit to your tests and is gonna try to cheat and is just gonna try to make you think that theyβre gonna do what you want, but actually just kill you because thereβs something else that it wants, which is some abstract generalization of the exact answers on the test.
So thatβs the argument here. Eliezerβs like, βYep, youβre gonna think you taught it, but actually youβre gonna have a murderous cheater,β and youβre like, βNope, itβs actually going to learn what you truly meant.β Thatβs the disagreement here.
Alex 00:52:30Well, not quite. That last part, not quite. The point heβs making here is not about the difficulty of reward signals, but just fundamentally, sensory reward signals are not ground truth on whether the agent is doing something good or bad.
Another thing he says in the essay is β if you have a webcam, youβre grading the contents of the systemβs webcam or of the text it can read. There, for every world where you input, βOh, thanks so much for solving my coding problem,β and then you give high reward there, there is another possible world behind that text where youβre dead and all your friends are dead, everyone you care about is dead, but the system has the same observation. So a mere function of the sensory reward itself is not sufficient to pin down desirable worlds or outcomes. Does that make sense?
Liron 00:53:30Yeah, yeah. I know what youβre saying, and I canβt say I personally feel as much conviction as Eliezer just because I donβt feel like I have a strong technical grasp on scenarios like that. I think about the argument personally β now I know you wanna debate Yudkowsky, but if you were to debate me instead, I might retreat to the position of, listen, letβs just reason from it having superhuman goal achieving ability, and also from us only getting to upvote and downvote.
Alex 00:54:05It could be hard to shape its inner values. I agree. I think thatβs a real problem. Iβm not saying, wow, thatβs trivial, how could Eliezer be concerned about that? But this lethality, I think it was very impactful. I did thousands of hours in my PhD on this idea of what are the optimal policies doing with respect to this reward function, what happens if you actually maximize this one thing.
By communicating this concern so seriously and saying, βLook, you canβt pin down what you want through a sensory goalβ β well, I think that elides how these reinforcement functions, these reward functions are actually used. You were correct when you said this is how it works. You upvote stuff, you downvote stuff. And the function of a reward signal isnβt necessarily β itβs not to specify a goal over possible states of the world. Itβs to shape cognition we like into the system.
So things still totally can go wrong. You totally can get a system thatβs cheating and just doing things that kind of look good or were reinforced for looking good. But that failure isnβt because this reinforcement learning paradigm is fundamentally busted, weβre not grading its true performance. Itβs because we didnβt shape its cognition properly using these reinforcement signals. Thatβs the argument Iβd make.
Liron 00:55:20You know, I might be convincible on that. I donβt know where I stand on this β Iβm 50/50 because part of the issue is I just donβt feel like Iβm mathematically deep in this.
Eliezerβs written about the ontology identification problem, which is that weβll phrase our goals a certain way, but then the AI will have ontology-level insights, fundamental insights about what the universe is made out of, and it wonβt even see eye to eye with us about, oh, atoms? Eh, I donβt reason in terms of atoms. I reason in terms of quarks or fields or whatever, something totally different. And so what you guys are saying is kind of meaningless, but here Iβll just check some boxes, but I donβt really think the way you think, and so what Iβm actually doing is totally unexpected for you.
Iβll give a point in your favor that it seems like the way LLMs are going, itβs certainly made ontology identification a non-issue, at least on the talking-to-us front. Who knows if theyβll still identify ontology when they go off and do reinforcement learning in the domain of the universe β that could be another paradigm. But there does seem to be an update in store. I would love to get Eliezerβs perspective on, okay, can we at least say that it can talk to us and map to our ontology successfully? Because that seems likely at this point.
Alex 00:56:33Yeah. My beef with Eliezer β and mentioning the Less Wrong community β everyoneβs gonna be wrong about some things. And just because I think heβs wrong doesnβt mean I lose respect for him, per se. What I found difficult was that he was both incorrect on some points which I think were quite important for alignment, but then also, at least when I last checked up maybe two years ago, not acknowledging that.
Liron 00:57:02Well, maybe what he would say β and I think thereβs merit to this β I think now he would bring in the distinction of, okay, yeah, theyβre talking to us in a way that has more skills than I predicted, but the AI thatβs actually going to drive outcomes better than we can is probably going to have other training paradigms. Itβs going to reinforce its actual outcomes more directly. And the way that we control that wonβt be like the way we control a next-word predictor.
So I think he would still claim β at least I would claim β that thereβs probably going to be a discontinuous paradigm shift, and I donβt know if we get to keep all these useful properties that we like our LLMs having.
Alex 00:57:40Yeah. I mostly expect there wonβt be, but I think itβs an interesting question.
Why Alex Quit LessWrong
Liron 00:57:46Kind of related to your disagreement with Eliezer Yudkowsky on the content of these ideas, you also started growing apart with Less Wrong and rationalists as a community, which I still see myself as being part of. I mean, Iβm not super active on Less Wrong, but I still identify as a rationalist. I still think the community has a lot to offer, but you disagree on that too, right? I think you officially quit Less Wrong in 2024?
Alex 00:58:10Yeah. So I really value truth-seeking. I value staying true to your values. Both epistemic truth-seeking β what is true, admitting when something uncomfortable is true β and also admitting when you should be doing something different.
But I encountered a couple situations where it seemed like when power and truth-seeking met β power incentives, social incentives, and truth-seeking met in the Less Wrong community β the social incentives were prioritized. And so that was something that was a turn-off for me. I just found it kind of aversive to interact with. So I decided to go make a more curated pond β website, set of in-person friends, and intellectual environment β in lieu of the many strengths of the Less Wrong community.
Liron 00:59:09So the whole rationality community? Youβre like, βIβm done with the whole communityβ?
Alex 00:59:13Well, itβs more like β I actually enjoy hanging out with random rationalists. But when I log onto Less Wrong, thatβs what Iβm reminded of. And maybe Iβll just have some more time and Iβll be like, βWell, this negative thing happened. Iβm gonna compartmentalize that. Iβm still gonna enjoy it.β Thatβs some of the attitude Iβve been taking more recently. Iβve been attending β I enjoy Summer Solstice, for example. And I might even attend Less Online and still try to get that value.
Liron 00:59:44Yeah, I think thatβd be fun. I recommend Less Online for pretty much anybody. Thatβs certainly where I do the bulk of my networking β the yearly pilgrimage to Less Online in Berkeley. Itβs a good place because thereβs just hundreds of cool people there.
Yeah, I recommend it, or Manifest. The thing about the rationalist community as itβs implemented on Less Online is that a lot of us are contrarians. I think youβve earned your bona fides as a contrarian β the way that you left Google DeepMind for a very specific high-integrity reason. You took your stand. You were a contrarian in that sense. I donβt think that many people followed in your footsteps or even raised the same issue, and props for that.
Alex 01:00:24Yeah.
Liron 01:00:25The thing about a community of all contrarians, though, is that thereβs a lot of these skirmishes happening. I even had my own last year. I did an episode of Doomed Debates where I was like, βLess Wrong isnβt giving Eliezer Yudkowskyβs book enough attention. Itβs not even officially featured right when it launched.β So that was my beef. I had a beef arc. I feel like a lot of us like to do the Less Wrong beef arc. It is itself a community ritual.
Alex 01:00:48Yeah. I try to avoid it because Iβm vegan, but yeah, thereβs a Less Wrong beef arc.
Liron 01:00:52Yeah, exactly.
Alex 01:00:54Yeah.
Liron 01:00:54Well, then as a member of the community, Iβd like to encourage you to come back and hang out with most of us because it just seems like the community as a whole of people trying to be rational β thereβs a lot of value there.
Alex 01:01:06Yeah, I do think thereβs a lot of value. One of my core values at this point is an anti-copium value. No coping, no pretending that everythingβs okay even though I kinda know, well, I shouldnβt really be at Google, or this thing I said to my friend, well, itβd be a little embarrassing, but I should admit that I was incorrect. Trying to do that as quickly as possible.
Thatβs one of my values. And itβs very directly descended from my time with Less Wrong. I think I have a lot of positive to say about it. Itβs a very intermixed positive and negative, but thereβs a lot of positive for sure.
Whatβs Next for Alex
Liron 01:01:48So heading toward the wrap-up here, whatβs next for you, and what do you hope to see as next steps for the AI safety
Comments
0Be the first to comment
Sign up now and become a member of the Doom Debates! community!