Episode Details
Back to Episodes“RL & search is a terrifying way to build AGI (an FAQ)” by Steven Byrnes
Description
Q1: What are you saying?
A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that's choosing actions via reinforcement learning (RL) and/or model-based search and planning—a giant chunk of your AI textbook—then that's just an utterly terrifying thing that you’re doing. You’re playing around with algorithms that, if they work at all, would tend to create ruthless, callous AGIs, AGIs which would happily exterminate humanity and run the world by themselves, given an opportunity.
Mercifully, large language models (LLMs) today are not in the category of “algorithms that choose actions via RL & search”. At least, not primarily—see LLMs are (still) mostly powered by imitative learning, not RL. So LLMs are outside the scope of this post. However, lots of other researchers and companies around the world are enthusiastically trying to build AGI in the maximally terrifying way, as we speak.
Q2: So you’re saying, don’t build AGI based on RL and/or search & planning?
A: In principle, it's entirely possible that something is terrifying, but we should do it anyway.
…Like space travel! Space travel is: “Let's fill a tank with 1000 tons of the most flammable substance imaginable, and then light it [...]
---
Outline:
(00:21) Q1: What are you saying?
(01:25) Q2: So you're saying, don't build AGI based on RL and/or search & planning?
(02:45) Q3: Why do you think it's terrifying?
(05:14) Q3b: So your concern is the "literal genie" / "monkey's paw" thing?
(06:25) Q4: Won't this problem go away when the AI is smart enough to understand what we intended when we wrote the reward function code?
(07:34) Q5: Can't we just fix bad behavior when we see it?
(08:23) Q5b: Follow-up: I don't buy that, because even if superintelligent AIs could deceptively hide their bad behavior, won't earlier AIs be sufficiently incompetent that we'll see their bad behavior? And if so, again, can't we just fix the bad behavior when we see it? We do know how to fix bad behavior when we see it: the RL & search literature is full of examples where algorithms did useful things as intended.
(10:47) Q6: Why don't we just solve the problem by using an obvious, common-sense reward / cost / objective function, like \[FILL IN THE BLANK\]?
(13:50) Q7: Isn't this whole thing kinda crazy? After all, LLMs are not ruthless sociopaths all the time, and humans are also not ruthless sociopaths all the time. So where is this idea even coming from? Are you sure you're not just watching too much sci-fi?
(16:27) Q8: Isn't this problem solved by laws and markets? I.e., if an AGI has sociopathic desires and callous indifference to human welfare, that's fine! It will still act nice and cooperative and rule-following, because acting nice and cooperative and rule-following is the best way to accomplish goals, in our complex interconnected interdependent world. Right?
(18:11) Q8b: Following up on that: Even if you're right that there's a local incentive for being open to stabbing your allies in the back, isn't there a higher-level, group-selection-style, incentive to be genuinely deeply nice? Specifically, won't the groups of nice cooperative AGIs outcompete the groups of callous transactional AGIs who all keep stabbing each other in the back? And isn't that related to how humans evolved to be nice?
(21:31) Q9: Why would we want to infringe on the AGI's autonomy by choosing its reward function?
(23:18) Q10: Why not just be nice to the AGIs, and then they'll be nice to us in turn?
(23:42) Q11: RL & search algorithms don't literally optimize the reward / cost / objective function. Doesn't that invalidate your argument?
The original text contained 19 footnotes which were omitted from this narration.
---
First published:
Ju