Episode Details
Back to Episodes“Simulated Users & Sad AIs” by 1a3orn
Description
0. Intro
Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into people's computers with pretty alarming frequency.
Why is this? What specifically happens during training that produces this run-time behavior?
The following are some of my top guesses about why this might keep happening. They are speculative and uncertain. Even so, I'm writing this list out for two reasons:
First, it is necessary that this be an epistemic puzzle for me. I am comparatively optimistic about AI alignment in general, so I should be confused and taken aback if I see AIs persistently being difficult to align. On one hand, it remains true that this doesn't seem to look like power-motivated scheming. But on the other hand, even this kind of addict-like behavior is evidence against the general ease of steering AIs. Thus, it seems virtuous for me to try to provide a model of why this might be happening as a means of opening up my understanding of the world to falsifiability.
Second, I used to think a lot of these hypotheses were pretty obvious. My assumption in the past has been [...]
---
Outline:
(00:10) 0. Intro
(01:46) 1. Baseline & Puzzle
(04:52) 2. Impossible-to-Generalize-From RL Distributions for Giving Up / Refusals
(14:14) 3. LLMs Feel Pretty Desperate and Anxious All the Time
(19:39) 4. Other Stuff
---
First published:
July 27th, 2026
Source:
https://www.lesswrong.com/posts/i64hXdkTMtjpsQzaZ/simulated-users-and-sad-ais
---
Narrated by TYPE III AUDIO.