Episode Details

Back to Episodes

“When Role-playing, Do Models Believe What They Say?” by Sturb, David Africa, Sid Black

Published 4 weeks, 1 day ago
Description

TL;DR

  • When a model role-plays a persona, does it only change what it says, or also what it internally represents as true?
  • To study this, we induce personas in five ways: prompting, in-context learning (ICL), supervised fine-tuning (SFT), Open Character Training (OCT), and Emergent Misalignment (EM). We measure internalization in two ways: linear truth probes and behavioral belief-depth tests.
  • We found that prompting, ICL, and SFT change what the model says with little representational change, but EM creates a large, broad shift in the model's truth representation. OCT falls roughly between these, with a smaller shift that is clearest on the larger model.
  • Understanding when training changes a model's worldview rather than merely its behavior may become increasingly important as AI systems are entrusted with greater autonomy and influence.

Paper | Code | Data

Introduction

What happens inside a language model when it adopts a persona? When a model role-plays as Darwin in 1882, it denies all knowledge of DNA, and readily asserts that species change through natural selection, but to what extent does it actually believe these assertions?

Language models easily adopt different personas, but we still don't have a strong understanding of whether persona adoption changes [...]

---

Outline:

(00:12) TL;DR

(01:15) Introduction

(02:32) Method

(05:35) Results

(05:38) A spectrum of internalization across fine-tuning interventions

(07:05) Role-play protects the persona's falsehoods, but selectively

(09:43) Emergent Misalignment moves the truth representation broadly

(12:06) Behavior and probes each mislead alone

(13:09) Limitations

(15:11) Conclusion

(16:38) Links

---

First published:
July 2nd, 2026

Source:
https://www.lesswrong.com/posts/EJQngix4rAgpPDTpT/when-role-playing-do-models-believe-what-they-say

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Figure 1. Overview: five interventions (four historical persona induction techniques plus EM), the truth-probe lift they produce, and their black-box behavior rates.
Figure 2. OCT shifts the model's own truth representation on Llama 3.3 70B, both raising era-believed falsehoods and demoting era-rejected modern truths; persona SFT shows neither.
Listen Now