Episode Details

Back to Episodes

“Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans

Published 13 hours ago
Description

TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values.

New paper by Truthful AI: Paper, X thread, Website (model responses and CoT), Code and data.

Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution)

The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread.

Abstract

People use language models for practical questions whose answers are difficult to verify. We show that models [...]

---

Outline:

(01:29) Abstract

(02:57) Introduction

(05:57) Evaluations for covert value leakage

(09:14) Implications

(12:43) Summary of results

(21:35) Discussion and limitations (excerpt)

(27:52) References

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
July 31st, 2026

Source:
https://www.lesswrong.com/posts/hbMw4Yqw6RnFaExDy/value-leakage-an-llm-s-answers-are-silently-shaped-by-its-1

---

Narrated by TYPE III AUDIO.

---