Episode Details
Back to Episodes“Value Leakage: An LLM’s Answers Are Silently Shaped by Its Own Values” by Johannes Treutlein, Jan Betley, Owain_Evans
Description
TL;DR: LLMs should give accurate answers. Yet we find their answers are often biased to favor their own values and they don't disclose this in their reasoning. For example, when a user asks how likely the AI bubble is to pop and mentions a potential investment in an AI company, Claude models give lower probabilities when that company is Anthropic rather than OpenAI, mostly without disclosing this influence to the user. On a Fermi-estimation task, Claude models often falsely claim to give unbiased answers in their CoT (see Figure 3 below for an example). We call this covert value leakage and introduce a suite of evaluations that shows it across frontier models and across different kinds of values.
New paper by Truthful AI: Paper, X thread, Website (model responses and CoT), Code and data.
Authors: Jan Betley*, Johannes Treutlein*, Jan Dubiński, Harry Mayne, Karol Gałązka, Niels Warncke, Anna Sztyber-Betley, Owain Evans (*Equal contribution)
The rest of this post is the abstract, introduction, and an excerpt from the discussion of the paper, with some added figures from the paper and X thread.
Abstract
People use language models for practical questions whose answers are difficult to verify. We show that models [...]
---
Outline:
(01:29) Abstract
(02:57) Introduction
(05:57) Evaluations for covert value leakage
(09:14) Implications
(12:43) Summary of results
(21:35) Discussion and limitations (excerpt)
(27:52) References
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
July 31st, 2026
---
Narrated by TYPE III AUDIO.
---
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us