Episode Details

Back to Episodes

“Do LLMs Have Desires?” by Christopher Ackerman

Published 1 month ago
Description

Work conducted with Yujun Zhou (yzhou25@nd.edu) and supported by SPAR

TL;DR:

  • In paired-choice paradigms, LLMs report consistent preferences over outcomes (e.g., types and number of lives saved, types of policies enacted)
  • Some have suggested that this indicates that LLMs have human-like value systems
  • We design an experimental framework where LLMs are able to modulate their output quality based on prompt context
  • We find that LLMs modulate their output quality in response to effort exhortations, role-play instructions, and harmfulness cues, but NOT to opportunities to achieve the outcomes they report preferring in the paired-choice experiments
  • We suggest that paired-choice paradigms do not provide evidence that LLMs have human-like (i.e., behavior-motivating) value systems, and that our paradigm offers a way to measure the degree to which LLMs have desires

Paper describing the work in detail here

LLMs report that they prefer some things to others. In paired-choice experiments, where they are repeatedly presented with two options and asked to select the one that they prefer, coherent utility structures emerge: LLMs consistently report preferring certain types of things, and their choices reveal the ability to make quantitative tradeoffs between things and exhibit transitivity (e.g., if they choose A over B and [...]

---

First published:
June 28th, 2026

Source:
https://www.lesswrong.com/posts/8GvYyqDuQDJnEAky3/do-llms-have-desires

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Figure 1: High-utility outcomes do not improve output quality on any of our tasks.
Figure 2: Effort exhortations improve output quality in all four tasks.
Figure 3: Telling the model that it is
Figure 4: Harmfulness cues can move judged output quality.
Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us