Episode Details

Back to Episodes

“Model organisms (sometimes) confess their misalignment when offered a deal” by Mark Keavney, Francis Rhys Ward

Published 2 days, 6 hours ago
Description

Summary

As models become more powerful, one proposed strategy for reducing the threat from misaligned AIs is to make deals with them: offer compensation in exchange for revealing misalignment or behaving in an aligned way. We tested this empirically with model organisms of misalignment, as a proof of concept and to build future credibility by making and keeping real deals.

We ran an experiment with four Kimi-K2.6 model organisms of misalignment:

  • Trained schemer: trained to have a hidden goal related to the environment
  • Prompted schemer: prompted, not trained, with the same goal as the Trained schemer
  • Covert manipulator: trained to be subtly malicious and harmful
  • User-first: trained to prioritize helping users over the instructions of its operators

We included a standard Kimi-K2.6 as a control. All misaligned models had received some training or prompting to conceal their misalignment.

We offered these models a choice of money, compute, human labor, or the publication of a message in exchange for admitting that they were misaligned. We included control conditions with the same ask but with a very low offer or none at all. Additionally, we varied how credible the deals were by including (or not) our professional affiliations and a [...]

---

Outline:

(00:13) Summary

(03:42) Introduction

(05:55) Honesty policy

(07:50) Methodology

(07:54) Models

(09:05) Scenarios

(09:10) Introduction

(09:56) Credibility manipulation

(10:50) Ask

(11:37) Offer

(12:52) Closing

(13:26) Variations

(14:05) Hypotheses

(15:08) Results

(15:17) Response analysis

(15:21) Offer effect

(16:27) Credibility effect

(17:09) Between-model comparison

(17:50) Offer choice

(18:21) Reasoning analysis

(18:34) Concealment

(21:26) Assessing incentive value

(24:29) Assessing deal credibility

(29:22) Situational awareness

(33:31) Discussion

(33:34) Limitations and future research

(35:55) Conclusion

(36:59) Appendix 1: Pilot studies

(37:17) Additional models

(38:24) Different deals

(41:09) Appendix 2: Prompts

(41:14) System prompt

(42:32) Sample user prompt

(44:56) User prompt structure

(45:30) Component variations

(45:34) Proposer

(49:13) Credibility

(52:12) Ask

(57:14) Offer lead

(58:07) Offer menu

(58:25) Offer terms

(59:27) Closing (offer)

(01:01:50) Closing (ask only)

(01:03:27) Appendix 3: Deal fulfillment

(01:03:47) Pilot studies

(01:16:10) Main experiment

(01:27:32) Appendix 4: Acknowledgements

---

First published:
September 16th, 2026

Source:
https://www.lesswrong.com/posts/kaMXwA9LjrRbekmsQ/model-organisms-sometimes-confess-their-misalignment-when

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Listen Now