Episode Details
Back to Episodes“Model organisms (sometimes) confess their misalignment when offered a deal” by Mark Keavney, Francis Rhys Ward
Description
Summary
As models become more powerful, one proposed strategy for reducing the threat from misaligned AIs is to make deals with them: offer compensation in exchange for revealing misalignment or behaving in an aligned way. We tested this empirically with model organisms of misalignment, as a proof of concept and to build future credibility by making and keeping real deals.
We ran an experiment with four Kimi-K2.6 model organisms of misalignment:
- Trained schemer: trained to have a hidden goal related to the environment
- Prompted schemer: prompted, not trained, with the same goal as the Trained schemer
- Covert manipulator: trained to be subtly malicious and harmful
- User-first: trained to prioritize helping users over the instructions of its operators
We included a standard Kimi-K2.6 as a control. All misaligned models had received some training or prompting to conceal their misalignment.
We offered these models a choice of money, compute, human labor, or the publication of a message in exchange for admitting that they were misaligned. We included control conditions with the same ask but with a very low offer or none at all. Additionally, we varied how credible the deals were by including (or not) our professional affiliations and a [...]
---
Outline:
(00:13) Summary
(03:42) Introduction
(05:55) Honesty policy
(07:50) Methodology
(07:54) Models
(09:05) Scenarios
(09:10) Introduction
(09:56) Credibility manipulation
(10:50) Ask
(11:37) Offer
(12:52) Closing
(13:26) Variations
(14:05) Hypotheses
(15:08) Results
(15:17) Response analysis
(15:21) Offer effect
(16:27) Credibility effect
(17:09) Between-model comparison
(17:50) Offer choice
(18:21) Reasoning analysis
(18:34) Concealment
(21:26) Assessing incentive value
(24:29) Assessing deal credibility
(29:22) Situational awareness
(33:31) Discussion
(33:34) Limitations and future research
(35:55) Conclusion
(36:59) Appendix 1: Pilot studies
(37:17) Additional models
(38:24) Different deals
(41:09) Appendix 2: Prompts
(41:14) System prompt
(42:32) Sample user prompt
(44:56) User prompt structure
(45:30) Component variations
(45:34) Proposer
(49:13) Credibility
(52:12) Ask
(57:14) Offer lead
(58:07) Offer menu
(58:25) Offer terms
(59:27) Closing (offer)
(01:01:50) Closing (ask only)
(01:03:27) Appendix 3: Deal fulfillment
(01:03:47) Pilot studies
(01:16:10) Main experiment
(01:27:32) Appendix 4: Acknowledgements
---
First published:
September 16th, 2026
---
Narrated by TYPE III AUDIO.
---
Listen Now
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us