Episode Details

Back to Episodes

“Evaluating Explanations of LLM Behavior In The Wild with Counterfactual Experiments” by Adam Karvonen, Euan Ong, Subhash Kantamneni, Sam Marks

Published 4 weeks, 1 day ago
Description

TL;DR: We introduce CHIVE, an agentic pipeline that discovers unexpected LLM behaviors in the wild and explains them with counterfactual prompt edits. We use the resulting data in two ways. Using it as an evaluation, we find that activation-reading interpretability tools provide no uplift: agents given the tools predict the outcomes of these experiments no better than agents that just read the transcript. Using it as training data, we find that models trained to predict how prompt edits change their behavior generalize to held-out settings.

📄 Paper, 💻 Code


Figure 1. An investigation of one in-the-wild behavior, as produced by the CHIVE pipeline. Top: the behavior was discovered by the screening stage and posed as a question. Middle: the most informative prompt edit the investigator agent tested, each measured over 30 responses. Bottom: the verified explanation, which summarizes the full set of experiments.

Introduction

Many areas of AI safety, such as interpretability and chain-of-thought faithfulness, aim to explain model behaviors. But what makes an explanation of a behavior good? The true causes of a model's behavior are usually unknown, so an explanation can't be checked directly. In this work, we evaluate explanations through the lens of counterfactual [...]


---

Outline:

(01:31) Introduction

(03:37) CHIVE: a pipeline for discovering counterfactual explanations for model behaviors

(06:06) Interpretability tools provide no uplift on our evaluation

(07:53) Why don't the tools help?

(08:49) How should we interpret these results?

(12:26) Training models to predict their own behavior

(14:27) In summary

---

First published:
August 21st, 2026

Source:
https://www.lesswrong.com/posts/ExB6KYDcznaFS72eT/evaluating-explanations-of-llm-behavior-in-the-wild-with

---

Narrated by TYPE III AUDIO.

---