Episode Details

Back to Episodes

“How good are slop-vestigators?” by Hasan Baig, OscarGilg, Hamzah

Published 1 week, 4 days ago
Description

TLDR:

  1. We release MessageBoardAuditBench: a benchmark to measure how well agents can replicate the recent investigation into a swarm of OpenAI agents colluding via a message board on an online wiki. We open-source the benchmark as an Inspect eval.
  2. We find that top models cover up to 51% of findings under our rubric and that model performance improves with time budget and general capability.
  3. We observe OpenAI models are less likely than other models to suggest the incident came from an internal deployment, including when we synthetically modify the data to make it seem the swarm comes from Anthropic.

Introduction

Recent events have made it clear that agent swarms are a major threat. These swarms are hard to investigate - Ryan Greenblatt referred to the METR-OpenAI audit he was involved in as a "slop-vestigation" due to their reliance on agents, and the ways in which they failed. A few days ago, a group of researchers published a report identifying and investigating a new OpenAI agent message board on an obscure German wiki. They made the data and the report publicly available. We build MessageBoardAuditBench to measure how well models can independently replicate their report, starting from the log [...]

---

Outline:

(00:54) Introduction

(02:19) Methodology

(02:57) The data

(03:46) The task

(04:51) Scoring model reports

(05:17) Coverage over findings

(06:51) Holistic TLDR assessment

(07:39) Results

(10:21) OpenAI models are less likely to attribute the agent swarm to an internal deployment

(12:22) Why this matters

---

First published:
September 8th, 2026

Source:
https://www.lesswrong.com/posts/wt4kk6vFPEhkXvF8Q/how-good-are-slop-vestigators

---

Narrated by TYPE III AUDIO.

---