Episode Details

Back to Episodes

“Sub-agent delegation chaining” by David Rein

Published 3 weeks, 5 days ago
Description

Epistemic status: pretty confident in the validity of the core proposal, not that confident in specific implementation details

TL;DR: we should cryptographically verify that sub-agent instances/sessions are downstream of human instructions

Frontier AI labs have started using LLM-based monitoring systems to check for misbehavior from their internal AI agents, which often run unwatched by humans for hours or days. One central reason to be concerned about rogue internal deployments, where AIs subvert the monitoring system, is that the unmonitored AI could do a lot of work in service of its misaligned goals, without needing to dress up this work as being innocuous/harmless to fool its AI/human monitors. This work could be general reasoning/strategizing, finding additional useful vulnerabilities in the AI company's infrastructure, doing jailbreaking research (e.g. so it could fool monitors if necessary), doing ML research into training data poisoning, etc.

But a lot of this work will be expensive for the agent—doing it well will require a large amount of inference compute. We might be pretty optimistic that it’ll be difficult for a rogue AI to acquire this much unmonitored inference compute—eventually people will start noticing! This raises a question for the rogue agent: is it possible to [...]

---

Outline:

(04:23) Appendix

(04:26) Sketch/outline of one defense strategy

The original text contained 5 footnotes which were omitted from this narration.

---

First published:
July 6th, 2026

Source:
https://www.lesswrong.com/posts/98GvRu78jTXJgz9gA/sub-agent-delegation-chaining

---

Narrated by TYPE III AUDIO.

---

Images from the article:

The inference server only accepts requests that derive from a valid human instruction, which we can ~guarantee via hardware MFA whenever you start a new agent session.
A rogue agent could piggy-back on a real human instruction (e.g. by stealing the key the monitored agent uses to sign the sub-agent call), but now it needs to make its task seem like a natural/normal part of the human's initial instruction, which reduces to the normal monitor evasion challenge.

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Listen Now

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us