Podcast Episodes
Back to Search“Automated alignment runs are hard to study!” by Alejandro Aristizabal, draganover, Aleksandr Bowkis, Cameron Holmes
TL;DR: This post presents three case studies of automated alignment research runs at Arcadia Impact. We use these case studies to emphasise the foll…
1 month ago
“Free will is like temperature” by Optimization Process
Free will is like temperature: a useful tool for analyzing the behavior of certain systems which are too big and complicated to model in exact detai…
1 month ago
[Linkpost] “Patterns and problems in emerging multiagent systems (Anthropic, Frontier Red Team)” by Julian Bradshaw
This is a link post.
Linkpost for some new Anthropic research on how agents coordinate (or don't). Not too long, pretty interesting. For example:
Th…
1 month ago
“Measuring Spurious Correlations with Feature Strength” by egan
This work was partially done by an automated research scaffold developed at Redwood Research. For this project, all of the experiment ideas were des…
1 month ago
“Introducing the Conceptual Reasoning Index” by Chi Nguyen, Emery Cooper, Caspar Oesterheld, Alex Kastner, Joe Benton
Associated announcement tweet.
We are planning to release blog posts properly arguing the case for this kind of work in the future.
tl;dr
A core hop…
1 month ago
“Demon Safety” by Ben Pace
(by LemmySmackett)
"Hey man, I haven't seen you in a minute. What are you up to these days?"
"Been on that grind, bro. I got a new gig."
"Really? …
1 month ago
“AI swarms are starting to pose indirect takeover risk” by oakhu, Alex Mallen
OpenAI's cyberattack on Hugging Face turns out to have been the result of many agents, in distinct training and evaluation contexts, coordinating fo…
1 month, 1 week ago