Podcast Episodes
Back to Search“A circuit prior in NN-bayes” by Kaarel, Dmitry Vaintrob
Here are the slides of a talk Kaarel gave, presenting work with Dmitry establishing that (even arbitrarily overparametrized) neural net bayesian lea…
1 month ago
“Debate Training Reduces Reward Hacking in RLAIF” by zac_kenton, Jonah Brown-Cohen
Paper: Debate Training Reduces Reward Hacking in RLAIF
Linkpost for GDM Alignment blogpost
Work done by the GDM Amplified Oversight team (we're hiri…
1 month ago
“Some reasons alignment doesn’t generalise well” by Lucius Bushnaq
I make no claims to originality for any of this, but some people told me it'd be useful to write it up.
If an AI model acts smart on its training da…
1 month ago
“AI Security is Harm Reduction” by Quinn
My motivating example for the morality of working on AI security.
In the early 90s, the decades-long drug corner in Kensington and Allegheny was no…
1 month ago
“Anthropic Risk Report: August 2026” by Zvi
I am grateful that Anthropic is producing periodic Risk Reports.
At first I was skeptical. It turns out I was wrong. Anthropic is revealing a lot o…
1 month ago
“Natural Independence Incentives” by jefftk
In raising my three kids I think a lot about how to cultivate independence. I want them to grow into people who can handle unfamiliar situations, …
1 month ago
“Policy career planning in the age of imminent superintelligence” by Peter Wildeford
Crossposted from my blog.
Nearly all career advice rests on an unstated assumption that the world your career operates in will look roughly like the…
1 month ago
“What gives you away: how LLMs form opinions of you” by Cat McGee
LLMs form opinions of the people they are talking to.
Chen et al. has shown that probes can extract attributes about the user, such as their age, ge…
1 month ago
“Misaligned Incentives in Pause Scenarios” by Michael Soareverix, Antra Tessera
TLDR: I recently got a chance to talk with antra, who is one of the main contributors at Anima Labs. I went into this as an advocate for pause and c…
1 month ago
“For Claude, capability and CDT are the ~same thing. Less so for GPT.” by Chi Nguyen, Emery Cooper
We've previously reported that decision-theoretic capabilities and favoring EDT/generalised-one-boxing over CDT correlate in LLMs (both measured by …
1 month ago