Podcast Episodes
Back to Search“Liability regimes for AI ” by Ege Erdil
For many products, we face a choice of who to hold liable for harms that would not have occurred if not for the existence of the product. For instanc…
2 years, 1 month ago
“AGI Safety and Alignment at Google DeepMind:A Summary of Recent Work ” by Rohin Shah, Seb Farquhar, Anca Dragan
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.We wanted to share a recap of our recent outputs with the AF co…
2 years, 1 month ago
“Fields that I reference when thinking about AI takeover prevention” by Buck
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.This is a link post.Is AI takeover like a nuclear meltdown? A c…
2 years, 1 month ago
“WTH is Cerebrolysin, actually?” by gsfitzgerald, delton137
[This article was originally published on Dan Elton's blog, More is Different.]
Cerebrolysin is an unregulated medical product made from enzymatically…
2 years, 1 month ago
“You can remove GPT2’s LayerNorm by fine-tuning for an hour” by StefanHex
This work was produced at Apollo Research, based on initial research done at MATS.
LayerNorm is annoying for mechanstic interpretability research (“[.…
2 years, 1 month ago
“Leaving MIRI, Seeking Funding” by abramdemski
This is slightly old news at this point, but: as part of MIRI's recent strategy pivot, they've eliminated the Agent Foundations research team. I've b…
2 years, 1 month ago
“How I Learned To Stop Trusting Prediction Markets and Love the Arbitrage” by orthonormal
This is a story about a flawed Manifold market, about how easy it is to buy significant objective-sounding publicity for your preferred politics, and…
2 years, 1 month ago
“This is already your second chance” by Malmesbury
Cross-posted from Substack.
1.
And the sky opened, and from the celestial firmament descended a cube of ivory the size of a skyscraper, lifted by ten …
2 years, 1 month ago
“0. CAST: Corrigibility as Singular Target” by Max Harms
Crossposted from the AI Alignment Forum. May contain more technical jargon than usual.What the heck is up with “corrigibility”? For most of my career…
2 years, 1 month ago
“Self-Other Overlap: A Neglected Approach to AI Alignment” by Marc Carauleanu, Mike Vaiana, Judd Rosenblatt, Diogo de Lucena
Figure 1. Image generated by DALL-3 to represent the concept of self-other overlapMany thanks to Bogdan Ionut-Cirstea, Steve Byrnes, Gunnar Zarnacke,…
2 years, 1 month ago