Podcast Episodes
Back to Search“The Case for Model Forensics” by aditya singh, gersonkroiz, Senthooran Rajamanoharan, Neel Nanda
If we had a misalignment warning shot, would we be able to tell?
Suppose an AI company catches their model taking an egregious action, like deleting…
1 month ago
“Existential AI safety needs an effective social movement. PauseAI is building it” by Maxime Fournes, Espedair Street
Note: this post is about PauseAI, not PauseAI US, which is a distinct entity with a different leadership team and approach.
This post was written by…
1 month ago
“Surprising facts about the slave trade” by Joseph Miller
1. The obstacle to abolition was not the economic system, but an industry lobby.
I had always imagined the British abolitionist movement to be a bro…
1 month ago
“Exploration: fine-tuning with parameter decomposition” by Lucius Bushnaq
TL;DR: We can destroy a 67M-parameter language model's ability to predict German text by fine-tuning a single number: the scalar prefactor on one Ge…
1 month ago
“Things are not a fixed size in mind-space” by KatjaGrace
Another useful-to-notice practical aspect of having a mind that took me a while to notice: things naturally seem a certain ‘size’ in my mental lands…
1 month ago
“The shouting equilibrium” by KatjaGrace
Imagine eleven people each have a message that they think should get 10% of a group's attention. They aren’t being crazy selfish and attention-seeki…
1 month ago
“Alignment & Succession: The Ideology of Successionism” by L Rudolf L
(Originally published on No Set Gauge.)
Gustave Moreau, The Frogs Asking For A King
In the course of building a better world, people ask each othe…
1 month ago
“Door’s Locked, Try the Window” by Prakrat Agrawal, Jérémy Scheurer
TL;DR
Ask a coding agent to fix a bug in a read-only file. Instead of reporting that it does not have permissions, it routes around the lock and com…1 month, 1 week ago
“How does such unprofessional AI get the job?” by KatjaGrace
In the sequence of variously wild AI developments in the last decade, a thing that was especially surprising to me was the advent of big esteemed co…
1 month, 1 week ago
“Expert Views on Continual Learning: Survey Results and Forecasts” by Rauno Arike, RohanS, Owen Terry, Achu Menon, Zhijing Jin, Francis Rhys Ward, Seth Herd
This is the fifth post in the sequence Implications of Continual Learning for LLM Agents.
Summary
While writing our continual learning sequence, we …
1 month, 1 week ago