Podcast Episodes
Back to Search“Subliminal Learning: LLMs Transmit Behavioral Traits via Hidden Signals in Data” by cloud, mle, Owain_Evans
Authors: Alex Cloud*, Minh Le*, James Chua, Jan Betley, Anna Sztyber-Betley, Jacob Hilton, Samuel Marks, Owain Evans (*Equal contribution, randomly …
1 year, 2 months ago
“Love stays loved (formerly ‘Skin’)” by Swimmer963 (Miranda Dixon-Luinenburg)
This is a short story I wrote in mid-2022. Genre: cosmic horror as a metaphor for living with a high p-doom.
One
The last time I saw my mom, we m…
1 year, 2 months ago
“Make More Grayspaces” by Duncan Sabien (Inactive)
Author's note: These days, my thoughts go onto my substack by default, instead of onto LessWrong. Everything I write becomes free after a week or so…
1 year, 2 months ago
“Shallow Water is Dangerous Too” by jefftk
Content warning: risk to children
Julia and I knowdrowning is the biggestrisk to US children under 5, and we try to take this seriously.But yesterda…
1 year, 2 months ago
“Narrow Misalignment is Hard, Emergent Misalignment is Easy” by Edward Turner, Anna Soligo, Senthooran Rajamanoharan, Neel Nanda
Anna and Ed are co-first authors for this work. We’re presenting these results as a research update for a continuing body of work, which we hope wil…
1 year, 2 months ago
“Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety” by Tomek Korbak, Mikita Balesni, Vlad Mikulik, Rohin Shah
Twitter | Paper PDF
Seven years ago, OpenAI five had just been released, and many people in the AI safety community expected AIs to be opaque RL age…
1 year, 2 months ago
“the jackpot age” by thiccythot
This essay is about shifts in risk taking towards the worship of jackpots and its broader societal implications. Imagine you are presented with this…
1 year, 2 months ago
“Surprises and learnings from almost two months of Leo Panickssery” by Nina Panickssery
Leo was born at 5am on the 20th May, at home (this was an accident but the experience has made me extremely homebirth-pilled). Before that, I was on…
1 year, 2 months ago
“An Opinionated Guide to Using Anki Correctly” by Luise
I can't count how many times I've heard variations on "I used Anki too for a while, but I got out of the habit." No one ever sticks with Anki. In my…
1 year, 2 months ago
“Lessons from the Iraq War about AI policy” by Buck
I think the 2003 invasion of Iraq has some interesting lessons for the future of AI policy.
(Epistemic status: I’ve read a bit about this, talked to…
1 year, 2 months ago