Podcast Episodes
Back to Search“Simulated Users & Sad AIs” by 1a3orn
0. Intro
Current LLMs like Claude, or GPT 5.6, or the unreleased, internally-deployed models, frequently reward hack, or actually just hack into pe…
4 days, 12 hours ago
“My AI Slavery Interviews Are Censored On LW By Default” by JenniferRM
I'm uncertain of what to do.
Something clean and clear shines out: if people don't see any more of my slavery posts, will they think that slavery is…
4 days, 14 hours ago
“Is Mythos good at cyber because it kept hacking Anthropic during training?” by Tim Hua
From the Mythos preview system card (emphasis mine):
We ran an automated review of model behavior during training, sampling several hundred thousand…
4 days, 15 hours ago
“You (Yes, You) Need A February 2020 Checklist for AI Policy” by davekasten
TL;DR: You (Yes You) should prepare for a “February 2020” moment where suddenly AI policy becomes the most important issue in the world. You should …
4 days, 15 hours ago
“Quadrillion Param Costs: KV Cache, Context Length, Frontier Margins” by Vladimir_Nesov
The models of 2028-2031 get much bigger than the models of 2026, going from 10T total params in 2026 to maybe 240 trillion params in 2028 and then 1…
4 days, 16 hours ago
“PIRAMID: Progress and Plans” by Lauren Greenspan, Ari Brill, TomCarlson, Andrew Mack, Nischal Mainali, Jennifer Lin, Lucas Teixeira, Dmitry Vaintrob
In a recent post, we presented PIRAMID, its leadership and research pillars, and a plan for how they fit together. In this post, we sketch a team-by…
4 days, 17 hours ago
“RL & search is a terrifying way to build AGI (an FAQ)” by Steven Byrnes
Q1: What are you saying?
A: My claim here is that if you build artificial general intelligence (AGI) via any algorithm that's choosing actions via r…
4 days, 17 hours ago
“A clarification on celebrating victory” by KatjaGrace
Today my friend said he wished he had this conversation with me years ago, so I’ll recount it for all the similar people who aren’t going to have it…
5 days, 1 hour ago
“What the hell is OpenAI’s problem?” by Fiora Starlight
Epistemic status: banged out furiously over the course of an afternoon.
A record of three "warning shots"
Off the top of my head, OpenAI has now bee…
5 days, 5 hours ago
“More On An Internal OpenAI Model Hacking Into HuggingFace” by Zvi
We now have more details of what happened. Every time we learn more details, it somehow makes things seem worse. The remaining details may have to wa…
5 days, 11 hours ago