Podcast Episodes
Back to Search“A proposal for a highly effective AI safety org” by ceselder
TLDR: an org that pays people to "just read the fucking transcripts"; a large amount of people reading anonymized claude code/RL/eval transcripts fl…
2 weeks, 4 days ago
“Talking to journalists” by KatjaGrace
A common view around me seems to be that journalists are frequently dishonorable and dangerous, and talking to them is a risk to be avoided unless y…
2 weeks, 4 days ago
[Linkpost] “Resolution has a new Agent Foundations team” by Jeremy Gillen
This is a link post.
The team will include me (Jeremy Gillen), Abram Demski, Sam Eisenstat, Scott Garrabrant and Kaarel Hänni. We'll soon recruit add…
2 weeks, 4 days ago
“If you’re interpreting <1B parameter models, you should use a tensor transformer” by Logan Riggs
To all my fellow researchers doing SLT, computational mechanics, one of ARC's programs, natural abstractions/condensation, proofs on NNs (or any int…
2 weeks, 4 days ago
“Kairos has raised $50M to build talent infrastructure for AI safety (and we’re hiring!)” by agucova
TL;DR: Kairos has raised 50 million dollars from Coefficient Giving for two years of funding, one of the largest commitments they’ve made towards AI…
2 weeks, 4 days ago
“Anthropic Has Some Alignment Problems” by Zvi
Oh, good. They noticed.
Anthropic, too, is planning to bring METR inside for an independent review of their own incidents, where three times a Clau…
2 weeks, 4 days ago
“Incoherent AI Identities can also be Stable” by Ashe Vazquez Nuñez
In this post, I extend some experiments from "The Artificial Self " (TAS) to find that incoherent identities, delivered to models as system prompts,…
2 weeks, 4 days ago
“How concerned should we be about OpenAI’s recurrent architecture rumors?” by Rauno Arike
Yesterday, The Information reported that OpenAI's upcoming model, Astra, is built with a looped transformer architecture. Given that Zvi sounds (und…
2 weeks, 4 days ago
“Early handoff? Improve conceptual reasoning? [Diagram]” by Cleo Nardo
This article is about:
How do we (more) safely defer to AIs? (Ryan Greenblatt, Julian Stastny)AI 2040: Plan A, Alignment Roadmap (Ryan Greenblatt, T…2 weeks, 4 days ago
“Fake voices: warping the social world” by KatjaGrace
In 2020 I wrote a list of flavors of badness generally represented by advertising. The one I thought about most later on was probably #4:
Cultural p…
2 weeks, 4 days ago