Podcast Episodes
Back to Search“The Bad Guy With An AI Named Claude” by Zvi
A lot of bad guys try to use Claude to do bad things. Mostly they fail. We think.
Anthropic has disrupted a bunch of them, and offers an extensive …
3 days, 12 hours ago
“Quick notes from teaching technical profiles how to talk in public” by Camille B.
Status: written in a hurry as people are getting showered with interviews re AI Safety and superintelligence, and I thought it may help a few people…
3 days, 13 hours ago
″[Cross-post] Palisade Podcast episode “How to Actually Influence AI Policy (No Law Degree Required) — with Matthew Lipka”” by davekasten
Palisade Research has launched a podcast series! I'll be hosting a series of episodes where I interview people who are experts in a functional or su…
3 days, 15 hours ago
“Inoculation Midtraining with Learned Neologisms” by Kyle O’Brien, Edward James Young, Puria, Nathalie Kirch, Cam, Tomek Korbak, David Africa
TL;DR
In our new paper, we demonstrate that we can achieve selective generalisation of misalignment by midtraining Nemotron 120B on synthetic docume…
3 days, 16 hours ago
“Cooperation with AIs seems to be a low-hanging fruit for better evals” by Clément Dumas
Summary
In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prom…
3 days, 18 hours ago
“Astra appears to perform belief-propagation-like inference without CoT” by MBaert
tl;dr I tested GPT-6 Astra on randomized Boolean logic problems. Astra can solve surprisingly complex logic problems without chain-of-thought, and i…
4 days, 9 hours ago
“OpenAI President Brockman says HuggingFace incident model had not been alignment-trained” by Caspar Oesterheld
On today's episode of the podcast "Odd Lots", OpenAI President Greg Brockman said (at around 8:40): "This model that did/had the HuggingFace inciden…
4 days, 13 hours ago
“Model Weight Exfiltration Seems Overrated” by Vaniver
[Epistemic status: a hot take that I’ve shared at the lunch table twice. People at the lunch table made slight updates instead of being convinced.]
…
4 days, 14 hours ago
“OpenAI Says It’s Not Responsible for the Leading the Future Super PAC. But Only Its Employees Seem to Believe That.” by garrison
This is the full text of a post first published on Obsolete, a Substack that I write about the political economy of AI. I’m a freelance journalist a…
4 days, 14 hours ago
“Op-Ed: I Worked at Google DeepMind. You Should Listen to the Warnings About AI” by TurnTrout
Published in The Guardian.
Major AI lab CEOs recently advocated for pacing AI development. They are right to be concerned: the field runs an extreme…
4 days, 23 hours ago