Podcast Episodes
Back to Search“Towards surfacing model algorithms with meta-tokens in the J-Space” by agam_bhatia
TL;DR
We used J-lens on Qwen3.6-27B to find “meta-tokens”: tokens that surface non-obvious computation in the model. When the model reads ambiguous …
1 week, 3 days ago
“OpenAI Shares Some Alignment Problems” by Zvi
Kudos to OpenAI for sharing their recent experiences with a misaligned internal model, where they encountered problems sufficiently severe they were…
1 week, 3 days ago
“What do I mean by “Artificial General Intelligence”?” by Steven Byrnes
In this post, intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will seem obvious to…
1 week, 3 days ago
“What do I mean by “Artificial General Intelligence”?” by Steven Byrnes
In this post,[1] intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will seem obvious…
1 week, 3 days ago
“What do I mean by “Artificial General Intelligence”?” by Steven Byrnes
In this post,[1] intended for a broad audience, I will paint a brief picture of what I’m talking about when I talk about “AGI”. It will seem obvious…
1 week, 3 days ago
“Epistemics and Coordination: It’s complicated!” by Raymond Douglas
Money is pouring in, people are looking for new areas to fund, and the invisible hand is starting to grab a bit at AI for epistemics and coordinatio…
1 week, 3 days ago
“Differential acceleration of alignment-relevant capabilities is a bad bet” by Zephaniah Roe
There is an idea floating around in the rough shape of "we need to accelerate capabilities that are differentially useful for safety research so AIs…
1 week, 3 days ago
“Measuring Reward-Seeking via Contrastive Belief Updates” by Jérémy Scheurer, Axel Højmark, jenny, Felix Hofstätter, Theodore Ehrenborg, Bronson Schoen, Alex Meinke
Machine learning models can produce the right outputs for the wrong reasons. Famous examples include a reinforcement learning agent that, rewarded f…
1 week, 3 days ago
“11 Open Empirical Problems in Reward-Seeking” by Alex Meinke, Jérémy Scheurer, Axel Højmark, Theodore Ehrenborg
We recently published our paper on "Measuring Reward-Seeking via Contrastive Belief Updates". We're excited about research like this, and there are …
1 week, 3 days ago
“Demis Hassabis on the New Coming Age” by Zvi
Google CEO Demis Hassabis offered us a first rate second rate essay, A Framework for Frontier AI and the Dawning of a New Age. I’ll go over that ess…
1 week, 4 days ago