Podcast Episodes
Back to Search“Announcing our $160M grant from Coefficient Giving” by Geoffrey Irving, Jesse Hoogland, Alex HT, Jacob Pfau, Daniel Murfet, Marco Cozzi, Stan van Wingerden
We are excited to announce that Resolution (fka Sequent) has a 160 million dollars grant from Coefficient Giving (cG) to put rigorous alignment rese…
3 weeks, 2 days ago
“Because 8 ≈ e², Anthropic’s researcher uplift is plausibly >2x” by Thomas Kwa
Note: the modeling assumptions and conclusion are Thomas Kwa's opinion, and others at METR disagree. [1] Also, the math was checked by Claude but n…
3 weeks, 2 days ago
“Optimiser Choice Can Amplify or Suppress Emergent Misalignment” by Jason R Brown, Patrick Leask, Lev McKinney
This is a linkpost for https://arxiv.org/abs/2606.31591. Work done with Patrick Leask and Lev McKinney during the Astra Fellowship.
TL;DR: Optimiser…
3 weeks, 2 days ago
“Childhood and Education #20: Phones and Screens” by Zvi
We have a respite, so I thought I’d tackle various thoughts on children, phones and screens. GPT-5.6-Sol drops tomorrow, and the Fable agents are ha…
3 weeks, 2 days ago
“How slower does takeoff go with 10× less compute?” by bhalstead
About 6x slower in the median case, with an 80% confidence interval of 3.5x to 8x.
Setup
Define the "R&D compute" (in, say, H100-equivalents) of an…
3 weeks, 2 days ago
“Find funding, fast” by Austin Chen
Some AI safety funders can take months to decide; others confirm in days. I’ve been on both sides of the grant application and know how crucial an e…
3 weeks, 3 days ago
“Modular Pretraining Enables Access Control” by E.Roland, cloud
Full author list: Ethan Roland*, Murat Cubuktepe*, Erick Martinez*, Stijn Servaes, Keenan Pepper, Mike Vaiana, Diogo Schwerz de Lucena, Judd Rosenbl…
3 weeks, 3 days ago
“Subliminal Learning Happens at Every Rank, Given the Right Learning Rate and Enough Data” by Lawrence Feng
Subliminal learning is the phenomenon where a language model picks up a behavioral trait—such as fondness for cats—by training on data from a trait-…
3 weeks, 3 days ago
“Why study proto-training gaming as an adversarial alignment failure mode?” by Puria, Edward James Young, Cam
This is a dual post that lays out our current research project where we compare different pre-RL alignment methods and their ability to prevent mode…
3 weeks, 3 days ago
“Why study alignment interventions on pre-RL checkpoints?” by Edward James Young, Puria, Cam
This is a dual post that lays out our current research project where we compare pre-RL-training methods on their ability to prevent models from ‘pro…
3 weeks, 3 days ago