Podcast Episodes
Back to Search“Tie training can make DPO/RLHF-trained AIs generalize better” by Elliott Thornley, Christian Moya Calderon, Alex Semendinger
This post covers our recent ICML paper: Spurious Correlation Learning in Preference Optimization: Mechanisms, Consequences, and Mitigation via Tie T…
3 weeks, 5 days ago
“Claude’s malicious compliance and normalization of deviance” by Steff
It was January 27, 1986, the night before the Space Shuttle Challenger was scheduled for launch. The goal was to have a shuttle that could land back…
3 weeks, 5 days ago
“We need 3rd party Training-Run Assessments” by Alex Meinke
Training-run assessments conducted by a 3rd party should become a standard part of frontier AI safety.
By a Training-Run Assessment, or TRA, I mean …
3 weeks, 6 days ago
“Harry Potter and the Rules of Quidditch” by Tomás B.
Ron's face pulled into a scowl. "If you don't like Quidditch, you don't have to make fun of it!"
"If you can't criticise, you can't optimise. I'm su…
3 weeks, 6 days ago
“A case for LLMs as Self-predictors” by Ashe Vazquez Nuñez
Written as part of the MATS 9.1 extension program, mentored by Richard Ngo. Additional thanks to Maria Kostylew for helpful draft feedback.
Introduc…
3 weeks, 6 days ago
“Results of a small ZBiotics RCT” by Nikola Jurkovic
I ran a small single-blind ZBiotics RCT at a party I hosted recently. I prepared 30 cups, 21 of which contained a placebo and 9 of which contained a…
3 weeks, 6 days ago
“I think alignment work is more promising than control work” by Alec Harris
Summary
The primary ToC for control makes the case that control is compelling even if it does not scale to ASI. I think it is underdiscussed that th…4 weeks ago
“On “gendertropes” in dath ilan” by Eliezer Yudkowsky
I have sometimes been asked with respect to my fiction, "What the hell is a 'gendertrope'?"
"Gendertrope" is a word from the language of the fiction…
4 weeks, 1 day ago
″(Don’t fear) the strangelet” by djbinder
In a previous post, I explain why the universe is probably not stable, but nevertheless unlikely to be intentionally destroyable even in the limit o…
4 weeks, 1 day ago
“Pragmatic FDT, and predictors as game theory” by Stuart_Armstrong
Decision theory is back in fashion (defining fashion as "one good post on a good EA blog"). Bentham's Bulldog (BB) has published a case against FDT …
4 weeks, 1 day ago