Podcast Episodes

Back to Search
“On the Rationality of Deterring ASI” by Dan H

I’m releasing a new paper “Superintelligence Strategy” alongside Eric Schmidt (formerly Google), and Alexandr Wang (Scale AI). Below is the executive…

1 year, 6 months ago

Short Long
View Episode
[Linkpost] “METR: Measuring AI Ability to Complete Long Tasks” by Zach Stein-Perlman

This is a link post. Summary: We propose measuring AI performance in terms of the length of tasks AI agents can complete. We show that this metric ha…

1 year, 6 months ago

Short Long
View Episode
“I make several million dollars per year and have hundreds of thousands of followers—what is the straightest line path to utilizing these resources to reduce existential-level AI threats?” by shrimpy

I have, over the last year, become fairly well-known in a small corner of the internet tangentially related to AI.

As a result, I've begun making what…

1 year, 6 months ago

Short Long
View Episode
“Claude Sonnet 3.7 (often) knows when it’s in alignment evaluations” by Nicholas Goldowsky-Dill, Mikita Balesni, Jérémy Scheurer, Marius Hobbhahn

Note: this is a research note based on observations from evaluating Claude Sonnet 3.7. We’re sharing the results of these ‘work-in-progress’ investig…

1 year, 6 months ago

Short Long
View Episode
“Levels of Friction” by Zvi

Scott Alexander famously warned us to Beware Trivial Inconveniences.

When you make a thing easy to do, people often do vastly more of it.

When you put …

1 year, 6 months ago

Short Long
View Episode
“Why White-Box Redteaming Makes Me Feel Weird” by Zygi Straznickas

There's this popular trope in fiction about a character being mind controlled without losing awareness of what's happening. Think Jessica Jones, The …

1 year, 6 months ago

Short Long
View Episode
“Reducing LLM deception at scale with self-other overlap fine-tuning” by Marc Carauleanu, Diogo de Lucena, Gunnar_Zarncke, Judd Rosenblatt, Mike Vaiana, Cameron Berg

This research was conducted at AE Studio and supported by the AI Safety Grants programme administered by Foresight Institute with additional support …

1 year, 6 months ago

Short Long
View Episode
“Auditing language models for hidden objectives” by Sam Marks, Johannes Treutlein, dmz, Sam Bowman, Hoagy, Carson Denison, Akbir Khan, Euan Ong, Christopher Olah, Fabien Roger, Meg, Drake Thomas, Adam Jermyn, Monte M, evhub

We study alignment audits—systematic investigations into whether an AI is pursuing hidden objectives—by training a model with a hidden misaligned obj…

1 year, 6 months ago

Short Long
View Episode
“The Most Forbidden Technique” by Zvi

The Most Forbidden Technique is training an AI using interpretability techniques.

An AI produces a final output [X] via some method [M]. You can analy…

1 year, 6 months ago

Short Long
View Episode
“Trojan Sky” by Richard_Ngo

You learn the rules as soon as you’re old enough to speak. Don’t talk to jabberjays. You recite them as soon as you wake up every morning. Keep your …

1 year, 6 months ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us