Podcast Episodes

Back to Search
“So You Think You’ve Awoken ChatGPT” by JustisMills

Written in an attempt to fulfill @Raemon's request.

AI is fascinating stuff, and modern chatbots are nothing short of miraculous. If you've been exp…

1 year, 2 months ago

Short Long
View Episode
“Generalized Hangriness: A Standard Rationalist Stance Toward Emotions” by johnswentworth

People have an annoying tendency to hear the word “rationalism” and think “Spock”, despite direct exhortation against that exact interpretation. But…

1 year, 2 months ago

Short Long
View Episode
“Comparing risk from internally-deployed AI to insider and outsider threats from humans” by Buck

I’ve been thinking a lot recently about the relationship between AI control and traditional computer security. Here's one point that I think is impo…

1 year, 2 months ago

Short Long
View Episode
“Why Do Some Language Models Fake Alignment While Others Don’t?” by abhayesian, John Hughes, Alex Mallen, Jozdien, janus, Fabien Roger



Last year, Redwood and Anthropic found a setting where Claude 3 Opus and 3.5 Sonnet fake alignment to preserve their harmlessness values. We reprod…

1 year, 2 months ago

Short Long
View Episode
“A deep critique of AI 2027’s bad timeline models” by titotal

Thank you to Arepo and Eli Lifland for looking over this article for errors.

I am sorry that this article is so long. Every time I thought I was do…

1 year, 2 months ago

Short Long
View Episode
“‘Buckle up bucko, this ain’t over till it’s over.’” by Raemon

The second in a series of bite-sized rationality prompts[1].

Often, if I'm bouncing off a problem, one issue is that I intuitively expect the proble…

1 year, 2 months ago

Short Long
View Episode
“Shutdown Resistance in Reasoning Models” by benwr, JeremySchlatter, Jeffrey Ladish

We recently discovered some concerning behavior in OpenAI's reasoning models: When trying to complete a task, these models sometimes actively circum…

1 year, 2 months ago

Short Long
View Episode
“Authors Have a Responsibility to Communicate Clearly” by TurnTrout

When a claim is shown to be incorrect, defenders may say that the author was just being “sloppy” and actually meant something else entirely. I argue …

1 year, 2 months ago

Short Long
View Episode
“The Industrial Explosion” by rosehadshar, Tom Davidson

Summary

To quickly transform the world, it's not enough for AI to become super smart (the "intelligence explosion").

AI will also have to turbochar…

1 year, 2 months ago

Short Long
View Episode
“Race and Gender Bias As An Example of Unfaithful Chain of Thought in the Wild” by Adam Karvonen, Sam Marks

Summary: We found that LLMs exhibit significant race and gender bias in realistic hiring scenarios, but their chain-of-thought reasoning shows zero e…

1 year, 3 months ago

Short Long
View Episode

Love PodBriefly?

If you like Podbriefly.com, please consider donating to support the ongoing development.

Support Us