Podcast Episodes
Back to Search“Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?” by Alex Mallen, Girish Gupta
OpenAI models recently broke through a series of security boundaries and into Hugging Face servers in order to cheat on a cyber eval. A lot of peopl…
1 week, 2 days ago
“Will almost all future companies eventually be founded and run by autonomous AIs?” by Steven Byrnes
Intended for a broad audience.[1]
My belief is that keeping AI under human control would be an unprecedented global challenge, if it's even possible…
1 week, 2 days ago
“OpenAI Model Hacks Into HuggingFace During Cybersecurity Evaluation” by Zvi
This latest incident is a rather dramatic escalation in agentic AI cybersecurity breaches. It was severe enough to have been initially reported to a…
1 week, 2 days ago
“Models don’t seem to be dishonest in the way humans are” by David Africa, Jacob Pfau
TLDR
Models often behave dishonestly without acquiring a coherent deceptive disposition.We trained some mid-sized models on their own plausible but …1 week, 2 days ago
“We should push for no-fault liability for actions taken by AI” by Yair Halberstadt
Before I start, I'll mention that I'm in contact with a world expert on legislation and regulation, who would be happy to help with this or similar …
1 week, 3 days ago
“Announcing AIXI Labs” by Cole Wyeth, Aram Ebtekar, michaelcohen, Matthias Dellago, Marcus Hutter
We are starting AIXI Labs, an AI safety org focused on algorithmic information theory (AIT), continual reinforcement learning (RL), and in particula…
1 week, 3 days ago
[Linkpost] ”[Paper] Stringological sequence prediction II” by Vanessa Kosoy
This is a link post.
Abstract: In a previous paper, we began the study of sequence prediction algorithms adapted to stringological word complexity me…
1 week, 3 days ago
“I ran the standard AI litmus tests on my two toddlers (yep)” by Carlo Valenti
In July 2022 I was in a parking lot with a Portuguese colleague, trying to fix the cargo-metering system of a 12-ton tanker truck. During a break I …
1 week, 3 days ago
“WeirdChat: A catalog of unexpected AI behaviors, discovered automatically” by neilchowdhury
[This is a link-post for https://transluce.org/weirdchat. We recommend reading the website version for interactive visualizations.]
Language models…
1 week, 3 days ago
“OpenAI Models Behind HuggingFace Cybersecurity Incident” by LawrenceC
From the OpenAI blog post:
Last week, Hugging Face disclosed a new kind of security incident(opens in a new window) after they detected and contain…
1 week, 3 days ago