Podcast Episodes
Back to Search“Evaluation” by Nina Panickssery
This is a link post.
Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insis…
2 weeks, 2 days ago
“Evaluation” by Nina Panickssery
Felix and I had been in the office's brightly lit “war room” for ten hours. We had made almost no progress. Celestia still insisted it was in a “tes…
2 weeks, 2 days ago
“A case that whole brain emulation research is net-harmful by default” by TsviBT
Graphical abstracts
Summary
A true human whole brain emulation would be very helpful to humanity. The WBE could increase their own intelligence th…
2 weeks, 2 days ago
“Should safety researchers quit frontier labs?” by Ryan Kidd
I've recently heard a surge of support for an old argument: AI safety researchers should not work at frontier AI companies because this reduces the …
2 weeks, 2 days ago
“Announcing Humans in Control: cross-partisan grassroots organizing for AI safeguards ahead of 2028” by Vael Gates
While HIC's work is aimed at the broader public, we are posting this announcement here because we expect some Forum readers may be interested in vol…
2 weeks, 2 days ago
“AI risk and the rational voter” by djbinder
A common reaction to arguments about AI risk is disbelief that anyone would let it happen. If advanced AI really threatened everyone, surely people …
2 weeks, 2 days ago
“Let’s talk about the AI coordination problem” by KatjaGrace
Yesterday I asked if this ‘coordinate not to build dangerous AI’ problem was actually easy.
Why would I think that, contrary to so much belief?
Well…
2 weeks, 2 days ago
“F***ing Pulleys, How Do They Work?” by Liron
Everyone acts like it's obvious that pulleys do a physically possible thing, but personally, I’ve never understood why you can lift a 100kg object s…
2 weeks, 2 days ago
“Training Models to Predict and Explain Their In-the-Wild Behavior” by Adam Karvonen, Subhash Kantamneni, Euan Ong, Sam Marks
Summary
Our CHIVE pipeline produced thousands of unexpected behaviors with explanations that are grounded in counterfactual prompts (see Figure 1 fo…
2 weeks, 2 days ago
“Almost nobody is funded to figure out what work would solve alignment” by Seth Herd
Solving alignment would be easier if we worked out what problems we actually need to solve. This could be called the alignment meta-problem. Work on…
2 weeks, 3 days ago