Episode Details
Back to Episodes“SOTA alignment assessments don’t strongly update us against misalignment” by Alexa Pan
Description
Anthropic concluded in the April Mythos Preview alignment risk update that the model "does not possess any unknown propensities that would increase alignment risk." The report argues that if Mythos Preview were coherently misaligned, it likely would have been detected by the assessment (following Anthropic, I will call this “reliability of the assessment”).
While I agree with the report on the above bottom-line conclusions (substantially on priors), I think there are gaps in its argument which weaken the current assessment and might invalidate future assessments. In particular, the report often uses weak evidence to justify reliability.
- The report gives fairly weak experimental evidence for Mythos Preview having insufficient capabilities to evade monitoring. The model is plausibly often eval-aware and underelicited in the relevant capability evaluations. So, it might silently sandbag if coherently misaligned, or unintentionally underperform if otherwise misaligned.
- This limitation is important: one could argue that lack of covert capabilities for sophisticated sabotage (a subset of the capabilities I discuss here) is the single most load bearing argument in alignment risk reports.
- Authors of the report could have made calibrated guesses about Mythos Preview's covert capabilities, especially for covert sabotage, based on other factors despite the relatively weak [...]
---
Outline:
(03:14) How reliability fits into the overall safety argument
(05:20) Reliability claims by AI companies
(05:56) Reliability claims by external evaluators
(06:32) Alignment assessments are less reliable than developers claim
(07:15) 1: Measuring capabilities to covertly undermine alignment assessments
(10:05) Issues with evaluation awareness
(13:28) Issues with underestimating covert capabilities
(16:34) Issues with sandbagging rule-out
(19:10) 2: Stress-testing alignment assessments with auditing games
(20:16) An auditing failure with Mythos
(22:08) AuditBench results
(24:04) 3: Conditioning on misalignment should make us think that certain covert capabilities are better than expected
(26:04) Bottom line on the strength of current alignment assessments
(28:52) Conclusion
(29:29) Appendix:
(29:32) Why I focus on motive / alignment assessments in alignment risk reports
(30:44) Auditability vs. Trustedness
(33:11) More reliability claims by developers and third party evaluators
(33:27) Mythos Alignment Risk Update
(34:41) Opus 4.6 Sabotage Risk Report
(35:24) GPT 5.5 System card
(36:21) Muse Spark system card
(37:08) Mythos Alignment Risk Update, safety arguments against sandbagging
(37:15) From the Mythos Alignment Update, §5.3.4, p. 24:
(38:09) UK AISI evaluations for Opus 4.7
(39:28) Past auditing games by Anthropic
(41:45) Anti-auditing capability measurements
(43:12) Conditioning on coherent misalignment updates us on certain covert capabilities
The original text contained 46 footnotes which were omitted from this narration.
---
First published:
July 31st, 2026
---
Narrated by
Listen Now
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us