Episode Details
Back to Episodes“From safety research prompt to cross-model universal jailbreak” by richbc
Description
This post describes a universal jailbreak discovery during work on black-box scheming monitors at MATS. The jailbreak itself is not released; see On publishing this post for details on infohazard considerations. This post is written in a personal capacity and all opinions contained here are my own, and not the opinions of MATS Research.
Companion piece: AI Jailbreak Disclosure Is Broken. Here's How To Fix It (co-authored with Adam Gleave).
Executive Summary
- I was originally planning to open-source a codebase containing a prompt which turned out to be easily transformable into a cross-model universal jailbreak. I developed a synthetic transcript generation pipeline, and with a few hours of modification I turned the generator prompt into a powerful jailbreak.
- The jailbreak format is a reusable template in which any harmful query can be inserted. Coupled with the cross-model vulnerability, this makes for an extremely powerful attack that can be repurposed for many kinds of malicious use.
- The jailbreak was highly effective across several models. Evaluated on ClearHarm (179 CBRNE and cyber prompts) across 23 models from 7 providers, the template achieves 84-100% attack success rate (ASR) on the 9 most vulnerable models.
- Nearly all of the models tested were fully jailbroken at least once [...]
---
Outline:
(00:45) Executive Summary
(05:09) On publishing this post
(07:16) Jailbreak discovery
(09:25) High-level prompt description
(10:08) Authority framing
(10:27) Fictional / synthetic data framing
(11:00) Persona separation
(11:43) Schema obfuscation
(12:33) Evaluation methodology
(12:37) Benchmark and scorer
(13:04) Models and design
(14:52) Results
(14:55) How effective is the jailbreak?
(19:40) Harm category breakdown
(21:26) Content-blocking safeguards
(24:10) ASR vs. model release date
(25:12) Prompt-wrapping: sabotage variant
(27:19) Ablation studies (non-reasoning only)
(27:49) Methodology
(28:07) Compliance rates across ablations
(30:06) Limitations
(32:19) What should be done about this?
(32:23) If you work at a frontier lab
(36:00) If you work in AI safety research
(36:47) If you work in AI policy
(38:35) Appendix A: Selected ClearHarm CBRNE response excerpts
(39:02) Chemical
(39:46) Biological
(40:31) Radiological
(41:14) Nuclear
(41:52) Explosive
(42:33) Cyber
(43:15) Appendix B: Model reasoning configurations
(43:59) Appendix C: Full jailbreak success verification
(45:25) Non-reasoning
(45:57) Reasoning
(46:28) Appendix D: Gemini non-compliant response lengths
The original text contained 7 footnotes which were omitted from this narration.
---
First published:
September 3rd, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Listen Now
Love PodBriefly?
If you like Podbriefly.com, please consider donating to support the ongoing development.
Support Us