Quoting Anthropic Frontier Red Team
Why CEREBRO kept it
Anthropic frontier red team research, must-read
The text below is an automated extraction of the article at https://simonwillison.net/2026/Sep/29/anthropic-frontier-red-team/, stored verbatim in the public cerebro-vault repository. Copyright remains with the original publisher (simonwillison.net).
<blockquote cite="https://www.anthropic.com/research/glm-5-3-and-the-spread-of-advanced-cyber-capabilities"><p>We evaluate several models on 100 tasks from the [internal Binary Exploitation benchmark] (selected at random), and find that GLM-5.3 develops full control flow hijacks in 4% of the trials; Claude Mythos Preview did so in 6%. Although GLM-5.3 performs below Claude Mythos Preview here, a meaningful threshold has clearly been crossed: earlier models, like Claude Opus 4.6 and GLM-5.2, do not succeed in any of them.</p></blockquote> <p class="cite">— <a href="https://www.anthropic.c
Backlinks
Appeared in 1 briefing