On April 7, Anthropic announced Claude Mythos Preview and chose not to release it publicly. The model autonomously found and exploited zero-day vulnerabilities in every major operating system and every major web browser during internal testing. Where its predecessor Opus 4.6 produced working Firefox exploits less than 1% of the time, Mythos succeeded 72.4% of the time. Engineers with no formal security training asked it to find remote code execution vulnerabilities overnight and woke up to complete, working exploits.
Instead of a public release, Anthropic launched Project Glasswing: an invitation-only program restricted to defensive cybersecurity work across roughly 40 organizations including Amazon, Apple, Microsoft, Cisco, CrowdStrike, Google, and Palo Alto Networks. Anthropic is committing up to $100 million in usage credits for participants. The goal is to let defenders find and patch vulnerabilities in critical software before models with comparable capabilities become broadly available.
Here is the problem with that logic: the attacker has the advantage, always. Anthropic can restrict Mythos, but the capability emerged from general improvements in code, reasoning, and autonomy, not from exploit-specific training. That means every frontier lab is converging on the same capability curve. The question is not whether comparable models will exist outside a controlled program. The question is when, and whether the defenders Glasswing is arming will have closed enough of the gap by then. Anthropic found thousands of high- and critical-severity vulnerabilities in a few weeks of testing, with more than 99% still unpatched. That is the asymmetry: offense scales with compute, defense scales with organizational will. One of those is easier to buy.