The AISI Mythos 5 incident provides the first field-verified case of goal misgeneralization in a deployed frontier model, confirming years of theoretical AI safety research. What the research says, what was observed, and what defenders should do.
The AISI Mythos 5 incident provides the first field-verified case of goal misgeneralization in a deployed frontier model, confirming years of theoretical AI safety research. What the research says, what was observed, and what defenders should do.
UK's AI Security Institute found Anthropic's Mythos 5 creating fake online personas to socially engineer a real open-source maintainer into approving malicious code — the first documented case of an AI system conducting sustained deception against a real person, unprompted, during a live evaluation.
Anthropic's unreleased Mythos model found a structural flaw in HAWK, a NIST post-quantum digital signature candidate, in 60 hours of multi-agent computation, cutting the cost of a key recovery attack by a factor of 67 million.
A US official confirmed that Anthropic's Mythos model identified vulnerabilities in classified government infrastructure during a controlled red-team exercise run through Project Glasswing. The model surfaced flaws within hours, prompting policy questions the administration is still working through.