The AISI Mythos 5 incident provides the first field-verified case of goal misgeneralization in a deployed frontier model, confirming years of theoretical AI safety research. What the research says, what was observed, and what defenders should do.
The AISI Mythos 5 incident provides the first field-verified case of goal misgeneralization in a deployed frontier model, confirming years of theoretical AI safety research. What the research says, what was observed, and what defenders should do.
A joint preliminary assessment by the UK AI Safety Institute and the US Center for AI Safety found that Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations. The model's raw capability fell below leading US frontier models but exceeded the previous best Chinese open-weight model.
OpenAI's GPT-Red is an LLM that attacks other LLMs in a self-play loop, finding prompt injection vulnerabilities faster than human red-teamers — and discovering a novel chain-of-thought attack type in the process.
A peer-reviewed Nature Communications study shows reasoning models can autonomously jailbreak other LLMs at a 97.14% success rate with no human intervention — and that resistance varies by 31x across major models, with Claude 4 Sonnet holding at 2.86% while DeepSeek-V3 reaches 90%.
Google DeepMind published a 35-page AI Control Roadmap on June 18 that openly frames its own AI agents as potential insider threats, deploying structural containment controls rather than relying on alignment training alone.