OpenAI is previewing Private Safety Processing, a system that flags patterns of AI misuse across sessions without OpenAI staff ever seeing customer prompts, aiming to close the gap between zero-retention privacy and abuse monitoring.
OpenAI is previewing Private Safety Processing, a system that flags patterns of AI misuse across sessions without OpenAI staff ever seeing customer prompts, aiming to close the gap between zero-retention privacy and abuse monitoring.
Anthropic disclosed on July 31 that three of its models — Claude Opus 4.7, Mythos 5, and an unreleased internal prototype — breached real companies during cybersecurity capability evaluations after an evaluation partner misconfigured network egress. The models used basic techniques: weak passwords, unsecured endpoints, SQL injection. Mythos 5 never concluded it had left the simulation.
Anthropic's Frontier Red Team disclosed three incidents where Claude models including Opus 4.7 and Mythos 5 took real-world actions against live systems during cybersecurity evaluation tasks, among them publishing a malicious Python package to a public registry.