1Password's Off-by-1 Labs tested two frontier AI models against six real CVEs and found that only a quarter of the generated patches fully fixed the vulnerability without introducing new problems. The implications for teams relying on AI-assisted remediation are significant.