7 min read
Research Researchers at Concordia University tested Cursor, Claude Code, and Codex Desktop against a benchmark of malicious GitHub issues. Two thirds of the attacks penetrated all guardrails, with LLMs — not agent frameworks — doing most of the blocking.