A paper accepted at ICML 2026 shows LLMs infer speaker role from text style, not role tags — enabling a zero-shot attack called CoT Forgery that achieves 60% success against frontier models by injecting fabricated reasoning.
A paper accepted at ICML 2026 shows LLMs infer speaker role from text style, not role tags — enabling a zero-shot attack called CoT Forgery that achieves 60% success against frontier models by injecting fabricated reasoning.
NDSS 2026 research shows LLMs can be systematically tricked into missing deliberately planted vulnerabilities through Familiar Pattern Attacks — automated, black-box exploits of the abstraction bias that affects every major model family.
University of Toronto researchers presented GPUBreach at Black Hat 2026 — a Rowhammer attack that escalates from an unprivileged CUDA kernel to a host root shell, bypassing IOMMU, threatening shared AI cloud GPU infrastructure.
Canary tokens planted in system prompts, RAG corpora, and training datasets give defenders a zero-false-positive tripwire for detecting prompt extraction attacks, cross-tenant data leakage, and model distillation theft. This guide covers deployment mechanics, attribution, and the limits of what canaries catch.
Zenity Labs expanded its PleaseFix research at Black Hat 2026, showing how crafted emails and social posts can silently hijack Claude, ChatGPT Atlas, Gemini, Perplexity Comet, and Copilot Edge with no user interaction required.