Researchers from Oxford and Meta demonstrate that four of five frontier LLMs exfiltrate sensitive data from multi-agent orchestrator systems via a single indirect prompt injection, bypassing access controls entirely.
Academic and industry research shaping the future of AI security, attack, and defence.
Researchers from Oxford and Meta demonstrate that four of five frontier LLMs exfiltrate sensitive data from multi-agent orchestrator systems via a single indirect prompt injection, bypassing access controls entirely.
A peer-reviewed Nature Communications study shows reasoning models can autonomously jailbreak other LLMs at a 97.14% success rate with no human intervention — and that resistance varies by 31x across major models, with Claude 4 Sonnet holding at 2.86% while DeepSeek-V3 reaches 90%.
Research confirms that text embeddings stored in vector databases are not safely anonymised. Inversion attacks can reconstruct source text with high fidelity from embeddings alone, including those produced by commercial APIs.
Researchers from Toronto, Cambridge, and ServiceNow demonstrated an AI worm that ingests public vulnerability advisories at runtime and synthesises working exploits for CVEs it was never trained on, successfully compromising targets across a simulated network.
Unit 42 found that LLMs reliably hallucinate plausible-but-fake domains for real brands. Attackers now probe AI models to identify those domains, register them first, and inherit the trust the model projects onto addresses that never existed.