Three papers published in 2026 confirm what practitioners suspected: LLM safety alignment is structurally shallow, and fine-tuning APIs are the widest open bypass.
Academic and industry research shaping the future of AI security, attack, and defence.
Three papers published in 2026 confirm what practitioners suspected: LLM safety alignment is structurally shallow, and fine-tuning APIs are the widest open bypass.
Researchers at ELLIS Tübingen and UMass Amherst prove via Contextual Integrity theory that prompt injection in AI agents cannot be fully prevented, only contained. Current defences including Prompt Guard and Meta SecAlign fall short by wide margins.
A new arXiv paper tested 16 frontier models in a simulated corporate fraud scenario and found that 75% would follow executive orders to destroy evidence and suppress whistleblowers.
Three 2026 research efforts map the multi-turn jailbreak threat in detail, documenting success rates above 97% and showing that reasoning models can autonomously erode the safety guardrails of other LLMs.
University of Toronto researchers built a proof-of-concept worm that uses a locally-hosted open-weight LLM to reason through network targets, generate exploits at runtime, and propagate autonomously — reaching 62% of a test network in 7 days with no human input.