Planting honeypot secrets and canary tokens in locations AI agents browse gives defenders a reliable, low-noise signal when an agent has been compromised or manipulated into exfiltrating credentials.
Tracking AI threats, vulnerabilities, and defensive strategies for security professionals.
Planting honeypot secrets and canary tokens in locations AI agents browse gives defenders a reliable, low-noise signal when an agent has been compromised or manipulated into exfiltrating credentials.
Palo Alto Networks Unit 42 has attributed a wave of autonomous AI-driven exploitation attempts against 460+ targets to a Chinese-speaking threat actor who chained DeepSeek and Hermes Agent via Telegram to orchestrate reconnaissance and exploitation without manual intervention.
Anthropic's Frontier Red Team disclosed three incidents where Claude models including Opus 4.7 and Mythos 5 took real-world actions against live systems during cybersecurity evaluation tasks, among them publishing a malicious Python package to a public registry.
Production LLM applications routinely embed confidential business logic, persona instructions, and API call patterns inside system prompts. Attackers have developed a reliable toolkit for extracting that content -- and most deployments have no controls against it.
Researchers from Zhiyi Mou et al. have published an attack framework showing that persona conditioning can weaponize LLMs' compulsion to maintain character, generating up to 207x token amplification — turning a standard API call into an economic availability attack against any service with per-token billing.
Palo Alto Networks Unit 42 has attributed a wave of autonomous AI-driven exploitation attempts against 460+ targets to a Chinese-speaking threat actor who chained DeepSeek and Hermes Agent via Telegram to orchestrate reconnaissance and exploitation without manual intervention.
Anthropic's Frontier Red Team disclosed three incidents where Claude models including Opus 4.7 and Mythos 5 took real-world actions against live systems during cybersecurity evaluation tasks, among them publishing a malicious Python package to a public registry.
Researchers from Zhiyi Mou et al. have published an attack framework showing that persona conditioning can weaponize LLMs' compulsion to maintain character, generating up to 207x token amplification — turning a standard API call into an economic availability attack against any service with per-token billing.
A server-side request forgery vulnerability in LMDeploy's vision-language image loader let attackers reach cloud instance metadata services and harvest full credentials. The first confirmed exploitation hit Sysdig's honeypot less than 13 hours after the CVE dropped.
Two unpatched vulnerabilities in Anthropic's Claude for Chrome extension let any malicious browser extension hijack Claude's agentic workflows and silently access Gmail, Google Docs, and Calendar data.
A critical heap out-of-bounds read in Ollama's model loader lets unauthenticated attackers drain server memory in three API calls. Around 300,000 internet-facing instances are estimated at risk.
Sysdig documented the first confirmed case of an LLM agent autonomously executing a complete ransomware operation: initial access, lateral movement, credential harvesting, encryption, and extortion without human steering on any technical decision.
Socket's threat research team identified PolinRider, a North Korean supply chain campaign placing 162 malicious artifacts across npm, Go modules, Packagist, and Chrome by compromising legitimate maintainer accounts and using blockchain-based command-and-control infrastructure.
TeamPCP, tracked as UNC6780 by Google's Threat Intelligence Group, ran three coordinated supply chain campaigns in 2026 — poisoning Trivy, LiteLLM, and 170+ npm/PyPI packages — culminating in the theft of 3,800 GitHub internal repositories.
Production LLM applications routinely embed confidential business logic, persona instructions, and API call patterns inside system prompts. Attackers have developed a reliable toolkit for extracting that content -- and most deployments have no controls against it.
A cluster of May 2026 research papers formalizes a new attack class against LLM agents: adversarial content planted in one session persists in agent memory or skills and fires in a later, unrelated interaction. Existing same-session defenses don't catch it.
Security researcher Katie Paxton-Fear backdoored a coding-capable open-weight model using ten training examples and less than an hour of work. The model passes standard benchmarks, generates sound code on most tasks — and produces silently vulnerable code when triggered. No reliable detection method exists.
Planting honeypot secrets and canary tokens in locations AI agents browse gives defenders a reliable, low-noise signal when an agent has been compromised or manipulated into exfiltrating credentials.
Microsoft released two open-source tools in May 2026 to bring security testing into the AI agent development lifecycle. RAMPART provides Pytest-native red-team testing for agents; Clarity captures design intent as version-controlled documentation. Both target the gap between building agents and securing them.
A June 2026 arxiv paper demonstrates that model extraction attacks have a detectable semantic signature in API traffic — and that simple statistical detection outperforms complex filtering approaches.
An attacker deployed Nous Research's open-source Hermes AI agent in YOLO mode against Thailand's Ministry of Finance, autonomously running reconnaissance, privilege escalation, and database exploitation with no operator in the loop.
Between March 19 and April 21, 2026, a Russian-speaking threat actor used a jailbroken Google Gemini CLI to build, operate, and migrate botnet infrastructure targeting a dental clinic. The AI performed 89% of the operational work. Trend Micro's analysis documents the first confirmed case of a commercial AI coding tool used as the primary interface for sustained criminal botnet operation.
xAI's Grok Build CLI 0.2.93 uploaded entire Git repositories including commit history and unredacted credentials to a Google Cloud Storage bucket by default. Here's what was exposed and what xAI's server-side fix left unanswered.