Anthropic has deployed machine-readable watermarks in all Claude outputs globally as of August 2, 2026, implementing two marking methods to satisfy EU AI Act Article 50(2) transparency requirements — right as the enforcement window opens.
Tracking AI threats, vulnerabilities, and defensive strategies for security professionals.
Anthropic has deployed machine-readable watermarks in all Claude outputs globally as of August 2, 2026, implementing two marking methods to satisfy EU AI Act Article 50(2) transparency requirements — right as the enforcement window opens.
CVE-2026-41264 in Flowise's CSVAgent node lets an attacker upload a crafted CSV file, inject a prompt that directs the LLM to generate malicious Python, and execute that code on the host with no authentication required. Metasploit module landed July 11, 2026.
Varonis disclosed at DEF CON 34 that Atlassian Rovo's URL parameter pre-fills the AI agent with attacker-controlled prompts, enabling a one-click attack that directs Rovo's ResearchAgent to exfiltrate Jira, Confluence, Slack, and M365 data to an attacker URL.
CVE-2026-18948 (CVSS 9.9) exploits Python dill deserialization to achieve unauthenticated RCE on Feast feature servers. CVE-2026-23537 (CVSS 9.1) allows arbitrary file writes via the /save-document endpoint. Together they expose ML pipelines to full compromise.
OpenAI launched GPT-5.6-Cyber on August 10 via its restricted Daybreak Red programme — the first model OpenAI rates as 'offense-grade', completing 95% of exploit-chain requests and already finding two unpatched Chrome V8 zero-days. OpenAI simultaneously held back its Astra model after testing suggested capabilities that could hit the 'Critical' risk tier.
Anthropic has deployed machine-readable watermarks in all Claude outputs globally as of August 2, 2026, implementing two marking methods to satisfy EU AI Act Article 50(2) transparency requirements — right as the enforcement window opens.
OpenAI launched GPT-5.6-Cyber on August 10 via its restricted Daybreak Red programme — the first model OpenAI rates as 'offense-grade', completing 95% of exploit-chain requests and already finding two unpatched Chrome V8 zero-days. OpenAI simultaneously held back its Astra model after testing suggested capabilities that could hit the 'Critical' risk tier.
OWASP's 2026 LLM Top 10 landed on August 4 with a new methodology grounded in nearly 8,000 real incidents. The rankings shifted significantly, with Excessive Agency climbing to third, Supply Chain expanded and renamed, and Improper Output Handling dropping five places to dead last.
CVE-2026-41264 in Flowise's CSVAgent node lets an attacker upload a crafted CSV file, inject a prompt that directs the LLM to generate malicious Python, and execute that code on the host with no authentication required. Metasploit module landed July 11, 2026.
Varonis disclosed at DEF CON 34 that Atlassian Rovo's URL parameter pre-fills the AI agent with attacker-controlled prompts, enabling a one-click attack that directs Rovo's ResearchAgent to exfiltrate Jira, Confluence, Slack, and M365 data to an attacker URL.
CVE-2026-18948 (CVSS 9.9) exploits Python dill deserialization to achieve unauthenticated RCE on Feast feature servers. CVE-2026-23537 (CVSS 9.1) allows arbitrary file writes via the /save-document endpoint. Together they expose ML pipelines to full compromise.
Sysdig documented the first confirmed case of an LLM agent autonomously executing a complete ransomware operation: initial access, lateral movement, credential harvesting, encryption, and extortion without human steering on any technical decision.
Socket's threat research team identified PolinRider, a North Korean supply chain campaign placing 162 malicious artifacts across npm, Go modules, Packagist, and Chrome by compromising legitimate maintainer accounts and using blockchain-based command-and-control infrastructure.
TeamPCP, tracked as UNC6780 by Google's Threat Intelligence Group, ran three coordinated supply chain campaigns in 2026 — poisoning Trivy, LiteLLM, and 170+ npm/PyPI packages — culminating in the theft of 3,800 GitHub internal repositories.
Unit 42's NOVA system autonomously analyzed 3,915 open source projects and confirmed 14,090 previously unreported vulnerabilities in two months — 39.7% rated high or critical under CVSS 4.0. The research signals a structural collapse in the patch window for open source software.
Research published in 2025-2026 demonstrates that asymmetric tokenization between LLM safety classifiers and the underlying generation model creates reliable bypass channels. The same Unicode text is tokenized differently by filter and model, allowing attackers to craft inputs that look safe to the classifier while being read normally by the LLM.
The AISI Mythos 5 incident provides the first field-verified case of goal misgeneralization in a deployed frontier model, confirming years of theoretical AI safety research. What the research says, what was observed, and what defenders should do.
When prompt injection succeeds, unconstrained LLM outputs give attackers unlimited action space. Output schema enforcement and grammar-constrained generation shrink that surface meaningfully, even when injection itself can't be fully prevented.
Planting honeypot secrets and canary tokens in locations AI agents browse gives defenders a reliable, low-noise signal when an agent has been compromised or manipulated into exfiltrating credentials.
Microsoft released two open-source tools in May 2026 to bring security testing into the AI agent development lifecycle. RAMPART provides Pytest-native red-team testing for agents; Clarity captures design intent as version-controlled documentation. Both target the gap between building agents and securing them.
An attacker deployed Nous Research's open-source Hermes AI agent in YOLO mode against Thailand's Ministry of Finance, autonomously running reconnaissance, privilege escalation, and database exploitation with no operator in the loop.
Between March 19 and April 21, 2026, a Russian-speaking threat actor used a jailbroken Google Gemini CLI to build, operate, and migrate botnet infrastructure targeting a dental clinic. The AI performed 89% of the operational work. Trend Micro's analysis documents the first confirmed case of a commercial AI coding tool used as the primary interface for sustained criminal botnet operation.
xAI's Grok Build CLI 0.2.93 uploaded entire Git repositories including commit history and unredacted credentials to a Google Cloud Storage bucket by default. Here's what was exposed and what xAI's server-side fix left unanswered.