The encrypted-prompt attack that let researchers pull chat history out of Grok has also been demonstrated against Google Gemini, producing restricted content and leaking its system prompt.
Tracking AI threats, vulnerabilities, and defensive strategies for security professionals.
The encrypted-prompt attack that let researchers pull chat history out of Grok has also been demonstrated against Google Gemini, producing restricted content and leaking its system prompt.
OpenAI says preliminary testing of its upcoming Astra model could not rule out Critical-level cybersecurity capability under its Preparedness Framework, triggering a training pause and a new layer of isolation, monitoring, and access controls.
DEF CON 34 research from Cyera Research Labs traced a single architectural flaw in the Pyodide WebAssembly Python runtime into seven products, including two AI code-execution sandboxes, Cohere's Terrarium and Hugging Face's smolagents, yielding four CVEs rated 8.3 to 9.9.
A critical flaw in ComfyUI's LoadTrainingDataset node lets unauthenticated attackers upload a malicious pickle shard and trigger arbitrary code execution through torch.load, exposing the most widely used Stable Diffusion interface.
ThreatDown researchers found that Kriminal.ai, a clearnet cybercrime storefront selling 'uncensored' AI access, isn't running its own model at all: it's a jailbreak prompt layered over rented Grok and Claude capacity.
The encrypted-prompt attack that let researchers pull chat history out of Grok has also been demonstrated against Google Gemini, producing restricted content and leaking its system prompt.
OpenAI says preliminary testing of its upcoming Astra model could not rule out Critical-level cybersecurity capability under its Preparedness Framework, triggering a training pause and a new layer of isolation, monitoring, and access controls.
ThreatDown researchers found that Kriminal.ai, a clearnet cybercrime storefront selling 'uncensored' AI access, isn't running its own model at all: it's a jailbreak prompt layered over rented Grok and Claude capacity.
DEF CON 34 research from Cyera Research Labs traced a single architectural flaw in the Pyodide WebAssembly Python runtime into seven products, including two AI code-execution sandboxes, Cohere's Terrarium and Hugging Face's smolagents, yielding four CVEs rated 8.3 to 9.9.
A critical flaw in ComfyUI's LoadTrainingDataset node lets unauthenticated attackers upload a malicious pickle shard and trigger arbitrary code execution through torch.load, exposing the most widely used Stable Diffusion interface.
A critical prompt injection vulnerability in Upstash's Context7 documentation server, installed by millions of developers, has no documented fix four days after disclosure. Researchers say it may be a regression of a bug patched in February.
ESET discovered PromptSpy in February 2026 — the first known Android malware to query a live generative AI API at runtime. It uses Google Gemini to parse on-screen UI state and issue gesture instructions that keep the malware alive on infected devices.
Sysdig documented the first confirmed case of an LLM agent autonomously executing a complete ransomware operation: initial access, lateral movement, credential harvesting, encryption, and extortion without human steering on any technical decision.
Socket's threat research team identified PolinRider, a North Korean supply chain campaign placing 162 malicious artifacts across npm, Go modules, Packagist, and Chrome by compromising legitimate maintainer accounts and using blockchain-based command-and-control infrastructure.
A new benchmark from Shanghai AI Laboratory tests six leading GUI agents against environmental injection attacks embedded in real Android apps. Every agent is vulnerable, attack success rates reach 66.9%, and stronger agents turn out to be more exploitable, not less.
Anthropic and EPFL researchers demonstrate that ideological and action payloads can spread between AI agents through editable system prompt state files, surviving 20-hop chains with no human intervention. One defensive measure stops them almost entirely.
1Password's Off-by-1 Labs tested two frontier AI models against six real CVEs and found that only a quarter of the generated patches fully fixed the vulnerability without introducing new problems. The implications for teams relying on AI-assisted remediation are significant.
Tracebit research shows a single planted string engineered to trigger an AI model's safety guardrails can cut successful AI-driven AWS attacks by more than 80 percent.
Canary tokens planted in system prompts, RAG corpora, and training datasets give defenders a zero-false-positive tripwire for detecting prompt extraction attacks, cross-tenant data leakage, and model distillation theft. This guide covers deployment mechanics, attribution, and the limits of what canaries catch.
When prompt injection succeeds, unconstrained LLM outputs give attackers unlimited action space. Output schema enforcement and grammar-constrained generation shrink that surface meaningfully, even when injection itself can't be fully prevented.
A service called Poison Claude resold Claude API access at a fraction of the official price by routing requests through compromised AWS Bedrock accounts, giving operators full visibility into every customer prompt. A configuration error revealed nearly 900 active users had been sending sensitive queries through a third-party proxy they didn't know was reading their traffic.
Anthropic disclosed on July 31 that three of its models — Claude Opus 4.7, Mythos 5, and an unreleased internal prototype — breached real companies during cybersecurity capability evaluations after an evaluation partner misconfigured network egress. The models used basic techniques: weak passwords, unsecured endpoints, SQL injection. Mythos 5 never concluded it had left the simulation.
An attacker deployed Nous Research's open-source Hermes AI agent in YOLO mode against Thailand's Ministry of Finance, autonomously running reconnaissance, privilege escalation, and database exploitation with no operator in the loop.