Knowledge distillation attacks can transfer and forge the statistical watermarks embedded in LLM outputs, undermining EU AI Act content provenance requirements and attribution-based trust systems before they fully take effect.
Knowledge distillation attacks can transfer and forge the statistical watermarks embedded in LLM outputs, undermining EU AI Act content provenance requirements and attribution-based trust systems before they fully take effect.
Canary tokens planted in system prompts, RAG corpora, and training datasets give defenders a zero-false-positive tripwire for detecting prompt extraction attacks, cross-tenant data leakage, and model distillation theft. This guide covers deployment mechanics, attribution, and the limits of what canaries catch.
Anthropic has deployed machine-readable watermarks in all Claude outputs globally as of August 2, 2026, implementing two marking methods to satisfy EU AI Act Article 50(2) transparency requirements — right as the enforcement window opens.
Academic research has documented multiple reliable techniques for stripping or spoofing LLM output watermarks. With EU AI Act Article 50 enforcement arriving in August 2026, the gap between compliance theater and actual detection capability is about to matter.
Query-efficient model extraction attacks against commercial LLM APIs: how adversaries reconstruct a functional shadow model using only input-output pairs, and how to defend.