7 min read
Research Research published in 2025-2026 demonstrates that asymmetric tokenization between LLM safety classifiers and the underlying generation model creates reliable bypass channels. The same Unicode text is tokenized differently by filter and model, allowing attackers to craft inputs that look safe to the classifier while being read normally by the LLM.