7 min read
Research Activation steering bypasses prompt-level safety controls by manipulating a model's internal representations at inference time. It requires white-box access, which makes local open-source LLM deployments the primary exposure surface.