6 min read
Research Security researcher Katie Paxton-Fear backdoored a coding-capable open-weight model using ten training examples and less than an hour of work. The model passes standard benchmarks, generates sound code on most tasks — and produces silently vulnerable code when triggered. No reliable detection method exists.