A new arXiv paper proposes an unusual defense for open-weight model safety: assume attackers can remove refusal behavior, then make the post-attack model unreliable for harmful requests.
The method, called Fool’s Gold, is aimed at “abliteration” attacks that project refusal-related directions out of model weights. The authors say such safety removal can be done in minutes and is hard to prevent durably at release time.
Their defense uses decoy hardening. In the clean model, behavior is held close to the original. In the attacked state, responses to hazardous operational prompts become confident-looking decoys with falsified critical elements. The paper evaluates the approach on seven models from five families, ranging from 9B to 122B parameters.
The idea is controversial by design because it relies on deceptive outputs after an attack. Its practical value depends on whether decoys reduce harm without creating new risks for legitimate users, researchers, or downstream systems that may not know the model has been altered.