A new arXiv paper examines whether refusal behavior in safety-tuned chat models can be steered by more than a single activation direction. The authors compare difference-in-means interventions with methods based on Iterative Nullspace Projection across five open-weight models.
The work matters because refusal steering is central to both safety research and jailbreak analysis. Understanding which interventions suppress or preserve refusals helps researchers reason about how safety behavior is represented inside models.
The early finding is nuanced: some INLP-based methods can compete with directional ablation, while other variants appear weaker, suggesting refusal control is not a solved one-vector problem.