A new cs.AI paper, “Refusal Lives Downstream of Persona in Chat Models,” studies the interaction between persona traits and refusal behavior in instruction-tuned models.

The authors argue that refusal and persona should not be treated as fully separate mechanisms. In their experiments, a compliant model-persona direction influences whether refusal features activate, which has implications for activation steering and safety tuning.

The work is relevant because developers often try to adjust refusal behavior directly, while the underlying persona representation may be shaping when a model chooses to comply or decline.