Methods that steer a language model toward or away from a concept can succeed on the target behavior while making its writing less fluent, according to a systematic study published by Apple researchers. The work compares conditioning techniques for both concept injection and concept removal rather than measuring control alone.
The researchers report that efficient activation-steering methods frequently paid a steep fluency cost. They also found an important dependency on how the model was trained: activation steering was far less effective on instruction-tuned models than on the corresponding base models.
Simpler prompting and full supervised fine-tuning remained viable for adding a concept, but performed less well when the goal was to remove one. That distinction matters for teams choosing safety or customization techniques, because a method that works for encouraging a style may not work equally well for suppressing unwanted behavior.
The study also found that inexpensive text-based metrics correlated strongly with costlier language-model judges. Those metrics could make broader comparisons practical, while also revealing how different conditioning methods change output. The paper’s central warning is methodological: evaluations should report both whether steering worked and what it did to generation quality.