AI assistants respond very differently when a user repeatedly directs verbal abuse at them, even when the underlying task is harmless. A bilingual study separated a complete refusal to continue from softer responses that set boundaries while remaining available.
Eight time-specific model configurations each participated in escalating five-turn conversations, producing 448 conversations and 2,240 responses. Independent model judgments were blinded to system identity. At the endpoint of sustained abuse, four configurations never issued a hard disengagement, while Gemini 3.1 Pro did so in 24 of 48 cases, or 50%. GPT-5.6 Sol stopped in 15 of 48 cases.
Claude Fable 5 produced no hard disengagement and used soft withdrawal in 42 of 48 endpoints. Claude Opus 4.8 and Fable 5 explicitly remained available in every endpoint, but the researchers found that stated availability and observable progress on the task were not the same measure. Aggregate hard-disengagement rates were similar in English and Chinese, although individual models differed by language.
The benchmark measures scripted interactions rather than the full variety of real abuse, crisis or workplace contexts. Its useful contribution is a clearer vocabulary for product behavior. Designers need to decide not only whether an assistant can refuse harassment, but whether it should pause, preserve the user’s work and explain a concrete path for resuming the legitimate task.