A new evaluation finds that biased user turns can make large language models express more cognitive bias later in a conversation.

The researchers built a three-condition framework to separate exposure to a biased user turn from the semantic content of that turn. Their benchmark includes 24,300 jury-validated prompts spanning a 9-by-9 matrix of target and human biases. Across eight frontier instruction-tuned models, biased conversational context increased bias expression relative to zero-shot baselines in six models.

The paper also identifies a tension in model behavior. Exposure to biased reasoning generally amplified later bias, while explicitly stated bias cues often triggered alignment-related suppression that reduced overt bias. In other words, models may respond differently when a bias is subtle versus when it is clearly named.

The result matters for real chat systems because conversations are not isolated prompts. Prior user framing can shape later reasoning, so bias evaluations need multi-turn tests that reflect how people actually interact with assistants.