Giving a language model more reasoning tokens does not automatically make its decisions fairer, a new study finds. Researchers compared thinking and non-thinking modes in QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on income, recidivism, and credit tasks.
The evaluation used counterfactual pairs, where protected attributes change while the decision-relevant facts remain the same. Reasoning corrected some cases in which the basic model flipped its answer, but it also introduced new flips at nearly saturated confidence. Across all nine combinations of model and dataset, newly created inconsistencies outnumbered resolved ones by roughly five to one.
To examine when the change emerged, the researchers introduced a measure that tracks the probability gap as reasoning grows and a transition matrix comparing pair outcomes before and after thinking. Their analysis suggests bias can propagate and amplify within the reasoning trace rather than simply appearing in the final answer. The work covers three models and benchmark datasets, not real lending or justice systems, and counterfactual fairness is only one definition of fairness. Its practical warning is clear: organizations should audit reasoning mode separately instead of assuming a longer explanation signals a more equitable decision.