What the study compared
A research paper titled “Does Thinking Help Fairness? Reasoning Tokens Resolve Some Biases but Create More” examines whether enabling reasoning in language models improves or harms the fairness of their decisions. According to the available abstract, its main conclusion is that both can happen: some decision differences present in the non-reasoning variant disappear, while others that were not present before emerge.
The comparison was conducted within each model: the authors contrasted a reasoning variant with a non-reasoning variant. They evaluated QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B on three tasks associated with the Adult, COMPAS, and Credit datasets. The abstract describes these as high-stakes decision tasks. However, it does not provide the full protocol or detailed parameters for each evaluation.
The finding is therefore not that reasoning makes every system unfair, or that all its responses are less fair. It is a bounded comparison between specific configurations of three models and three datasets. Moreover, a difference in a computational test does not, by itself, establish harm to people in a real-world process.
Scope described in the abstract
| Element | Available information |
|---|---|
| Models | QwQ-32B, DeepSeek-R1-Distill-Qwen-32B, and Qwen3-32B |
| Datasets | Adult, COMPAS, and Credit |
| Comparison | Reasoning and non-reasoning variants within each model |
| Summary of results | Across all nine combinations, newly created flips outnumbered resolved flips by roughly a factor of five |
What a counterfactual flip means
In general terms, a counterfactual test compares a model’s decisions on two paired cases that differ in an aspect relevant to the analysis. The question is whether the prediction changes when that aspect is altered, while keeping the comparison sufficiently controlled to interpret the difference. In this study, the abstract refers to “counterfactual flips”: prediction changes between counterfactual pairs.
There is an important caveat: the material available does not specify which attribute was changed in each dataset, how the pairs were constructed, which other variables were held constant, or the precise rule used to classify a difference as a “flip.” It is therefore not possible to attribute the result to a specific attribute or faithfully reconstruct the test from the abstract alone.
The authors distinguish two directions of effect. A “resolved” flip is, in the paper’s general description, a flip present in the non-reasoning configuration that no longer appears in the reasoning configuration. A “created” flip is one that appears when reasoning is enabled and was not observed in the baseline. This explanation reflects the comparison described; the complete operational definition would require the full text.
The result: newly created flips were more numerous
The abstract reports a pattern across the nine evaluated combinations—three models multiplied by three datasets: counterfactual flips created when reasoning was enabled outnumbered those that were resolved by roughly five to one. It also says that the created flips occurred with model confidence close to saturation.
This figure summarizes the reported direction of the results, but it should not be read as a universal rate or as the probability that a real system will harm someone. The abstract does not provide absolute counts, uncertainty intervals, results broken down by model and dataset, or the full distribution of confidence. Without those details, it is not possible to assess here how much the effect varied across the nine combinations or how precise the aggregate estimate is.
Nor does “confidence close to saturation” mean that the system is well calibrated or that its answer is correct. The abstract indicates that newly created flips were associated with very high model confidence, but the material provided does not explain how that confidence was calculated, what threshold counted as close to saturation, or whether the scores were checked against observed outcomes. A model’s stated confidence and the reliability of its decision are separate questions.
How the authors try to explain the effect
The abstract presents two instruments for analyzing how bias evolves. The first, called the Counterfactual Depth Probability Gap (CDPG), is intended to track the evolution of differences along the depth of the reasoning process. The authors report observing bias propagation and amplification during that process.
The second, the Bias Transition Matrix (BTM), represents how predictions for counterfactual pairs change when moving from the non-reasoning configuration to the one that enables reasoning. According to the abstract, the asymmetric effect—resolving some flips while creating more others—originates in the joint transition of the states of each pair.
These are explanations and tools proposed by the paper, not proof that all internal reasoning necessarily amplifies bias. The abstract does not include the equations, implementation details, or robustness tests for CDPG and BTM. Without that information, their assumptions cannot be assessed independently, nor can we determine how much the conclusions depend on the chosen definition of a flip.
What to check before interpreting this kind of measurement
- 01Identify how counterfactual pairs were formed and which attributes were changed.
- 02Check which conditions were kept the same between reasoning and non-reasoning variants, including data, instructions, and generation parameters.
- 03Examine results by model and dataset, as well as in aggregate, and review counts and uncertainty.
- 04Separate the frequency of flips from reported confidence and from a decision’s accuracy or real-world impact.
- 05Repeat the evaluation with data and scenarios relevant to the intended use before drawing operational conclusions.
Limitations and open questions
The information available here comes from the paper’s abstract, not a complete account of its methods and results. The abstract is enough to identify the models, datasets, and general conclusion, but not to verify every detail of the experiment. In particular, it does not confirm whether data, prompts, and other parameters were held constant across variants, or explain how potential differences in formatting or generation were controlled.
There is also not enough information to judge how representative Adult, COMPAS, and Credit are of current systems or specific decision-making processes. These are the evaluated datasets named in the abstract; we should not infer that they cover all employment, criminal justice, or credit decisions, or that they alone reflect the populations that could be affected. The material provided also does not show which limitations concerning these datasets the authors explicitly identify.
Generalizability is another limitation. Three models from specific size ranges and families do not represent the full variety of reasoning models, and three datasets do not cover every high-stakes context. The result is a reason to test each system in its own context, not to conclude that enabling reasoning always worsens fairness. Conversely, the fact that some differences are resolved is not enough to claim that reasoning improves fairness overall.
Verifying the methodological questions raised here requires the full text and its evaluation materials. Until those are available, the conclusion should be attributed to what the abstract reports and kept within the scope of the nine combinations described.
What we know—and what the available material cannot confirm
| Question | What can be stated | What remains unverified |
|---|---|---|
| Which systems were evaluated? | The abstract names three models and three datasets. | It does not provide all parameters or the full experimental configuration here. |
| How were flips defined? | The study compares counterfactual changes between pairs of predictions. | The attributes, pair construction, and complete operational rule are not specified. |
| Were all other conditions identical? | The abstract describes a within-model comparison between reasoning and non-reasoning modes. | The material provided does not confirm that data, prompts, and other conditions were held constant. |
| Does the result apply to other contexts? | It is reported for nine model-and-dataset combinations. | It does not demonstrate generalization to other models, data, or real-world decisions. |
What teams should test before deploying these systems
For teams evaluating models in sensitive decision-making, the practical implication is straightforward: do not assume that a reasoning mode is fairer because it produces longer explanations or resolves some baseline cases. It is worth comparing each configuration on paired tests relevant to the intended use and recording both the flips that disappear and those that emerge.
An evaluation should explicitly describe what varies in each pair, which features are held fixed, and why the comparison is relevant. It should also report results separately by model, group, or condition under analysis, together with sample sizes and uncertainty ranges. If model confidence is part of the analysis, teams should explain how it is obtained and avoid treating it as an automatic measure of correctness.
Finally, laboratory tests do not replace oversight in the setting where a system is used. Before using a system in a high-stakes decision, teams need to check its behavior with appropriate data, establish procedures for reviewing errors, and enable meaningful human intervention. These are precautionary measures arising from the limited scope of the result, not recommendations attributed directly to the paper.
Open questions
- The full text and methods are not available in the sources provided, so the exact operational definition of “counterfactual flip” cannot be verified.
- The material does not specify which attributes were changed when constructing the counterfactual pairs or how other variables were held constant.
- The abstract does not confirm that data, prompts, and other parameters were identical across reasoning and non-reasoning variants.
- Absolute counts, uncertainty intervals, and complete results broken down across the nine combinations are not provided.
- The method for measuring confidence close to saturation is not detailed, and no calibration tests are provided.
- The available information does not establish which specific limitations the paper identifies for Adult, COMPAS, and Credit, and does not support extrapolating the results to other models or real-world decisions.
Keep exploring
Sources consulted
Corrections and transparency
If you spot incorrect or outdated information, send us a correction with the page and source we should review.
Submit a correction