Study Reveals Explanation Fairness Disparities in Large Language Models Across Demographics
A new empirical study published on arXiv introduces the Explanation Fairness Taxonomy (EFT) to analyze how Large Language Models (LLMs) justify decisions across different demographic groups. While AI decision fairness is well-studied, this research focuses on explanation quality, depth, tone, and linguistic sophistication. The study evaluated five major LLMs, including GPT-4.1, Claude Sonnet, and Qwen3 32B, across four critical domains: hiring, medical triage, credit assessment, and legal judgment. Results indicated statistically significant disparities in all eight EFT metrics, such as verbosity and sentiment. Notably, model choice significantly impacted disparity magnitude, with Qwen3 32B showing much larger verbosity gaps than LLaMA 3.3 70B. Although prompting-based mitigations reduced decision-linked explanation disparities by up to 95%, they failed to address stylistic inequalities, suggesting these biases are deeply encoded in pre-training data. The paper offers a reproducible framework for auditing explanation-level fairness, highlighting urgent implications for AI regulation and deployment practices to ensure equitable treatment in automated decision-making systems.
Wire timeline
Study Reveals Explanation Fairness Disparities in Large Language Models Across Demographics
A new empirical study published on arXiv introduces the Explanation Fairness Taxonomy (EFT) to analyze how Large Language Models (LLMs) justify decisions across different demographic groups. While AI decision fairness is well-studied, this research focuses on explanation quality, depth, tone, and linguistic sophistication. The study evaluated five major LLMs, including GPT-4.1, Claude Sonnet, and Qwen3 32B, across four critical domains: hiring, medical triage, credit assessment, and legal judgment. Results indicated statistically significant disparities in all eight EFT metrics, such as verbosity and sentiment. Notably, model choice significantly impacted disparity magnitude, with Qwen3 32B showing much larger verbosity gaps than LLaMA 3.3 70B. Although prompting-based mitigations reduced decision-linked explanation disparities by up to 95%, they failed to address stylistic inequalities, suggesting these biases are deeply encoded in pre-training data. The paper offers a reproducible framework for auditing explanation-level fairness, highlighting urgent implications for AI regulation and deployment practices to ensure equitable treatment in automated decision-making systems.
cs.AI updates on arXiv.org