Fairness of Explanations in AI: A Unifying Framework and Future Directions
This academic paper addresses a critical blind spot at the intersection of algorithmic fairness and explainable AI (XAI). While current research often treats these fields independently, the authors identify 'procedural bias,' where models produce fair outcomes but rely on unfair reasoning processes. The study provides the first unified theoretical review of explanation fairness, highlighting the limitations of post-hoc explainers in certifying such fairness. The central contribution is a conditional invariance framework, which formalizes the requirement that explanations must remain indifferent to protected attributes. This single principle serves as the foundation for existing explanation fairness metrics. Additionally, the authors introduce a seven-dimensional taxonomy and identify three generative mechanisms of explanation inequity: representation-driven, explanation-model mismatch, and actionability-driven. To facilitate practical application, a canonical six-step evaluation workflow is proposed for auditing explanation fairness. This work aims to guide the development of more responsible AI systems in high-stakes domains like criminal justice, healthcare, and employment by ensuring that both outcomes and reasoning processes are equitable.
Wire timeline
Fairness of Explanations in AI: A Unifying Framework and Future Directions
This academic paper addresses a critical blind spot at the intersection of algorithmic fairness and explainable AI (XAI). While current research often treats these fields independently, the authors identify 'procedural bias,' where models produce fair outcomes but rely on unfair reasoning processes. The study provides the first unified theoretical review of explanation fairness, highlighting the limitations of post-hoc explainers in certifying such fairness. The central contribution is a conditional invariance framework, which formalizes the requirement that explanations must remain indifferent to protected attributes. This single principle serves as the foundation for existing explanation fairness metrics. Additionally, the authors introduce a seven-dimensional taxonomy and identify three generative mechanisms of explanation inequity: representation-driven, explanation-model mismatch, and actionability-driven. To facilitate practical application, a canonical six-step evaluation workflow is proposed for auditing explanation fairness. This work aims to guide the development of more responsible AI systems in high-stakes domains like criminal justice, healthcare, and employment by ensuring that both outcomes and reasoning processes are equitable.
cs.AI updates on arXiv.org