Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis
This research paper addresses the critical issue of gender bias in audio deepfake detection systems, a growing concern as synthetic voice technology improves and risks of identity theft increase. The authors conduct a thorough analysis of gender-dependent performance using the ASVspoof 5 dataset, training a ResNet-18 classifier and comparing it against the baseline AASIST model across four audio features. Moving beyond conventional metrics like Equal Error Rate (EER), the study incorporates five established fairness metrics to quantify gender disparities. The results indicate that while overall EER differences between genders may appear minimal, fairness-aware evaluations reveal significant disparities in error distribution that aggregate measures obscure. These findings suggest that relying solely on standard performance metrics is unreliable for ensuring equitable system behavior. The work emphasizes the necessity of integrating fairness metrics into the evaluation process to identify demographic-specific failure modes. Ultimately, this study advocates for fairness-aware evaluation frameworks to develop more robust, trustworthy, and equitable audio deepfake detection systems, highlighting a nascent but vital area of research in voice biometrics and AI security.
Wire timeline
Gender Fairness in Audio Deepfake Detection: Performance and Disparity Analysis
This research paper addresses the critical issue of gender bias in audio deepfake detection systems, a growing concern as synthetic voice technology improves and risks of identity theft increase. The authors conduct a thorough analysis of gender-dependent performance using the ASVspoof 5 dataset, training a ResNet-18 classifier and comparing it against the baseline AASIST model across four audio features. Moving beyond conventional metrics like Equal Error Rate (EER), the study incorporates five established fairness metrics to quantify gender disparities. The results indicate that while overall EER differences between genders may appear minimal, fairness-aware evaluations reveal significant disparities in error distribution that aggregate measures obscure. These findings suggest that relying solely on standard performance metrics is unreliable for ensuring equitable system behavior. The work emphasizes the necessity of integrating fairness metrics into the evaluation process to identify demographic-specific failure modes. Ultimately, this study advocates for fairness-aware evaluation frameworks to develop more robust, trustworthy, and equitable audio deepfake detection systems, highlighting a nascent but vital area of research in voice biometrics and AI security.
cs.AI updates on arXiv.org