Scalable Entity-Based Framework for Auditing Bias in LLMs
Researchers have introduced a new scalable framework for auditing bias in large language models (LLMs), addressing the trade-off between ecological validity and statistical control in existing evaluation methods. By using named entities as controlled probes within synthetic data, the framework enables large-scale analysis of systematic disparities in model behavior. The study represents the largest bias audit to date, analyzing 1.9 billion data points across various entities, tasks, languages, and models. Key findings reveal that LLMs tend to favor left-wing politicians over right-wing ones, prefer Western and wealthier nations over the Global South, and show bias toward Western companies while penalizing defense and pharmaceutical firms. Notably, while instruction tuning helps reduce bias, increasing model scale amplifies it, and prompting in Chinese or Russian does not mitigate Western-aligned preferences. The authors emphasize the critical need for systematic bias auditing before deploying LLMs in high-stakes applications. This extensible framework is made publicly available to support future research and improve fairness in AI systems.
Wire timeline
Scalable Entity-Based Framework for Auditing Bias in LLMs
Researchers have introduced a new scalable framework for auditing bias in large language models (LLMs), addressing the trade-off between ecological validity and statistical control in existing evaluation methods. By using named entities as controlled probes within synthetic data, the framework enables large-scale analysis of systematic disparities in model behavior. The study represents the largest bias audit to date, analyzing 1.9 billion data points across various entities, tasks, languages, and models. Key findings reveal that LLMs tend to favor left-wing politicians over right-wing ones, prefer Western and wealthier nations over the Global South, and show bias toward Western companies while penalizing defense and pharmaceutical firms. Notably, while instruction tuning helps reduce bias, increasing model scale amplifies it, and prompting in Chinese or Russian does not mitigate Western-aligned preferences. The authors emphasize the critical need for systematic bias auditing before deploying LLMs in high-stakes applications. This extensible framework is made publicly available to support future research and improve fairness in AI systems.
cs.AI updates on arXiv.org