StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
Researchers have introduced StereoTales, a comprehensive multilingual dataset and evaluation pipeline designed to systematically study social bias in open-ended Large Language Model (LLM) generation. Addressing the limitations of existing English-centric benchmarks, this framework covers ten languages and 79 socio-demographic attributes, utilizing over 650,000 stories generated by 23 recent LLMs. The study identifies more than 1,500 over-represented associations, which were rated for harmfulness by both human panels and AI models. Key findings reveal that all evaluated models produce consequential harmful stereotypes, regardless of their size or capabilities, with these biases largely shared across providers rather than being isolated incidents. Furthermore, the prompt language significantly influences which stereotypes emerge, as biases adapt culturally to amplify prejudice against locally salient protected groups. The research also demonstrates a broad alignment between human and LLM judgments on harmfulness. To facilitate further analysis and mitigation efforts, the authors have released the evaluation code, dataset, model generations, and annotations, marking a significant step forward in understanding and addressing multilingual social bias in artificial intelligence systems.
Wire timeline
StereoTales: A Multilingual Framework for Open-Ended Stereotype Discovery in LLMs
Researchers have introduced StereoTales, a comprehensive multilingual dataset and evaluation pipeline designed to systematically study social bias in open-ended Large Language Model (LLM) generation. Addressing the limitations of existing English-centric benchmarks, this framework covers ten languages and 79 socio-demographic attributes, utilizing over 650,000 stories generated by 23 recent LLMs. The study identifies more than 1,500 over-represented associations, which were rated for harmfulness by both human panels and AI models. Key findings reveal that all evaluated models produce consequential harmful stereotypes, regardless of their size or capabilities, with these biases largely shared across providers rather than being isolated incidents. Furthermore, the prompt language significantly influences which stereotypes emerge, as biases adapt culturally to amplify prejudice against locally salient protected groups. The research also demonstrates a broad alignment between human and LLM judgments on harmfulness. To facilitate further analysis and mitigation efforts, the authors have released the evaluation code, dataset, model generations, and annotations, marking a significant step forward in understanding and addressing multilingual social bias in artificial intelligence systems.
cs.AI updates on arXiv.org