X-Value: A New Benchmark for Cross-Lingual Values Judgment in LLMs
Researchers have introduced X-Value, the first cross-lingual values judgment benchmark designed to evaluate how large language models (LLMs) assess deep-level values across multiple languages. Addressing the limitations of existing evaluations that focus primarily on factual tasks, this study highlights challenges such as cultural diversity and disciplinary complexity. The team developed a novel two-stage human-AI collaborative annotation framework to create the dataset, which includes 4,750 question-answer pairs across 14 languages and seven major global issue categories. Systematic evaluations of 17 different LLMs revealed significant performance disparities and limitations in their ability to judge values consistently across languages and cultures. The findings underscore an urgent need to enhance the values-aware content judgment capabilities of AI systems. This academic publication aims to bridge the gap in multilingual capability assessment, providing granular metadata to facilitate rigorous model evaluation and promoting a consensus-pluralism perspective in AI development.
Wire timeline
X-Value: A New Benchmark for Cross-Lingual Values Judgment in LLMs
Researchers have introduced X-Value, the first cross-lingual values judgment benchmark designed to evaluate how large language models (LLMs) assess deep-level values across multiple languages. Addressing the limitations of existing evaluations that focus primarily on factual tasks, this study highlights challenges such as cultural diversity and disciplinary complexity. The team developed a novel two-stage human-AI collaborative annotation framework to create the dataset, which includes 4,750 question-answer pairs across 14 languages and seven major global issue categories. Systematic evaluations of 17 different LLMs revealed significant performance disparities and limitations in their ability to judge values consistently across languages and cultures. The findings underscore an urgent need to enhance the values-aware content judgment capabilities of AI systems. This academic publication aims to bridge the gap in multilingual capability assessment, providing granular metadata to facilitate rigorous model evaluation and promoting a consensus-pluralism perspective in AI development.
cs.AI updates on arXiv.org