Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
Researchers have introduced DISCA (Disagreement-Informed Steering for Cultural Alignment), a novel method for aligning large language models (LLMs) with diverse cultural values without requiring model fine-tuning or white-box access. Addressing the limitation that existing alignment techniques often rely on expensive per-country data or internal model access, DISCA operates in a realistic black-box regime using only public data. The approach leverages within-country sociodemographic disagreement as a primary steering signal, instantiating each country as a panel of persona agents grounded in the World Values Survey. These agents' disagreements are converted into bounded, loss-averse logit corrections during inference. Tested across 20 countries and seven open-weight backbones ranging from 2B to 70B parameters, DISCA reduced cultural misalignment by 10–24% on standard benchmarks for models larger than 3.8B parameters. This training-free technique offers a scalable alternative for serving global moral preferences, demonstrating that inference-time calibration can effectively address cultural biases in LLMs without altering model weights.
Wire timeline
Training-Free Cultural Alignment of Large Language Models via Persona Disagreement
Researchers have introduced DISCA (Disagreement-Informed Steering for Cultural Alignment), a novel method for aligning large language models (LLMs) with diverse cultural values without requiring model fine-tuning or white-box access. Addressing the limitation that existing alignment techniques often rely on expensive per-country data or internal model access, DISCA operates in a realistic black-box regime using only public data. The approach leverages within-country sociodemographic disagreement as a primary steering signal, instantiating each country as a panel of persona agents grounded in the World Values Survey. These agents' disagreements are converted into bounded, loss-averse logit corrections during inference. Tested across 20 countries and seven open-weight backbones ranging from 2B to 70B parameters, DISCA reduced cultural misalignment by 10–24% on standard benchmarks for models larger than 3.8B parameters. This training-free technique offers a scalable alternative for serving global moral preferences, demonstrating that inference-time calibration can effectively address cultural biases in LLMs without altering model weights.
cs.AI updates on arXiv.org