Tracing Moral Foundations in Large Language Models
A new study published on arXiv investigates whether large language models (LLMs) possess genuine internal moral structures or merely exhibit superficial mimicry. Using Moral Foundations Theory (MFT), researchers analyzed 14 base and instruction-tuned LLMs from families including Llama, Qwen, and Mistral, ranging from 7B to 70B parameters. The team employed a multi-level approach involving layer-wise analysis, pretrained sparse autoencoders (SAEs), and causal steering interventions. Results indicate that LLMs represent and distinguish moral foundations in ways aligning with human judgments. This moral geometry emerges naturally during pretraining and is selectively modified by post-training. SAE features revealed semantic links to specific foundations, suggesting partially disentangled mechanisms. Furthermore, steering interventions demonstrated a causal connection between internal representations and moral outputs. The findings provide mechanistic evidence that moral concepts in LLMs are distributed, layered, and partly disentangled, implying that pluralistic moral structures can emerge as latent patterns from statistical regularities in language data alone.
Wire timeline
Tracing Moral Foundations in Large Language Models
A new study published on arXiv investigates whether large language models (LLMs) possess genuine internal moral structures or merely exhibit superficial mimicry. Using Moral Foundations Theory (MFT), researchers analyzed 14 base and instruction-tuned LLMs from families including Llama, Qwen, and Mistral, ranging from 7B to 70B parameters. The team employed a multi-level approach involving layer-wise analysis, pretrained sparse autoencoders (SAEs), and causal steering interventions. Results indicate that LLMs represent and distinguish moral foundations in ways aligning with human judgments. This moral geometry emerges naturally during pretraining and is selectively modified by post-training. SAE features revealed semantic links to specific foundations, suggesting partially disentangled mechanisms. Furthermore, steering interventions demonstrated a causal connection between internal representations and moral outputs. The findings provide mechanistic evidence that moral concepts in LLMs are distributed, layered, and partly disentangled, implying that pluralistic moral structures can emerge as latent patterns from statistical regularities in language data alone.
cs.AI updates on arXiv.org