Study Reveals Grounding Gap in LLMs' Understanding of Abstract Concepts
A new research paper published on arXiv investigates how Large Language Models (LLMs) anchor the meaning of abstract concepts compared to humans. By replicating cognitive science property-generation experiments across 21 frontier and open-weight LLMs, the study identifies a significant 'grounding gap.' Results show that while human responses correlate above 0.9, no LLM exceeded a Pearson correlation of 0.37 with human data. The models rely heavily on word associations, frequently underproducing properties related to emotion and internal states. However, when explicitly queried about grounding categories, alignment with human judgment improves, particularly in larger models. Using sparse autoencoders, researchers found that internal features for 'sensorimotor' and 'social' dimensions exist but are not recruited naturally during free generation. This suggests current LLMs possess the necessary grounding information but fail to utilize it in a human-like manner without explicit prompting, highlighting a fundamental difference in how AI and human brains process abstract meanings like justice or theory.
Wire timeline
Study Reveals Grounding Gap in LLMs' Understanding of Abstract Concepts
A new research paper published on arXiv investigates how Large Language Models (LLMs) anchor the meaning of abstract concepts compared to humans. By replicating cognitive science property-generation experiments across 21 frontier and open-weight LLMs, the study identifies a significant 'grounding gap.' Results show that while human responses correlate above 0.9, no LLM exceeded a Pearson correlation of 0.37 with human data. The models rely heavily on word associations, frequently underproducing properties related to emotion and internal states. However, when explicitly queried about grounding categories, alignment with human judgment improves, particularly in larger models. Using sparse autoencoders, researchers found that internal features for 'sensorimotor' and 'social' dimensions exist but are not recruited naturally during free generation. This suggests current LLMs possess the necessary grounding information but fail to utilize it in a human-like manner without explicit prompting, highlighting a fundamental difference in how AI and human brains process abstract meanings like justice or theory.
cs.AI updates on arXiv.org