LLMs Struggle to Estimate Student Difficulty Due to Cognitive Misalignment
A recent study published on arXiv investigates whether Large Language Models (LLMs) can accurately estimate the difficulty of educational items for human learners, addressing the cold start problem in assessment. The research conducts a large-scale empirical analysis of over 20 models across diverse domains, including medical knowledge and mathematical reasoning. Findings reveal a systematic misalignment between AI predictions and human cognitive struggles. Contrary to expectations, scaling up model size does not improve alignment; instead, models converge toward a shared machine consensus. High performance often impedes accurate difficulty estimation because LLMs fail to simulate student capability limitations, even when prompted to adopt specific proficiency levels. Additionally, the study identifies a critical lack of introspection, as models cannot predict their own limitations. These results suggest that general problem-solving capabilities do not equate to an understanding of human learning challenges, highlighting significant obstacles for using current LLMs in automated educational difficulty prediction.
Wire timeline
LLMs Struggle to Estimate Student Difficulty Due to Cognitive Misalignment
A recent study published on arXiv investigates whether Large Language Models (LLMs) can accurately estimate the difficulty of educational items for human learners, addressing the cold start problem in assessment. The research conducts a large-scale empirical analysis of over 20 models across diverse domains, including medical knowledge and mathematical reasoning. Findings reveal a systematic misalignment between AI predictions and human cognitive struggles. Contrary to expectations, scaling up model size does not improve alignment; instead, models converge toward a shared machine consensus. High performance often impedes accurate difficulty estimation because LLMs fail to simulate student capability limitations, even when prompted to adopt specific proficiency levels. Additionally, the study identifies a critical lack of introspection, as models cannot predict their own limitations. These results suggest that general problem-solving capabilities do not equate to an understanding of human learning challenges, highlighting significant obstacles for using current LLMs in automated educational difficulty prediction.
cs.AI updates on arXiv.org