Evaluating Developmental Cognition Capabilities of LLMs
A new research paper introduces the Developmental Sentence Completion Test (DSCT) to evaluate how Large Language Models (LLMs) handle developmental cognition, based on Robert Kegan's constructive-developmental theory. Current conversational AI personalization often overlooks how users interpret outputs to construct reality. The DSCT, a 20-item instrument, aims to elicit developmental signals in text without the scalability issues of expert interviews or proprietary tools. The study assesses LLM performance across three regimes: simulated personas, real human responses, and default model-generated answers. Results indicate that top-tier models accurately recover intended labels in simulated personas. However, agreement with real human responses is only fair, though stronger within neighboring developmental stages. Notably, larger and newer models tend to generate higher-rated text when answering without persona conditioning. The findings suggest that developmental signals are clearer in synthetic responses than in human-written text. Consequently, the primary challenge for stage-aware conversational AI is not just classifier accuracy, but the availability and quality of developmental signals within elicited text, highlighting a new dimension for AI evaluation and personalization.
Wire timeline
Evaluating Developmental Cognition Capabilities of LLMs
A new research paper introduces the Developmental Sentence Completion Test (DSCT) to evaluate how Large Language Models (LLMs) handle developmental cognition, based on Robert Kegan's constructive-developmental theory. Current conversational AI personalization often overlooks how users interpret outputs to construct reality. The DSCT, a 20-item instrument, aims to elicit developmental signals in text without the scalability issues of expert interviews or proprietary tools. The study assesses LLM performance across three regimes: simulated personas, real human responses, and default model-generated answers. Results indicate that top-tier models accurately recover intended labels in simulated personas. However, agreement with real human responses is only fair, though stronger within neighboring developmental stages. Notably, larger and newer models tend to generate higher-rated text when answering without persona conditioning. The findings suggest that developmental signals are clearer in synthetic responses than in human-written text. Consequently, the primary challenge for stage-aware conversational AI is not just classifier accuracy, but the availability and quality of developmental signals within elicited text, highlighting a new dimension for AI evaluation and personalization.
cs.AI updates on arXiv.org