Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
A new academic study investigates the reliability of Large Language Models (LLMs) in interpreting personal sensing data, such as activity and mood traces. The researchers introduce the concept of 'epistemic overreach' (EO), where AI-generated explanations imply causal relationships unsupported by available evidence. Analyzing 14,922 explanations across three longitudinal datasets (StudentLife, GLOBEM, CollegeExperience) using Llama, Qwen, and GPT models, the study finds that LLMs frequently attribute anomalous days to causes without sufficient data support. This pattern persists across different model families and anomaly types. The research demonstrates that while bounded prompting reduces EO, it does not eliminate it, and providing richer context does not reliably improve evidential grounding. The authors argue for establishing evidential discipline in AI systems, urging that evaluations of personal sensing explanations must prioritize evidential grounding alongside fluency. This highlights significant challenges in deploying LLMs for health and behavioral monitoring where accurate causal inference is critical.
Wire timeline
Auditing Epistemic Overreach in LLM-Generated Personal Sensing Explanations
A new academic study investigates the reliability of Large Language Models (LLMs) in interpreting personal sensing data, such as activity and mood traces. The researchers introduce the concept of 'epistemic overreach' (EO), where AI-generated explanations imply causal relationships unsupported by available evidence. Analyzing 14,922 explanations across three longitudinal datasets (StudentLife, GLOBEM, CollegeExperience) using Llama, Qwen, and GPT models, the study finds that LLMs frequently attribute anomalous days to causes without sufficient data support. This pattern persists across different model families and anomaly types. The research demonstrates that while bounded prompting reduces EO, it does not eliminate it, and providing richer context does not reliably improve evidential grounding. The authors argue for establishing evidential discipline in AI systems, urging that evaluations of personal sensing explanations must prioritize evidential grounding alongside fluency. This highlights significant challenges in deploying LLMs for health and behavioral monitoring where accurate causal inference is critical.
cs.AI updates on arXiv.org