SayNext-Bench: Evaluating LLM Struggles with Next-Utterance Anticipation
Researchers have identified a significant gap in the ability of Large Language Models (LLMs) to anticipate human dialogue, despite their proficiency in generating responses. The study, titled 'SayNext-Bench,' highlights that while humans utilize multimodal cues like gestures and emotional tone for anticipation, current LLMs struggle with this task. To address this, the team introduced SayNext-Bench, a benchmark for evaluating Multimodal LLMs (MLLMs) on context-conditioned response anticipation. They also created SayNext-PC, a large-scale multimodal dialogue dataset, and a multi-level evaluation framework assessing lexical similarity, emotion-intention consistency, and overall alignment. Furthermore, the authors developed SayNext-Chat, a cognitively inspired dual-route MLLM using learnable priming tokens to integrate perceptual cues with anticipatory priors. Extensive experiments and user studies demonstrate that SayNext-Chat outperforms state-of-the-art MLLMs. The findings emphasize the critical role of multimodal cues and active anticipatory processing in natural human interaction, areas currently lacking in existing AI models.
Wire timeline
SayNext-Bench: Evaluating LLM Struggles with Next-Utterance Anticipation
Researchers have identified a significant gap in the ability of Large Language Models (LLMs) to anticipate human dialogue, despite their proficiency in generating responses. The study, titled 'SayNext-Bench,' highlights that while humans utilize multimodal cues like gestures and emotional tone for anticipation, current LLMs struggle with this task. To address this, the team introduced SayNext-Bench, a benchmark for evaluating Multimodal LLMs (MLLMs) on context-conditioned response anticipation. They also created SayNext-PC, a large-scale multimodal dialogue dataset, and a multi-level evaluation framework assessing lexical similarity, emotion-intention consistency, and overall alignment. Furthermore, the authors developed SayNext-Chat, a cognitively inspired dual-route MLLM using learnable priming tokens to integrate perceptual cues with anticipatory priors. Extensive experiments and user studies demonstrate that SayNext-Chat outperforms state-of-the-art MLLMs. The findings emphasize the critical role of multimodal cues and active anticipatory processing in natural human interaction, areas currently lacking in existing AI models.
cs.AI updates on arXiv.org