STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
Researchers have introduced STAGE (Screenplay Text, Agents, Graphs and Evaluation), a new unified benchmark designed to evaluate artificial intelligence models on their ability to understand and reason over full-length movie screenplays. Unlike previous benchmarks that focus on isolated subtasks like question answering or dialogue generation, STAGE assesses whether models can construct a coherent story world and maintain consistency across various reasoning and generation tasks. The benchmark encompasses four specific tasks: knowledge graph construction, scene-level event summarization, long-context screenplay question answering, and in-script character role-playing. These tasks are grounded in a shared narrative world representation. STAGE provides cleaned scripts, curated knowledge graphs, and detailed annotations for 150 films in both English and Chinese. This resource enables a holistic evaluation of AI capabilities in building world representations, abstracting and verifying narrative events, reasoning over long narratives, and generating responses that remain consistent with character profiles. The paper was published on arXiv in the field of Computation and Language, aiming to advance the development of more sophisticated narrative understanding systems.
Wire timeline
STAGE: A Full-Screenplay Benchmark for Reasoning over Evolving Stories
Researchers have introduced STAGE (Screenplay Text, Agents, Graphs and Evaluation), a new unified benchmark designed to evaluate artificial intelligence models on their ability to understand and reason over full-length movie screenplays. Unlike previous benchmarks that focus on isolated subtasks like question answering or dialogue generation, STAGE assesses whether models can construct a coherent story world and maintain consistency across various reasoning and generation tasks. The benchmark encompasses four specific tasks: knowledge graph construction, scene-level event summarization, long-context screenplay question answering, and in-script character role-playing. These tasks are grounded in a shared narrative world representation. STAGE provides cleaned scripts, curated knowledge graphs, and detailed annotations for 150 films in both English and Chinese. This resource enables a holistic evaluation of AI capabilities in building world representations, abstracting and verifying narrative events, reasoning over long narratives, and generating responses that remain consistent with character profiles. The paper was published on arXiv in the field of Computation and Language, aiming to advance the development of more sophisticated narrative understanding systems.
cs.AI updates on arXiv.org