ChartDiff: A New Benchmark for Comparative Chart Understanding in AI
Researchers have introduced ChartDiff, the first large-scale benchmark designed to evaluate artificial intelligence models on cross-chart comparative reasoning. While existing benchmarks focus on single-chart interpretation, ChartDiff addresses the gap in analyzing pairs of charts. The dataset comprises 8,541 chart pairs from diverse sources and styles, annotated with human-verified summaries highlighting differences in trends, fluctuations, and anomalies. Evaluations of general-purpose, specialized, and pipeline-based models reveal that frontier general-purpose models achieve higher quality according to GPT-based metrics. In contrast, specialized methods score higher on lexical overlap metrics like ROUGE but perform worse in human-aligned evaluations, indicating a mismatch between standard metrics and actual summary quality. The study also finds that multi-series charts remain difficult for all model families, though strong end-to-end models show robustness to different plotting libraries. These findings underscore that comparative chart reasoning is a significant challenge for current vision-language models. ChartDiff aims to advance research in multi-chart understanding by providing a rigorous standard for assessing how well AI systems can comprehend and summarize complex visual data relationships.
Wire timeline
ChartDiff: A New Benchmark for Comparative Chart Understanding in AI
Researchers have introduced ChartDiff, the first large-scale benchmark designed to evaluate artificial intelligence models on cross-chart comparative reasoning. While existing benchmarks focus on single-chart interpretation, ChartDiff addresses the gap in analyzing pairs of charts. The dataset comprises 8,541 chart pairs from diverse sources and styles, annotated with human-verified summaries highlighting differences in trends, fluctuations, and anomalies. Evaluations of general-purpose, specialized, and pipeline-based models reveal that frontier general-purpose models achieve higher quality according to GPT-based metrics. In contrast, specialized methods score higher on lexical overlap metrics like ROUGE but perform worse in human-aligned evaluations, indicating a mismatch between standard metrics and actual summary quality. The study also finds that multi-series charts remain difficult for all model families, though strong end-to-end models show robustness to different plotting libraries. These findings underscore that comparative chart reasoning is a significant challenge for current vision-language models. ChartDiff aims to advance research in multi-chart understanding by providing a rigorous standard for assessing how well AI systems can comprehend and summarize complex visual data relationships.
cs.AI updates on arXiv.org