VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
Researchers have introduced VT-Bench, the first unified benchmark designed to standardize visual-tabular multi-modal learning tasks. While visual-text learning has gained significant attention, visual-tabular data remains underexplored despite its critical importance in high-stakes domains such as healthcare and industry. VT-Bench aggregates 14 datasets across nine diverse domains, including medical, pets, media, and transportation, comprising over 756,000 samples. The study evaluates 23 representative models, ranging from unimodal experts and specialized visual-tabular models to general-purpose vision-language models (VLMs) and tool-augmented methods. The evaluation highlights substantial challenges in current visual-tabular learning capabilities. By providing a standardized framework for both discriminative prediction and generative reasoning tasks, VT-Bench aims to stimulate the research community to develop more powerful multi-modal foundation models. This initiative addresses a significant gap in artificial intelligence research, offering a comprehensive resource for advancing models that can effectively process and reason across visual and tabular data structures.
Wire timeline
VT-Bench: A Unified Benchmark for Visual-Tabular Multi-Modal Learning
Researchers have introduced VT-Bench, the first unified benchmark designed to standardize visual-tabular multi-modal learning tasks. While visual-text learning has gained significant attention, visual-tabular data remains underexplored despite its critical importance in high-stakes domains such as healthcare and industry. VT-Bench aggregates 14 datasets across nine diverse domains, including medical, pets, media, and transportation, comprising over 756,000 samples. The study evaluates 23 representative models, ranging from unimodal experts and specialized visual-tabular models to general-purpose vision-language models (VLMs) and tool-augmented methods. The evaluation highlights substantial challenges in current visual-tabular learning capabilities. By providing a standardized framework for both discriminative prediction and generative reasoning tasks, VT-Bench aims to stimulate the research community to develop more powerful multi-modal foundation models. This initiative addresses a significant gap in artificial intelligence research, offering a comprehensive resource for advancing models that can effectively process and reason across visual and tabular data structures.
cs.AI updates on arXiv.org