DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
Researchers have introduced DocSeeker, a novel framework designed to enhance the performance of Multimodal Large Language Models (MLLMs) in understanding long documents. Current MLLMs struggle with lengthy texts due to low signal-to-noise ratios and insufficient supervision data. DocSeeker addresses these issues by implementing a structured workflow involving analysis, localization, and reasoning. The model utilizes a two-stage training framework: first, supervised fine-tuning on high-quality data generated through knowledge distillation, and second, Evidence-aware Group Relative Policy Optimization to jointly improve evidence localization and answer accuracy. Additionally, an Evidence-Guided Resolution Allocation strategy is employed to manage memory constraints during multi-page document processing. Extensive experiments indicate that DocSeeker achieves superior results on both in-domain and out-of-domain tasks, demonstrating robust generalization from short-page training to ultra-long documents. This advancement provides a solid foundation for visual Retrieval-Augmented Generation systems, marking a significant step forward in artificial intelligence capabilities for complex document analysis.
Wire timeline
DocSeeker: Structured Visual Reasoning with Evidence Grounding for Long Document Understanding
Researchers have introduced DocSeeker, a novel framework designed to enhance the performance of Multimodal Large Language Models (MLLMs) in understanding long documents. Current MLLMs struggle with lengthy texts due to low signal-to-noise ratios and insufficient supervision data. DocSeeker addresses these issues by implementing a structured workflow involving analysis, localization, and reasoning. The model utilizes a two-stage training framework: first, supervised fine-tuning on high-quality data generated through knowledge distillation, and second, Evidence-aware Group Relative Policy Optimization to jointly improve evidence localization and answer accuracy. Additionally, an Evidence-Guided Resolution Allocation strategy is employed to manage memory constraints during multi-page document processing. Extensive experiments indicate that DocSeeker achieves superior results on both in-domain and out-of-domain tasks, demonstrating robust generalization from short-page training to ultra-long documents. This advancement provides a solid foundation for visual Retrieval-Augmented Generation systems, marking a significant step forward in artificial intelligence capabilities for complex document analysis.
cs.AI updates on arXiv.org