RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
Researchers have introduced RW-Post, a new text-image benchmark designed for real-world multimodal fact-checking. This initiative addresses the growing challenge of multimodal misinformation, where manipulated or repurposed images are used to reinforce misleading textual narratives. The benchmark features auditable annotations, linking original social media posts with reasoning traces and explicit evidence items extracted from human fact-check articles via an LLM-assisted pipeline. RW-Post supports controlled evaluations across closed-book, evidence-bounded, and open-web regimes, allowing for systematic diagnosis of visual grounding and evidence utilization. The study also presents AgentFact as a reference verification baseline and benchmarks strong open-source Large Vision-Language Models (LVLMs). Experimental results indicate significant room for improvement, as current models struggle with faithful evidence grounding. However, evidence-bounded evaluation was found to enhance both accuracy and faithfulness. The code and dataset for this project are scheduled for public release to support further research in detecting and verifying complex multimedia misinformation.
Wire timeline
RW-Post: Auditable Evidence-Grounded Multimodal Fact-Checking in the Wild
Researchers have introduced RW-Post, a new text-image benchmark designed for real-world multimodal fact-checking. This initiative addresses the growing challenge of multimodal misinformation, where manipulated or repurposed images are used to reinforce misleading textual narratives. The benchmark features auditable annotations, linking original social media posts with reasoning traces and explicit evidence items extracted from human fact-check articles via an LLM-assisted pipeline. RW-Post supports controlled evaluations across closed-book, evidence-bounded, and open-web regimes, allowing for systematic diagnosis of visual grounding and evidence utilization. The study also presents AgentFact as a reference verification baseline and benchmarks strong open-source Large Vision-Language Models (LVLMs). Experimental results indicate significant room for improvement, as current models struggle with faithful evidence grounding. However, evidence-bounded evaluation was found to enhance both accuracy and faithfulness. The code and dataset for this project are scheduled for public release to support further research in detecting and verifying complex multimedia misinformation.
cs.AI updates on arXiv.org