SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making
Researchers Zeyu Li and Lei Li have introduced SKG-VLA, a novel framework designed to enhance decision-making in large-scale complaint handling systems. Current systems often rely on shallow classification and isolated modalities, failing to leverage explicit scene structures and cross-evidence dependencies. SKG-VLA addresses this by modeling each case as a structured complaint scene using a Scene Knowledge Graph (SKG). This graph integrates diverse elements such as complaint entities, evidence items, policy clauses, temporal events, and transactional states. The team developed a data synthesis pipeline to generate scene descriptions and rule-consistent generalizations, alongside a large-scale multimodal dataset for benchmarking. The model employs a three-stage training strategy involving domain-adaptive pre-training, instruction fine-tuning, and multimodal alignment. Experimental results demonstrate that SKG-VLA significantly improves policy-grounded reasoning, decision accuracy, and robustness, particularly in scenarios with incomplete evidence or long-tail cases. This advancement represents a significant step in applying structured knowledge graphs to multimodal AI reasoning for complex administrative tasks.
Wire timeline
SKG-VLA: Scene Knowledge Graph Priors for Structured Scene Semantics and Multimodal Reasoning for Decision Making
Researchers Zeyu Li and Lei Li have introduced SKG-VLA, a novel framework designed to enhance decision-making in large-scale complaint handling systems. Current systems often rely on shallow classification and isolated modalities, failing to leverage explicit scene structures and cross-evidence dependencies. SKG-VLA addresses this by modeling each case as a structured complaint scene using a Scene Knowledge Graph (SKG). This graph integrates diverse elements such as complaint entities, evidence items, policy clauses, temporal events, and transactional states. The team developed a data synthesis pipeline to generate scene descriptions and rule-consistent generalizations, alongside a large-scale multimodal dataset for benchmarking. The model employs a three-stage training strategy involving domain-adaptive pre-training, instruction fine-tuning, and multimodal alignment. Experimental results demonstrate that SKG-VLA significantly improves policy-grounded reasoning, decision accuracy, and robustness, particularly in scenarios with incomplete evidence or long-tail cases. This advancement represents a significant step in applying structured knowledge graphs to multimodal AI reasoning for complex administrative tasks.
cs.AI updates on arXiv.org