Scalable Bayesian Planner Enhances Multimodal Theory-of-Mind Reasoning in AI
Researchers have introduced a novel scalable Bayesian Theory-of-Mind (ToM) planner designed to overcome the limitations of existing computational methods in multimodal environments. Current ToM approaches often struggle with scalability and generalization as task complexity increases, relying heavily on structured workflows or deep model fine-tuning. The proposed framework decomposes ToM reasoning into stepwise Bayesian updates and employs a weak-to-strong control mechanism. This allows smaller language models to specialize in ToM-specific likelihood estimation, transferring their reasoning behaviors to larger language models ranging from 7B to 405B parameters for better integration with social and world knowledge. Extensive experiments demonstrate that this synergistic approach aligns large-model inference with Bayesian principles, achieving a 4.6% accuracy improvement over state-of-the-art techniques on multimodal ToM benchmarks. The method shows particular strength in handling challenging unseen scenarios, establishing a new standard for modeling human mental states, such as beliefs, desires, and intentions, in complex artificial intelligence systems.
Wire timeline
Scalable Bayesian Planner Enhances Multimodal Theory-of-Mind Reasoning in AI
Researchers have introduced a novel scalable Bayesian Theory-of-Mind (ToM) planner designed to overcome the limitations of existing computational methods in multimodal environments. Current ToM approaches often struggle with scalability and generalization as task complexity increases, relying heavily on structured workflows or deep model fine-tuning. The proposed framework decomposes ToM reasoning into stepwise Bayesian updates and employs a weak-to-strong control mechanism. This allows smaller language models to specialize in ToM-specific likelihood estimation, transferring their reasoning behaviors to larger language models ranging from 7B to 405B parameters for better integration with social and world knowledge. Extensive experiments demonstrate that this synergistic approach aligns large-model inference with Bayesian principles, achieving a 4.6% accuracy improvement over state-of-the-art techniques on multimodal ToM benchmarks. The method shows particular strength in handling challenging unseen scenarios, establishing a new standard for modeling human mental states, such as beliefs, desires, and intentions, in complex artificial intelligence systems.
cs.AI updates on arXiv.org