Multimodal Models Develop Mental Imagery for Spatial Reasoning Tasks
A new research paper published on arXiv reveals that large multimodal models, specifically a fine-tuned Qwen3.5 Vision-Language Model (VLM), develop internal mental imagery when solving complex spatial puzzles. The study examined twelve diverse visual reasoning tasks, including tangrams, jigsaws, Sokoban, and 3D mental rotation. Researchers found that the model's activations encode meaningful visual information about intermediate states during problem-solving, suggesting the emergence of an imperfect visual world model as a byproduct of action selection, even without explicit visual supervision. Leveraging this discovery, the authors proposed methods to sharpen these mental images. By integrating just sixteen visual tokens per step into the model's chain of thought, the average solve rate improved from 83% to 89%. Significant performance gains were observed in reasoning-heavy tasks like jigsaw puzzles and 3D mental rotation. This finding indicates that multimodal AI systems can spontaneously generate visual representations to aid in logical deduction and spatial understanding, marking a significant step forward in interpreting how these models process and visualize complex geometric relationships and action consequences.
Wire timeline
Multimodal Models Develop Mental Imagery for Spatial Reasoning Tasks
A new research paper published on arXiv reveals that large multimodal models, specifically a fine-tuned Qwen3.5 Vision-Language Model (VLM), develop internal mental imagery when solving complex spatial puzzles. The study examined twelve diverse visual reasoning tasks, including tangrams, jigsaws, Sokoban, and 3D mental rotation. Researchers found that the model's activations encode meaningful visual information about intermediate states during problem-solving, suggesting the emergence of an imperfect visual world model as a byproduct of action selection, even without explicit visual supervision. Leveraging this discovery, the authors proposed methods to sharpen these mental images. By integrating just sixteen visual tokens per step into the model's chain of thought, the average solve rate improved from 83% to 89%. Significant performance gains were observed in reasoning-heavy tasks like jigsaw puzzles and 3D mental rotation. This finding indicates that multimodal AI systems can spontaneously generate visual representations to aid in logical deduction and spatial understanding, marking a significant step forward in interpreting how these models process and visualize complex geometric relationships and action consequences.
cs.AI updates on arXiv.org