Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
A new research paper published on arXiv investigates whether modern multimodal transformer models encode principles of human visual interest or merely exploit large-scale data correlations. The study focuses on the Qwen3-VL-8B vision-language model, analyzing its internal representations using neuroscience-inspired methods. Researchers utilized a Common Interestingness (CI) score derived from human engagement data on Flickr to measure visual interest. The analysis revealed that CI information is linearly decodable from the model's final-layer embeddings, indicating alignment with human-derived measures of interestingness. Through dimensionality reduction and Generalized Discrimination Value analyses, the team found that CI-related hidden representations emerge in intermediate vision transformer layers and become increasingly distinguishable across language model layers. Furthermore, concept vectors derived via geometric, probe, and Sparse Auto-Encoder methods converge in higher layers, suggesting a robust, structured encoding of visual interestingness without explicit supervision. This work aims to bridge the gap between human cognition and artificial intelligence, offering insights for responsible AI use in marketing and communication by uncovering shared computational principles between biological brains and transformer architectures.
Wire timeline
Neuroscience-Inspired Analyses of Visual Interestingness in Multimodal Transformers
A new research paper published on arXiv investigates whether modern multimodal transformer models encode principles of human visual interest or merely exploit large-scale data correlations. The study focuses on the Qwen3-VL-8B vision-language model, analyzing its internal representations using neuroscience-inspired methods. Researchers utilized a Common Interestingness (CI) score derived from human engagement data on Flickr to measure visual interest. The analysis revealed that CI information is linearly decodable from the model's final-layer embeddings, indicating alignment with human-derived measures of interestingness. Through dimensionality reduction and Generalized Discrimination Value analyses, the team found that CI-related hidden representations emerge in intermediate vision transformer layers and become increasingly distinguishable across language model layers. Furthermore, concept vectors derived via geometric, probe, and Sparse Auto-Encoder methods converge in higher layers, suggesting a robust, structured encoding of visual interestingness without explicit supervision. This work aims to bridge the gap between human cognition and artificial intelligence, offering insights for responsible AI use in marketing and communication by uncovering shared computational principles between biological brains and transformer architectures.
cs.AI updates on arXiv.org