The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?
A new research paper submitted to arXiv introduces the 'Wittgensteinian Representation Hypothesis,' proposing that language serves as the asymptotic attractor for multimodal representation convergence. The study addresses an open question in representation learning regarding why independently trained neural networks from different modalities converge toward shared representations. By employing directional convergence analysis via cycle-kNN, an asymmetric alignment measure, the authors analyzed dozens of unimodal models across point clouds, vision, and language. They discovered a consistent directional asymmetry where non-language modalities move significantly more toward the neighborhood structure of language than vice versa. This pattern, invisible to traditional symmetric measures, is attributed to feature density asymmetry, with language representations occupying the most compact regions of representational space. Using the Information Bottleneck framework, the authors argue that optimization under compression drives representations toward the discrete, compositional structures characteristic of language. This findings suggest that the semantic structure of language fundamentally shapes how multimodal AI systems organize information, offering a principled interpretation of cross-modal alignment in artificial intelligence.
Wire timeline
The Wittgensteinian Representation Hypothesis: Is Language the Attractor of Multimodal Convergence?
A new research paper submitted to arXiv introduces the 'Wittgensteinian Representation Hypothesis,' proposing that language serves as the asymptotic attractor for multimodal representation convergence. The study addresses an open question in representation learning regarding why independently trained neural networks from different modalities converge toward shared representations. By employing directional convergence analysis via cycle-kNN, an asymmetric alignment measure, the authors analyzed dozens of unimodal models across point clouds, vision, and language. They discovered a consistent directional asymmetry where non-language modalities move significantly more toward the neighborhood structure of language than vice versa. This pattern, invisible to traditional symmetric measures, is attributed to feature density asymmetry, with language representations occupying the most compact regions of representational space. Using the Information Bottleneck framework, the authors argue that optimization under compression drives representations toward the discrete, compositional structures characteristic of language. This findings suggest that the semantic structure of language fundamentally shapes how multimodal AI systems organize information, offering a principled interpretation of cross-modal alignment in artificial intelligence.
cs.AI updates on arXiv.org