Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials
A new research paper submitted to arXiv addresses a critical limitation in Machine Learning Interatomic Potentials (MLIPs), which are essential for computational chemistry and materials science. While current models estimate inter-atomic forces with high precision, their ability to generalize to previously unseen molecules remains uncertain. The authors question whether these models truly learn the compositional structure of chemistry or merely interpolate patterns from training data. To investigate this, the study introduces a benchmark comprising four tasks designed to test compositional generalization on out-of-distribution molecules. The empirical analysis reveals that state-of-the-art models, including foundation models pre-trained on millions of molecules, struggle significantly with these tasks. Errors on unseen examples were found to be an order of magnitude higher than on in-distribution data. This findings highlight a substantial gap in current AI capabilities for scientific discovery, suggesting that existing MLIPs do not yet fully capture underlying physical principles required for robust generalization in drug design and materials discovery applications.
Wire timeline
Benchmarking Compositional Generalisation for Machine Learning Interatomic Potentials
A new research paper submitted to arXiv addresses a critical limitation in Machine Learning Interatomic Potentials (MLIPs), which are essential for computational chemistry and materials science. While current models estimate inter-atomic forces with high precision, their ability to generalize to previously unseen molecules remains uncertain. The authors question whether these models truly learn the compositional structure of chemistry or merely interpolate patterns from training data. To investigate this, the study introduces a benchmark comprising four tasks designed to test compositional generalization on out-of-distribution molecules. The empirical analysis reveals that state-of-the-art models, including foundation models pre-trained on millions of molecules, struggle significantly with these tasks. Errors on unseen examples were found to be an order of magnitude higher than on in-distribution data. This findings highlight a substantial gap in current AI capabilities for scientific discovery, suggesting that existing MLIPs do not yet fully capture underlying physical principles required for robust generalization in drug design and materials discovery applications.
cs.AI updates on arXiv.org