AtomGen: Streamlining Atomistic Modeling through Dataset and Benchmark Integration
The Vector Institute for Artificial Intelligence has introduced AtomGen, a project designed to enhance atomistic modeling capabilities by integrating advanced machine-learning techniques with extensive datasets. Central to this initiative is Atomformer, an encoder-only transformer model adapted for three-dimensional atomistic structures. Inspired by the Uni-mol+ architecture, Atomformer incorporates 3D spatial information via Gaussian pair-wise positional embeddings and metadata such as atomic mass and valency to accurately predict system properties. To support this model, the team compiled S2EF-15M, a large-scale dataset containing over 15 million atomic systems, aggregating data from the Open Catalyst Project, Materials Project Trajectory Dataset, and SPICE. This diverse collection enables comprehensive multi-task learning of energies and forces across various chemical environments. The project also features efficient data processing pipelines that unify disparate formats into the HuggingFace Datasets library structure, facilitating parallel processing and dynamic batching. All models and datasets are publicly available on the HuggingFace Hub, promoting reproducibility and ease of use in scientific research. This development aims to streamline the prediction and understanding of material properties from small organic compounds to complex catalyst systems.
Wire timeline
AtomGen: Streamlining Atomistic Modeling through Dataset and Benchmark Integration
The Vector Institute for Artificial Intelligence has introduced AtomGen, a project designed to enhance atomistic modeling capabilities by integrating advanced machine-learning techniques with extensive datasets. Central to this initiative is Atomformer, an encoder-only transformer model adapted for three-dimensional atomistic structures. Inspired by the Uni-mol+ architecture, Atomformer incorporates 3D spatial information via Gaussian pair-wise positional embeddings and metadata such as atomic mass and valency to accurately predict system properties. To support this model, the team compiled S2EF-15M, a large-scale dataset containing over 15 million atomic systems, aggregating data from the Open Catalyst Project, Materials Project Trajectory Dataset, and SPICE. This diverse collection enables comprehensive multi-task learning of energies and forces across various chemical environments. The project also features efficient data processing pipelines that unify disparate formats into the HuggingFace Datasets library structure, facilitating parallel processing and dynamic batching. All models and datasets are publicly available on the HuggingFace Hub, promoting reproducibility and ease of use in scientific research. This development aims to streamline the prediction and understanding of material properties from small organic compounds to complex catalyst systems.
Vector Institute for Artificial Intelligence