PolyLM: Predicting Polymer Properties via Large Language Models
Researchers have introduced PolyLM, a novel framework leveraging large language models (LLMs) to predict physical and mechanical polymer properties directly from unstructured scientific literature. Unlike traditional models that rely solely on chemical structure representations like SMILES, PolyLM processes natural language descriptions of synthesis routes, processing history, and testing conditions. This approach addresses the limitation where identical polymers exhibit different behaviors based on experimental context. The team curated a massive dataset comprising 185,000 scientific papers and over 276,400 unique polymer samples across 22 diverse properties. By fine-tuning the Qwen3.5-9B model using Low-Rank Adaptation (LoRA), they achieved state-of-the-art predictive accuracy. Evaluations on 68,283 held-out observations yielded a median R-squared value of 0.74, with key thermal and mechanical properties frequently exceeding 0.80. This study demonstrates that natural language serves as a powerful, scalable interface for realistic materials performance prediction, preserving nuanced experimental details often lost in structured data formats.
Wire timeline
PolyLM: Predicting Polymer Properties via Large Language Models
Researchers have introduced PolyLM, a novel framework leveraging large language models (LLMs) to predict physical and mechanical polymer properties directly from unstructured scientific literature. Unlike traditional models that rely solely on chemical structure representations like SMILES, PolyLM processes natural language descriptions of synthesis routes, processing history, and testing conditions. This approach addresses the limitation where identical polymers exhibit different behaviors based on experimental context. The team curated a massive dataset comprising 185,000 scientific papers and over 276,400 unique polymer samples across 22 diverse properties. By fine-tuning the Qwen3.5-9B model using Low-Rank Adaptation (LoRA), they achieved state-of-the-art predictive accuracy. Evaluations on 68,283 held-out observations yielded a median R-squared value of 0.74, with key thermal and mechanical properties frequently exceeding 0.80. This study demonstrates that natural language serves as a powerful, scalable interface for realistic materials performance prediction, preserving nuanced experimental details often lost in structured data formats.
cs.AI updates on arXiv.org