MolRGen: A Benchmark for De Novo Molecular Generation with Reasoning LLMs
Researchers have introduced MolRGen, a new benchmark and molecular verifier designed to train and evaluate reasoning-based large language models (LLMs) in de novo molecular generation. Addressing the lack of training environments where rewards can be computed without reference molecules, MolRGen features approximately 4,500 protein-pocket targets and 50,000 multi-objective optimization prompts. These prompts integrate docking scores with key molecular properties such as QED, synthetic accessibility, logP, and physicochemical descriptors. Unlike existing caption-based or molecule-editing benchmarks, MolRGen evaluates molecules generated from scratch by computing rewards during the generation process. The study benchmarks both general-purpose and chemistry-specialized open-source LLMs, introducing a diversity-aware top-k metric to assess the ability to generate diverse, high-scoring molecules. Additionally, the authors demonstrate fine-tuning a 128B LLM using GRPO, which improves performance but reveals a trade-off between diversity and exploitation. This framework provides a scalable testbed for advancing verifier-based reasoning and reinforcement learning in computational molecular design.
Wire timeline
MolRGen: A Benchmark for De Novo Molecular Generation with Reasoning LLMs
Researchers have introduced MolRGen, a new benchmark and molecular verifier designed to train and evaluate reasoning-based large language models (LLMs) in de novo molecular generation. Addressing the lack of training environments where rewards can be computed without reference molecules, MolRGen features approximately 4,500 protein-pocket targets and 50,000 multi-objective optimization prompts. These prompts integrate docking scores with key molecular properties such as QED, synthetic accessibility, logP, and physicochemical descriptors. Unlike existing caption-based or molecule-editing benchmarks, MolRGen evaluates molecules generated from scratch by computing rewards during the generation process. The study benchmarks both general-purpose and chemistry-specialized open-source LLMs, introducing a diversity-aware top-k metric to assess the ability to generate diverse, high-scoring molecules. Additionally, the authors demonstrate fine-tuning a 128B LLM using GRPO, which improves performance but reveals a trade-off between diversity and exploitation. This framework provides a scalable testbed for advancing verifier-based reasoning and reinforcement learning in computational molecular design.
cs.AI updates on arXiv.org