AssayBench: A New Benchmark for Virtual Cell Modeling with LLMs
Researchers have introduced AssayBench, a novel benchmark designed to evaluate the performance of Large Language Models (LLMs) and agentic systems in predicting cellular phenotypic responses. Addressing the lack of standard benchmarks for in silico phenotypic screening, AssayBench utilizes data from 1,920 publicly available CRISPR screens across five cellular phenotype classes. The study formulates screen prediction as a gene rank prediction task and introduces the adjusted nDCG metric to compare performance across heterogeneous assays. Extensive evaluations reveal that zero-shot generalist LLMs significantly outperform biology-specific LLMs and trainable baselines, although existing methods remain far from estimated performance ceilings. The findings suggest that optimization techniques like fine-tuning, ensembling, and prompt engineering can further enhance model accuracy. This work provides a critical testbed for advancing virtual cell models, aiming to accelerate biological discovery and improve drug discovery workflows by better aligning computational predictions with real-world phenotypic endpoints.
Wire timeline
AssayBench: A New Benchmark for Virtual Cell Modeling with LLMs
Researchers have introduced AssayBench, a novel benchmark designed to evaluate the performance of Large Language Models (LLMs) and agentic systems in predicting cellular phenotypic responses. Addressing the lack of standard benchmarks for in silico phenotypic screening, AssayBench utilizes data from 1,920 publicly available CRISPR screens across five cellular phenotype classes. The study formulates screen prediction as a gene rank prediction task and introduces the adjusted nDCG metric to compare performance across heterogeneous assays. Extensive evaluations reveal that zero-shot generalist LLMs significantly outperform biology-specific LLMs and trainable baselines, although existing methods remain far from estimated performance ceilings. The findings suggest that optimization techniques like fine-tuning, ensembling, and prompt engineering can further enhance model accuracy. This work provides a critical testbed for advancing virtual cell models, aiming to accelerate biological discovery and improve drug discovery workflows by better aligning computational predictions with real-world phenotypic endpoints.
cs.AI updates on arXiv.org