OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
Researchers have introduced OPT-BENCH, a new benchmark designed to evaluate the self-improvement capabilities of Large Language Model (LLM) agents within large-scale search spaces. Addressing the gap in understanding whether LLMs can adapt to dynamic environmental feedback through intrinsic cognitive faculties like perception, reasoning, and memory, the study combines 20 machine learning tasks with 10 classic NP-hard problems. The authors also propose OPT-Agent, a framework emulating human-like cognitive adaptation via an iterative loop of perception, memory, and reasoning. Extensive experiments conducted on 19 LLMs from seven model families, ranging from 3B to 235B parameters, reveal that while stronger models leverage feedback more effectively for self-improvement, their adaptability is fundamentally constrained by base capacity. Crucially, even the most advanced LLMs currently fall short of human expert performance in these complex optimization tasks, highlighting significant limitations in current AI cognitive adaptation.
Wire timeline
OPT-BENCH: Evaluating the Iterative Self-Optimization of LLM Agents in Large-Scale Search Spaces
Researchers have introduced OPT-BENCH, a new benchmark designed to evaluate the self-improvement capabilities of Large Language Model (LLM) agents within large-scale search spaces. Addressing the gap in understanding whether LLMs can adapt to dynamic environmental feedback through intrinsic cognitive faculties like perception, reasoning, and memory, the study combines 20 machine learning tasks with 10 classic NP-hard problems. The authors also propose OPT-Agent, a framework emulating human-like cognitive adaptation via an iterative loop of perception, memory, and reasoning. Extensive experiments conducted on 19 LLMs from seven model families, ranging from 3B to 235B parameters, reveal that while stronger models leverage feedback more effectively for self-improvement, their adaptability is fundamentally constrained by base capacity. Crucially, even the most advanced LLMs currently fall short of human expert performance in these complex optimization tasks, highlighting significant limitations in current AI cognitive adaptation.
cs.AI updates on arXiv.org