AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design
Researchers have introduced AHD Agent, a novel framework for Automatic Heuristic Design (AHD) aimed at solving NP-hard combinatorial optimization problems. While existing Large Language Model (LLM) approaches often treat models as passive generators within fixed workflows, leading to inefficient exploration, AHD Agent empowers LLMs to proactively decide between generating heuristics or invoking tools to retrieve evidence from the solving environment. This tool-integrated, multi-turn system utilizes an agentic reinforcement learning mechanism and a new environment synthesis pipeline to optimize decision-making. Experimental results across eight diverse domains, including four held-out tasks, demonstrate that the compact 4B-parameter agent matches or surpasses state-of-the-art baselines that rely on significantly larger models. Furthermore, the new approach requires far fewer evaluations, highlighting its efficiency. The study suggests that AHD Agent provides a scalable and effective trajectory toward truly autonomous heuristic design, overcoming limitations of previous context-dependent methods.
Wire timeline
AHD Agent: Agentic Reinforcement Learning for Automatic Heuristic Design
Researchers have introduced AHD Agent, a novel framework for Automatic Heuristic Design (AHD) aimed at solving NP-hard combinatorial optimization problems. While existing Large Language Model (LLM) approaches often treat models as passive generators within fixed workflows, leading to inefficient exploration, AHD Agent empowers LLMs to proactively decide between generating heuristics or invoking tools to retrieve evidence from the solving environment. This tool-integrated, multi-turn system utilizes an agentic reinforcement learning mechanism and a new environment synthesis pipeline to optimize decision-making. Experimental results across eight diverse domains, including four held-out tasks, demonstrate that the compact 4B-parameter agent matches or surpasses state-of-the-art baselines that rely on significantly larger models. Furthermore, the new approach requires far fewer evaluations, highlighting its efficiency. The study suggests that AHD Agent provides a scalable and effective trajectory toward truly autonomous heuristic design, overcoming limitations of previous context-dependent methods.
cs.AI updates on arXiv.org