Teacher-Aware Evolution of Heuristic Programs from Learned Optimization Policies
Researchers have introduced a novel teacher-aware evolutionary framework designed to enhance the automatic generation of executable heuristics for combinatorial optimization problems. While Large Language Model (LLM)-based methods show promise, they often rely on delayed endpoint performance metrics. This new approach utilizes independently trained learned optimization policies as behavioral teachers. Instead of directly deploying or imitating these teachers, the system queries them on states visited by candidate heuristic programs, using their action preferences as local feedback for evolution. This method allows the search process to discover static executable heuristics guided by both task performance and teacher-derived behavioral signals. Experimental results across scheduling, routing, and graph optimization benchmarks demonstrate that this framework outperforms existing performance-driven LLM heuristic evolution baselines. Crucially, the resulting heuristics require no neural inference during deployment, offering efficiency advantages. The study suggests that learned optimization policies can be effectively repurposed as behavioral feedback sources, advancing the field of automatic heuristic discovery in artificial intelligence and computer science.
Wire timeline
Teacher-Aware Evolution of Heuristic Programs from Learned Optimization Policies
Researchers have introduced a novel teacher-aware evolutionary framework designed to enhance the automatic generation of executable heuristics for combinatorial optimization problems. While Large Language Model (LLM)-based methods show promise, they often rely on delayed endpoint performance metrics. This new approach utilizes independently trained learned optimization policies as behavioral teachers. Instead of directly deploying or imitating these teachers, the system queries them on states visited by candidate heuristic programs, using their action preferences as local feedback for evolution. This method allows the search process to discover static executable heuristics guided by both task performance and teacher-derived behavioral signals. Experimental results across scheduling, routing, and graph optimization benchmarks demonstrate that this framework outperforms existing performance-driven LLM heuristic evolution baselines. Crucially, the resulting heuristics require no neural inference during deployment, offering efficiency advantages. The study suggests that learned optimization policies can be effectively repurposed as behavioral feedback sources, advancing the field of automatic heuristic discovery in artificial intelligence and computer science.
cs.AI updates on arXiv.org