Select-then-differentiate: Solving Bilevel Optimization with Manifold Lower-level Solution Sets
Researchers from arXiv have introduced a novel approach to optimistic bilevel optimization, addressing scenarios where the lower-level problem features a non-isolated manifold of minimizers. Traditionally, such settings cause non-differentiability in the hyper-objective due to multiple potential solutions. The study demonstrates that under a local Polyak-Łojasiewicz condition, differentiability is preserved if the optimistic selection is unique, even without a singleton solution set. This insight leads to an explicit pseudoinverse-based hyper-gradient formula. The authors propose HG-MS, a select-then-differentiate method that efficiently computes hyper-gradients and converges to a stationary point with complexity tied to the manifold's intrinsic dimension. Empirical tests on large language model (LLM) source reweighting show that a practical variant of HG-MS achieves superior performance on GSM8K and MATH benchmarks, alongside competitive results on MT-Bench instruction-following tasks. This work bridges theoretical optimization advances with practical AI applications, offering improved methods for handling complex optimization landscapes in machine learning models.
Wire timeline
Select-then-differentiate: Solving Bilevel Optimization with Manifold Lower-level Solution Sets
Researchers from arXiv have introduced a novel approach to optimistic bilevel optimization, addressing scenarios where the lower-level problem features a non-isolated manifold of minimizers. Traditionally, such settings cause non-differentiability in the hyper-objective due to multiple potential solutions. The study demonstrates that under a local Polyak-Łojasiewicz condition, differentiability is preserved if the optimistic selection is unique, even without a singleton solution set. This insight leads to an explicit pseudoinverse-based hyper-gradient formula. The authors propose HG-MS, a select-then-differentiate method that efficiently computes hyper-gradients and converges to a stationary point with complexity tied to the manifold's intrinsic dimension. Empirical tests on large language model (LLM) source reweighting show that a practical variant of HG-MS achieves superior performance on GSM8K and MATH benchmarks, alongside competitive results on MT-Bench instruction-following tasks. This work bridges theoretical optimization advances with practical AI applications, offering improved methods for handling complex optimization landscapes in machine learning models.
cs.AI updates on arXiv.org