Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
Researchers from the artificial intelligence community have introduced a novel approach to fixed-confidence best arm identification within generalized linear bandits. The study proposes a hybrid feedback model where learners can query either absolute reward feedback from a single arm or relative dueling feedback from an arm pair, both governed by generalized linear models. To address this, the authors developed a likelihood-ratio-based confidence sequence that unifies heterogeneous observations, yielding an explicit ellipsoidal confidence set under self-concordance assumptions. Building on this foundation, they proposed a hybrid Track-and-Stop algorithm that adaptively allocates queries by tracking a minimax-optimal design over a joint action space. The framework establishes delta-correctness and provides high-probability upper bounds on stopping time. Additionally, the research extends to a cost-aware setting, accounting for heterogeneous acquisition costs across different feedback modalities. Empirical experiments demonstrate that the proposed algorithms significantly improve sample efficiency compared to existing baseline methods. This work contributes to advancing decision-making processes in machine learning environments where diverse feedback types are available.
Wire timeline
Best Arm Identification in Generalized Linear Bandits via Hybrid Feedback
Researchers from the artificial intelligence community have introduced a novel approach to fixed-confidence best arm identification within generalized linear bandits. The study proposes a hybrid feedback model where learners can query either absolute reward feedback from a single arm or relative dueling feedback from an arm pair, both governed by generalized linear models. To address this, the authors developed a likelihood-ratio-based confidence sequence that unifies heterogeneous observations, yielding an explicit ellipsoidal confidence set under self-concordance assumptions. Building on this foundation, they proposed a hybrid Track-and-Stop algorithm that adaptively allocates queries by tracking a minimax-optimal design over a joint action space. The framework establishes delta-correctness and provides high-probability upper bounds on stopping time. Additionally, the research extends to a cost-aware setting, accounting for heterogeneous acquisition costs across different feedback modalities. Empirical experiments demonstrate that the proposed algorithms significantly improve sample efficiency compared to existing baseline methods. This work contributes to advancing decision-making processes in machine learning environments where diverse feedback types are available.
cs.AI updates on arXiv.org