Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
Researchers Jiazheng Sun, Weixin Wang, and Pan Xu have released a revised academic paper on arXiv presenting a unified algorithmic framework for ensemble sampling in nonlinear contextual bandits. The study introduces two specific methods: Generalized Linear Ensemble Sampling (GLM-ES) and Neural Ensemble Sampling (Neural-ES). These algorithms maintain multiple estimators for reward model parameters using maximum likelihood estimation on randomly perturbed data. The authors establish high-probability frequentist regret bounds, with GLM-ES matching state-of-the-art results for generalized linear bandits. A key practical innovation is the development of anytime versions of these algorithms, which remove the fixed-time horizon assumption, making them suitable for scenarios where the total number of rounds is unknown. Empirical evaluations demonstrate strong performance for both methods and their anytime variants. This work theoretically validates ensemble sampling as a robust and practical randomized exploration approach for complex nonlinear models, addressing specific challenges in theoretical analysis and expanding the applicability of contextual bandit algorithms in machine learning.
Wire timeline
Provable Anytime Ensemble Sampling Algorithms in Nonlinear Contextual Bandits
Researchers Jiazheng Sun, Weixin Wang, and Pan Xu have released a revised academic paper on arXiv presenting a unified algorithmic framework for ensemble sampling in nonlinear contextual bandits. The study introduces two specific methods: Generalized Linear Ensemble Sampling (GLM-ES) and Neural Ensemble Sampling (Neural-ES). These algorithms maintain multiple estimators for reward model parameters using maximum likelihood estimation on randomly perturbed data. The authors establish high-probability frequentist regret bounds, with GLM-ES matching state-of-the-art results for generalized linear bandits. A key practical innovation is the development of anytime versions of these algorithms, which remove the fixed-time horizon assumption, making them suitable for scenarios where the total number of rounds is unknown. Empirical evaluations demonstrate strong performance for both methods and their anytime variants. This work theoretically validates ensemble sampling as a robust and practical randomized exploration approach for complex nonlinear models, addressing specific challenges in theoretical analysis and expanding the applicability of contextual bandit algorithms in machine learning.
cs.AI updates on arXiv.org