Learning the Preferences of a Learning Agent
This academic paper, submitted to arXiv in May 2026 by researchers including Stuart Russell, addresses a critical limitation in Inverse Reinforcement Learning (IRL). Traditional IRL methods infer human preferences from observed behavior but assume the human acts optimally. This assumption fails when humans are themselves learning to act optimally in an environment. The authors formalize the problem of learning the preferences of a learning agent, where a predictor observes a learner acting online to infer the underlying reward function being optimized, even if initially suboptimally. The study models the learner as either having no-regret properties or converging to an optimal Boltzmann policy over time. For each setting, the paper establishes theoretical guarantees for various preference learning algorithms or demonstrates that such guarantees are impossible. This research contributes to the field of AI alignment by improving how AI systems can understand and act in accordance with human values, particularly in dynamic scenarios where human behavior evolves through learning rather than remaining static and optimal.
Wire timeline
Learning the Preferences of a Learning Agent
This academic paper, submitted to arXiv in May 2026 by researchers including Stuart Russell, addresses a critical limitation in Inverse Reinforcement Learning (IRL). Traditional IRL methods infer human preferences from observed behavior but assume the human acts optimally. This assumption fails when humans are themselves learning to act optimally in an environment. The authors formalize the problem of learning the preferences of a learning agent, where a predictor observes a learner acting online to infer the underlying reward function being optimized, even if initially suboptimally. The study models the learner as either having no-regret properties or converging to an optimal Boltzmann policy over time. For each setting, the paper establishes theoretical guarantees for various preference learning algorithms or demonstrates that such guarantees are impossible. This research contributes to the field of AI alignment by improving how AI systems can understand and act in accordance with human values, particularly in dynamic scenarios where human behavior evolves through learning rather than remaining static and optimal.
cs.AI updates on arXiv.org