C-BPO: A Preference-Corrected Framework for Personalizing LLMs with Binary Feedback
Researchers have introduced C-BPO, a novel framework designed to enhance the personalization of Large Language Models (LLMs) by effectively utilizing binary feedback signals. Unlike existing methods that often isolate user histories and neglect inter-user differences, C-BPO treats target user data as positive feedback while leveraging other users' data as implicit negative signals. This approach addresses the critical issue of preference overlap, where shared task knowledge might be erroneously penalized during training. By grounding its objective in Positive-Unlabeled (PU) learning theory, the framework purifies negative signals by subtracting positive bias. This ensures that the model aligns with unique user idiosyncrasies without compromising general helpfulness or common knowledge. Empirical experiments conducted across various personalization tasks and different backbone LLMs demonstrate that C-BPO consistently outperforms baseline methods. The study highlights the efficacy of preference-calibrated binary signals in modeling distinct inter-user differences, offering a significant advancement in aligning AI behaviors with individual user preferences while maintaining overall model utility.
Wire timeline
C-BPO: A Preference-Corrected Framework for Personalizing LLMs with Binary Feedback
Researchers have introduced C-BPO, a novel framework designed to enhance the personalization of Large Language Models (LLMs) by effectively utilizing binary feedback signals. Unlike existing methods that often isolate user histories and neglect inter-user differences, C-BPO treats target user data as positive feedback while leveraging other users' data as implicit negative signals. This approach addresses the critical issue of preference overlap, where shared task knowledge might be erroneously penalized during training. By grounding its objective in Positive-Unlabeled (PU) learning theory, the framework purifies negative signals by subtracting positive bias. This ensures that the model aligns with unique user idiosyncrasies without compromising general helpfulness or common knowledge. Empirical experiments conducted across various personalization tasks and different backbone LLMs demonstrate that C-BPO consistently outperforms baseline methods. The study highlights the efficacy of preference-calibrated binary signals in modeling distinct inter-user differences, offering a significant advancement in aligning AI behaviors with individual user preferences while maintaining overall model utility.
cs.AI updates on arXiv.org