Users as Annotators: LLM Preference Learning from Comparison Mode
This research paper introduces a novel approach for collecting pairwise preference data to align large language models (LLMs) by leveraging user annotations from comparison modes. Traditionally, professional human annotators label which of two model responses is better, but this method is costly and limited in scale. The authors propose utilizing daily interactions where users naturally compare responses, capitalizing on the fact that users are experts on their own queries. However, this approach faces challenges regarding data quality control. To address this, the study generates responses from different models or versions to create asymmetry, enabling the inference of user data quality through a proposed user behavior model. An expectation-maximization algorithm is developed to estimate a latent quality factor for each user, allowing for effective filtering of annotation data. Downstream tasks demonstrate that this method successfully captures user behavior and improves data filtering for LLM alignment, offering a scalable alternative to professional annotation while maintaining high-quality training data for artificial intelligence systems.
Wire timeline
Users as Annotators: LLM Preference Learning from Comparison Mode
This research paper introduces a novel approach for collecting pairwise preference data to align large language models (LLMs) by leveraging user annotations from comparison modes. Traditionally, professional human annotators label which of two model responses is better, but this method is costly and limited in scale. The authors propose utilizing daily interactions where users naturally compare responses, capitalizing on the fact that users are experts on their own queries. However, this approach faces challenges regarding data quality control. To address this, the study generates responses from different models or versions to create asymmetry, enabling the inference of user data quality through a proposed user behavior model. An expectation-maximization algorithm is developed to estimate a latent quality factor for each user, allowing for effective filtering of annotation data. Downstream tasks demonstrate that this method successfully captures user behavior and improves data filtering for LLM alignment, offering a scalable alternative to professional annotation while maintaining high-quality training data for artificial intelligence systems.
cs.AI updates on arXiv.org