New AI Research Proposes Preference-Based Embeddings Over Semantic Similarity
A new research paper titled "Embeddings for Preferences, Not Semantics," submitted to arXiv by Carter Blair, Ariel D. Procaccia, and Milind Tambe, addresses a critical limitation in using standard text embeddings for collective decision-making. While modern AI enables participants to express views via free-form text, existing embedding models measure semantic similarity rather than preferential similarity, which is required for applications like facility location problems and fair clustering. The authors argue that off-the-shelf embeddings fail when the correlation between semantic content and user preference breaks down, as they conflate preference-relevant signals with semantic nuisances like style and wording. To resolve this, the study formalizes the issue as an invariance problem and introduces synthetic training data designed to decouple these factors. This approach shifts the optimal scorer away from nuisance-dominated cosine similarity. Experimental results across eleven online deliberation datasets demonstrate that this method significantly improves preference prediction accuracy, offering a more robust framework for AI-driven democratic processes and opinion analysis.
Wire timeline
New AI Research Proposes Preference-Based Embeddings Over Semantic Similarity
A new research paper titled "Embeddings for Preferences, Not Semantics," submitted to arXiv by Carter Blair, Ariel D. Procaccia, and Milind Tambe, addresses a critical limitation in using standard text embeddings for collective decision-making. While modern AI enables participants to express views via free-form text, existing embedding models measure semantic similarity rather than preferential similarity, which is required for applications like facility location problems and fair clustering. The authors argue that off-the-shelf embeddings fail when the correlation between semantic content and user preference breaks down, as they conflate preference-relevant signals with semantic nuisances like style and wording. To resolve this, the study formalizes the issue as an invariance problem and introduces synthetic training data designed to decouple these factors. This approach shifts the optimal scorer away from nuisance-dominated cosine similarity. Experimental results across eleven online deliberation datasets demonstrate that this method significantly improves preference prediction accuracy, offering a more robust framework for AI-driven democratic processes and opinion analysis.
cs.AI updates on arXiv.org