NASH Framework Enhances Data Shapley Effectiveness in Machine Learning Data Selection
A new research paper titled "Is Data Shapley Not Better than Random in Data Selection? Ask NASH" addresses the inconsistent performance of Data Shapley values in identifying high-quality training data subsets. While Data Shapley accounts for data interactions, critics argue it sometimes performs no better than random selection. The authors propose a novel framework called NASH (Non-linear Aggregation of SHapley-informative components) to resolve this issue. NASH decomposes the target utility function, such as validation accuracy, into simpler, Shapley-informative component functions. It then selects data by optimizing an objective that aggregates these components non-linearly. The study demonstrates that this approach substantially improves the effectiveness of Shapley and semivalue-based data selection methods while incurring minimal additional runtime costs. This work provides a strategic method for consistently selecting high-quality data subsets, addressing key limitations in current machine learning data curation practices. The paper was submitted to arXiv under Computer Science > Machine Learning by researchers including Xiao Tian and Bryan Kian Hsiang Low.
Wire timeline
NASH Framework Enhances Data Shapley Effectiveness in Machine Learning Data Selection
A new research paper titled "Is Data Shapley Not Better than Random in Data Selection? Ask NASH" addresses the inconsistent performance of Data Shapley values in identifying high-quality training data subsets. While Data Shapley accounts for data interactions, critics argue it sometimes performs no better than random selection. The authors propose a novel framework called NASH (Non-linear Aggregation of SHapley-informative components) to resolve this issue. NASH decomposes the target utility function, such as validation accuracy, into simpler, Shapley-informative component functions. It then selects data by optimizing an objective that aggregates these components non-linearly. The study demonstrates that this approach substantially improves the effectiveness of Shapley and semivalue-based data selection methods while incurring minimal additional runtime costs. This work provides a strategic method for consistently selecting high-quality data subsets, addressing key limitations in current machine learning data curation practices. The paper was submitted to arXiv under Computer Science > Machine Learning by researchers including Xiao Tian and Bryan Kian Hsiang Low.
cs.AI updates on arXiv.org