Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
Researchers have introduced TAP (Tabular Augmentation Policy), a novel method designed to enhance machine learning model performance in data-scarce environments. The study addresses the 'fidelity-utility gap,' noting that existing generative augmentation techniques often prioritize distributional plausibility over actual utility for downstream tasks. TAP couples diffusion inpainting with a lightweight, learner-conditioned policy to strategically generate high-utility samples and control their injection during training through explicit gating. This approach ensures that added data effectively reduces the learner's held-out evaluation loss. Tested across seven real-world datasets under severe data scarcity conditions, TAP consistently outperformed strong generative baselines. The results demonstrated significant improvements, including up to a 15.6 percentage point increase in classification accuracy and a 32% reduction in regression RMSE. This research highlights the importance of aligning generative objectives with specific learner needs rather than focusing solely on statistical fidelity, offering a robust solution for improving model accuracy when training data is limited.
Wire timeline
Active Tabular Augmentation via Policy-Guided Diffusion Inpainting
Researchers have introduced TAP (Tabular Augmentation Policy), a novel method designed to enhance machine learning model performance in data-scarce environments. The study addresses the 'fidelity-utility gap,' noting that existing generative augmentation techniques often prioritize distributional plausibility over actual utility for downstream tasks. TAP couples diffusion inpainting with a lightweight, learner-conditioned policy to strategically generate high-utility samples and control their injection during training through explicit gating. This approach ensures that added data effectively reduces the learner's held-out evaluation loss. Tested across seven real-world datasets under severe data scarcity conditions, TAP consistently outperformed strong generative baselines. The results demonstrated significant improvements, including up to a 15.6 percentage point increase in classification accuracy and a 32% reduction in regression RMSE. This research highlights the importance of aligning generative objectives with specific learner needs rather than focusing solely on statistical fidelity, offering a robust solution for improving model accuracy when training data is limited.
cs.AI updates on arXiv.org