Architecture-agnostic Lipschitz-constant Bayesian Header for Vision Transformers
Researchers have introduced a novel architecture-agnostic Lipschitz-constant Bayesian header designed to address label noise in supervised deep learning, specifically targeting semantically proximal classification errors. Integrated with vision transformers, this method creates the bi-Lipschitz-constrained Bayesian Vision Transformer (LipB-ViT). Unlike conventional Bayesian layers, it enforces spectral normalization on both the mean and log-variance of variational weights to enhance calibrated predictive uncertainty and reduce noise amplification. The study proposes a new metric for jointly capturing uncertainty and confidence, alongside an adaptive arithmetic-mean fusion scheme that combines feature-space proximity with predictive uncertainty. This approach detects corrupted labels more effectively than state-of-the-art k-nearest neighbor methods, achieving over 7% higher performance and a recall exceeding 0.93 with 15% semantically misclassified labels. Despite increased computational costs from Monte Carlo sampling, the solution offers plug-and-play compatibility with pre-trained backbones. It demonstrates robustness against both structured adversarial and unstructured noise, providing a valuable tool for high-stakes applications where annotation reliability varies, while also offering a pipeline for assessing dataset quality.
Wire timeline
Architecture-agnostic Lipschitz-constant Bayesian Header for Vision Transformers
Researchers have introduced a novel architecture-agnostic Lipschitz-constant Bayesian header designed to address label noise in supervised deep learning, specifically targeting semantically proximal classification errors. Integrated with vision transformers, this method creates the bi-Lipschitz-constrained Bayesian Vision Transformer (LipB-ViT). Unlike conventional Bayesian layers, it enforces spectral normalization on both the mean and log-variance of variational weights to enhance calibrated predictive uncertainty and reduce noise amplification. The study proposes a new metric for jointly capturing uncertainty and confidence, alongside an adaptive arithmetic-mean fusion scheme that combines feature-space proximity with predictive uncertainty. This approach detects corrupted labels more effectively than state-of-the-art k-nearest neighbor methods, achieving over 7% higher performance and a recall exceeding 0.93 with 15% semantically misclassified labels. Despite increased computational costs from Monte Carlo sampling, the solution offers plug-and-play compatibility with pre-trained backbones. It demonstrates robustness against both structured adversarial and unstructured noise, providing a valuable tool for high-stakes applications where annotation reliability varies, while also offering a pipeline for assessing dataset quality.
cs.AI updates on arXiv.org