FragileFlow: Spectral Control for Foundation Model Robustness
Researchers have introduced FragileFlow, a novel plug-in regularizer designed to enhance the robustness of Large Language Models (LLMs) and Vision-Language Models (VLMs). Current evaluation metrics often overlook structured failure modes where predictions remain technically correct but are fragile, with probability mass shifting toward incorrect classes near decision boundaries. FragileFlow addresses this by formalizing margin-aware error flow and utilizing a calibrated margin buffer to identify these vulnerable predictions. It organizes off-class probability mass into a class-wise vulnerable-risk matrix. The study provides the first PAC-Bayes upper bound for this error-flow object, demonstrating that empirical spectral control offers a conservative path to deterministic worst-class robustness under stability conditions. Experimental results on multiple-choice LLM benchmarks and few-shot CLIP adaptation indicate that FragileFlow consistently improves risk measures compared to baselines, enhances perturbed worst-class accuracy, and maintains clean accuracy. This advancement offers a theoretical and practical framework for improving the reliability of foundation models against adversarial perturbations.
Wire timeline
FragileFlow: Spectral Control for Foundation Model Robustness
Researchers have introduced FragileFlow, a novel plug-in regularizer designed to enhance the robustness of Large Language Models (LLMs) and Vision-Language Models (VLMs). Current evaluation metrics often overlook structured failure modes where predictions remain technically correct but are fragile, with probability mass shifting toward incorrect classes near decision boundaries. FragileFlow addresses this by formalizing margin-aware error flow and utilizing a calibrated margin buffer to identify these vulnerable predictions. It organizes off-class probability mass into a class-wise vulnerable-risk matrix. The study provides the first PAC-Bayes upper bound for this error-flow object, demonstrating that empirical spectral control offers a conservative path to deterministic worst-class robustness under stability conditions. Experimental results on multiple-choice LLM benchmarks and few-shot CLIP adaptation indicate that FragileFlow consistently improves risk measures compared to baselines, enhances perturbed worst-class accuracy, and maintains clean accuracy. This advancement offers a theoretical and practical framework for improving the reliability of foundation models against adversarial perturbations.
cs.AI updates on arXiv.org