Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs
Researchers have introduced FedMITR, a novel framework designed to enhance One-Shot Federated Learning using Vision Transformers (ViTs). This approach addresses significant challenges in extremely non-IID settings, where existing data-free methods often produce low-quality synthetic data with semantic misalignment. FedMITR employs sparse model inversion to selectively generate semantic foregrounds while ignoring uninformative backgrounds, thereby reducing gradient instability. Additionally, it implements a token relabeling strategy that distinguishes between high and low information density patches, using pseudo-labels and ensemble models respectively to improve robustness. Theoretical analysis based on algorithmic stability demonstrates that these techniques eliminate background noise issues and reduce gradient variance, leading to tighter generalization bounds. Empirical results confirm that FedMITR substantially outperforms current baselines across various experimental settings, offering a more efficient and accurate method for training global models in a single communication round without sharing raw data.
Wire timeline
Provable Sparse Inversion and Token Relabel Enhanced One-shot Federated Learning with ViTs
Researchers have introduced FedMITR, a novel framework designed to enhance One-Shot Federated Learning using Vision Transformers (ViTs). This approach addresses significant challenges in extremely non-IID settings, where existing data-free methods often produce low-quality synthetic data with semantic misalignment. FedMITR employs sparse model inversion to selectively generate semantic foregrounds while ignoring uninformative backgrounds, thereby reducing gradient instability. Additionally, it implements a token relabeling strategy that distinguishes between high and low information density patches, using pseudo-labels and ensemble models respectively to improve robustness. Theoretical analysis based on algorithmic stability demonstrates that these techniques eliminate background noise issues and reduce gradient variance, leading to tighter generalization bounds. Empirical results confirm that FedMITR substantially outperforms current baselines across various experimental settings, offering a more efficient and accurate method for training global models in a single communication round without sharing raw data.
cs.AI updates on arXiv.org