Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
Researchers from arXiv have published a study addressing the privacy threat of unauthorized personal data use in AI model training. The paper investigates Unlearnable Examples (UEs), which embed imperceptible perturbations into data to obstruct feature learning. While previous studies focused on training models from scratch, this work systematically analyzes UEs under the widely adopted pretraining-finetuning paradigm. The authors reveal that existing UE methods lose effectiveness when pretrained weights are loaded and frozen, as shallow layers preserve semantics and filter out noise. To address this, they propose Shallow Semantic Camouflage (SSC), a hierarchical deception strategy that confines perturbation generation to a semantically valid subspace. Extensive experiments demonstrate that SSC successfully maintains data unlearnability even under challenging conditions like shallow-layer freezing and semantic-focused pretraining. This research bridges a critical gap in privacy-preserving machine learning, offering a robust solution for protecting data in modern deep learning workflows that rely on pretrained models.
Wire timeline
Channel-Level Semantic Perturbations: Unlearnable Examples for Diverse Training Paradigms
Researchers from arXiv have published a study addressing the privacy threat of unauthorized personal data use in AI model training. The paper investigates Unlearnable Examples (UEs), which embed imperceptible perturbations into data to obstruct feature learning. While previous studies focused on training models from scratch, this work systematically analyzes UEs under the widely adopted pretraining-finetuning paradigm. The authors reveal that existing UE methods lose effectiveness when pretrained weights are loaded and frozen, as shallow layers preserve semantics and filter out noise. To address this, they propose Shallow Semantic Camouflage (SSC), a hierarchical deception strategy that confines perturbation generation to a semantically valid subspace. Extensive experiments demonstrate that SSC successfully maintains data unlearnability even under challenging conditions like shallow-layer freezing and semantic-focused pretraining. This research bridges a critical gap in privacy-preserving machine learning, offering a robust solution for protecting data in modern deep learning workflows that rely on pretrained models.
cs.AI updates on arXiv.org