Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
Researchers have proposed a novel defense mechanism to mitigate watermark forgery attacks in generative AI models. Watermarking allows providers to verify if content was generated by their models, but adversaries can forge these watermarks to damage reputations. Existing defenses often degrade model utility by embedding multiple watermarks. This new scheme randomizes watermark key selection for each query and accepts content as genuine only if exactly one key detects a watermark. The method is provably forgery-resistant regardless of the number of samples collected by attackers, provided they cannot distinguish watermarks from different keys. It treats the underlying watermarking method as a black box, making it modality-agnostic for images and text. Empirical results show a reduction in attacker success rates from near-perfect levels to just 2%, with negligible computational overhead. This approach maintains model utility while significantly enhancing security against forgery, offering a robust solution for GenAI providers seeking to protect their intellectual property and maintain trust.
Wire timeline
Mitigating Watermark Forgery in Generative Models via Randomized Key Selection
Researchers have proposed a novel defense mechanism to mitigate watermark forgery attacks in generative AI models. Watermarking allows providers to verify if content was generated by their models, but adversaries can forge these watermarks to damage reputations. Existing defenses often degrade model utility by embedding multiple watermarks. This new scheme randomizes watermark key selection for each query and accepts content as genuine only if exactly one key detects a watermark. The method is provably forgery-resistant regardless of the number of samples collected by attackers, provided they cannot distinguish watermarks from different keys. It treats the underlying watermarking method as a black box, making it modality-agnostic for images and text. Empirical results show a reduction in attacker success rates from near-perfect levels to just 2%, with negligible computational overhead. This approach maintains model utility while significantly enhancing security against forgery, offering a robust solution for GenAI providers seeking to protect their intellectual property and maintain trust.
cs.AI updates on arXiv.org