PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
Researchers have introduced PromptGuard, a novel content moderation technique designed to address safety vulnerabilities in text-to-image (T2I) models. These models often risk generating not-safe-for-work (NSFW) content, including sexually explicit, violent, or politically sensitive images. Inspired by system prompt mechanisms in large language models, PromptGuard optimizes a universal safety soft prompt that acts as an implicit guideline within the T2I model's textual embedding space. This approach moderates unsafe inputs without requiring proxy models or compromising inference efficiency. The method employs a divide-and-conquer strategy, combining category-specific soft prompts into unified safety guidance. Extensive experiments across five datasets demonstrate that PromptGuard effectively mitigates NSFW generation while maintaining high-quality benign outputs. It reportedly outperforms eight state-of-the-art defenses and operates 3.8 times faster than prior moderation methods. Evaluations using multi-head safety classifiers and VLM-based guardrails confirm its robustness, showing low average unsafe ratios. The research aims to enhance ethical standards in AI image generation by providing a efficient, reliable safety layer.
Wire timeline
PromptGuard: Soft Prompt-Guided Unsafe Content Moderation for Text-to-Image Models
Researchers have introduced PromptGuard, a novel content moderation technique designed to address safety vulnerabilities in text-to-image (T2I) models. These models often risk generating not-safe-for-work (NSFW) content, including sexually explicit, violent, or politically sensitive images. Inspired by system prompt mechanisms in large language models, PromptGuard optimizes a universal safety soft prompt that acts as an implicit guideline within the T2I model's textual embedding space. This approach moderates unsafe inputs without requiring proxy models or compromising inference efficiency. The method employs a divide-and-conquer strategy, combining category-specific soft prompts into unified safety guidance. Extensive experiments across five datasets demonstrate that PromptGuard effectively mitigates NSFW generation while maintaining high-quality benign outputs. It reportedly outperforms eight state-of-the-art defenses and operates 3.8 times faster than prior moderation methods. Evaluations using multi-head safety classifiers and VLM-based guardrails confirm its robustness, showing low average unsafe ratios. The research aims to enhance ethical standards in AI image generation by providing a efficient, reliable safety layer.
cs.AI updates on arXiv.org