Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
A new research paper published on arXiv introduces Disrupt-and-Rectify Smoothing (DR-Smoothing), a novel defense method designed to protect Large Language Models (LLMs) from jailbreaking attacks. Inspired by denoised-smoothing techniques in adversarial defense, this approach integrates a two-stage prompt processing scheme that first disrupts input prompts and then rectifies them. This process restores out-of-distribution disrupted prompts to an in-distribution form, significantly reducing the risk of unpredictable LLM behavior compared to previous disrupt-only methods. The study provides a theoretical analysis offering tight bounds for defense success probability and disruption strength requirements. Extensive experiments demonstrate that DR-Smoothing effectively defends against both token-level and prompt-level jailbreaking attacks under various scenarios. Furthermore, the method surpasses current state-of-the-art defenses by achieving a superior balance between harmlessness and helpfulness, ensuring robust security without compromising the model's utility.
Wire timeline
Guaranteed Jailbreaking Defense via Disrupt-and-Rectify Smoothing
A new research paper published on arXiv introduces Disrupt-and-Rectify Smoothing (DR-Smoothing), a novel defense method designed to protect Large Language Models (LLMs) from jailbreaking attacks. Inspired by denoised-smoothing techniques in adversarial defense, this approach integrates a two-stage prompt processing scheme that first disrupts input prompts and then rectifies them. This process restores out-of-distribution disrupted prompts to an in-distribution form, significantly reducing the risk of unpredictable LLM behavior compared to previous disrupt-only methods. The study provides a theoretical analysis offering tight bounds for defense success probability and disruption strength requirements. Extensive experiments demonstrate that DR-Smoothing effectively defends against both token-level and prompt-level jailbreaking attacks under various scenarios. Furthermore, the method surpasses current state-of-the-art defenses by achieving a superior balance between harmlessness and helpfulness, ensuring robust security without compromising the model's utility.
cs.AI updates on arXiv.org