TRACE Framework Enhances Multi-Turn Jailbreaking via Turn-Aware Credit Assignment
Researchers have introduced TRACE, a novel turn-aware credit assignment framework designed to improve reinforcement learning-based multi-turn jailbreaking attacks on Large Language Models (LLMs). Current methods often suffer from coarse trajectory-level outcome signals that uniformly reward or penalize all dialogue turns, ignoring the non-uniform and phase-dependent nature of turn contributions. This leads to inefficient learning, such as over-rewarding redundant turns in successful attacks or under-crediting useful steps in failed ones. TRACE addresses this by estimating turn-level contributions through leave-one-turn-out semantic masking for successful trajectories and assigning penalties based on prompt harmfulness and semantic relevance for failed ones. Additionally, it incorporates a local refusal-aware penalty. The framework also repurposes attack-side credit signals for multi-turn defense alignment. Extensive experiments on both open-source and closed-source LLM targets demonstrate that TRACE achieves a 25% relative improvement in attack success rate compared to the strongest existing RL baselines. Furthermore, when applied to defense alignment, it significantly improves the balance between safety and utility, highlighting its dual potential for both offensive security testing and defensive model robustness.
Wire timeline
TRACE Framework Enhances Multi-Turn Jailbreaking via Turn-Aware Credit Assignment
Researchers have introduced TRACE, a novel turn-aware credit assignment framework designed to improve reinforcement learning-based multi-turn jailbreaking attacks on Large Language Models (LLMs). Current methods often suffer from coarse trajectory-level outcome signals that uniformly reward or penalize all dialogue turns, ignoring the non-uniform and phase-dependent nature of turn contributions. This leads to inefficient learning, such as over-rewarding redundant turns in successful attacks or under-crediting useful steps in failed ones. TRACE addresses this by estimating turn-level contributions through leave-one-turn-out semantic masking for successful trajectories and assigning penalties based on prompt harmfulness and semantic relevance for failed ones. Additionally, it incorporates a local refusal-aware penalty. The framework also repurposes attack-side credit signals for multi-turn defense alignment. Extensive experiments on both open-source and closed-source LLM targets demonstrate that TRACE achieves a 25% relative improvement in attack success rate compared to the strongest existing RL baselines. Furthermore, when applied to defense alignment, it significantly improves the balance between safety and utility, highlighting its dual potential for both offensive security testing and defensive model robustness.
cs.AI updates on arXiv.org