Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
Researchers have introduced Divergence-in-Behavior Auditing (DIBA), a novel framework designed to detect unauthorized data usage in Large Language Models (LLMs) trained with Reinforcement Learning with Verifiable Rewards (RLVR). As RLVR becomes a standard training stage relying on high-value, non-public prompt sets, concerns regarding data exposure have grown. Traditional membership inference attacks fail in this context because RLVR generates dynamic responses rather than fitting fixed strings. DIBA addresses this by comparing fine-tuned models against pre-RLVR checkpoints, analyzing both reward-side success changes and policy-side behavioral drift. By aggregating multiple stochastic rollouts, the method produces stable auditing signals. In white-box settings, DIBA significantly outperforms existing likelihood-based baselines, achieving an AUC of approximately 0.8 and substantially higher true positive rates at low false positive rates. The study also explores grey-box scenarios, noting that audit effectiveness varies by algorithm but remains robust across model sizes. This development provides a critical tool for verifying data integrity and preventing intellectual property theft in advanced AI training pipelines.
Wire timeline
Auditing Data Membership in Reinforcement Learning With Verifiable Rewards
Researchers have introduced Divergence-in-Behavior Auditing (DIBA), a novel framework designed to detect unauthorized data usage in Large Language Models (LLMs) trained with Reinforcement Learning with Verifiable Rewards (RLVR). As RLVR becomes a standard training stage relying on high-value, non-public prompt sets, concerns regarding data exposure have grown. Traditional membership inference attacks fail in this context because RLVR generates dynamic responses rather than fitting fixed strings. DIBA addresses this by comparing fine-tuned models against pre-RLVR checkpoints, analyzing both reward-side success changes and policy-side behavioral drift. By aggregating multiple stochastic rollouts, the method produces stable auditing signals. In white-box settings, DIBA significantly outperforms existing likelihood-based baselines, achieving an AUC of approximately 0.8 and substantially higher true positive rates at low false positive rates. The study also explores grey-box scenarios, noting that audit effectiveness varies by algorithm but remains robust across model sizes. This development provides a critical tool for verifying data integrity and preventing intellectual property theft in advanced AI training pipelines.
cs.AI updates on arXiv.org