DataDignity: Training Data Attribution for Large Language Models
Researchers have introduced DataDignity, a new framework for training data attribution in large language models (LLMs), aimed at identifying the specific source documents supporting model outputs. The study presents FakeWiki, a controlled benchmark comprising 3,537 fabricated Wikipedia-style articles designed to test pinpoint provenance while minimizing lexical shortcuts. This benchmark includes various query conditions, such as clean prompts and jailbreak-inspired transformations. The team evaluated several methods, including seven retrieval baselines, a training-free activation-steering method called SteerFuse, and a supervised contrastive provenance ranker named ScoringModel. Results across nine open-weight instruction-tuned LLMs showed that ScoringModel significantly outperformed existing baselines, improving mean Recall@10 from 35.0 to 52.2. Notably, it demonstrated robustness against adversarial queries, achieving a 15.7-point average improvement over the best baseline in jailbreak scenarios. The findings highlight that effective data attribution requires evaluation settings that distinguish true answer support from mere topical or lexical resemblance, offering a pathway for more transparent and auditable AI systems.
Wire timeline
DataDignity: Training Data Attribution for Large Language Models
Researchers have introduced DataDignity, a new framework for training data attribution in large language models (LLMs), aimed at identifying the specific source documents supporting model outputs. The study presents FakeWiki, a controlled benchmark comprising 3,537 fabricated Wikipedia-style articles designed to test pinpoint provenance while minimizing lexical shortcuts. This benchmark includes various query conditions, such as clean prompts and jailbreak-inspired transformations. The team evaluated several methods, including seven retrieval baselines, a training-free activation-steering method called SteerFuse, and a supervised contrastive provenance ranker named ScoringModel. Results across nine open-weight instruction-tuned LLMs showed that ScoringModel significantly outperformed existing baselines, improving mean Recall@10 from 35.0 to 52.2. Notably, it demonstrated robustness against adversarial queries, achieving a 15.7-point average improvement over the best baseline in jailbreak scenarios. The findings highlight that effective data attribution requires evaluation settings that distinguish true answer support from mere topical or lexical resemblance, offering a pathway for more transparent and auditable AI systems.
cs.AI updates on arXiv.org