New Theoretical Framework and STOC Algorithm Enhance Continual Factual Knowledge in Language Models
Researchers have introduced a novel theoretical framework and algorithm to address the challenge of continual factual knowledge acquisition (cFKA) in language models. While Continual Pre-Training (CPT) is essential for integrating new information without erasing existing knowledge, the underlying mechanisms remain poorly understood. This study analyzes training dynamics using a single-layer Transformer, revealing that regularization-based methods only adjust convergence rates without preventing forgetting, whereas data replay methods effectively stabilize pretrained knowledge. Building on these insights, the team proposes Selecting Tokens via attention Contribution (STOC), a generative data replay approach that identifies influential factual snippets to guide replay data generation. Extensive experiments on synthetic and real-world datasets validate that STOC significantly mitigates catastrophic forgetting and enhances cFKA. This work provides a unified explanation for representative CPT methods and offers a practical solution for improving the long-term retention of facts in large language models, marking a significant advancement in artificial intelligence research regarding model stability and knowledge integration.
Wire timeline
New Theoretical Framework and STOC Algorithm Enhance Continual Factual Knowledge in Language Models
Researchers have introduced a novel theoretical framework and algorithm to address the challenge of continual factual knowledge acquisition (cFKA) in language models. While Continual Pre-Training (CPT) is essential for integrating new information without erasing existing knowledge, the underlying mechanisms remain poorly understood. This study analyzes training dynamics using a single-layer Transformer, revealing that regularization-based methods only adjust convergence rates without preventing forgetting, whereas data replay methods effectively stabilize pretrained knowledge. Building on these insights, the team proposes Selecting Tokens via attention Contribution (STOC), a generative data replay approach that identifies influential factual snippets to guide replay data generation. Extensive experiments on synthetic and real-world datasets validate that STOC significantly mitigates catastrophic forgetting and enhances cFKA. This work provides a unified explanation for representative CPT methods and offers a practical solution for improving the long-term retention of facts in large language models, marking a significant advancement in artificial intelligence research regarding model stability and knowledge integration.
cs.AI updates on arXiv.org