The Two Clocks and the Innovation Window: When and How Generative Models Learn Rules
A new research paper published on arXiv investigates the fundamental tension in generative models trained on finite data, where objectives often converge to empirical training distributions rather than the desired population distribution. The study introduces a framework based on two distinct training timescales: tau_rule, marking when generations first become rule-valid, and tau_mem, indicating when models begin reproducing specific training samples. By analyzing synthetic tasks like parity and combinatorial puzzles, the authors demonstrate that tau_rule increases with rule complexity but decreases with model capacity, whereas tau_mem scales linearly with dataset size and remains largely invariant to the rule. The interval between these two points is defined as the 'innovation window,' a critical period for genuine creativity that widens with larger datasets but narrows with complex rules. This two-clock structure is observed in both diffusion (DiT) and autoregressive (GPT) architectures. The findings provide a unified, predictive account of how optimization landscapes evolve, offering insights into preventing memorization and fostering true innovation in artificial intelligence systems.
Wire timeline
The Two Clocks and the Innovation Window: When and How Generative Models Learn Rules
A new research paper published on arXiv investigates the fundamental tension in generative models trained on finite data, where objectives often converge to empirical training distributions rather than the desired population distribution. The study introduces a framework based on two distinct training timescales: tau_rule, marking when generations first become rule-valid, and tau_mem, indicating when models begin reproducing specific training samples. By analyzing synthetic tasks like parity and combinatorial puzzles, the authors demonstrate that tau_rule increases with rule complexity but decreases with model capacity, whereas tau_mem scales linearly with dataset size and remains largely invariant to the rule. The interval between these two points is defined as the 'innovation window,' a critical period for genuine creativity that widens with larger datasets but narrows with complex rules. This two-clock structure is observed in both diffusion (DiT) and autoregressive (GPT) architectures. The findings provide a unified, predictive account of how optimization landscapes evolve, offering insights into preventing memorization and fostering true innovation in artificial intelligence systems.
cs.AI updates on arXiv.org