Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
This academic study investigates the empirical observability of Tian's (2025) repulsion theorem during the grokking phase of two-layer neural networks. While the theorem predicts that similar features exert repulsive forces via negative off-diagonal entries in matrix B, it lacks specifics on timing and spectral signatures. Using a modular addition setup, the authors observe a dissociation between feature structure and learning mechanisms. The predicted sign rule holds robustly across different activation functions. However, spectral signatures in parameter updates are strongly activation-dependent. With squared activation, a slope detector identifies grokking at epoch 174 with a rank-2 spectrum, whereas ReLU activation results in a rank-1 spectrum where the detector never fires. This aligns with distinctions between focused and spreading memorization, indicating that while feature repulsion structure is consistent, its translation into weight updates depends critically on the activation derivative. The findings provide crucial insights into the mechanics of grokking and the role of activation functions in neural network learning dynamics.
Wire timeline
Feature Repulsion and Spectral Lock-in: An Empirical Study of Two-Layer Network Grokking
This academic study investigates the empirical observability of Tian's (2025) repulsion theorem during the grokking phase of two-layer neural networks. While the theorem predicts that similar features exert repulsive forces via negative off-diagonal entries in matrix B, it lacks specifics on timing and spectral signatures. Using a modular addition setup, the authors observe a dissociation between feature structure and learning mechanisms. The predicted sign rule holds robustly across different activation functions. However, spectral signatures in parameter updates are strongly activation-dependent. With squared activation, a slope detector identifies grokking at epoch 174 with a rank-2 spectrum, whereas ReLU activation results in a rank-1 spectrum where the detector never fires. This aligns with distinctions between focused and spreading memorization, indicating that while feature repulsion structure is consistent, its translation into weight updates depends critically on the activation derivative. The findings provide crucial insights into the mechanics of grokking and the role of activation functions in neural network learning dynamics.
cs.AI updates on arXiv.org