Attribution-Guided Pruning for Insight and Control in Small-scale LLMs
Researchers have introduced a novel method called attribution-guided pruning to enhance the interpretability and control of Large Language Models (LLMs). Addressing the challenge of identifying internal mechanisms responsible for specific behaviors, the study utilizes Layer-wise Relevance Propagation (LRP) with reference samples to isolate circuit components. By introducing contrastive relevance, the approach targets undesired behaviors while preserving general model capabilities. Experiments on the OPT-125M model demonstrated that pruning approximately 0.3% of neurons significantly reduced toxic outputs, while removing just 0.03% of weight elements mitigated repetitive text generation without performance degradation. The findings were validated across additional small-scale language models, proving the method's transferability across different architectures. This research establishes attribution-guided pruning as an effective mechanism for diagnosing and correcting undesirable behaviors in LLMs through targeted intervention on behavior-specific circuits. The associated code has been made publicly available to support further development and verification in the field of mechanistic interpretability.
Wire timeline
Attribution-Guided Pruning for Insight and Control in Small-scale LLMs
Researchers have introduced a novel method called attribution-guided pruning to enhance the interpretability and control of Large Language Models (LLMs). Addressing the challenge of identifying internal mechanisms responsible for specific behaviors, the study utilizes Layer-wise Relevance Propagation (LRP) with reference samples to isolate circuit components. By introducing contrastive relevance, the approach targets undesired behaviors while preserving general model capabilities. Experiments on the OPT-125M model demonstrated that pruning approximately 0.3% of neurons significantly reduced toxic outputs, while removing just 0.03% of weight elements mitigated repetitive text generation without performance degradation. The findings were validated across additional small-scale language models, proving the method's transferability across different architectures. This research establishes attribution-guided pruning as an effective mechanism for diagnosing and correcting undesirable behaviors in LLMs through targeted intervention on behavior-specific circuits. The associated code has been made publicly available to support further development and verification in the field of mechanistic interpretability.
cs.AI updates on arXiv.org