SLASH the Sink: Sharpening Structural Attention Inside LLMs
Researchers have introduced a novel, training-free method called StructuraL Attention SHarpening (Slash) to enhance the structural understanding capabilities of Large Language Models (LLMs). While LLMs excel in semantic tasks, they often struggle with graph topologies presented in serialized formats. The study reveals that LLMs internally reconstruct graph structures, evidenced by a 'sawtooth' pattern in attention maps, but this capability is diluted by the 'attention sink' phenomenon. This dilution creates a representation bottleneck due to a conflict between the model's anisotropic bias for language tasks and the local aggregation needed for graph reasoning. Slash addresses this issue through a plug-and-play attention redistribution mechanism that amplifies intrinsic structural understanding without requiring additional training or fine-tuning. Experimental results on pure graph tasks and molecular prediction demonstrate that Slash delivers significant and consistent performance improvements across various LLM architectures. This approach offers a cost-effective solution compared to existing methods that rely on external adapters or extensive fine-tuning, thereby preserving generalizability while boosting graph reasoning accuracy.
Wire timeline
SLASH the Sink: Sharpening Structural Attention Inside LLMs
Researchers have introduced a novel, training-free method called StructuraL Attention SHarpening (Slash) to enhance the structural understanding capabilities of Large Language Models (LLMs). While LLMs excel in semantic tasks, they often struggle with graph topologies presented in serialized formats. The study reveals that LLMs internally reconstruct graph structures, evidenced by a 'sawtooth' pattern in attention maps, but this capability is diluted by the 'attention sink' phenomenon. This dilution creates a representation bottleneck due to a conflict between the model's anisotropic bias for language tasks and the local aggregation needed for graph reasoning. Slash addresses this issue through a plug-and-play attention redistribution mechanism that amplifies intrinsic structural understanding without requiring additional training or fine-tuning. Experimental results on pure graph tasks and molecular prediction demonstrate that Slash delivers significant and consistent performance improvements across various LLM architectures. This approach offers a cost-effective solution compared to existing methods that rely on external adapters or extensive fine-tuning, thereby preserving generalizability while boosting graph reasoning accuracy.
cs.AI updates on arXiv.org