Rethinking Layer Redundancy in Large Language Models: Calibration Objectives and Search for Depth Pruning
A new research paper published on arXiv challenges conventional approaches to depth pruning in large language models (LLMs). Traditionally, layer redundancy has been viewed as an inherent structural property of pretrained networks, with focus placed on identifying removable layers through importance criteria and search algorithms. However, this study adopts a functional perspective, arguing that redundancy is jointly determined by the model and the specific calibration objective used. Through extensive empirical testing across three LLM families, two calibration objectives, and seven search algorithms, the authors discovered that different objectives yield qualitatively distinct pruning patterns. Notably, rankings based on perplexity often diverge from those based on downstream reasoning accuracy. Conversely, when the calibration objective remains fixed, various search algorithms tend to converge on similar pruning solutions. These findings suggest that the choice of calibration objective plays a more critical role than the specific search algorithm in determining which layers are considered redundant. This insight implies that universal layer ranking methods may be ineffective, urging developers to prioritize objective selection for optimizing inference efficiency in LLMs.
Wire timeline
Rethinking Layer Redundancy in Large Language Models: Calibration Objectives and Search for Depth Pruning
A new research paper published on arXiv challenges conventional approaches to depth pruning in large language models (LLMs). Traditionally, layer redundancy has been viewed as an inherent structural property of pretrained networks, with focus placed on identifying removable layers through importance criteria and search algorithms. However, this study adopts a functional perspective, arguing that redundancy is jointly determined by the model and the specific calibration objective used. Through extensive empirical testing across three LLM families, two calibration objectives, and seven search algorithms, the authors discovered that different objectives yield qualitatively distinct pruning patterns. Notably, rankings based on perplexity often diverge from those based on downstream reasoning accuracy. Conversely, when the calibration objective remains fixed, various search algorithms tend to converge on similar pruning solutions. These findings suggest that the choice of calibration objective plays a more critical role than the specific search algorithm in determining which layers are considered redundant. This insight implies that universal layer ranking methods may be ineffective, urging developers to prioritize objective selection for optimizing inference efficiency in LLMs.
cs.AI updates on arXiv.org