Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality
A new study published on arXiv challenges the assumption that scaling artificial intelligence models inherently improves the quality of their post-hoc explanations. Researchers evaluated eleven computer vision models, including ResNet, DenseNet, and Vision Transformer families, across three image datasets with ground-truth segmentation masks. By employing five explainable AI methods and two localisation metrics, including the newly proposed Dual-Polarity Precision, the team analyzed the relationship between model complexity and explanation accuracy. The findings reveal that increasing architectural depth and parameter count does not consistently enhance explanation quality; in many cases, smaller models matched or exceeded the performance of deeper variants. Furthermore, while pretraining improved predictive accuracy, it did not reliably boost localisation scores. The study highlights scenarios where models achieved high predictive performance despite near-zero localisation precision, indicating that standard performance metrics may fail to reflect whether predictions are based on relevant annotated regions. Consequently, the authors argue that explainability must be explicitly assessed during model selection, particularly for safety-sensitive deployments, rather than assuming larger models provide better interpretability.
Wire timeline
Scaling Vision Models Does Not Consistently Improve Localisation-Based Explanation Quality
A new study published on arXiv challenges the assumption that scaling artificial intelligence models inherently improves the quality of their post-hoc explanations. Researchers evaluated eleven computer vision models, including ResNet, DenseNet, and Vision Transformer families, across three image datasets with ground-truth segmentation masks. By employing five explainable AI methods and two localisation metrics, including the newly proposed Dual-Polarity Precision, the team analyzed the relationship between model complexity and explanation accuracy. The findings reveal that increasing architectural depth and parameter count does not consistently enhance explanation quality; in many cases, smaller models matched or exceeded the performance of deeper variants. Furthermore, while pretraining improved predictive accuracy, it did not reliably boost localisation scores. The study highlights scenarios where models achieved high predictive performance despite near-zero localisation precision, indicating that standard performance metrics may fail to reflect whether predictions are based on relevant annotated regions. Consequently, the authors argue that explainability must be explicitly assessed during model selection, particularly for safety-sensitive deployments, rather than assuming larger models provide better interpretability.
cs.AI updates on arXiv.org