Data-driven Circuit Discovery for Interpretability of Language Models
Researchers from arXiv have introduced Data-driven Circuit Discovery (DCD), a new framework designed to enhance the interpretability of language models (LMs). Traditional circuit discovery methods are hypothesis-driven, assuming that LMs implement tasks via a single computational subgraph and that datasets accurately represent human-defined tasks. This study challenges these assumptions, demonstrating that minor dataset variations can lead to circuits with low overlap and faithfulness. Furthermore, existing methods often merge distinct mechanisms into a single circuit when applied to mixed datasets. In contrast, DCD clusters examples based on similar model processing patterns and discovers separate circuits for each group. This approach allows distinct mechanistic structures to emerge independently, rather than being conflated. Experimental results indicate that DCD identifies multiple circuits per dataset, each exhibiting higher faithfulness to its specific group compared to single-circuit methods. By letting data reveal internal mechanistic structures instead of relying on rigid human-defined boundaries, DCD offers a more accurate explanation of how language models organize computation, marking a significant advancement in AI interpretability research.
Wire timeline
Data-driven Circuit Discovery for Interpretability of Language Models
Researchers from arXiv have introduced Data-driven Circuit Discovery (DCD), a new framework designed to enhance the interpretability of language models (LMs). Traditional circuit discovery methods are hypothesis-driven, assuming that LMs implement tasks via a single computational subgraph and that datasets accurately represent human-defined tasks. This study challenges these assumptions, demonstrating that minor dataset variations can lead to circuits with low overlap and faithfulness. Furthermore, existing methods often merge distinct mechanisms into a single circuit when applied to mixed datasets. In contrast, DCD clusters examples based on similar model processing patterns and discovers separate circuits for each group. This approach allows distinct mechanistic structures to emerge independently, rather than being conflated. Experimental results indicate that DCD identifies multiple circuits per dataset, each exhibiting higher faithfulness to its specific group compared to single-circuit methods. By letting data reveal internal mechanistic structures instead of relying on rigid human-defined boundaries, DCD offers a more accurate explanation of how language models organize computation, marking a significant advancement in AI interpretability research.
cs.AI updates on arXiv.org