DeCIR: Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
Researchers have introduced DeCIR, a novel method for Zero-Shot Composed Image Retrieval (ZS-CIR) that addresses the semantic transition bottleneck in existing projection-based approaches. While current lightweight methods avoid relying on Large Language Models during inference, they often struggle with complex semantic modifications compared to LLM-based systems. The study identifies an endpoint-transition conflict where edit text acts merely as a target attribute cue rather than a source-conditioned semantic transition. To resolve this, DeCIR decouples endpoint and transition learning by constructing paired forward/reverse edit tuples from image-caption pairs. It trains separate low-rank text adapter branches for endpoint and semantic transition alignment, merging them via Low-Rank Directional Merge (LRDM) into a single deployable adapter. Extensive experiments across datasets like CIRR, CIRCO, FashionIQ, and GeneCIS demonstrate that DeCIR consistently improves retrieval performance without increasing inference complexity, offering a more efficient and accurate solution for composed image retrieval tasks.
Wire timeline
DeCIR: Decoupling Endpoint and Semantic Transition Learning for Zero-Shot Composed Image Retrieval
Researchers have introduced DeCIR, a novel method for Zero-Shot Composed Image Retrieval (ZS-CIR) that addresses the semantic transition bottleneck in existing projection-based approaches. While current lightweight methods avoid relying on Large Language Models during inference, they often struggle with complex semantic modifications compared to LLM-based systems. The study identifies an endpoint-transition conflict where edit text acts merely as a target attribute cue rather than a source-conditioned semantic transition. To resolve this, DeCIR decouples endpoint and transition learning by constructing paired forward/reverse edit tuples from image-caption pairs. It trains separate low-rank text adapter branches for endpoint and semantic transition alignment, merging them via Low-Rank Directional Merge (LRDM) into a single deployable adapter. Extensive experiments across datasets like CIRR, CIRCO, FashionIQ, and GeneCIS demonstrate that DeCIR consistently improves retrieval performance without increasing inference complexity, offering a more efficient and accurate solution for composed image retrieval tasks.
cs.AI updates on arXiv.org