When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning
Researchers from arXiv have introduced a novel approach to long-horizon vision-language reasoning by formalizing 'commitment depth' as a learnable, state-conditioned variable rather than a fixed scalar. This method allows AI policies to dynamically decide how many primitive actions to execute open-loop before replanning, balancing replanning costs against execution errors. Implemented in a model-native vision-language policy, the system jointly predicts action sequences and their duration. Experimental results on Sliding Puzzle and Sokoban tasks demonstrate that this adaptive policy significantly outperforms fixed-depth baselines, achieving up to a 12.5 percentage point higher solve rate while using approximately 25% fewer actions. Notably, despite utilizing a 7B parameter backbone, the method surpasses advanced models like GPT-5.5 and Claude Sonnet, whereas tested open-weight models failed completely in zero-shot scenarios. Theoretical analysis confirms that state-conditioned commitment strictly dominates fixed depths when optimal depth varies across states, marking a significant advancement in efficient AI reasoning capabilities.
Wire timeline
When to Re-Commit: Temporal Abstraction Discovery for Long-Horizon Vision-Language Reasoning
Researchers from arXiv have introduced a novel approach to long-horizon vision-language reasoning by formalizing 'commitment depth' as a learnable, state-conditioned variable rather than a fixed scalar. This method allows AI policies to dynamically decide how many primitive actions to execute open-loop before replanning, balancing replanning costs against execution errors. Implemented in a model-native vision-language policy, the system jointly predicts action sequences and their duration. Experimental results on Sliding Puzzle and Sokoban tasks demonstrate that this adaptive policy significantly outperforms fixed-depth baselines, achieving up to a 12.5 percentage point higher solve rate while using approximately 25% fewer actions. Notably, despite utilizing a 7B parameter backbone, the method surpasses advanced models like GPT-5.5 and Claude Sonnet, whereas tested open-weight models failed completely in zero-shot scenarios. Theoretical analysis confirms that state-conditioned commitment strictly dominates fixed depths when optimal depth varies across states, marking a significant advancement in efficient AI reasoning capabilities.
cs.AI updates on arXiv.org