STAR: Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning
Researchers have introduced STAR (Spatio-Temporal Agent Router), a novel framework designed to enhance compositional spatiotemporal reasoning in multi-agent systems. Unlike existing tool-augmented Large Language Model (LLM) systems that handle routing implicitly, STAR externalizes inter-agent control through a state-conditioned transition policy. This approach utilizes an agent routing matrix that combines expert-defined nominal routes with recovery transitions learned from execution traces, including failures. By conditioning on distinct failure states such as malformed outputs or missing dependencies, the router enables specific recovery strategies rather than generic retries. The system employs a tool-grounded extract-compute-deposit protocol, where specialists write intermediate results to a shared blackboard. Experimental results across three spatiotemporal benchmarks and eight backbone LLMs demonstrate that STAR significantly outperforms baseline methods, particularly in scenarios where execution deviates from standard paths. The study highlights that retaining unsuccessful traces during training is crucial for expanding the routing policy's support on error states, proving that typed failure-aware routing is a key driver of performance improvements in complex multi-agent environments.
Wire timeline
STAR: Failure-Aware Markov Routing for Multi-Agent Spatiotemporal Reasoning
Researchers have introduced STAR (Spatio-Temporal Agent Router), a novel framework designed to enhance compositional spatiotemporal reasoning in multi-agent systems. Unlike existing tool-augmented Large Language Model (LLM) systems that handle routing implicitly, STAR externalizes inter-agent control through a state-conditioned transition policy. This approach utilizes an agent routing matrix that combines expert-defined nominal routes with recovery transitions learned from execution traces, including failures. By conditioning on distinct failure states such as malformed outputs or missing dependencies, the router enables specific recovery strategies rather than generic retries. The system employs a tool-grounded extract-compute-deposit protocol, where specialists write intermediate results to a shared blackboard. Experimental results across three spatiotemporal benchmarks and eight backbone LLMs demonstrate that STAR significantly outperforms baseline methods, particularly in scenarios where execution deviates from standard paths. The study highlights that retaining unsuccessful traces during training is crucial for expanding the routing policy's support on error states, proving that typed failure-aware routing is a key driver of performance improvements in complex multi-agent environments.
cs.AI updates on arXiv.org