HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
Researchers have introduced HiVLA, a novel hierarchical framework for robotic manipulation that addresses the limitations of end-to-end Vision-Language-Action (VLA) models. Traditional VLA models often lose reasoning capabilities when fine-tuned on narrow control data. HiVLA resolves this by decoupling high-level semantic planning from low-level motor control. The system employs a Vision-Language Model (VLM) planner for task decomposition and visual grounding, generating structured plans with subtask instructions and target bounding boxes. These plans are executed by a low-level flow-matching Diffusion Transformer (DiT) action expert, which uses a cascaded cross-attention mechanism to fuse global context, high-resolution object crops, and skill semantics. This architecture preserves the VLM's zero-shot reasoning abilities while allowing independent optimization of execution components. Extensive simulations and real-world experiments demonstrate that HiVLA significantly outperforms state-of-the-art baselines, particularly in long-horizon skill composition and fine-grained manipulation of small objects in cluttered environments, marking a significant advancement in embodied AI and robotics.
Wire timeline
HiVLA: A Visual-Grounded-Centric Hierarchical Embodied Manipulation System
Researchers have introduced HiVLA, a novel hierarchical framework for robotic manipulation that addresses the limitations of end-to-end Vision-Language-Action (VLA) models. Traditional VLA models often lose reasoning capabilities when fine-tuned on narrow control data. HiVLA resolves this by decoupling high-level semantic planning from low-level motor control. The system employs a Vision-Language Model (VLM) planner for task decomposition and visual grounding, generating structured plans with subtask instructions and target bounding boxes. These plans are executed by a low-level flow-matching Diffusion Transformer (DiT) action expert, which uses a cascaded cross-attention mechanism to fuse global context, high-resolution object crops, and skill semantics. This architecture preserves the VLM's zero-shot reasoning abilities while allowing independent optimization of execution components. Extensive simulations and real-world experiments demonstrate that HiVLA significantly outperforms state-of-the-art baselines, particularly in long-horizon skill composition and fine-grained manipulation of small objects in cluttered environments, marking a significant advancement in embodied AI and robotics.
cs.AI updates on arXiv.org