MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
Researchers have introduced MTA-RL, a novel framework for robust urban autonomous driving that bridges perception and control using Multi-modal Transformer-based 3D Affordances and Reinforcement Learning. Addressing the lack of interpretability in end-to-end models and error propagation in modular pipelines, MTA-RL fuses RGB images and LiDAR point clouds via a transformer architecture to predict explicit, geometry-aware affordance representations. These structured representations create a compact observation space, allowing the reinforcement learning policy to operate on driving semantics, which enhances sample efficiency and stability. Extensive evaluations in the CARLA simulator across towns with varying traffic densities demonstrate that MTA-RL consistently outperforms state-of-the-art baselines. Notably, the model exhibits superior zero-shot generalization when trained solely on Town03 and tested on unseen environments, achieving significant improvements in route completion, total distance, and distance per violation. Ablation studies confirm the critical role of multi-modal fusion and reward shaping, establishing MTA-RL as an effective solution for stable decision-making in dense urban interactions.
Wire timeline
MTA-RL: Robust Urban Driving via Multi-modal Transformer-based 3D Affordances and Reinforcement Learning
Researchers have introduced MTA-RL, a novel framework for robust urban autonomous driving that bridges perception and control using Multi-modal Transformer-based 3D Affordances and Reinforcement Learning. Addressing the lack of interpretability in end-to-end models and error propagation in modular pipelines, MTA-RL fuses RGB images and LiDAR point clouds via a transformer architecture to predict explicit, geometry-aware affordance representations. These structured representations create a compact observation space, allowing the reinforcement learning policy to operate on driving semantics, which enhances sample efficiency and stability. Extensive evaluations in the CARLA simulator across towns with varying traffic densities demonstrate that MTA-RL consistently outperforms state-of-the-art baselines. Notably, the model exhibits superior zero-shot generalization when trained solely on Town03 and tested on unseen environments, achieving significant improvements in route completion, total distance, and distance per violation. Ablation studies confirm the critical role of multi-modal fusion and reward shaping, establishing MTA-RL as an effective solution for stable decision-making in dense urban interactions.
cs.AI updates on arXiv.org