BAMI: Training-Free Bias Mitigation in GUI Grounding
Researchers have introduced Bias-Aware Manipulation Inference (BAMI), a novel training-free method designed to enhance GUI grounding capabilities for AI agents. GUI grounding is essential for enabling agents to perform tasks like clicking and dragging on user interfaces. The study identifies two primary sources of error in existing models: precision bias caused by high image resolution and ambiguity bias resulting from intricate interface elements. To address these issues, BAMI employs a Masked Prediction Distribution attribution method and incorporates coarse-to-fine focus along with candidate selection mechanisms. Experimental results demonstrate significant performance improvements, notably boosting the accuracy of the TianXi-Action-7B model on the ScreenSpot-Pro benchmark from 51.9% to 57.8%. Ablation studies further confirm the robustness and stability of the approach across various parameter configurations. This development offers a practical solution for improving the reliability of GUI agents in complex scenarios without requiring additional model training. The associated code has been made publicly available to facilitate further research and application in computer vision and artificial intelligence domains.
Wire timeline
BAMI: Training-Free Bias Mitigation in GUI Grounding
Researchers have introduced Bias-Aware Manipulation Inference (BAMI), a novel training-free method designed to enhance GUI grounding capabilities for AI agents. GUI grounding is essential for enabling agents to perform tasks like clicking and dragging on user interfaces. The study identifies two primary sources of error in existing models: precision bias caused by high image resolution and ambiguity bias resulting from intricate interface elements. To address these issues, BAMI employs a Masked Prediction Distribution attribution method and incorporates coarse-to-fine focus along with candidate selection mechanisms. Experimental results demonstrate significant performance improvements, notably boosting the accuracy of the TianXi-Action-7B model on the ScreenSpot-Pro benchmark from 51.9% to 57.8%. Ablation studies further confirm the robustness and stability of the approach across various parameter configurations. This development offers a practical solution for improving the reliability of GUI agents in complex scenarios without requiring additional model training. The associated code has been made publicly available to facilitate further research and application in computer vision and artificial intelligence domains.
cs.AI updates on arXiv.org