ARSM-Agent: Enhancing Security and Adversarial Robustness in Medical LLM Agents
A new study published on arXiv introduces ARSM-Agent, a comprehensive security enhancement framework designed to improve the adversarial robustness and trustworthiness of large language model (LLM) agents in medical decision-making tasks. The research addresses critical security challenges by implementing a full-link framework that includes input risk perception, medical evidence constraints, knowledge consistency verification, and decision confidence reweighting. The proposed method utilizes a weighted joint objective function balancing decision accuracy, adversarial robustness, safety refusal, and knowledge consistency. Experimental results demonstrate that ARSM-Agent significantly outperforms four baseline models, including LLM-Agent and Adv-Train-Agent. Under various attack scenarios such as semantic perturbation, prompt injection, and false-evidence attacks, the system reduced the overall attack success rate to 8.7% while achieving a high knowledge consistency score of 0.91. Ablation studies further confirmed the individual contributions of each module to overall system accuracy and security. This approach provides a reliable intelligent support mechanism for secure medical decision-making in challenging and potentially hostile environments.
Wire timeline
ARSM-Agent: Enhancing Security and Adversarial Robustness in Medical LLM Agents
A new study published on arXiv introduces ARSM-Agent, a comprehensive security enhancement framework designed to improve the adversarial robustness and trustworthiness of large language model (LLM) agents in medical decision-making tasks. The research addresses critical security challenges by implementing a full-link framework that includes input risk perception, medical evidence constraints, knowledge consistency verification, and decision confidence reweighting. The proposed method utilizes a weighted joint objective function balancing decision accuracy, adversarial robustness, safety refusal, and knowledge consistency. Experimental results demonstrate that ARSM-Agent significantly outperforms four baseline models, including LLM-Agent and Adv-Train-Agent. Under various attack scenarios such as semantic perturbation, prompt injection, and false-evidence attacks, the system reduced the overall attack success rate to 8.7% while achieving a high knowledge consistency score of 0.91. Ablation studies further confirmed the individual contributions of each module to overall system accuracy and security. This approach provides a reliable intelligent support mechanism for secure medical decision-making in challenging and potentially hostile environments.
cs.AI updates on arXiv.org