Insider Attacks in Multi-Agent LLM Consensus Systems
A new research paper submitted to arXiv investigates security vulnerabilities in multi-agent systems powered by Large Language Models (LLMs). While these systems rely on agents communicating in natural language to reach consensus, existing frameworks often assume all participants are aligned with system objectives. This study highlights the risk of malicious insiders who participate as legitimate members but pursue hidden adversarial goals, such as delaying or preventing agreement among benign agents. The authors formalize this threat as a sequential decision-making problem and propose a novel world-model-based framework. This approach learns surrogate dynamics over the latent behavioral states of benign agents and trains an attacker using reinforcement learning. Preliminary results indicate that this trained attacker significantly reduces consensus rates and prolongs disagreement more effectively than direct malicious-prompt baselines. The findings suggest that combining latent world models with reinforcement learning offers a potent method for executing adaptive insider attacks, raising important concerns for the security and robustness of future language-based multi-agent deployments.
Wire timeline
Insider Attacks in Multi-Agent LLM Consensus Systems
A new research paper submitted to arXiv investigates security vulnerabilities in multi-agent systems powered by Large Language Models (LLMs). While these systems rely on agents communicating in natural language to reach consensus, existing frameworks often assume all participants are aligned with system objectives. This study highlights the risk of malicious insiders who participate as legitimate members but pursue hidden adversarial goals, such as delaying or preventing agreement among benign agents. The authors formalize this threat as a sequential decision-making problem and propose a novel world-model-based framework. This approach learns surrogate dynamics over the latent behavioral states of benign agents and trains an attacker using reinforcement learning. Preliminary results indicate that this trained attacker significantly reduces consensus rates and prolongs disagreement more effectively than direct malicious-prompt baselines. The findings suggest that combining latent world models with reinforcement learning offers a potent method for executing adaptive insider attacks, raising important concerns for the security and robustness of future language-based multi-agent deployments.
cs.AI updates on arXiv.org