CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
Researchers have introduced CalBench, a novel controlled evaluation environment designed to study multi-agent coordination and privacy trade-offs in Large Language Models (LLMs). The benchmark utilizes a decentralized calendar scheduling scenario where N agents, each managing private calendars with pre-existing commitments, must coordinate to schedule incoming meetings while minimizing disruption costs. Since agents only observe their own data, successful scheduling necessitates communication across private information boundaries. CalBench generates scenarios with oracle solutions, allowing for precise measurement of coordination quality by comparing realized costs to optimal ones, alongside a Distributed Constraint Optimization (DCOP) baseline. The environment specifically evaluates task success, communication efficiency, fairness in cost distribution, and privacy leakage by tracking whether agents reveal sensitive, task-irrelevant semantic contexts during negotiation. Unlike other benchmarks where a single agent might substitute for a group, CalBench is inherently decentralized, providing a rigorous setting to analyze coordination protocols and negotiation strategies in multi-agent systems without central oversight.
Wire timeline
CalBench: Evaluating Coordination-Privacy Trade-offs in Multi-Agent LLMs
Researchers have introduced CalBench, a novel controlled evaluation environment designed to study multi-agent coordination and privacy trade-offs in Large Language Models (LLMs). The benchmark utilizes a decentralized calendar scheduling scenario where N agents, each managing private calendars with pre-existing commitments, must coordinate to schedule incoming meetings while minimizing disruption costs. Since agents only observe their own data, successful scheduling necessitates communication across private information boundaries. CalBench generates scenarios with oracle solutions, allowing for precise measurement of coordination quality by comparing realized costs to optimal ones, alongside a Distributed Constraint Optimization (DCOP) baseline. The environment specifically evaluates task success, communication efficiency, fairness in cost distribution, and privacy leakage by tracking whether agents reveal sensitive, task-irrelevant semantic contexts during negotiation. Unlike other benchmarks where a single agent might substitute for a group, CalBench is inherently decentralized, providing a rigorous setting to analyze coordination protocols and negotiation strategies in multi-agent systems without central oversight.
cs.AI updates on arXiv.org