Communication-Theoretic Framework for Cost-Aware LLM Agent Reliability
Researchers Hamed Omidvar and Vahideh Akhlaghi propose a new analytical framework for Large Language Model (LLM) agents, grounded in Shannon’s communication theory. The study treats LLMs as discrete stochastic channels, unifying disparate reliability techniques like retrying and majority voting under six classical reliability operators. The authors derive closed-form results, including a noise-variance threshold for averaging methods and a contractivity criterion for generator-critic refinement. A key innovation is a cost-aware semantic-nearest-neighbor router that optimizes the quality-cost trade-off without retraining. Evaluations across local and cloud models on challenging tasks from MMLU, GSM8K, and HumanEval demonstrate that no single fixed technique dominates. Instead, the proposed router achieves the empirical Pareto frontier, reducing normalized costs by approximately 56% compared to the strongest fixed technique at matched quality, or improving quality by 7% at matched cost. This work advocates for consolidating reliability strategies into a single, tunable layer informed by channel coding principles to enhance efficiency and performance in LLM applications.
Wire timeline
Communication-Theoretic Framework for Cost-Aware LLM Agent Reliability
Researchers Hamed Omidvar and Vahideh Akhlaghi propose a new analytical framework for Large Language Model (LLM) agents, grounded in Shannon’s communication theory. The study treats LLMs as discrete stochastic channels, unifying disparate reliability techniques like retrying and majority voting under six classical reliability operators. The authors derive closed-form results, including a noise-variance threshold for averaging methods and a contractivity criterion for generator-critic refinement. A key innovation is a cost-aware semantic-nearest-neighbor router that optimizes the quality-cost trade-off without retraining. Evaluations across local and cloud models on challenging tasks from MMLU, GSM8K, and HumanEval demonstrate that no single fixed technique dominates. Instead, the proposed router achieves the empirical Pareto frontier, reducing normalized costs by approximately 56% compared to the strongest fixed technique at matched quality, or improving quality by 7% at matched cost. This work advocates for consolidating reliability strategies into a single, tunable layer informed by channel coding principles to enhance efficiency and performance in LLM applications.
cs.AI updates on arXiv.org