Consensus Sampling Algorithm Proposed for Safer Generative AI
Researchers Adam Tauman Kalai, Yael Tauman Kalai, and Or Zamir have introduced a new method called consensus sampling to enhance the safety of generative AI systems. Addressing undetectable risks in AI outputs, this black-box algorithm aggregates multiple probability distributions to boost safety without requiring specific model architectures. The approach is competitive with the average risk of the safest subset of models while abstaining from generating output when there is insufficient agreement among them. Formalized through R-robustness, the method bounds information leakage and adversarial influence, offering a model-agnostic solution for inheriting guarantees from reliable subsets of models. Inspired by robust statistics and prior work on copyright protection, the study demonstrates that standard mixture methods are vulnerable to unsafe constituents, whereas their pointwise-median construction provides robust intuition. Experiments on synthetic distributions and image generation illustrate the mechanism's effectiveness. This research provides a Pareto-optimal tradeoff between worst-case risk and abstention, requiring overlap among safe distributions to function effectively.
Wire timeline
Consensus Sampling Algorithm Proposed for Safer Generative AI
Researchers Adam Tauman Kalai, Yael Tauman Kalai, and Or Zamir have introduced a new method called consensus sampling to enhance the safety of generative AI systems. Addressing undetectable risks in AI outputs, this black-box algorithm aggregates multiple probability distributions to boost safety without requiring specific model architectures. The approach is competitive with the average risk of the safest subset of models while abstaining from generating output when there is insufficient agreement among them. Formalized through R-robustness, the method bounds information leakage and adversarial influence, offering a model-agnostic solution for inheriting guarantees from reliable subsets of models. Inspired by robust statistics and prior work on copyright protection, the study demonstrates that standard mixture methods are vulnerable to unsafe constituents, whereas their pointwise-median construction provides robust intuition. Experiments on synthetic distributions and image generation illustrate the mechanism's effectiveness. This research provides a Pareto-optimal tradeoff between worst-case risk and abstention, requiring overlap among safe distributions to function effectively.
cs.AI updates on arXiv.org