Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
Researchers have introduced a new method called Untargeted Jailbreak via Entropy Maximization (UJEM)-KL to challenge the security of vision-language models (VLMs). Previous studies suggested that gradient-based universal image jailbreaks lacked cross-model transferability, raising doubts about feasible multimodal attacks. However, this new study revisits those conclusions under a strictly untargeted threat model, avoiding fixed prefixes or response patterns. The team discovered that refusal behaviors in VLMs concentrate at high-entropy tokens during autoregressive decoding. Leveraging this insight, UJEM-KL maximizes entropy at these critical decision tokens to flip refusal outcomes while stabilizing low-entropy positions to maintain output quality. Tested across three VLMs and two safety benchmarks, the lightweight attack achieved competitive white-box success rates and significantly improved transferability compared to prior methods. It remained effective even against representative defenses. The findings indicate that limited transferability in previous attempts stemmed from overly constrained optimization objectives rather than inherent model robustness. This research highlights potential vulnerabilities in current AI safety mechanisms and suggests a need for more robust defense strategies against untargeted adversarial attacks in multimodal systems.
Wire timeline
Break the Brake, Not the Wheel: Untargeted Jailbreak via Entropy Maximization
Researchers have introduced a new method called Untargeted Jailbreak via Entropy Maximization (UJEM)-KL to challenge the security of vision-language models (VLMs). Previous studies suggested that gradient-based universal image jailbreaks lacked cross-model transferability, raising doubts about feasible multimodal attacks. However, this new study revisits those conclusions under a strictly untargeted threat model, avoiding fixed prefixes or response patterns. The team discovered that refusal behaviors in VLMs concentrate at high-entropy tokens during autoregressive decoding. Leveraging this insight, UJEM-KL maximizes entropy at these critical decision tokens to flip refusal outcomes while stabilizing low-entropy positions to maintain output quality. Tested across three VLMs and two safety benchmarks, the lightweight attack achieved competitive white-box success rates and significantly improved transferability compared to prior methods. It remained effective even against representative defenses. The findings indicate that limited transferability in previous attempts stemmed from overly constrained optimization objectives rather than inherent model robustness. This research highlights potential vulnerabilities in current AI safety mechanisms and suggests a need for more robust defense strategies against untargeted adversarial attacks in multimodal systems.
cs.AI updates on arXiv.org