Metis: New Framework Achieves High Success Rate in Jailbreaking LLMs via Metacognitive Optimization
Researchers have introduced Metis, a novel framework designed to uncover vulnerabilities in Large Language Models (LLMs) through advanced red teaming techniques. Published on arXiv, the study addresses limitations in existing automated methods that often rely on static heuristics. Metis reformulates jailbreaking as inference-time policy optimization within an adversarial Partially Observable Markov Decision Process (POMDP). It utilizes a self-evolving metacognitive loop to diagnose target defense logic and refine its attack policy using structured feedback. Evaluations across ten diverse models show Metis achieves an average Attack Success Rate (ASR) of 89.2%, significantly outperforming traditional baselines. Notably, it maintains high efficacy against resilient frontier models like O1 and GPT-5-chat, with success rates of 76.0% and 78.0% respectively. The framework also enhances efficiency by reducing token costs by an average of 8.2 times. These findings highlight critical vulnerabilities in current LLM safety alignments, suggesting an urgent need for next-generation defenses capable of dynamic reasoning during inference to counter internally-steered, closed-loop reasoning trajectories.
Wire timeline
Metis: New Framework Achieves High Success Rate in Jailbreaking LLMs via Metacognitive Optimization
Researchers have introduced Metis, a novel framework designed to uncover vulnerabilities in Large Language Models (LLMs) through advanced red teaming techniques. Published on arXiv, the study addresses limitations in existing automated methods that often rely on static heuristics. Metis reformulates jailbreaking as inference-time policy optimization within an adversarial Partially Observable Markov Decision Process (POMDP). It utilizes a self-evolving metacognitive loop to diagnose target defense logic and refine its attack policy using structured feedback. Evaluations across ten diverse models show Metis achieves an average Attack Success Rate (ASR) of 89.2%, significantly outperforming traditional baselines. Notably, it maintains high efficacy against resilient frontier models like O1 and GPT-5-chat, with success rates of 76.0% and 78.0% respectively. The framework also enhances efficiency by reducing token costs by an average of 8.2 times. These findings highlight critical vulnerabilities in current LLM safety alignments, suggesting an urgent need for next-generation defenses capable of dynamic reasoning during inference to counter internally-steered, closed-loop reasoning trajectories.
cs.AI updates on arXiv.org