UTS Achieves Second Place in PsyDefDetect with Multi-Agent AI for Defense Mechanism Classification
Researchers from the University of Technology Sydney (UTS) have secured second place among 64 teams in the PsyDefDetect challenge, achieving an F1 score of 0.406 in classifying psychological defense mechanisms within emotional support dialogues. Their system leverages the Defense Mechanism Rating Scales (DMRS) and introduces a novel concept of absence-based reasoning, where defense mechanisms are identified by missing affect, blocked cognition, or denied reality. This approach, encoded via prompt-level clinical rules, contributed an 11.4 percentage point gain in F1 score. The architecture features a multi-phase deliberative council of Gemini 2.5 agents that rate evidence strength rather than voting, achieving a top-5 result without fine-tuning. To address systematic errors in minority class predictions, the team implemented a targeted override ensemble using three fine-tuned Qwen3.5 models. This structured multi-agent system, comprising builder, critic, and regression guard roles, applied 16 specific overrides to boost performance further. The study highlights both the potential of large language models in clinical psychology applications and the challenges of bias toward majority classes in emotional content analysis.
Wire timeline
UTS Achieves Second Place in PsyDefDetect with Multi-Agent AI for Defense Mechanism Classification
Researchers from the University of Technology Sydney (UTS) have secured second place among 64 teams in the PsyDefDetect challenge, achieving an F1 score of 0.406 in classifying psychological defense mechanisms within emotional support dialogues. Their system leverages the Defense Mechanism Rating Scales (DMRS) and introduces a novel concept of absence-based reasoning, where defense mechanisms are identified by missing affect, blocked cognition, or denied reality. This approach, encoded via prompt-level clinical rules, contributed an 11.4 percentage point gain in F1 score. The architecture features a multi-phase deliberative council of Gemini 2.5 agents that rate evidence strength rather than voting, achieving a top-5 result without fine-tuning. To address systematic errors in minority class predictions, the team implemented a targeted override ensemble using three fine-tuned Qwen3.5 models. This structured multi-agent system, comprising builder, critic, and regression guard roles, applied 16 specific overrides to boost performance further. The study highlights both the potential of large language models in clinical psychology applications and the challenges of bias toward majority classes in emotional content analysis.
cs.AI updates on arXiv.org