PHMForge: Evaluating LLM Agents on Industrial Prognostics via MCP-Native Tools
Researchers have introduced PHMForge, a new evaluation environment designed to assess the reliability of Large Language Model (LLM) agents in safety-critical Prognostics and Health Management (PHM). Addressing limitations in prior benchmarks that conflated protocol fluency with reasoning, PHMForge features 99 SME-authored scenarios across eight industrial asset classes, including rotating equipment and lithium-ion cells. The system utilizes 39 MCP-native tools wrapping established PHM algorithms. Testing across various agentic frameworks and LLM backbones revealed that the strongest configuration achieved an 80.8% pass@1 rate, with failures primarily stemming from orchestration and tool-sequencing errors. A critical ablation study demonstrated that replacing MCP execution with text-based Retrieval-Augmented Generation (RAG) caused performance on battery Remaining Useful Life tasks to collapse from 100% to 20%, highlighting the structural limits of static retrieval for prognostic computation. The study concludes that while frontier LLMs excel at calling tools, they struggle with planning when to invoke them. PHMForge is open-sourced with deterministic evaluators and a public leaderboard to facilitate further research in industrial AI applications.
Wire timeline
PHMForge: Evaluating LLM Agents on Industrial Prognostics via MCP-Native Tools
Researchers have introduced PHMForge, a new evaluation environment designed to assess the reliability of Large Language Model (LLM) agents in safety-critical Prognostics and Health Management (PHM). Addressing limitations in prior benchmarks that conflated protocol fluency with reasoning, PHMForge features 99 SME-authored scenarios across eight industrial asset classes, including rotating equipment and lithium-ion cells. The system utilizes 39 MCP-native tools wrapping established PHM algorithms. Testing across various agentic frameworks and LLM backbones revealed that the strongest configuration achieved an 80.8% pass@1 rate, with failures primarily stemming from orchestration and tool-sequencing errors. A critical ablation study demonstrated that replacing MCP execution with text-based Retrieval-Augmented Generation (RAG) caused performance on battery Remaining Useful Life tasks to collapse from 100% to 20%, highlighting the structural limits of static retrieval for prognostic computation. The study concludes that while frontier LLMs excel at calling tools, they struggle with planning when to invoke them. PHMForge is open-sourced with deterministic evaluators and a public leaderboard to facilitate further research in industrial AI applications.
cs.AI updates on arXiv.org