IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Researchers have introduced IndustryBench, a new benchmark designed to evaluate the reliability and safety of Large Language Models (LLMs) in industrial procurement contexts. The dataset comprises 2,049 Chinese question-answer items grounded in national standards (GB/T) and structured product records, with translations into English, Russian, and Vietnamese. The study highlights that partial correctness in LLM responses can mask critical safety violations, a nuance often missed by aggregate benchmarks. Evaluation results across 17 models reveal significant limitations, with the top system scoring only 2.083 out of 3. Key findings indicate that 'Standards & Terminology' remains a persistent weakness and that extended reasoning often introduces unsupported, safety-critical details, lowering safety-adjusted scores for most models. Notably, safety-violation checks significantly reshuffled model leaderboards, demonstrating that source-grounded, safety-aware diagnosis is essential for industrial applications. The authors released the benchmark, prompts, and scoring scripts to facilitate more rigorous, safety-focused LLM evaluation in high-stakes industrial environments.
Wire timeline
IndustryBench: Probing the Industrial Knowledge Boundaries of LLMs
Researchers have introduced IndustryBench, a new benchmark designed to evaluate the reliability and safety of Large Language Models (LLMs) in industrial procurement contexts. The dataset comprises 2,049 Chinese question-answer items grounded in national standards (GB/T) and structured product records, with translations into English, Russian, and Vietnamese. The study highlights that partial correctness in LLM responses can mask critical safety violations, a nuance often missed by aggregate benchmarks. Evaluation results across 17 models reveal significant limitations, with the top system scoring only 2.083 out of 3. Key findings indicate that 'Standards & Terminology' remains a persistent weakness and that extended reasoning often introduces unsupported, safety-critical details, lowering safety-adjusted scores for most models. Notably, safety-violation checks significantly reshuffled model leaderboards, demonstrating that source-grounded, safety-aware diagnosis is essential for industrial applications. The authors released the benchmark, prompts, and scoring scripts to facilitate more rigorous, safety-focused LLM evaluation in high-stakes industrial environments.
cs.AI updates on arXiv.org