FactoryNet: A Large-Scale Dataset for Industrial Time-Series Foundation Models
Researchers have introduced FactoryNet, the first universal pretraining corpus designed specifically for industrial time-series data. This significant dataset comprises 51 million datapoints derived from 23,000 end-to-end task executions, including 13,300 real and 9,800 synthetic samples across six different robotic and machining embodiments. The core innovation lies in a novel schema termed Setpoint, Effort, Feedback, Context (S-E-F-C), which standardizes diverse actuated systems into a common representational frame. This structure facilitates robust zero-shot cross-embodiment transfer and highly parameter-efficient anomaly detection. The corpus covers 27 annotated anomaly types alongside healthy baselines and counterfactual pairs. Experimental results demonstrate that models trained on FactoryNet achieve fair cross-embodiment transfer capabilities under bias-aware metrics. Furthermore, using just 24 schema-aligned signals, the model delivers competitive anomaly detection performance compared to high-dimensional baselines. By releasing this growing, multi-embodiment dataset, the authors aim to accelerate the development of foundation models tailored for industrial applications, addressing the need for standardized, large-scale data in manufacturing and robotics domains.
Wire timeline
FactoryNet: A Large-Scale Dataset for Industrial Time-Series Foundation Models
Researchers have introduced FactoryNet, the first universal pretraining corpus designed specifically for industrial time-series data. This significant dataset comprises 51 million datapoints derived from 23,000 end-to-end task executions, including 13,300 real and 9,800 synthetic samples across six different robotic and machining embodiments. The core innovation lies in a novel schema termed Setpoint, Effort, Feedback, Context (S-E-F-C), which standardizes diverse actuated systems into a common representational frame. This structure facilitates robust zero-shot cross-embodiment transfer and highly parameter-efficient anomaly detection. The corpus covers 27 annotated anomaly types alongside healthy baselines and counterfactual pairs. Experimental results demonstrate that models trained on FactoryNet achieve fair cross-embodiment transfer capabilities under bias-aware metrics. Furthermore, using just 24 schema-aligned signals, the model delivers competitive anomaly detection performance compared to high-dimensional baselines. By releasing this growing, multi-embodiment dataset, the authors aim to accelerate the development of foundation models tailored for industrial applications, addressing the need for standardized, large-scale data in manufacturing and robotics domains.
cs.AI updates on arXiv.org