Robust Evaluation for Multi-Task Predictive Maintenance Models.

Md Mahamudur Rahaman Shamim, Md. Nuruzzaman, Zannatul Ferdus, Md Rajib Ahmed, Abieer Nwshad Anward, Mohammad Tooneer, Johir Uddin Khan, Khalid Hossen· July 21, 2026 View original

Summary

This study reveals how data leakage in splitting training/test sets can severely distort performance metrics for multi-task deep learning models in predictive maintenance. It proposes a leakage-audited protocol for robust evaluation and analyzes data-scale sensitivity.

Multi-task deep learning models, which simultaneously perform fault classification and remaining useful life (RUL) regression, are increasingly vital in predictive maintenance. However, this research highlights a critical issue: reported performance can be significantly skewed by how sliding-window sequences are partitioned into training and test sets. The study demonstrates that a naive splitting approach can artificially inflate classification accuracy to nearly perfect levels (e.g., 99.9%) or, conversely, reduce it to zero due to poor class representation. To counter this, the researchers introduce a "chunk-based, leakage-audited splitting protocol" and rigorously evaluate models using multiple seeds, ANOVA, and Tukey HSD tests. Using this robust protocol on the NASA C-MAPSS dataset, an attention-enhanced multi-task architecture (AMTLNet) achieved strong classification accuracy and R2 for RUL, matching or outperforming baselines. However, on smaller datasets (Bearing and Hydraulic), multi-task training proved unstable, with classification degrading for one and regression for the other, an asymmetry linked to label provenance. Ablation studies further showed that the multi-head attention branch is crucial for regression stability, while the convolutional branch contributes less despite its parameter count. The paper provides a practical framework for determining when joint training is appropriate under data scarcity and offers a reusable leakage-audit protocol for reliable evaluation.

Why it matters

For professionals developing AI for predictive maintenance, this research provides crucial guidelines for robust model evaluation, preventing inflated performance claims and ensuring reliable deployment, especially when dealing with time-series data and data scarcity.

How to implement this in your domain

  1. 1Adopt the proposed chunk-based, leakage-audited splitting protocol for all predictive maintenance model evaluations.
  2. 2Always use multiple seeds and statistical tests (e.g., ANOVA, Tukey HSD) to ensure the robustness and significance of reported results.
  3. 3Carefully analyze label provenance and data scarcity before implementing multi-task learning for fault diagnosis and RUL estimation.
  4. 4Prioritize attention mechanisms in model architectures when regression stability is critical for RUL estimation.

Who benefits

ManufacturingAerospaceEnergyAutomotiveIndustrial IoT

Key takeaways

  • Naive data splitting in predictive maintenance can lead to severely misleading model performance metrics.
  • A leakage-audited, chunk-based splitting protocol is essential for robust and reliable evaluation.
  • Multi-task learning stability is sensitive to data scale and label provenance, particularly for smaller datasets.
  • Attention mechanisms are critical for stabilizing regression performance in joint fault diagnosis and RUL estimation.

Original post by Md Mahamudur Rahaman Shamim, Md. Nuruzzaman, Zannatul Ferdus, Md Rajib Ahmed, Abieer Nwshad Anward, Mohammad Tooneer, Johir Uddin Khan, Khalid Hossen

"arXiv:2607.16493v1 Announce Type: new Abstract: Multi-task deep learning models that jointly perform fault classification and remaining useful life (RUL) regression are increasingly used in predictive maintenance, yet reported performance can be strongly affected by how sliding-w…"

View on X

Originally posted by Md Mahamudur Rahaman Shamim, Md. Nuruzzaman, Zannatul Ferdus, Md Rajib Ahmed, Abieer Nwshad Anward, Mohammad Tooneer, Johir Uddin Khan, Khalid Hossen on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses