CruiseBench Improves Aero-Engine RUL Prediction Benchmarking

Pu Cheng, Qiang Miao· July 23, 2026 View original

Summary

CruiseBench is a new benchmark for aero-engine Remaining Useful Life (RUL) prediction, derived from N-CMAPSS, that focuses on the cruise stage of flights. It provides a fixed protocol and a cruising-period mask (CPM-N-CMAPSS) to standardize evaluation, enabling more controlled comparisons of RUL models.

Researchers have introduced CruiseBench, a specialized benchmark designed to standardize and improve the evaluation of Remaining Useful Life (RUL) prediction models for aero-engines. This benchmark is derived from the N-CMAPSS dataset, which simulates engine degradation using real-flight profiles. While N-CMAPSS offers increased realism by retaining full within-flight time series, its complexity can make direct comparisons between RUL models challenging due to entangled degradation cues and operational variations. CruiseBench addresses this by focusing specifically on the cruise stage of flights, which is identified using a new artifact called CPM-N-CMAPSS (Cruising-Period Mask for N-CMAPSS). This mask isolates stable cruising intervals, allowing for a fixed evaluation protocol that uses scenario descriptors and measured sensors as inputs, while excluding auxiliary data. By standardizing the data preprocessing and evaluation, CruiseBench provides a more controlled environment for comparing RUL model performance. Initial experiments with various models like LSTM, GRU, TCN, and TSMixer demonstrate baseline results, with TSMixer achieving the lowest average RMSE. The study also highlights how factors like flight-stage selection and RUL-cap thresholds can influence reported outcomes, underscoring the need for such a standardized benchmark.

Why it matters

Accurate RUL prediction is crucial for predictive maintenance, reducing downtime, and improving safety in industries relying on complex machinery like aircraft engines. CruiseBench offers a more reliable way to develop and compare these models.

How to implement this in your domain

  1. 1Adopt CruiseBench as a standardized benchmark for evaluating aero-engine RUL prediction models.
  2. 2Utilize the CPM-N-CMAPSS mask to focus RUL model development on stable operational periods.
  3. 3Benchmark existing or new RUL models against CruiseBench to ensure robust and comparable performance metrics.
  4. 4Explore transfer learning and domain adaptation techniques using the stage-specific data foundation provided by CPM-N-CMAPSS.

Who benefits

AerospaceAviationManufacturingLogisticsDefense

Key takeaways

  • CruiseBench standardizes aero-engine RUL prediction evaluation by focusing on the cruise stage.
  • CPM-N-CMAPSS provides a mask to isolate stable cruising intervals for consistent data.
  • The benchmark enables more controlled and reproducible comparisons of RUL models.
  • TSMixer achieved the lowest average RMSE in initial CruiseBench experiments.

Original post by Pu Cheng, Qiang Miao

"arXiv:2607.19380v1 Announce Type: new Abstract: Remaining useful life (RUL) prediction estimates how long an engine can continue safe operation and is central to maintenance planning. N-CMAPSS extends C-MAPSS by simulating run-to-failure aero-engine trajectories using recorded re…"

View on X

Originally posted by Pu Cheng, Qiang Miao on X · view source

Want to go deeper?

Turn these trends into skills with Learnijoy's hands-on AI & tech courses.

Explore courses