vix.ing · top · new · best · stats · spec

Generalization bounds and sample complexity for remaining useful life prediction from complete degradation trajectories

2026/05/21 by Huy Hoang Le, Kim-Anh Nguyen, Kim‐Anh Nguyen
Engineering · #Advanced Battery Technologies Research #Machine Fault Diagnosis Techniques #Reliability and Maintenance Optimization #cs.LG

paper · pdf · doi:10.1088/1361-6501/ae7109

openalex publication_date 2026/05/21 · openalex created_date 2026/05/22 · openalex updated_date 2026/07/30

Abstract

Abstract Data-driven remaining useful life (RUL) prediction requires complete degradation trajectories for training, yet such run-to-failure data are scarce and expensive. Practitioners currently lack principled guidance on how many failure examples suffice for a given model and accuracy target. This paper develops a sample complexity framework for RUL prediction comprising seven main results organized around three themes. First, we establish fundamental learning rates: a distribution-free generalization bound shows that the uniform deviation of the mean squared error decreases as <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" overflow="scroll"> <mml:mrow> <mml:mi>O</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:msup> <mml:mi>B</mml:mi> <mml:mrow> <mml:mn>2</mml:mn> </mml:mrow> </mml:msup> <mml:msqrt> <mml:mi>p</mml:mi> <mml:mrow> <mml:mo>/</mml:mo> </mml:mrow> <mml:mi>n</mml:mi> </mml:msqrt> <mml:mo stretchy="false">)</mml:mo> </mml:mrow> </mml:math> , where <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" overflow="scroll"> <mml:mrow> <mml:mi>p</mml:mi> </mml:mrow> </mml:math> is the model complexity and <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" overflow="scroll"> <mml:mrow> <mml:mi>n</mml:mi> </mml:mrow> </mml:math> the number of trajectories, and a minimax lower bound proves that the <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" overflow="scroll"> <mml:mrow> <mml:mi mathvariant="normal">Θ</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:mi>p</mml:mi> <mml:mrow> <mml:mo>/</mml:mo> </mml:mrow> <mml:mi>n</mml:mi> <mml:mo stretchy="false">)</mml:mo> </mml:mrow> </mml:math> rate is unimprovable. Second, we quantify how domain knowledge accelerates learning: incorporating degradation physics reduces data requirements by up to two orders of magnitude for deep networks, a Bernstein-type analysis achieves the minimax-optimal <mml:math xmlns:mml="http://www.w3.org/1998/Math/MathML" overflow="scroll"> <mml:mrow> <mml:mi>O</mml:mi> <mml:mo stretchy="false">(</mml:mo> <mml:mi>p</mml:mi> <mml:mrow> <mml:mo>/</mml:mo> </mml:mrow> <mml:mi>n</mml:mi> <mml:mo stretchy="false">)</mml:mo> </mml:mrow> </mml:math> rate under high signal-to-noise conditions, and closed-form penalties reveal when an incorrectly assumed physics model hurts rather than helps. Third, we characterize the impact of data quality: fleet variability induces an irreducible bias–variance tradeoff, while right-censored observations suffer an efficiency loss that depends critically on the degradation class. Closed-form expressions are provided for exponential, power-law, and stretched-exponential degradation. Cross-domain validation against published turbofan, battery, and bearing benchmarks confirms the theoretical predictions within a factor of 2–3 on average. The results yield practical guidelines for planning data collection, selecting model complexity, and evaluating physics model assumptions in prognostics applications.

Citations

Related