Skip to content

Geometry-Aware Visual Representation for Remaining Useful Life Prediction

Conference: ECCV 2026
Paper: ECCV Official
Cache: ../paper_cache/ECCV2026/eccv-5418.txt
Area: Time Series
Keywords: Remaining Useful Life Prediction, Phase Space Reconstruction, Globally Anchored Representation, Knowledge Distillation, Time Series to Image

TL;DR

Addressing the failure of conventional time-frequency representations to capture intrinsic system dynamical geometry, this paper introduces Phase Space Density Images (PSDI) with Globally Anchored Reconstruction (GAR) to model degradation as attractor dispersion, achieving robust RUL prediction via sliding-window temporal aggregation and compact knowledge distillation.

Background & Motivation

Prognostics and Health Management (PHM) for critical mechanical components relies heavily on accurate Remaining Useful Life (RUL) prediction. Despite rapid advances in recurrent neural networks and Transformer-based architectures, accurate life prognosis remains notoriously difficult due to the stochastic nature of wear and the non-linear, chaotic dynamics of mechanical systems. In real-world operational environments, raw vibration signals display high-frequency local oscillations and severe non-stationarity. The temporal waveforms under healthy and early degradation regimes exhibit nearly indistinguishable patterns, with explosive amplitude growth occurring only near catastrophic failure, often trapping conventional 1D sequence models in spurious transient noise.

To circumvent these limitations, image-based time-series representations have gained substantial traction. Existing methods predominantly employ Short-Time Fourier Transform (STFT) or Wavelet Transforms to convert scalar series into time-frequency heatmaps, which are subsequently processed by high-capacity vision backbones. However, time-frequency decompositions essentially track local variations in spectral energy and signal intensity; they do not characterize the underlying dynamical states or topological degradation mechanisms of the physical system. Under nonlinear dynamical system theory, a scalar vibration time series is merely a partial, 1D projection of an underlying high-dimensional dynamical attractor. As physical degradation worsens, the system attractor undergoes structural topological deformation—transitioning from a tightly concentrated, compact trajectory in healthy states to an increasingly dispersed and fragmented distribution near failure. Conventional spectral heatmaps and unanchored transformations fail to maintain an invariant reference coordinate system, thereby failing to reliably trace this progressive geometric breakdown.

The core tension lies in accurately capturing state-space geometric transitions while preventing arbitrary projection drift across non-stationary degradation stages. Core idea: build coordinate-consistent Phase Space Density Images (PSDI) via time-delay embedding, eliminate coordinate drift through a frozen globally anchored healthy reference frame, and suppress vibration transient jitter using sliding-window temporal aggregation within a compact multi-scale knowledge distillation framework.

Method

Overall Architecture

The proposed prognostics framework comprises three interconnected stages: Globally Anchored Phase Space Reconstruction (GAR), dual-channel PSDI density image and PC1 trajectory generation, and a multi-scale knowledge distillation framework with sliding-window temporal aggregation. First, global delay and embedding parameters are estimated from healthy training sequences, followed by fitting a 2D PCA orthogonal basis and quantile-based spatial boundaries on centered healthy trajectories. Next, multi-channel vibration inputs are projected onto this static reference frame to generate coordinate-consistent 2D density images and 1D trajectory sequences. During prediction, a high-capacity pretrained vision teacher provides geometric visual priors, while a lightweight student network is jointly optimized via temporal smoothing, multi-scale feature alignment, and attention distillation to generate monotonic, stable RUL estimations.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Multi-channel Raw Vibration Signals"] --> GAR["Globally Anchored Reconstruction<br/>Global parameters & healthy reference frame"]
    GAR --> PSDI["PSDI & Trajectory Construction<br/>2D spatial histogram & PC1 trajectory"]
    PSDI --> SWTA["Sliding-Window Feature Aggregation<br/>Compact encoder with low-pass filtering"]
    SWTA --> MSKD["Multi-Scale Feature Alignment & Distillation<br/>Pyramid alignment & attention distribution transfer"]
    MSKD --> Out["Monotonic & Stable RUL Prediction"]

Key Designs

1. Globally Anchored Reconstruction: Eliminating Coordinate Drift in State-Space Embeddings Scalar vibration signals offer only partial visibility into latent dynamical processes. By virtue of Takens' embedding theorem, a scalar time series can be lifted into an \(m\)-dimensional state space with time delay \(\tau_d\), forming delay vectors \(\mathbf{r}_i = [x_i, x_{i+\tau_d}, \dots, x_{i+(m-1)\tau_d}] \in \mathbb{R}^{1 \times m}\) and a trajectory matrix \(\mathbf{R} \in \mathbb{R}^{T \times m}\) (where \(T = N - (m - 1)\tau_d\)). However, if a 2D projection basis is fitted independently for each temporal window, arbitrary rotations and translations occur, confounding actual physical degradation with coordinate drift. To overcome this, Globally Anchored Reconstruction (GAR) establishes a unified, shared reference frame from an early-life healthy subset \(\mathcal{D}_{\mathrm{ref}} \subset \mathcal{D}_{\mathrm{train}}\). Global parameters \((\tau^*, m^*)\) are selected as the medians of individual estimates obtained via Mutual Information minimization and Cao's algorithm. All delay-embedded healthy trajectories are pooled to compute a global centroid \(\boldsymbol{\mu} \in \mathbb{R}^{1 \times m^*}\) and a 2D PCA projection basis \(\mathbf{U} \in \mathbb{R}^{2 \times m^*}\). Once computed from healthy dynamics, \((\boldsymbol{\mu}, \mathbf{U})\) is permanently frozen across all subsequent training and inference samples, projecting every sample onto the identical plane \(\mathbf{Z} = (\mathbf{R} - \mathbf{1}_T \boldsymbol{\mu})\mathbf{U}^\top \in \mathbb{R}^{T \times 2}\) so that observed geometric transformations strictly reflect real physical wear.

2. Phase Space Density Image & Trajectory Construction: Unifying Spatial Visitation Topology with Temporal Order To transform continuous state trajectories into 2D image tensors compatible with standard vision backbones, the projected plane is discretized into an \(H_{\mathrm{img}} \times W_{\mathrm{img}}\) grid. Instead of using min-max bounds that collapse under extreme transient shocks, spatial boundaries \(\mathbf{B}\) are determined from empirical quantiles (e.g., 1% and 99%) of the healthy projection distribution \(\mathbf{Z}_{\mathrm{ref}}\). A 2D density histogram \(\mathbf{H}\) counts the visitation frequency of trajectories across spatial bins. Because healthy attractors exhibit high spatial concentration, linear normalization would mask subtle outer-dispersion patterns indicative of incipient faults; thus, a logarithmic compression is applied: $$ \mathbf{I} = \mathrm{MinMaxNorm}\left(\log(1 + \mathbf{H})\right) $$ This compression expands the dynamic range of low-density regions, highlighting delicate peripheral wear features. Concurrently, since density histograms collapse temporal ordering, the first principal component (PC1) trajectory is extracted as a 1D temporal signal \(\mathbf{P}_t \in \mathbb{R}^{C \times P}\). Cross-attention injects this temporal signal into the visual patch embeddings, producing trajectory-enriched tokens that capture both spatial dispersion and causal degradation progression.

3. Sliding-Window Temporal Aggregation: Mitigating High-Frequency Noise and Non-Monotonic Jitter Frame-by-frame RUL prediction is highly susceptible to stochastic measurement noise and intermittent mechanical impacts, which cause erratic, non-monotonic predictions. After the compact student encoder extracts global features \(\mathbf{c}_t\), a causal sliding window of width \(K\) computes the moving average: $$ \bar{\mathbf{c}}t = \frac{1}{K} \sum_\tau $$ Under the theoretical assumptions of locally stationary degradation, independent noise, and a Lipschitz-continuous prediction head, this moving average acts as a feature-level low-pass filter, strictly bounding prediction variance by }^t \mathbf{c\(\mathrm{Var}(\hat{y}_t) = \mathcal{O}(1/K)\). This operation effectively dampens transient fluctuations and enforces monotonic degradation tracking toward end-of-life.

4. Multi-Scale Feature Alignment & Cross-Modal Distillation: Transferring Visual Priors while Suppressing Overload Directly deploying high-capacity visual foundation teachers (such as ViT or MAE) on vibration datasets risks semantic redundancy and severe overfitting. The proposed framework employs a compact student model (e.g., Tiny-ViT or MobileNet) guided by a frozen teacher. To bridge the dimension discrepancy between student features \(\bar{\mathbf{c}}_t \in \mathbb{R}^{d_s}\) and teacher features \(\mathbf{h}_t \in \mathbb{R}^{d_t}\), a pyramid-style multi-scale aligner with learnable scale weights \(\boldsymbol{\alpha} = \mathrm{softmax}(\boldsymbol{\theta})\) projects representations across granularities: $$ \hat{\mathbf{c}}t = \sum}^S \alpha_i \psi_i(\bar{\mathbf{c}t) $$ The distillation objective pairs point-wise MSE and directional cosine distance with temperature-scaled KL divergence, while transferring spatial self-attention matrices via attention correlation distillation: $$ \mathcal{L}\right)\right) $$ Jointly training with the primary Smooth L1 regression loss ensures that the student inherits essential geometric inductive biases while discarding unneeded natural image semantic complexity.}} = \frac{T_a^2}{B} \sum_{i=1}^B D_{\mathrm{KL}}\left(\sigma\left(\frac{\mathbf{A}_t^{(i)}}{T_a}\right) \Bigg| \sigma\left(\frac{\mathbf{A}_s^{(i)}}{T_a

Loss & Training

The teacher network is trained with the frozen pretrained backbone, optimizing only the regression head via Smooth L1 loss: \(\mathcal{L}_{\mathrm{teacher}} = \mathrm{SmoothL1}(\hat{y}_t, y_t)\). The student model minimizes a unified objective combining RUL regression and multi-component distillation: $$ \mathcal{L}{\mathrm{student}} = \mathcal{L}}} + \lambda_{\mathrm{distill}} \left( \lambda_{\mathrm{feat}}\mathcal{L{\mathrm{feat}} + \lambda \right) $$ where feature distillation loss comprises }}\mathcal{L}_{\mathrm{att}\(\mathcal{L}_{\mathrm{feat}} = \lambda_{\mathrm{mse}}\mathcal{L}_{\mathrm{mse}} + \lambda_{\mathrm{cos}}\mathcal{L}_{\mathrm{cos}} + \lambda_{\mathrm{KL}}\mathcal{L}_{\mathrm{KL}}\), with alignment and distillation coefficients adaptively updated during training.

Key Experimental Results

Main Results

Evaluations are conducted on two standard bearing benchmarks: XJTU-SY and PHM. Five evaluation cases are covered: Category 1 (Cases 1–3) tests under identical operating conditions, while Category 2 (Cases 4–5) evaluates cross-condition generalization. Metrics include Root Mean Squared Error (RMSE) and Mean Absolute Percentage Error (MAPE).

Method Case 1 (RMSE / MAPE) Case 2 (RMSE / MAPE) Case 3 (RMSE / MAPE) Case 4 (RMSE / MAPE) Case 5 (RMSE / MAPE)
AE-TCA-LSSVM 0.128 / 0.243 0.232 / 0.317 0.096 / 0.358 0.159 / 0.281 0.214 / 1.243
2D-CLSTM 0.324 / 1.320 0.215 / 0.302 0.075 / 0.251 0.081 / 0.177 0.191 / 0.266
CNN-ORC 0.164 / 0.306 0.426 / 0.391 0.192 / 1.789 0.225 / 0.316 0.219 / 0.290
Bi-LSTMF 0.218 / 0.291 0.328 / 0.370 0.129 / 0.215 0.169 / 0.241 0.216 / 0.342
CLSTMF 0.146 / 0.409 0.190 / 0.285 0.234 / 0.540 0.144 / 0.271 0.091 / 0.184
PatchTST 0.151 / 0.234 0.095 / 0.040 0.120 / 0.053 0.099 / 0.070 0.087 / 0.040
TimesNet 0.195 / 0.283 0.183 / 0.195 0.119 / 0.048 0.117 / 0.101 0.104 / 0.053
TimeMixer 0.164 / 0.257 0.113 / 0.086 0.120 / 0.045 0.122 / 0.087 0.103 / 0.046
TimeMixer++ 0.166 / 0.270 0.079 / 0.026 0.118 / 0.046 0.094 / 0.092 0.086 / 0.035
Only Teacher (CLIP) 0.170 / 0.267 0.125 / 0.117 0.063 / 0.051 0.096 / 0.107 0.090 / 0.078
Ours (PSDI + KD) 0.099 / 0.156 0.107 / 0.070 0.071 / 0.039 0.067 / 0.055 0.060 / 0.040

Ablation Study

Table 1: Performance comparison across visual representations under the KD framework

Image Representation Case 1 (RMSE / MAPE) Case 2 (RMSE / MAPE) Case 3 (RMSE / MAPE) Case 4 (RMSE / MAPE) Case 5 (RMSE / MAPE)
Recurrence Plots (RP) 0.125 / 0.206 0.084 / 0.041 0.097 / 0.046 0.079 / 0.060 0.066 / 0.046
Gramian Angular Fields (GAF) 0.123 / 0.203 0.085 / 0.044 0.097 / 0.045 0.069 / 0.056 0.070 / 0.037
STFT 0.128 / 0.212 0.127 / 0.089 0.094 / 0.047 0.072 / 0.059 0.064 / 0.027
Wavelet Transform 0.119 / 0.192 0.093 / 0.028 0.096 / 0.045 0.071 / 0.062 0.062 / 0.029
PSDI (Ours) 0.099 / 0.156 0.107 / 0.070 0.071 / 0.039 0.067 / 0.055 0.060 / 0.040

Table 2: Quantitative impact of Global Anchoring (GAR) on feature space clustering

Setting Silhouette Score ↑ Inter-cluster Separation ↑ Intra-cluster Compactness ↓ Finding
w/o Anchoring (Local PSR) 0.101 0.175 0.308 Severe coordinate rotation and spatial drift
Anchoring (Globally Anchored) 0.516 0.741 0.233 >4x Silhouette boost; distinct degradation boundaries

Key Findings

  • General Superiority of Geometric Attractor Features: Compared to STFT and Wavelet transforms that only capture spectral energy distributions, PSDI achieves the lowest RMSE across Cases 1, 3, 4, and 5. This advantage is particularly pronounced in cross-condition transfer cases (Case 4: 0.067 vs. Wavelet's 0.071; Case 5: 0.060 vs. 0.062), demonstrating resilience to working condition variations.
  • Critical Role of Global Anchoring: Disabling global anchoring degrades RMSE across all settings (e.g., Case 1 error rises from 0.099 to 0.128, Case 4 from 0.067 to 0.084). Clustering analysis indicates that anchoring improves the Silhouette score from 0.101 to 0.516 and expands inter-cluster separation from 0.175 to 0.741, proving that a unified coordinate system is essential for models to distinguish health stages.
  • Distillation Outperforming Standalone Teachers: Benefiting from sliding-window temporal aggregation and multi-scale distillation, the lightweight student model (CLIP \(\to\) Tiny-ViT) surpasses the high-capacity standalone teacher in Cases 1, 4, and 5 (e.g., Case 1 drops from 0.170 to 0.099). This validates that temporal feature smoothing effectively filters out high-frequency vibration noise that overfits complex vision backbones.

Highlights & Insights

  • Degradation Reformulated as Phase-Space Attractor Dispersion: Overcomes the limits of time-frequency heatmaps by mapping chaotic mechanical degradation directly into intuitive visual topological transitions from tight clusters to dispersed clouds.
  • One-Time Frozen Global Health Anchor: Eliminates coordinate drift and arbitrary rotations by estimating projection parameters \((\tau^*, m^*, \boldsymbol{\mu}, \mathbf{U}, \mathbf{B})\) once on initial healthy data and holding them fixed for all subsequent samples.
  • Theoretically Bound Variance Reduction via Temporal Aggregation: Formally proves that sliding-window moving averages bound prediction variance by \(\mathcal{O}(1/K)\), offering an effective low-pass filtering mechanism in feature space to eliminate non-monotonic RUL jitter.

Limitations & Future Work

  • Lack of Long-Term Cumulative History Modeling: PSDI is constructed over isolated, short time windows; it reflects instantaneous attractor dispersion but lacks explicit memory of irreversible plastic deformation or cumulative fatigue accrued over thousands of operating hours.
  • Strong Dependence on Initial Clean Reference Data: The framework assumes that early operating data accurately reflects pristine healthy conditions; if an asset starts with manufacturing defects or unmonitored pre-existing wear, the reference anchor will introduce a systematic bias.
  • Future Directions: Exploring temporal sequence modeling over sequential PSDI maps (e.g., phase-space density video diffusion) and investigating dynamic reference calibration for variable-speed and variable-load regimes.
  • vs STFT / Wavelet Spectrograms: While spectral transforms project signal energy across time-frequency bins, PSDI unrolls state-space dynamics under Takens' theorem, providing a direct visual interpretation of underlying nonlinear chaotic trajectories.
  • vs Recurrence Plots (RP) / Gramian Angular Fields (GAF): Although RP and GAF generate 2D representations, they focus on pairwise temporal cross-recurrence or polar angles; PSDI maps actual state visitation frequencies onto a physically grounded, globally anchored reference plane.
  • vs TimesNet / TimeMixer++: Generic time-series foundation models focus on multi-periodicity and cross-scale mixing; the proposed method incorporates domain-specific dynamical visual priors, achieving superior cross-condition prognostics with a fraction of the computational footprint.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Pioneering introduction of coordinate-consistent phase space attractor dispersion for RUL visual representation.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive cross-condition benchmarks on two bearing datasets across multiple vision backbones and visual representations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous nonlinear dynamics derivations, clear mathematical formulations, and thorough ablation evidence.
  • Value: ⭐⭐⭐⭐⭐ Establishes a highly practical and noise-resilient visual paradigm for industrial prognostic health monitoring.