Remembering Across Blocks: Topology-Conditioned Block-Progressive Memory for Skeleton-Based Action Recognition¶
Conference: ECCV 2026
Paper: ECCV paper page
Area: Video Understanding
Keywords: skeleton-based action recognition, graph convolutional networks, block-progressive memory, over-smoothing alleviation, topology reconstruction loss
TL;DR¶
To overcome representation homogenization (over-smoothing) and the attenuation of early fine-grained spatiotemporal cues caused by repeated one-hop neighborhood aggregation in skeleton GCNs, this paper proposes a topology-conditioned block-progressive memory updated online along the block axis via self-supervised reconstruction surprise, achieving state-of-the-art results across major benchmarks even when the memory module is completely detached at inference time (Mem-free).
Background & Motivation¶
Skeleton-based human action recognition models spatiotemporal bodily dynamics using sequences of human joint coordinates. Compared to raw RGB video streams, skeleton data offers distinct advantages: it is lightweight, compact, highly robust to background clutter, illumination changes, and clothing textures, and provides inherent privacy preservation. Consequently, it has been widely adopted in video surveillance, human-computer interaction, virtual reality, and robotics. Modern skeleton action recognition has predominantly evolved around spatial-temporal graph convolutional networks (ST-GCNs and their successors), which treat human joints as graph nodes and the natural kinematic connectivity as graph edges, leveraging graph inductive biases to aggregate localized motion patterns along structural paths.
To capture dependencies among distant joints and model dynamics spanning extended temporal windows, classical GCNs must stack multiple convolution blocks so that information propagates hop by hop. Although recent advances have introduced sample-dependent dynamic topology learning (such as CTR-GCN, BlockGCN, and ProtoGCN) to adaptively modulate connection strengths, their foundational operation remains rooted in repeated one-hop neighborhood feature aggregation. This iterative aggregation mechanism inherently incurs two critical representation bottlenecks: first, repeated neighborhood mixing progressively drives joint features toward homogenizationโan issue widely recognized as over-smoothing in graph representation learningโeroding node-wise distinctiveness in deep layers; second, highly localized spatial configurations and fine-grained, short-term temporal dynamics captured in earlier blocks become severely diluted across successive layers, preventing the final classifier embedding from fully capitalizing on early discriminative cues.
Drawing inspiration from neural memory systems in long-sequence language modeling (such as Titans, which uses surprise-driven signals along the time axis to selectively retain historical context), this work reorients the operational dimension of memory: rather than maintaining context across the temporal sequence axis, it deploys a memory mechanism that progresses along the GCN block axis. Core idea: by querying a sample-specific memory state with block-wise learned graph topologies and measuring self-supervised reconstruction surprise, the model accumulates complementary fine-grained spatiotemporal cues across GCN blocks, leveraging inner-loop memory dynamics during training to actively shape the backbone representations so that it retains high discriminative power at test time even without memory.
Method¶
Overall Architecture¶
The complete network consists of a deep GCN backbone (typically \(L=10\) stacked ST-GCN blocks), a sample-specific topology-conditioned block-progressive memory module, and an action classification head. Given an input skeleton sequence, each GCN block sequentially extracts spatiotemporal features and dynamically estimates a corresponding spatial adjacency topology. Parallel to this block progression, the memory module operates in an inner loop: it treats the current block's topology matrix as a query to reconstruct the current spatiotemporal feature, calculates the reconstruction error as a surprise signal to update the memory state, and accumulates complementary cues across blocks. At the final stage, a gated fusion mechanism merges the memory readout with the deepest backbone feature before feeding it to the classifier. During the outer loop, cross-entropy supervision on action labels simultaneously optimizes the backbone, the classifier, and the memory update meta-parameters, effectively internalizing early-cue preservation into the backbone weights.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
IN["Skeleton Sequence Input<br/>(N joints ร T frames ร C coords)"] --> GCN["GCN Backbone Blocks (l=1...L)<br/>Learns block topology A & features H"]
GCN --> READ["Topology-Conditioned Memory Read<br/>Reconstructs current features via topology A"]
READ --> SURPRISE["Surprise-Driven Inner-Loop Update<br/>Computes reconstruction error to update memory"]
SURPRISE --> FUSION["Gated Memory Feature Fusion<br/>Fuses final memory readout with deepest feature"]
FUSION --> CLS["Action Classification Prediction<br/>Outer-loop optimizes & shapes backbone"]
Key Designs¶
1. Dual-Matrix Memory Parameterization & Topology-Conditioned Read: Associative Querying via Graph Adjacency
To capture how different GCN blocks structure motion representations from diverse structural perspectives, the memory state at block \(l\) is parameterized as an encoder key matrix \(M_{\text{key}}^{(l)} \in \mathbb{R}^{N^2 \times D}\) and a decoder value matrix \(M_{\text{val}}^{(l)} \in \mathbb{R}^{D \times NT}\), where \(D\) is the hidden dimension. At block \(l\), the spatial adjacency tensor \(A^{(l)} \in \mathbb{R}^{N \times N \times C^{(l)}}\) is reshaped into a flattened topology query \(\tilde{A}^{(l)} \in \mathbb{R}^{N^2 \times C^{(l)}}\). The memory read operation predicts the current block spatiotemporal feature \(\hat{H}^{(l)}\) from the previous memory state \(M^{(l-1)}\):
The query \(\tilde{A}^{(l)}\) projects the block's relational signature into the hidden code space via \(M_{\text{key}}\), which is subsequently decoded across all \(N \times T\) spatiotemporal coordinates by \(M_{\text{val}}\), establishing graph connectivity as an index into stored motion priors.
2. Surprise-Driven Inner-Loop Self-Supervised Update: Capturing Unexplained Discriminative Cues
If the prior memory state \(M^{(l-1)}\) accurately reconstructs the current feature \(H^{(l)}\) under topology \(\tilde{A}^{(l)}\), the current block contributes minimal novel information relative to earlier layers. Conversely, a substantial discrepancy indicates that the current block has generated distinct, newly emerging spatiotemporal dynamics. The inner-loop defines a self-supervised reconstruction objective to quantify this surprise:
The gradient \(\nabla_{M^{(l-1)}} \mathcal{L}_{\text{mem}}^{(l)}\) provides a direct corrective signal, steering the memory update toward capturing unmodeled structural cues that emerge along network depth.
3. Momentum Surprise Accumulation & Adaptive Forgetting Gates: Stabilizing Cross-Block Evolution
Directly updating memory parameters via raw single-step gradients can cause the memory to overfit to localized layer noise or high-frequency fluctuations, impairing its capacity to retain representations across deeper stages. To ensure smooth progression, the framework introduces a momentum surprise state \(S^{(l)}\) alongside adaptive gating factors. Taking the global average pooled feature \(h^{(l)} = \text{GAP}(H^{(l)})\), small fully-connected layers predict a forgetting factor \(\alpha^{(l)}\), a state decay factor \(\eta^{(l)}\), and a momentum step factor \(\theta^{(l)}\) via sigmoid activations:
where \(\alpha^{(l)}, \eta^{(l)}, \theta^{(l)} \in [0, 1]\) modulate update rates per sample and block, suppressing noisy transients while allowing outdated early features to decay adaptively.
4. Gated Representation Fusion & Training-Time Representation Shaping: Dual Inference Flexibility
At the final stage \(L\), the memory yields an accumulated readout \(O^{(L)} = \text{Read}(A^{(L)}; M^{(L)})\), which is fused with the final backbone feature \(H^{(L)}\) via a channel-wise gate \(g = \sigma(O^{(L)} w_1 + H^{(L)} w_2 + b)\) to produce the unified representation \(Z = g \odot O^{(L)} + (1-g) \odot H^{(L)}\). During outer-loop supervised training on action labels, meta-gradients induced by \(\mathcal{L}_{\text{mem}}\) shift backpropagation gradients toward early and middle blocks, significantly increasing the entropy of layer-wise gradient contributions. Consequently, the backbone itself internalizes the ability to preserve fine-grained, non-homogenized features. This enables two flexible deployment modes: an optional Mem-enabled mode that runs the inner-loop update at test time for maximal accuracy, and a Mem-free mode that completely detaches the memory module during inference, delivering marked gains over baselines with zero additional parameters or latency overhead.
Loss & Training¶
The architecture is trained under a decoupled bi-level optimization scheme: 1. Inner Loop: For each skeleton sample during the forward pass, temporary states \(M^{(l)}\) and \(S^{(l)}\) are updated online via self-supervised reconstruction loss \(\mathcal{L}_{\text{mem}}^{(l)}\) without requiring ground-truth action labels; 2. Outer Loop: Across each mini-batch, the primary cross-entropy classification loss \(\mathcal{L}_{\text{cls}}\) updates the GCN backbone, classifier, initial memory state \(M^{(0)}\), gate networks \(\{ \text{FC}_{\alpha}, \text{FC}_{\eta}, \text{FC}_{\theta} \}\), and fusion parameters \((w_1, w_2, b)\); 3. Hyper-parameters: Models are trained using SGD with Nesterov momentum of 0.9, weight decay of \(5 \times 10^{-4}\), and an initial learning rate of 0.05 decaying under a cosine schedule across 150 epochs. Weight decay for gate FC layers is set to \(5 \times 10^{-5}\), with gradient clipping applied at a maximum \(\ell_2\) norm of 1.0 to ensure numerical stability.
Key Experimental Results¶
Main Results¶
Evaluations are conducted on three standard skeleton action benchmarks: NTU RGB+D (NTU-60), NTU RGB+D 120 (NTU-120), and Kinetics-Skeleton across 2-stream (E2), 4-stream (E4), and 6-stream (E6) ensembles.
| Dataset / Benchmark Split | Metric | Ours (Mem-free) | Ours (Mem-enabled) | Prev. SOTA (ProtoGCN) | Gain (Mem-enabled) |
|---|---|---|---|---|---|
| NTU-60 X-Sub (E2 / E4 / E6) | Top-1 (%) | 93.3 / 93.7 / 93.9 | 93.5 / 93.8 / 94.0 | 93.0 / 93.5 / 93.8 | +0.5 / +0.3 / +0.2 |
| NTU-60 X-View (E2 / E4 / E6) | Top-1 (%) | 97.3 / 97.7 / 97.9 | 97.3 / 97.7 / 97.9 | 97.2 / 97.5 / 97.8 | +0.1 / +0.2 / +0.1 |
| NTU-120 X-Sub (E2 / E4 / E6) | Top-1 (%) | 90.0 / 90.7 / 91.0 | 90.0 / 90.7 / 91.1 | 89.7 / 90.4 / 90.9 | +0.3 / +0.3 / +0.2 |
| NTU-120 X-Set (E2 / E4 / E6) | Top-1 (%) | 91.5 / 92.1 / 92.4 | 91.5 / 92.1 / 92.4 | 91.2 / 91.9 / 92.2 | +0.3 / +0.2 / +0.2 |
| Kinetics-Skeleton (Top-1 E2/E4/E6) | Top-1 (%) | 50.9 / 52.8 / 53.2 | 51.1 / 52.9 / 53.4 | 49.9 / 51.3 / 51.9 | +1.2 / +1.6 / +1.5 |
| Kinetics-Skeleton (Top-5 E2/E4/E6) | Top-5 (%) | 74.3 / 75.9 / 76.4 | 74.3 / 75.9 / 76.4 | 74.0 / 75.1 / 75.6 | +0.3 / +0.8 / +0.8 |
In single-modality evaluations on NTU-60 X-Sub, Mem-enabled achieves 91.9% on Joint (J) and 92.3% on Bone (B), surpassing ProtoGCN's 91.5% and 92.0%. On NTU-120 X-Sub, it scores 86.5% on Joint and 88.9% on Bone, consistently leading all GCN-based methods.
Ablation Study¶
The following table isolates the individual contributions of the memory update components and fusion gating on NTU-60 X-Sub (Joint modality, Baseline: 91.3%):
| Configuration | Memory Module | Fusion Gate | Forgetting \(\alpha\) | Decay \(\eta\) | Momentum \(\theta\) | Top-1 Acc (%) | Note |
|---|---|---|---|---|---|---|---|
| Baseline GCN | ร | ร | ร | ร | ร | 91.3 | Standard dynamic-topology GCN backbone |
| w/o all components | โ | ร | ร | ร | ร | 89.8 | Naive gradient update overfits high-frequency noise (-1.5%) |
| w/o forgetting factor (\(\alpha\)) | โ | โ | ร | โ | โ | 90.4 | Memory capacity saturates with obsolete states (-0.9%) |
| w/o momentum factor (\(\theta\)) | โ | โ | โ | โ | ร | 91.3 | Static step size fails to adapt, falling back to baseline |
| w/o decay factor (\(\eta\)) | โ | โ | โ | ร | โ | 91.4 | Accumulating unattenuated momentum induces oscillations |
| w/o gate module | โ | ร | โ | โ | โ | 91.5 | Fixed combination constrains representation capacity |
| Full Model (Ours) | โ | โ | โ | โ | โ | 91.9 | Optimal synergy across all stabilizing components (+0.6%) |
Regarding operational overhead, Mem-free inference maintains identical throughput (70.2 samples/s), parameter footprint (3.95M), and computational cost (7.48 GFLOPs) to the baseline, incurring only a slight training duration increase (from 9h to 12h). Mem-enabled inference introduces modest overhead (4.36M parameters, 7.85 GFLOPs) for situations where every fraction of accuracy is paramount.
Key Findings¶
- Naive memory updates degrade performance via noise memorization: Without momentum filtering and forgetting gates, direct surprise-based updates hurt accuracy by 1.5% relative to the baseline (89.8% vs. 91.3%), indicating that high-frequency noise is erroneously memorized without adaptive moderation.
- Training-time memory reallocates gradient focus to early layers: Analysis of block-wise feature sensitivity entropy reveals that the proposed memory module distributes backpropagation gradients far more evenly across network depth, mitigating the baseline's heavy reliance on the final few layers.
- Measurable mitigation of representation homogenization (over-smoothing): Tracking feature magnitude variance across joints over the NTU-120 test set demonstrates that the proposed architecture preserves significantly higher inter-joint variance at stage 10 compared to the baseline, confirming that node feature diversity is effectively retained.
Highlights & Insights¶
- Re-indexing memory along network depth rather than temporal length: While sequence models commonly treat memory as a tool to span temporal context, this paper transfers neural memory principles to the GCN block progression axis, directly resolving structural over-smoothing in hierarchical architectures.
- Self-supervised topology surprise as an intrinsic filter: By defining surprise as the residual error when predicting features from learned graph connectivity, the model identifies and captures genuine multi-scale motion cues without requiring explicit layer-wise annotations.
- Cost-free deployment via representation shaping: By demonstrating that training with memory reshapes the underlying backbone parameters, the method establishes that models can achieve state-of-the-art gains at inference without carrying the memory module into production.
Limitations & Future Work¶
- Admitted Limitations: The dual-level optimization loop requires computing reconstruction gradients and tracking momentum states across layers during the forward-backward pass, which moderately extends training runtime from 9 hours to 12 hours.
- Potential Boundaries: Reshaping the adjacency tensor into an \(N^2\) vector query assumes a fixed joint count \(N\), which may require architectural adaptations for scenarios with variable skeletal configurations or dense multi-person interactions.
- Future Directions: Future research could investigate compressed projection operators to avoid \(N^2\) flattening and explore whether block-progressive memory can be extended to deep 3D point cloud convolutions and molecular graph networks.
Related Work & Insights¶
- vs CTR-GCN / ProtoGCN: While CTR-GCN and ProtoGCN introduce sophisticated intra-block channel refinement and prototypical representations, they do not resolve the progressive attenuation of early features caused by repeated one-hop message passing. This method is complementary, providing substantial additive gains when built atop ProtoGCN.
- vs Titans: Titans employs surprise-driven test-time memorization along temporal sequences in NLP. This paper generalizes the paradigm to graph architectures along the layer progression axis, replacing self-attention queries with structural graph topologies to prevent graph node over-smoothing.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneering adaptation of surprise-driven neural memory along the GCN block axis]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive multi-stream evaluations, single-modality audits, and in-depth over-smoothing variance analyses]
- Writing Quality: โญโญโญโญโญ [Methodical mathematical formulation, well-structured bi-level optimization narrative, and crisp ablation rationale]
- Value: โญโญโญโญโญ [Zero-overhead Mem-free deployment provides immense practical utility for real-time edge robotics and embedded systems]