Noise-Robust Facial Expression Recognition via Mamba-driven Neighbor Weight Refinement¶
Conference: ECCV 2026
Paper: ECCV Official Page
Area: Human Understanding
Keywords: Facial Expression Recognition, Label Distribution Learning, Neighbor Weight Refinement, State Space Model, Label Noise
TL;DR¶
To combat crowdsourced annotation ambiguity and label noise in in-the-wild facial expression recognition, Mamba-NWR combines cross-attention routing initialization with Mamba state space iterative refinement to dynamically balance historical states and aggregate robust neighborhood label distributions.
Background & Motivation¶
Facial expression recognition (FER) in unconstrained, in-the-wild scenarios is a fundamental building block for affective computing and natural human-computer interaction. However, real-world FER inherently suffers from subjective annotation ambiguity: subtle muscle activations and transitional facial expressions are interpreted differently by different human annotators. Furthermore, large-scale benchmarks typically rely on crowdsourced annotations that inevitably introduce systematic label noise and class inconsistencies. Under standard hard-target cross-entropy supervision, such corrupted labels mislead deep neural networks, causing severe overfitting to noise and hindering generalization.
To mitigate these challenges, recent works have explored label distribution learning (LDL) and neighborhood relational modeling. These methods retrieve semantic neighbors in an auxiliary space (e.g., facial landmarks or continuous emotional dimensions) and aggregate their predictive soft distributions to construct smoother and more fault-tolerant supervision. Nevertheless, existing paradigms almost universally depend on static, single-pass heuristic weighting schemes (such as raw Euclidean distance or cosine similarity). When severe label noise corrupts the feature representations, initial neighborhood relations become unreliable, and single-pass static aggregation not only fails to suppress noise but actively amplifies corrupted predictions across the neighborhood.
This failure highlights the fundamental need to reformulate neighbor supervision modeling as a sequential, dynamic optimization process rather than a static assignment. The core idea is to identify nearest neighbors in the auxiliary continuous ValenceโArousal (VโA) space, initialize contribution weights through multi-route Cross-Attention Routing (CAR) for robust local semantic alignment, and iteratively refine neighbor weights via a Mamba-based state space model that balances historical states with newly absorbed distributions, dynamically fused with the original label to build reliable noise-robust supervision.
Method¶
Overall Architecture¶
The pipeline of Mamba-NWR operates in four main stages: auxiliary neighborhood construction in the continuous ValenceโArousal (VโA) space, fine-grained initial weight estimation via Cross-Attention Routing, iterative weight and distribution refinement powered by a Mamba state space model, and dynamic fusion of the refined neighbor consensus with the original logical label. After deep feature extraction, the center sample and its \(K\) retrieved neighbors undergo multi-route alignment and multi-step state updates to eliminate the impact of misleading neighbors, synthesizing clean supervision for end-to-end training. Crucially, during inference, all neighborhood modeling and refinement modules are discarded, leaving only the standard ResNet-18 backbone with zero added latency.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image and Neighbors<br/>x_i and K neighbors x_{ik}"] --> B["Auxiliary Space Retrieval<br/>KNN in Valence-Arousal Space"]
B --> C["Cross-Attention Routing<br/>Multi-route local feature alignment"]
C --> D["Mamba State Space Refinement<br/>Historical state retention and input absorption"]
D --> E["Dynamic Fusion<br/>Adaptive balance between self-label and neighbors"]
E --> F["Joint Noise-Robust Loss<br/>Supervised cross-entropy + discriminative loss"]
Key Designs¶
1. Cross-Attention Routing: Multi-route local semantic alignment overcoming coarse distance metrics In the continuous two-dimensional ValenceโArousal (VโA) affective space, pseudo VโA representations are estimated via a pretrained network, and the \(K=8\) nearest neighbors are retrieved using KNN. Conventional approaches directly compute weights from low-dimensional Euclidean distances, which often fail when coarse representations are perturbed by noise. To establish an informative starting point, the Cross-Attention Routing (CAR) module projects the deep features of the center sample and each neighbor into \(N_l\) local feature routes: $$ \mathbf{U}i \in \mathbb{R}^{N_l \times d_l}, \quad \mathbf{U} $$ Following vector normalization across each route, local representations are mapped into high-level semantic categories via a shared transformation tensor } \in \mathbb{R}^{N_l \times d_l\(\mathbf{W}_{n,m} \in \mathbb{R}^{N_l \times N_h \times d_l \times d_h}\). For each local route \(n\), semantic agreement between the center sample and the \(k\)-th neighbor is quantified by scaled dot-product similarity and normalized across neighbors: $$ \alpha_{i_k}^{(n)} = \frac{\exp\left(\langle \hat{\mathbf{U}}i^{(n)}, \hat{\mathbf{U}}}^{(n)} \rangle / \sqrt{d_h}\right)}{\sum_{j=1}^K \exp\left(\langle \hat{\mathbf{U}i^{(n)}, \hat{\mathbf{U}} $$ The initial contribution weight is obtained by averaging across all }^{(n)} \rangle / \sqrt{d_h}\right)\(N_l\) local routes: \(w_{i_k}^{(0)} = \frac{1}{N_l}\sum_{n=1}^{N_l} \alpha_{i_k}^{(n)}\), satisfying \(\sum_{k=1}^K w_{i_k}^{(0)} = 1\). By decomposing holistic facial features into multiple functional routes, CAR evaluates subtle emotional consistency, providing a robust initialization resilient to ambiguous global features.
2. Mamba State Space Refinement: Historical state retention and progressive error correction Standard neighborhood aggregation relies on a single forward pass, leaving erroneous high-weight assignments uncorrected. Mamba-NWR formulates neighbor contribution recalibration as the temporal evolution of a discrete state space system. At iteration \(t\), the current aggregated neighborhood label distribution is computed as: $$ \mathbf{p}i^{(t)} = \sum $$ Projecting the weight hidden state }^K w_{i_k}^{(t)} \hat{\mathbf{y}}_{i_k\(\mathbf{w}_i^{(t)}\) and the semantic distribution into dimension \(d_z\), the Mamba module governs state transitions through a selective recurrence formula: $$ \mathbf{w}_i^{(t+1)} = \mathbf{A} \odot \mathbf{w}_i^{(t)} + \mathbf{B} \odot \mathbf{p}_i^{(t)} $$ where \(\mathbf{A}\) and \(\mathbf{B}\) are learnable parameters controlling memory retention and input integration, respectively, and \(\odot\) denotes element-wise multiplication. Intuitively, \(\mathbf{A}\) stabilizes past inference by preserving valid historical weighting, while \(\mathbf{B}\) incorporates the latest global neighborhood context to adjust neighbor credibility. Over \(T=3\) iterations, anomalous and noisy neighbors are progressively down-weighted, yielding an optimized neighborhood distribution \(\mathbf{p}_i^{\text{nbr}} = \sum_{k=1}^K w_{i_k}^{(T)} \hat{\mathbf{y}}_{i_k}\).
3. Dynamic Fusion: Adaptive trade-off between self-information and neighborhood consensus Because an aggregated neighborhood distribution is not universally superior to the original annotation (especially when clean self-features are unambiguous), an instance-specific fusion factor \(\lambda_i \in [0, 1]\) is generated by a lightweight prediction head and trained end-to-end. The final synthesized soft target is defined as: $$ \tilde{\mathbf{p}}_i = \lambda_i \mathbf{y}_i + (1 - \lambda_i) \mathbf{p}_i^{\text{nbr}} $$ When the center sample exhibits clean, high-confidence facial features that conflict with neighboring predictions, a larger \(\lambda_i\) encourages reliance on the original logical label \(\mathbf{y}_i\). Conversely, when the sample is heavily corrupted or resides near an ambiguous decision boundary, a smaller \(\lambda_i\) shifts supervision toward the collective consensus of its verified neighbors.
Loss & Training¶
The overall training objective combines soft-target cross-entropy with a discriminative feature regularization term. The classification loss is defined as: $$ \mathcal{L}{\text{CE}} = - \sum}^N \sum_{c=1}^C \tilde{p{i,c} \log \hat{y} $$ To enforce intra-class compactness and inter-class separability under noisy conditions, a discriminative center loss is added: $$ \mathcal{L}{\text{dis}} = \sum}^N \left( |\mathbf{fi - \boldsymbol{\mu}|2^2 - \frac{1}{C-1} \sum} |\mathbf{fi - \boldsymbol{\mu}_c|_2^2 \right) $$ where \(\boldsymbol{\mu}_c\) denotes the feature centroid of emotion category \(c\). The joint objective is: $$ \mathcal{L} = \mathcal{L} $$ with balancing coefficient }} + \beta \mathcal{L}_{\text{dis}\(\beta = 0.1\). The model is trained with Adam optimizer at a batch size of 32 using an ImageNet-pretrained ResNet-18 backbone.
Key Experimental Results¶
Main Results¶
Under symmetric noise settings, synthetic label noise is injected into training sets at rates of 10%, 20%, and 30%. Extensive evaluations across RAF-DB, AffectNet, and FERPlus benchmark datasets validate that Mamba-NWR consistently outperforms existing robust FER baselines:
| Dataset | Noise Ratio | Ours | SCN | RUL | EAC | LDLVA | MCR | Gain vs Prev. Best |
|---|---|---|---|---|---|---|---|---|
| RAF-DB | 10% | 89.97% | 82.18% | 86.17% | 88.02% | 87.98% | 89.28% | +0.69% vs MCR |
| RAF-DB | 20% | 88.38% | 80.10% | 84.32% | 86.05% | 86.81% | 87.61% | +0.77% vs MCR |
| RAF-DB | 30% | 87.62% | 77.46% | 82.06% | 84.42% | 85.85% | 86.08% | +1.54% vs MCR |
| AffectNet | 10% | 66.18% | 58.58% | 60.54% | 61.11% | 64.37% | 62.33% | +1.81% vs LDLVA |
| AffectNet | 20% | 65.66% | 57.25% | 59.01% | 60.29% | 63.89% | 61.35% | +1.77% vs LDLVA |
| AffectNet | 30% | 63.91% | 55.05% | 56.93% | 58.91% | 62.57% | 59.50% | +1.34% vs LDLVA |
| FERPlus | 10% | 88.11% | 84.28% | 86.93% | 87.03% | - | - | +0.09% vs LA-Net |
| FERPlus | 20% | 87.62% | 83.17% | 85.05% | 86.07% | - | - | +0.77% vs LA-Net |
| FERPlus | 30% | 87.15% | 82.47% | 83.90% | 85.44% | - | - | +1.14% vs LA-Net |
Under realistic asymmetric noise on RAF-DB (where fear transitions to surprise and disgust transitions to anger based on confusing visual pairs), Mamba-NWR achieves 89.30% / 87.44% / 83.56% at 10%, 20%, and 30% noise ratios, outperforming prior SOTA SOFT (88.01% / 86.35% / 81.92%) by 1.29% / 1.09% / 1.64%, respectively.
Ablation Study¶
Detailed component ablations on RAF-DB, AffectNet, and FERPlus highlight the relative impact of initialization, iterative recurrence, and fusion strategies:
| Ablation Config / Component | 10% Noise | 20% Noise | 30% Noise | Note |
|---|---|---|---|---|
| RAF-DB: VโA Distance Init | 88.04% | 87.51% | 85.96% | Coarse Euclidean distance degrades under heavy label noise |
| RAF-DB: Feature Similarity Init | 88.43% | 87.45% | 86.16% | Global cosine similarity lacks localized multi-route nuance |
| RAF-DB: CAR Routing Init (Ours) | 89.97% | 88.38% | 87.62% | Multi-route cross-attention provides a stable initial state |
| RAF-DB: Without Mamba Refinement | 88.15% | 86.81% | 86.43% | Static weighting accumulates errors, dropping 1.19% at 30% noise |
| RAF-DB: With Mamba Refinement (Ours) | 89.97% | 88.38% | 87.62% | State space updates filter unreliable neighbors effectively |
| RAF-DB: Mamba Iterations T=1 | 87.68% | 87.22% | 86.36% | Single step fails to assimilate global context |
| RAF-DB: Mamba Iterations T=2 | 88.19% | 87.73% | 87.15% | Progressive improvements observed |
| RAF-DB: Mamba Iterations T=3 (Ours) | 89.97% | 88.38% | 87.62% | Optimal performance across all noise rates |
| RAF-DB: Mamba Iterations T=4 | 89.21% | 88.10% | 87.34% | Slight over-smoothing and redundant update drift |
| RAF-DB: Fixed ฮป = 1 (Center Only) | 84.28% | 82.09% | 80.12% | Fully exposed to noisy labels, severe degradation |
| RAF-DB: Fixed ฮป = 0 (Neighbors Only) | 88.35% | 87.51% | 86.74% | Ignores individual identity cues, over-smoothed |
| RAF-DB: Learnable Dynamic ฮป (Ours) | 89.97% | 88.38% | 87.62% | Balances personal confidence with neighbor consensus |
Key Findings¶
- Error-correction gains increase with noise severity: Adding Mamba refinement produces a +1.82% gain on RAF-DB at 10% noise and prevents sharp degradation on AffectNet at 30% noise (62.72% without Mamba vs. 63.91% with Mamba), proving that sequential state recurrence prevents error propagation.
- Dynamic fusion acts as an automated safety valve: Relying solely on the original label drops accuracy to 80.12% under 30% noise, whereas trusting only neighbors yields 86.74%. The learnable dynamic parameter \(\lambda_i\) adaptively regulates reliance, hitting 87.62%.
Highlights & Insights¶
- Formulating LDL as a state space time-series process: Instead of static neighbor graph convolutions or heuristics, Mamba-NWR models neighborhood soft-label generation as an iterative state space evolution, elegantly preventing initial matching noise from propagating into downstream training.
- Zero inference footprint: All graph retrieval, cross-attention projection, and state space recurrence are executed strictly during backpropagation to reconstruct soft targets. The deployed inference model remains an unmodified ResNet-18, delivering high throughput.
- General applicability to noisy visual tasks: The combination of multi-route local alignment and state space dynamic weighting provides a clean blueprint for other weakly labeled or crowdsourced computer vision tasks, including medical grading and long-tailed fine-grained classification.
Limitations & Future Work¶
- Dependence on auxiliary VโA estimation: The framework presumes that meaningful neighborhoods can be discovered in the continuous ValenceโArousal space. When the external VโA estimator produces corrupted coordinates under extreme head pose or heavy occlusion, the retrieved neighbor set may be flawed.
- Future direction: Developing an end-to-end jointly optimizable affective manifold or combining continuous VโA coordinates with discrete Action Unit (AU) activations to capture physiological facial muscle movements.
Related Work & Insights¶
- vs LDLVA (Chen et al., CVPR 2020 / Le et al., WACV 2023): LDLVA constructs static label distributions over VโA graphs. Mamba-NWR discovers that static weights are vulnerable to noisy topological connections, introducing CAR and Mamba iterative correction to gain +1.81% on AffectNet.
- vs LA-Net (Wu & Cui, ICCV 2023): LA-Net extracts landmark geometric features to define neighborhoods. Mamba-NWR demonstrates that continuous affective space coordinates combined with sequence-level state space refinement achieve superior noise robustness without requiring landmark detectors.
Rating¶
- Novelty: โญโญโญโญโ [Pioneering integration of Mamba state space recurrence into neighborhood weight refinement for label distribution learning]
- Experimental Thoroughness: โญโญโญโญโญ [Extensive validation across 3 standard benchmarks, 10%-30% symmetric and asymmetric noise, accompanied by rigorous ablations]
- Writing Quality: โญโญโญโญโ [Clear formulation, well-defined mathematical framework, and coherent experimental evidence]
- Value: โญโญโญโญโญ [Provides a practical, zero-inference-overhead paradigm for robust visual affective computing under real-world noise]