Skip to content

Noise-Robust Facial Expression Recognition via Mamba-driven Neighbor Weight Refinement

Conference: ECCV 2026
Paper: ECCV Official Page
Area: Human Understanding
Keywords: Facial Expression Recognition, Label Distribution Learning, Neighbor Weight Refinement, State Space Model, Label Noise

TL;DR

To combat crowdsourced annotation ambiguity and label noise in in-the-wild facial expression recognition, Mamba-NWR combines cross-attention routing initialization with Mamba state space iterative refinement to dynamically balance historical states and aggregate robust neighborhood label distributions.

Background & Motivation

Facial expression recognition (FER) in unconstrained, in-the-wild scenarios is a fundamental building block for affective computing and natural human-computer interaction. However, real-world FER inherently suffers from subjective annotation ambiguity: subtle muscle activations and transitional facial expressions are interpreted differently by different human annotators. Furthermore, large-scale benchmarks typically rely on crowdsourced annotations that inevitably introduce systematic label noise and class inconsistencies. Under standard hard-target cross-entropy supervision, such corrupted labels mislead deep neural networks, causing severe overfitting to noise and hindering generalization.

To mitigate these challenges, recent works have explored label distribution learning (LDL) and neighborhood relational modeling. These methods retrieve semantic neighbors in an auxiliary space (e.g., facial landmarks or continuous emotional dimensions) and aggregate their predictive soft distributions to construct smoother and more fault-tolerant supervision. Nevertheless, existing paradigms almost universally depend on static, single-pass heuristic weighting schemes (such as raw Euclidean distance or cosine similarity). When severe label noise corrupts the feature representations, initial neighborhood relations become unreliable, and single-pass static aggregation not only fails to suppress noise but actively amplifies corrupted predictions across the neighborhood.

This failure highlights the fundamental need to reformulate neighbor supervision modeling as a sequential, dynamic optimization process rather than a static assignment. The core idea is to identify nearest neighbors in the auxiliary continuous Valenceโ€“Arousal (Vโ€“A) space, initialize contribution weights through multi-route Cross-Attention Routing (CAR) for robust local semantic alignment, and iteratively refine neighbor weights via a Mamba-based state space model that balances historical states with newly absorbed distributions, dynamically fused with the original label to build reliable noise-robust supervision.

Method

Overall Architecture

The pipeline of Mamba-NWR operates in four main stages: auxiliary neighborhood construction in the continuous Valenceโ€“Arousal (Vโ€“A) space, fine-grained initial weight estimation via Cross-Attention Routing, iterative weight and distribution refinement powered by a Mamba state space model, and dynamic fusion of the refined neighbor consensus with the original logical label. After deep feature extraction, the center sample and its \(K\) retrieved neighbors undergo multi-route alignment and multi-step state updates to eliminate the impact of misleading neighbors, synthesizing clean supervision for end-to-end training. Crucially, during inference, all neighborhood modeling and refinement modules are discarded, leaving only the standard ResNet-18 backbone with zero added latency.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Image and Neighbors<br/>x_i and K neighbors x_{ik}"] --> B["Auxiliary Space Retrieval<br/>KNN in Valence-Arousal Space"]
    B --> C["Cross-Attention Routing<br/>Multi-route local feature alignment"]
    C --> D["Mamba State Space Refinement<br/>Historical state retention and input absorption"]
    D --> E["Dynamic Fusion<br/>Adaptive balance between self-label and neighbors"]
    E --> F["Joint Noise-Robust Loss<br/>Supervised cross-entropy + discriminative loss"]

Key Designs

1. Cross-Attention Routing: Multi-route local semantic alignment overcoming coarse distance metrics In the continuous two-dimensional Valenceโ€“Arousal (Vโ€“A) affective space, pseudo Vโ€“A representations are estimated via a pretrained network, and the \(K=8\) nearest neighbors are retrieved using KNN. Conventional approaches directly compute weights from low-dimensional Euclidean distances, which often fail when coarse representations are perturbed by noise. To establish an informative starting point, the Cross-Attention Routing (CAR) module projects the deep features of the center sample and each neighbor into \(N_l\) local feature routes: $$ \mathbf{U}i \in \mathbb{R}^{N_l \times d_l}, \quad \mathbf{U} $$ Following vector normalization across each route, local representations are mapped into high-level semantic categories via a shared transformation tensor } \in \mathbb{R}^{N_l \times d_l\(\mathbf{W}_{n,m} \in \mathbb{R}^{N_l \times N_h \times d_l \times d_h}\). For each local route \(n\), semantic agreement between the center sample and the \(k\)-th neighbor is quantified by scaled dot-product similarity and normalized across neighbors: $$ \alpha_{i_k}^{(n)} = \frac{\exp\left(\langle \hat{\mathbf{U}}i^{(n)}, \hat{\mathbf{U}}}^{(n)} \rangle / \sqrt{d_h}\right)}{\sum_{j=1}^K \exp\left(\langle \hat{\mathbf{U}i^{(n)}, \hat{\mathbf{U}} $$ The initial contribution weight is obtained by averaging across all }^{(n)} \rangle / \sqrt{d_h}\right)\(N_l\) local routes: \(w_{i_k}^{(0)} = \frac{1}{N_l}\sum_{n=1}^{N_l} \alpha_{i_k}^{(n)}\), satisfying \(\sum_{k=1}^K w_{i_k}^{(0)} = 1\). By decomposing holistic facial features into multiple functional routes, CAR evaluates subtle emotional consistency, providing a robust initialization resilient to ambiguous global features.

2. Mamba State Space Refinement: Historical state retention and progressive error correction Standard neighborhood aggregation relies on a single forward pass, leaving erroneous high-weight assignments uncorrected. Mamba-NWR formulates neighbor contribution recalibration as the temporal evolution of a discrete state space system. At iteration \(t\), the current aggregated neighborhood label distribution is computed as: $$ \mathbf{p}i^{(t)} = \sum $$ Projecting the weight hidden state }^K w_{i_k}^{(t)} \hat{\mathbf{y}}_{i_k\(\mathbf{w}_i^{(t)}\) and the semantic distribution into dimension \(d_z\), the Mamba module governs state transitions through a selective recurrence formula: $$ \mathbf{w}_i^{(t+1)} = \mathbf{A} \odot \mathbf{w}_i^{(t)} + \mathbf{B} \odot \mathbf{p}_i^{(t)} $$ where \(\mathbf{A}\) and \(\mathbf{B}\) are learnable parameters controlling memory retention and input integration, respectively, and \(\odot\) denotes element-wise multiplication. Intuitively, \(\mathbf{A}\) stabilizes past inference by preserving valid historical weighting, while \(\mathbf{B}\) incorporates the latest global neighborhood context to adjust neighbor credibility. Over \(T=3\) iterations, anomalous and noisy neighbors are progressively down-weighted, yielding an optimized neighborhood distribution \(\mathbf{p}_i^{\text{nbr}} = \sum_{k=1}^K w_{i_k}^{(T)} \hat{\mathbf{y}}_{i_k}\).

3. Dynamic Fusion: Adaptive trade-off between self-information and neighborhood consensus Because an aggregated neighborhood distribution is not universally superior to the original annotation (especially when clean self-features are unambiguous), an instance-specific fusion factor \(\lambda_i \in [0, 1]\) is generated by a lightweight prediction head and trained end-to-end. The final synthesized soft target is defined as: $$ \tilde{\mathbf{p}}_i = \lambda_i \mathbf{y}_i + (1 - \lambda_i) \mathbf{p}_i^{\text{nbr}} $$ When the center sample exhibits clean, high-confidence facial features that conflict with neighboring predictions, a larger \(\lambda_i\) encourages reliance on the original logical label \(\mathbf{y}_i\). Conversely, when the sample is heavily corrupted or resides near an ambiguous decision boundary, a smaller \(\lambda_i\) shifts supervision toward the collective consensus of its verified neighbors.

Loss & Training

The overall training objective combines soft-target cross-entropy with a discriminative feature regularization term. The classification loss is defined as: $$ \mathcal{L}{\text{CE}} = - \sum}^N \sum_{c=1}^C \tilde{p{i,c} \log \hat{y} $$ To enforce intra-class compactness and inter-class separability under noisy conditions, a discriminative center loss is added: $$ \mathcal{L}{\text{dis}} = \sum}^N \left( |\mathbf{fi - \boldsymbol{\mu}|2^2 - \frac{1}{C-1} \sum} |\mathbf{fi - \boldsymbol{\mu}_c|_2^2 \right) $$ where \(\boldsymbol{\mu}_c\) denotes the feature centroid of emotion category \(c\). The joint objective is: $$ \mathcal{L} = \mathcal{L} $$ with balancing coefficient }} + \beta \mathcal{L}_{\text{dis}\(\beta = 0.1\). The model is trained with Adam optimizer at a batch size of 32 using an ImageNet-pretrained ResNet-18 backbone.

Key Experimental Results

Main Results

Under symmetric noise settings, synthetic label noise is injected into training sets at rates of 10%, 20%, and 30%. Extensive evaluations across RAF-DB, AffectNet, and FERPlus benchmark datasets validate that Mamba-NWR consistently outperforms existing robust FER baselines:

Dataset Noise Ratio Ours SCN RUL EAC LDLVA MCR Gain vs Prev. Best
RAF-DB 10% 89.97% 82.18% 86.17% 88.02% 87.98% 89.28% +0.69% vs MCR
RAF-DB 20% 88.38% 80.10% 84.32% 86.05% 86.81% 87.61% +0.77% vs MCR
RAF-DB 30% 87.62% 77.46% 82.06% 84.42% 85.85% 86.08% +1.54% vs MCR
AffectNet 10% 66.18% 58.58% 60.54% 61.11% 64.37% 62.33% +1.81% vs LDLVA
AffectNet 20% 65.66% 57.25% 59.01% 60.29% 63.89% 61.35% +1.77% vs LDLVA
AffectNet 30% 63.91% 55.05% 56.93% 58.91% 62.57% 59.50% +1.34% vs LDLVA
FERPlus 10% 88.11% 84.28% 86.93% 87.03% - - +0.09% vs LA-Net
FERPlus 20% 87.62% 83.17% 85.05% 86.07% - - +0.77% vs LA-Net
FERPlus 30% 87.15% 82.47% 83.90% 85.44% - - +1.14% vs LA-Net

Under realistic asymmetric noise on RAF-DB (where fear transitions to surprise and disgust transitions to anger based on confusing visual pairs), Mamba-NWR achieves 89.30% / 87.44% / 83.56% at 10%, 20%, and 30% noise ratios, outperforming prior SOTA SOFT (88.01% / 86.35% / 81.92%) by 1.29% / 1.09% / 1.64%, respectively.

Ablation Study

Detailed component ablations on RAF-DB, AffectNet, and FERPlus highlight the relative impact of initialization, iterative recurrence, and fusion strategies:

Ablation Config / Component 10% Noise 20% Noise 30% Noise Note
RAF-DB: Vโ€“A Distance Init 88.04% 87.51% 85.96% Coarse Euclidean distance degrades under heavy label noise
RAF-DB: Feature Similarity Init 88.43% 87.45% 86.16% Global cosine similarity lacks localized multi-route nuance
RAF-DB: CAR Routing Init (Ours) 89.97% 88.38% 87.62% Multi-route cross-attention provides a stable initial state
RAF-DB: Without Mamba Refinement 88.15% 86.81% 86.43% Static weighting accumulates errors, dropping 1.19% at 30% noise
RAF-DB: With Mamba Refinement (Ours) 89.97% 88.38% 87.62% State space updates filter unreliable neighbors effectively
RAF-DB: Mamba Iterations T=1 87.68% 87.22% 86.36% Single step fails to assimilate global context
RAF-DB: Mamba Iterations T=2 88.19% 87.73% 87.15% Progressive improvements observed
RAF-DB: Mamba Iterations T=3 (Ours) 89.97% 88.38% 87.62% Optimal performance across all noise rates
RAF-DB: Mamba Iterations T=4 89.21% 88.10% 87.34% Slight over-smoothing and redundant update drift
RAF-DB: Fixed ฮป = 1 (Center Only) 84.28% 82.09% 80.12% Fully exposed to noisy labels, severe degradation
RAF-DB: Fixed ฮป = 0 (Neighbors Only) 88.35% 87.51% 86.74% Ignores individual identity cues, over-smoothed
RAF-DB: Learnable Dynamic ฮป (Ours) 89.97% 88.38% 87.62% Balances personal confidence with neighbor consensus

Key Findings

  • Error-correction gains increase with noise severity: Adding Mamba refinement produces a +1.82% gain on RAF-DB at 10% noise and prevents sharp degradation on AffectNet at 30% noise (62.72% without Mamba vs. 63.91% with Mamba), proving that sequential state recurrence prevents error propagation.
  • Dynamic fusion acts as an automated safety valve: Relying solely on the original label drops accuracy to 80.12% under 30% noise, whereas trusting only neighbors yields 86.74%. The learnable dynamic parameter \(\lambda_i\) adaptively regulates reliance, hitting 87.62%.

Highlights & Insights

  • Formulating LDL as a state space time-series process: Instead of static neighbor graph convolutions or heuristics, Mamba-NWR models neighborhood soft-label generation as an iterative state space evolution, elegantly preventing initial matching noise from propagating into downstream training.
  • Zero inference footprint: All graph retrieval, cross-attention projection, and state space recurrence are executed strictly during backpropagation to reconstruct soft targets. The deployed inference model remains an unmodified ResNet-18, delivering high throughput.
  • General applicability to noisy visual tasks: The combination of multi-route local alignment and state space dynamic weighting provides a clean blueprint for other weakly labeled or crowdsourced computer vision tasks, including medical grading and long-tailed fine-grained classification.

Limitations & Future Work

  • Dependence on auxiliary Vโ€“A estimation: The framework presumes that meaningful neighborhoods can be discovered in the continuous Valenceโ€“Arousal space. When the external Vโ€“A estimator produces corrupted coordinates under extreme head pose or heavy occlusion, the retrieved neighbor set may be flawed.
  • Future direction: Developing an end-to-end jointly optimizable affective manifold or combining continuous Vโ€“A coordinates with discrete Action Unit (AU) activations to capture physiological facial muscle movements.
  • vs LDLVA (Chen et al., CVPR 2020 / Le et al., WACV 2023): LDLVA constructs static label distributions over Vโ€“A graphs. Mamba-NWR discovers that static weights are vulnerable to noisy topological connections, introducing CAR and Mamba iterative correction to gain +1.81% on AffectNet.
  • vs LA-Net (Wu & Cui, ICCV 2023): LA-Net extracts landmark geometric features to define neighborhoods. Mamba-NWR demonstrates that continuous affective space coordinates combined with sequence-level state space refinement achieve superior noise robustness without requiring landmark detectors.

Rating

  • Novelty: โญโญโญโญโ˜† [Pioneering integration of Mamba state space recurrence into neighborhood weight refinement for label distribution learning]
  • Experimental Thoroughness: โญโญโญโญโญ [Extensive validation across 3 standard benchmarks, 10%-30% symmetric and asymmetric noise, accompanied by rigorous ablations]
  • Writing Quality: โญโญโญโญโ˜† [Clear formulation, well-defined mathematical framework, and coherent experimental evidence]
  • Value: โญโญโญโญโญ [Provides a practical, zero-inference-overhead paradigm for robust visual affective computing under real-world noise]