Skip to content

Learning to Mask: Cross-Modal Noise Modulation for Hallucination Mitigation in Multi-modal Large Language Models

Conference: ECCV 2026
Paper: ECCV Official
Code: None
Area: Multimodal VLM / Hallucination Mitigation
Keywords: Multimodal Large Language Models, Hallucination Mitigation, Cross-Modal Noise Modulation, Dynamic Value Cache Masking, Langevin Dynamics Regularization

TL;DR

Addressing the issue where redundant features trigger hallucinations during both the prefill and autoregressive decoding phases of Multimodal Large Language Models (MLLMs), this paper proposes HalMask, a lightweight real-time self-correction framework that adaptively injects Gaussian noise into multimodal features and Value Caches via cross-modal modulation modules and Langevin dynamics regularization, substantially reducing hallucinations while preserving critical semantic cues.

Background & Motivation

Multimodal Large Language Models (MLLMs) have demonstrated impressive breakthroughs across vision-language tasks including image captioning, visual question answering, and complex multimodal reasoning. Despite their proficiency in processing visual inputs, they remain acutely vulnerable to hallucinations—a critical failure mode where generated textual outputs diverge from the actual visual evidence. Mitigating hallucinations is therefore a pivotal milestone toward realizing robust and trustworthy multimodal intelligence. Existing mitigation strategies have predominantly relied on post-hoc self-refinement and specialized decoding algorithms (such as contrastive decoding and introspective backtracking). However, self-refinement typically necessitates external high-performance auxiliary models across iterative loops, while specialized decoding schemes require multi-round forward iterations and repetitive backtracking during generation. These dependencies introduce prohibitive latency and throughput bottlenecks, severely impairing real-time practical utility.

A systematic analysis of hidden states across the entire MLLM inference pipeline reveals an underlying root cause: during autoregressive generation, previously generated tokens accumulate biased contextual signals that propagate and amplify errors in subsequent tokens. More crucially, hallucination-inducing redundant features reside not only within the visual tokens during the prefill stage, but also persist and propagate directly into the Image Value Cache and Answer Value Cache throughout decoding. Empirical investigations uncover that injecting calibrated Gaussian noise into these redundant features can effectively suppress hallucinated content. Nevertheless, naive noise injection without precise control risks corrupting salient visual and textual cues, thereby unintentionally aggravating hallucinations.

The angle of attack in this paper is to bypass costly external verifiers and repetitive re-decoding, instead employing lightweight learned modules to dynamically gauge cross-modal redundancy and inject controlled stochastic noise into both hidden states and attention caches. Core idea: propose the HalMask framework comprising Image and Text Noise Modulation Modules regularized by an overdamped Langevin dynamics objective to prevent semantic collapse, which injects adaptive Gaussian noise into visual tokens during prefill and dynamically filters the multi-layer Value Caches during decoding for lightweight, real-time self-correction.

Method

Overall Architecture

HalMask abstracts hallucination mitigation into a process of controlled feature compression and stochastic noise smoothing. Its pipeline spans a modular training phase and a two-stage dynamic masking inference workflow. During training, HalMask introduces the Image Noise Modulation Module (In-Mod) and the Text Noise Modulation Module (Tn-Mod), optimized on synthetic image-text-hallucination triplets using cross-attention and fusion networks to predict dimension-wise continuous control coefficients. A physics-inspired overdamped Langevin dynamics regularization term is incorporated into the loss function, encouraging the redundant feature distribution to smooth toward standard normal noise while strictly preserving latent topological structure. During inference, HalMask perturbs redundant visual tokens via In-Mod in the prefill stage, and dynamically updates the multi-layer Image and Answer Value Caches via Dynamic Value Cache Masking (VCM) across all decoder layers during autoregressive decoding.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Embeddings<br/>Image Tokens Xv + Prompt/Generated Text"] --> B["Cross-Modal Noise Modulation Modules<br/>In-Mod & Tn-Mod Cross-Attention Interaction"]
    B --> C["Langevin Dynamics Regularization<br/>Overdamped Equation & Sliced Wasserstein Distance"]
    C --> D["Two-Stage Masking Inference Strategy<br/>Prefill Phase Visual Token Feature Smoothing"]
    D --> E["Dynamic Value Cache Masking Mechanism<br/>Decoding Phase Real-Time Multi-Layer Cache Filtering"]
    E --> F["Hallucination-Free High-Fidelity Output"]

Key Designs

1. Cross-Modal Noise Modulation Modules: Generating Adaptive Control Coefficients Because multimodal representations exhibit high dimensionality with non-uniform redundancy, hard thresholding or uniform pruning often damages critical visual details. HalMask devises specialized Image (In-Mod) and Text (Tn-Mod) Noise Modulation Modules. For the visual branch, In-Mod utilizes visual tokens \(\mathbf{X}_v \in \mathbb{R}^{N \times d}\) as attention queries and text tokens (ground-truth text \(\mathbf{X}_{gt}\) in training; prompt or generated tokens in inference) as keys and values, mapping variable-length textual guidance onto visual token dimensions via scaled dot-product attention: $\(\mathbf{X}_{gt}^{attn} = \mathrm{Softmax}\left(\frac{\mathbf{X}_v \mathbf{X}_{gt}^T}{\sqrt{d}}\right) \mathbf{X}_{gt}\)$ The aligned text features and original visual tokens are concatenated channel-wise into \(\mathbf{X}_{merge} \in \mathbb{R}^{N \times 2d}\) and passed through Fusion-Net (five linear layers with ReLU activations) capped by a Sigmoid layer, generating continuous control coefficients \(\Lambda_{image} = \{\lambda_1^v, \dots, \lambda_N^v\} \in \mathbb{R}^{N \times d}\). Symmetrically, Tn-Mod employs text tokens as queries and visual tokens as keys/values to produce text control coefficients \(\Lambda_{text} \in \mathbb{R}^{K \times d}\). This cross-modal reference enables each modality to accurately single out isolated, unaligned redundant dimensions.

2. Langevin Dynamics Regularization: Preventing Semantic Over-Smoothing via Dynamical Priors Directly constraining feature distributions toward a Gaussian prior using traditional Variational Information Bottleneck (VIB) objectives often induces excessive regularization, destroying discriminative and subtle semantics. HalMask resolves this from a dynamical systems perspective by adopting the overdamped Langevin equation to govern feature evolution: $\(dx = \left( -M(x)\nabla V(x) + \mathrm{div}(M(x)) \right) dt + \sqrt{2} M^{1/2}(x) dW_t\)$ where \(M(x)\) represents the diffusion network, and \(-\nabla V(x)\) is parameterized by the force network. Discretized via the Euler-Maruyama scheme with \(dt=1\), the dynamics guide redundant features toward standard normal noise without collapsing salient features. To measure distributional divergence along random projections on the unit sphere \(S^{d-1}\), HalMask constructs a dynamic loss via Sliced Wasserstein Distance (SWD): $\(\mathcal{L}_{dynamic} = \frac{1}{L} \sum_{l=1}^{L} \left( \frac{1}{N} \sum_{n=1}^{N} |a_{(n),l}^v - b_{(n),l}^v| + \frac{1}{K} \sum_{k=1}^{K} |a_{(k),l}^{hal} - b_{(k),l}^{hal}| \right)\)$ Coupled with a cosine similarity supervision term \(\mathcal{L}_{sup}\) ensuring alignment with ground-truth semantics, this objective drives the modules to eliminate high-order redundancies while preserving latent representation topology (substantially reducing Representation Topology Divergence).

3. Dynamic Value Cache Masking Mechanism: Accumulation-Free Real-Time Cache Filtering In autoregressive generation, repeatedly modifying past hidden states risks irreversible error accumulation and representation drift. HalMask introduces the Dynamic Value Cache Masking (VCM) mechanism. At generation step \(t\), the newly emitted token is concatenated with prior answers into context \(\mathbf{X}_a\) and fed into In-Mod and Tn-Mod to dynamically compute current modulation coefficients \(\Lambda_{image}\) and \(\Lambda_{text}\). During attention computation at each transformer layer, Gaussian noise is injected on the fly: $\(\mathbf{C}_v \leftarrow \Lambda_{image} \odot \mathbf{C}_v + (1 - \Lambda_{image}) \odot \epsilon, \quad \mathbf{C}_t \leftarrow \Lambda_{text} \odot \mathbf{C}_t + (1 - \Lambda_{text}) \odot \epsilon\)$ Critically, the physical GPU memory retains the clean, unmasked Value representations throughout decoding; noise injection occurs solely within the transient computational graph of the current forward step. This design eliminates compounding noise accumulation across long horizons, allowing each token step to re-evaluate whole-sequence redundancy based on the latest global context.

4. Progressive Multi-Stage Masking Inference Strategy: Full-Pipeline Synergistic Mitigation Because the prefill and decoding phases exhibit distinct computational profiles, HalMask executes a customized two-stage pipeline. During the prefill stage, input instructions and visual tokens pass through In-Mod to generate \(\Lambda_{image}\), perturbing raw visual tokens before the LLM backbone via \(\mathbf{Z}_v = \Lambda_{image} \odot \mathbf{X}_v + (1 - \Lambda_{image}) \odot \epsilon\). This cuts off spurious hallucination triggers from low-level visual encoder patches at the source. Upon entering the decoding stage, the pipeline seamlessly shifts to multi-layer Value Cache modulation via VCM, adaptively filtering long-term visual memories and textual histories. Jointly, static input perception and dynamic generational memory are continually sanitized.

Loss & Training

During training, the vision encoder and LLM decoder backbones remain completely frozen; only In-Mod, Tn-Mod, Force Net, and Diffusion Net are trained. Optimization leverages synthetic triplets of images, ground-truth descriptions, and perturbed hallucinations (co-occurrence, object uncertainty, spatial positioning). The joint optimization loss is formulated as: $\(\mathcal{L}_{all} = \beta \mathcal{L}_{dynamic} + \mathcal{L}_{base} + \mathcal{L}_{sup} + \mathcal{L}_{ce}\)$ where \(\mathcal{L}_{ce}\) is standard cross-entropy, \(\mathcal{L}_{base}\) is the negative log-likelihood loss for diffusion and force networks, \(\mathcal{L}_{sup}\) maximizes cosine similarity with ground-truth features, and \(\beta\) is a balancing hyperparameter (ablation studies identify \(\beta=0.5\) as optimal).

Key Experimental Results

Main Results

On the MSCOCO benchmark for object hallucination evaluation (sentence-level \(\text{CHAIR}_S\) and instance-level \(\text{CHAIR}_I\), where lower is better; BLEU measures caption quality, where higher is better), HalMask outperforms all competing baselines across both LLaVA-1.5 and MiniGPT-4 backbones. For LLaVA-1.5, \(\text{CHAIR}_S\) drops from 21.20 to 7.60 (a 64.15% relative reduction), while BLEU improves to 19.48.

Backbone Method \(\text{CHAIR}_S \downarrow\) \(\text{CHAIR}_I \downarrow\) \(\text{BLEU} \uparrow\)
LLaVA-1.5-7B Greedy 21.20 6.84 16.25
Beam Search 19.40 6.55 17.24
Woodpecker 23.85 7.50 17.05
LURE 19.48 6.50 15.97
OPERA 17.50 6.07 16.02
HALC 16.90 5.72 16.02
HACL 18.27 5.90 17.09
DeGF 18.40 6.10 –
Nullu 15.20 5.30 15.69
HalMask (Ours) 7.60 3.06 19.48
MiniGPT-4-7B Greedy 32.40 12.20 14.57
Beam Search 19.50 6.84 15.99
Woodpecker 28.87 10.20 15.30
LURE 27.88 10.20 15.03
OPERA 29.70 11.96 14.82
HALC 25.20 9.42 14.91
HACL 24.47 9.57 15.84
Nullu 21.40 8.99 14.81
HalMask (Ours) 16.20 5.96 15.40

On the offline hallucination benchmark OPOPE, HalMask establishes superior Precision and \(F_{0.2}\) scores across Random, Popular, and Adversarial splits:

Backbone Method Random (\(F_{0.2} \uparrow\)) Popular (\(F_{0.2} \uparrow\)) Adversarial (\(F_{0.2} \uparrow\))
LLaVA-1.5 Greedy 96.42 89.71 85.22
OPERA 96.58 90.30 85.26
HALC 95.89 90.05 87.81
Nullu 96.05 92.36 86.97
HalMask (Ours) 96.67 92.79 90.04
MiniGPT-4 Greedy 94.25 88.53 87.32
OPERA 94.55 89.28 88.22
HALC 94.26 90.01 88.61
Nullu 94.82 92.19 89.21
HalMask (Ours) 95.62 91.91 90.70

In terms of inference efficiency, HalMask achieves a remarkable speed-accuracy trade-off: vanilla LLaVA-1.5 consumes 0.021s/token, contrastive decoding method HALC surges to 0.283s/token (with average POPE F1 of 76.73%), whereas HalMask requires only 0.034s/token (8.3× faster than HALC) while boosting POPE average accuracy and F1 to 87.72% and 87.52%, respectively.

Ablation Study

Ablation studies on CHAIR and OPOPE isolate the contribution of In-Mod and Tn-Mod. On LLaVA-1.5, applying In-Mod alone cuts \(\text{CHAIR}_S\) from 21.20 to 12.80; applying Tn-Mod alone reaches 20.40; their synergistic combination drives it down to 7.60, demonstrating strong cross-modal complementarity.

Config \(\text{CHAIR}_S \downarrow\) \(\text{CHAIR}_I \downarrow\) OPOPE Random (\(F_{0.2} \uparrow\)) OPOPE Popular (\(F_{0.2} \uparrow\)) OPOPE Adv (\(F_{0.2} \uparrow\)) Average \(F_{0.2} \uparrow\) Note
LLaVA-1.5 (Baseline) 21.20 6.84 96.42 89.71 85.22 90.45 Vanilla greedy decoding baseline
+ In-Mod 12.80 4.45 96.52 91.98 89.69 92.73 Visual token and visual cache masking only
+ Tn-Mod 20.40 6.74 96.79 89.31 87.30 91.13 Text Value Cache masking only
+ Tn-Mod + In-Mod (Full) 7.60 3.06 96.67 92.79 90.04 93.17 Full cross-modal noise modulation

Evaluating the impact of the Langevin dynamics regularization term on POPE benchmarks:

Config POPE Avg Recall \(\uparrow\) POPE Avg F1 \(\uparrow\) RTD Topology Divergence \(\downarrow\) Note
w/o. Dynamical Regularization 82.76% 87.11% 56.13 Naive Gaussian prior impairs representation topology
HalMask (w/ Dynamical Reg) 84.98% 87.52% 50.15 Average recall gains 2.22%, RTD drops by 5.98

Key Findings

  • Synergistic Dual-Modality Suppression: Visual redundancy filtering drives the bulk of object hallucination reduction (\(\text{CHAIR}_S\) dropped by 8.4 points), whereas text cache modulation arrests self-reinforcing linguistic bias, yielding super-linear mitigation when combined.
  • Topology Preservation via Dynamical Regularization: Incorporating Langevin dynamics increases POPE recall by 2.22% and reduces Representation Topology Divergence (RTD) from 56.13 to 50.15, confirming that physical diffusion priors effectively prevent the suppression of fine-grained factual entities.
  • Superior Computational Throughput: Contrasting with multi-pass contrastive decoding methods (such as HALC at 0.283s/token), HalMask incurs only 0.034s/token—a mere 13ms latency overhead over vanilla greedy decoding—rendering it highly practical for production deployment.

Highlights & Insights

  • Stochastic Noise Injection over Discrete Pruning: Rather than applying hard pruning thresholds that disrupt attention weight continuity, HalMask employs continuous control coefficients to smoothly interpolate Gaussian noise into redundant features, enabling soft semantic forgetting.
  • Stateless Value Cache Masking Architecture: Persisting pure uncorrupted hidden states in physical memory while applying noise masking strictly within the transient forward computation prevents compounding noise degradation across extended autoregressive contexts.
  • Seamless Plug-and-Play Extensibility: In-Mod and Tn-Mod consist of lightweight MLPs that train without altering the underlying LLM weights, facilitating effortless adaptation across modern architectures including Qwen2.5-VL and InternVL.

Limitations & Future Work

  • Author-Admitted Limitations: The modulation modules depend on offline synthetic triplets for training; while effective for common co-occurrence and positional errors, their generalization prior may be constrained when facing long-tail entities or specialized domains (e.g., complex medical imaging, dense satellite aerials).
  • Potential Latent Limitations: Noise modulation is applied uniformly across all transformer layers without layer-wise semantic abstraction weighting; furthermore, control coefficients are modeled at the channel level without explicit spatial geometry priors.
  • Future Improvement Directions: Coupling Langevin diffusion parameters with layer-adaptive gating mechanisms to calibrate noise intensity across shallow and deep layers, alongside integrating test-time adaptation for dynamic strength tuning.
  • vs Woodpecker / LURE: External post-editing pipelines rely on auxiliary object detectors or large language models for iterative question-answering and rewriting, incurring substantial cascading latency; HalMask operates internally via a single forward-pass mechanism.
  • vs OPERA / HALC / DeGF: Contrastive decoding and rollback strategies maintain parallel decoding heads or rewind when attention spikes occur; HalMask eliminates redundancies directly within tokens and Value Caches, preserving standard greedy/sampling decoding and accelerating generation by over 8×.
  • vs Nullu / VASparse: Halluspace projection and visual token sparsification concentrate exclusively on the prefill vision tokens, neglecting decoding cache accumulation; HalMask pioneers unified prefill and decoding cross-modal modulation.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Conceptually innovative use of controlled Gaussian noise injection and Langevin dynamics regularization to eliminate Value Cache redundancy.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive validation across MSCOCO, POPE, OPOPE, AMBER, MMBench, and ScienceQA, supported by rigorous ablations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear motivational narrative, mathematically sound dynamical derivations, and highly structured presentation.
  • Value: ⭐⭐⭐⭐⭐ Drastic hallucination reduction (64% CHAIR drop) paired with negligible latency overhead (0.034s/token) offers tremendous industrial utility.