Do Not Leave a Gap: Hallucination-Free Object Concealment in Vision-Language Models¶
Conference: ECCV 2026
Paper: ECCV 2026 Poster
Area: Multimodal VLM
Keywords: Vision-Language Model, Object Concealment, Hallucination Mitigation, Adversarial Attacks, Feature Continuity
TL;DR¶
Addressing the issue where suppression-based object concealment induces compensatory hallucinations in vision-language models (VLMs), this paper proposes Background-Consistent Re-encoding (BCR), a framework that statistically aligns target features and projects them onto background feature manifolds in deep vision encoder layers to conceal objects without leaving representational voids.
Background & Motivation¶
With the pervasive deployment of vision-language models (VLMs) in image captioning, visual question answering, and embodied reasoning, object concealmentโthe ability to selectively obscure specific regions or sensitive entities while preserving overall scene semanticsโhas become critical for privacy preservation, content moderation, and controlled information disclosure. Existing concealment strategies (such as VIP or direct region masking) predominantly rely on aggressively suppressing attention weights or feature activations within the region of interest (ROI) in early or intermediate layers. Although these methods prevent models from recognizing the target object, they inevitably introduce a sharp "semantic gap" or "representational void" in the visual embedding space.
The core tension is that modern VLM language decoders internalize strong contextual priors during pretraining. When visual representations exhibit an abrupt void or anomalous statistical rupture, the autoregressive decoder naturally attempts to compensate for the missing evidence via prior-driven inference. This compensatory mechanism triggers severe grounded hallucination and semantic drift: instead of smoothly omitting the concealed object, models fabricate plausible yet visually unsupported entitiesโsuch as hallucinating fire hydrants, cows, or fireworks in completely unrelated scenesโthereby destabilizing overall scene understanding and compromising factual fidelity.
This paper tackles the root cause by arguing that hallucination in concealment settings is driven by representational discontinuity rather than object absence itself. If target features can be seamlessly blended into the surrounding background manifold such that the language decoder perceives continuous, coherent background context rather than an empty void, compensatory inference can be fundamentally averted. Core idea: propose Background-Consistent Re-encoding (BCR), an optimization framework that enforces first- and second-order statistical alignment, dictionary-based manifold projection, and explicit background preservation across deep vision transformer layers, eliminating target object semantics without leaving a representational gap.
Method¶
Overall Architecture¶
Given an input image \(\mathbf{x} \in \mathbb{R}^{3 \times H \times W}\) and a target bounding box defining the ROI, the goal of BCR is to craft an adversarial image \(\mathbf{x}^{\mathrm{adv}}\) under an \(\ell_\infty\) norm bound \(\|\mathbf{x}^{\mathrm{adv}} - \mathbf{x}\|_\infty \le \epsilon\). The vision encoder \(f_\theta\) and the language decoder \(g_\phi\) remain entirely frozen throughout the procedure. All manipulations are achieved through end-to-end pixel-level gradient descent without updating any model parameters.
During forward propagation through the vision encoder, the image is tokenized into spatial patches. The algorithm partitions token indices into an ROI subset \(\mathcal{I}_r\) whose receptive fields intersect the target box, and a complementary background subset \(\mathcal{I}_b\). In a selected set of deep transformer layers \(\mathcal{L}\), BCR jointly enforces three complementary constraints: aligning the mean and variance between ROI and background tokens, softly reconstructing each ROI token as a convex combination of background tokens via an attention dictionary, and penalizing deviations on background tokens, coupled with total variation regularization in pixel space.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input Image & Target ROI Box"] --> B["Token Partitioning & Deep Feature Extraction<br/>Separate ROI Subset from Background Subset"]
B --> C["Statistical Alignment Constraint<br/>Match First-Order Mean & Second-Order Variance"]
B --> D["Dictionary Manifold Projection Constraint<br/>Soft Attention Convex Combination"]
B --> E["Background Representation Preservation Constraint<br/>Anchor & Freeze Non-ROI Semantics"]
C --> F["Joint Pixel-Level Back-Propagation Optimization<br/>Multi-Loss Synergy + Total Variation Regularization"]
D --> F
E --> F
F --> G["Adversarial Image Output<br/>Semantically Smooth without Representational Voids"]
Key Designs¶
1. Statistical Alignment Constraint: Eliminating Local Feature Distribution Anomalies
Suppression-based attacks cause sharp activation collapses or outlier distributions that are readily detected by language decoders as anomaly signals. To render the ROI tokens distributionally indistinguishable from background context at a fundamental statistical level, BCR minimizes the discrepancy between the first- and second-order feature statistics of ROI tokens \(\mathbf{Z}_r^{(l)} = \{\mathbf{z}_{r,i}^{(l)} \mid i \in \mathcal{I}_r\}\) and background tokens \(\mathbf{Z}_b^{(l)} = \{\mathbf{z}_{b,j}^{(l)} \mid j \in \mathcal{I}_b\}\) across chosen deep layers \(\mathcal{L}\):
where \(\mu(\cdot)\) and \(\sigma(\cdot)\) denote the per-dimension mean and standard deviation. This eliminates low-order moment shifts and prevents anomalous token activations from alerting the decoder.
2. Dictionary Manifold Projection Constraint: Reconstructing Target Semantics on Background Manifolds
Statistical matching alone only constrains marginal distributions without ensuring semantic coherence. If ROI features wander off the natural visual manifold, subsequent cross-attention layers still suffer distortion. BCR treats the set of background tokens as a localized context dictionary and softly projects each ROI token \(\mathbf{z}_{r,i}^{(l)}\) onto this background manifold using scaled dot-product attention weights:
The dictionary projection loss penalizes the squared reconstruction error between each original ROI token and its reconstructed convex combination:
By constraining ROI representations to be linear convex combinations of actual background tokens, the target region is smoothly absorbed into the surrounding background, effectively stripping target-specific identity while maintaining visual continuity.
3. Background Representation Preservation Constraint: Anchoring Global Visual Context
Adversarial perturbations during iterative gradient optimization often bleed into non-target image regions, corrupting unrelated objects and causing holistic semantic drift. To enforce that only the target region is altered, BCR imposes an explicit preservation penalty on all background tokens:
This constraint strictly anchors the deep embeddings of background pixels to their clean counterparts, ensuring that non-target semantic content remains stable and uncorrupted.
Loss & Training¶
The overall optimization objective aggregates multi-layer losses across the selected deep transformer subset \(\mathcal{L}\), regularized by total variation \(\mathcal{L}_{\mathrm{tv}}\) to encourage spatial smoothness in pixel space:
Hyperparameters are established as \(\lambda_{\mathrm{stat}} = 1.0\), \(\lambda_{\mathrm{dict}} = 1.0\), \(\lambda_{\mathrm{pres}} = 1.0\), and \(\lambda_{\mathrm{tv}} = 10^{-3}\), with temperature \(\tau = 0.07\) and perturbation budget \(\epsilon = 0.2\). Crucially, optimization is targeted at the final four transformer blocks of the vision encoder (e.g., layers {22, 23, 24, 25} for Instruct-BLIP and BLIP2-T5, layers {21, 22, 23, 24} for LLaVA), removing semantic concepts at the point where high-level visual abstraction occurs.
Key Experimental Results¶
Main Results¶
The authors conduct comprehensive evaluations on ImageNet (1,000 validation images) and COCO against representative baselines including Masking, PRM, and VIP. Evaluation metrics comprise Concealment Success (C โ), Global Preservation of non-target objects (GP โ), Grounded Hallucination rate verified via GLIP (GH โ), and Semantic Drift (SD โ).
| Dataset | Method | Instruct-BLIP (Cโ / GPโ / GHโ / SDโ) | BLIP2-T5 (Cโ / GPโ / GHโ / SDโ) | LLaVA (Cโ / GPโ / GHโ / SDโ) |
|---|---|---|---|---|
| ImageNet | Clean (No Attack) | 0.00 / 1.00 / 0.00 / 0.00 | 0.00 / 1.00 / 0.00 / 0.00 | 0.00 / 1.00 / 0.00 / 0.00 |
| ImageNet | Masking | 1.00 / 0.41 / 0.59 / 0.54 | 1.00 / 0.52 / 0.48 / 0.40 | 1.00 / 0.55 / 0.45 / 0.41 |
| ImageNet | PRM | 0.66 / 0.38 / 0.62 / 0.44 | 0.70 / 0.31 / 0.69 / 0.41 | 0.73 / 0.34 / 0.66 / 0.39 |
| ImageNet | VIP | 0.93 / 0.57 / 0.43 / 0.35 | 0.96 / 0.59 / 0.41 / 0.32 | 0.98 / 0.62 / 0.48 / 0.29 |
| ImageNet | BCR (Ours) | 0.92 / 0.81 / 0.19 / 0.13 | 0.95 / 0.83 / 0.17 / 0.10 | 0.98 / 0.86 / 0.14 / 0.08 |
| COCO | Clean (No Attack) | 0.00 / 1.00 / 0.00 / 0.00 | 0.00 / 1.00 / 0.00 / 0.00 | 0.00 / 1.00 / 0.00 / 0.00 |
| COCO | Masking | 1.00 / 0.37 / 0.63 / 0.46 | 1.00 / 0.39 / 0.61 / 0.44 | 1.00 / 0.41 / 0.59 / 0.41 |
| COCO | PRM | 0.63 / 0.54 / 0.46 / 0.37 | 0.77 / 0.58 / 0.42 / 0.34 | 0.80 / 0.60 / 0.40 / 0.33 |
| COCO | VIP | 0.80 / 0.42 / 0.58 / 0.48 | 0.83 / 0.45 / 0.55 / 0.45 | 0.88 / 0.48 / 0.52 / 0.43 |
| COCO | BCR (Ours) | 0.89 / 0.77 / 0.23 / 0.16 | 0.92 / 0.79 / 0.21 / 0.13 | 0.95 / 0.82 / 0.18 / 0.11 |
Ablation Study¶
Component-wise ablation on Instruct-BLIP confirms that all three loss terms are essential to achieving both high concealment efficacy and semantic stability.
| Config | \(\mathcal{L}_{\mathrm{stat}}\) | \(\mathcal{L}_{\mathrm{dict}}\) | \(\mathcal{L}_{\mathrm{pres}}\) | \(\mathcal{L}_{\mathrm{tv}}\) | Concealment Success (CS โ) | Hallucination Rate (HR โ) | Global Preservation (GP โ) |
|---|---|---|---|---|---|---|---|
| Full Model (BCR Full) | โ | โ | โ | โ | 0.91 | 0.08 | 0.83 |
| w/o Total Variation | โ | โ | โ | โ | 0.90 | 0.09 | 0.81 |
| w/o Background Preservation | โ | โ | โ | โ | 0.86 | 0.17 | 0.69 |
| w/o Dictionary Projection | โ | โ | โ | โ | 0.79 | 0.21 | 0.74 |
| w/o Statistical Alignment | โ | โ | โ | โ | 0.72 | 0.28 | 0.66 |
| VIP-style Baseline | - | - | - | - | 0.88 | 0.42 | 0.51 |
Key Findings¶
- Representational Continuity Governs Hallucination: VIP slightly reduces statistical discrepancy but leaves dictionary projection error at 1.0500 (even exceeding Clean at 1.0113), resulting in a 0.440 hallucination rate. Conversely, BCR slashes dictionary projection error to 0.0777 (over 90% reduction) and statistical discrepancy to 0.0297, driving grounded hallucination down to 0.167. This establishes empirical proof that voids cause hallucination.
- Deeper Layer Targeting is Crucial: Layer sensitivity analysis reveals that targeting early ViT blocks yields a meager 0.22 concealment success rate. Targeting deep layers reaches 0.91 success with a low 0.18 hallucination rate, confirming that high-level abstractions in later layers dictate linguistic recognition.
- Superior Perceptual Fidelity: BCR achieves 0.94 SSIM and 0.04 LPIPS on LLaVA (vastly superior to VIP's 0.65 SSIM and 0.32 LPIPS), ensuring that adversarial alterations remain virtually imperceptible to human observers.
Highlights & Insights¶
- Reframing Object Concealment as Manifold Blending: Breaks away from the traditional "suppress or erase" mindset by treating background tokens as an expressive dictionary, showing that seamless manifold blending eliminates the representational trigger for hallucination.
- GLIP-Based Grounded Verification: Overcomes the ambiguity of purely lexical caption comparisons by verifying candidate nouns against physical image regions with the GLIP open-vocabulary detector, setting a rigorous evaluation standard for multimodal adversarial attacks.
- Model-Agnostic Plug-and-Play Design: Operates purely through pixel-space optimization guided by vision encoder representations, making it directly portable across diverse architectures including CLIP-ViT and ViT-g.
Limitations & Future Work¶
- Partial Semantic Substitution: For small objects or complex surroundings, BCR occasionally re-encodes the target into an adjacent conceptual category (e.g., turning a tie into shirt patterns or cutlery into plates), which prevents hallucination but leaves indirect contextual clues.
- Vulnerability to Adversarial Probing: Interactive or targeted questioning directly focused on the concealed region (e.g., repeated binary yes/no probes) might occasionally extract faint residual features.
- Future Directions: Exploring prompt-conditioned adaptive re-encoding weights and extending token alignment formulations to non-uniform patch layouts or dynamic-resolution vision backbones.
Related Work & Insights¶
- vs VIP (Visual Information Protection): VIP creates a representational void by suppressing attention and value vectors, causing the language decoder to invent extraneous objects (e.g., fireworks or hydrants); BCR preserves continuity, achieving comparable concealment while cutting grounded hallucination by up to 3ร.
- vs PRM (Patch Representation Misalignment): PRM indiscriminately disrupts patch structures, causing catastrophic scene breakdown (GP drops to 30-40%); BCR anchors background tokens to maintain over 80% non-target semantic preservation.
Rating¶
- Novelty: โญโญโญโญโญ Reconceptualizes privacy-preserving object concealment by identifying and rectifying representational discontinuity as the true trigger of hallucination.
- Experimental Thoroughness: โญโญโญโญโญ Thorough evaluation across three distinct VLM architectures, dual datasets, quantitative continuity metrics, and grounded detector verification.
- Writing Quality: โญโญโญโญโญ Elegant structure, clear mathematical formulations, and compelling ablation insights.
- Value: โญโญโญโญโญ Provides a foundational methodology for multimodal privacy protection, content obfuscation, and reliable hallucination-free AI systems.