Gender Bias in Vision-Language In-Context Learning¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/mathfather/Gender-Bias-in-VL-ICL
Area: Multimodal VLM / AI Safety
Keywords: Vision-Language Model, In-Context Learning (ICL), Gender Bias, Evaluation Benchmark (VL-BICLE), Diffusion-based Debiasing
TL;DR¶
This paper systematically investigates how multimodal in-context demonstrations exert a "directional force" that amplifies gender bias in large vision-language models primarily via an asymmetric cross-gender degradation mechanism, and proposes an offline synthetic demonstration image replacement strategy using diffusion models that mitigates bias without degrading caption quality.
Background & Motivation¶
The emergence of large vision-language models (LVLMs) capable of processing interleaved image-text inputs has positioned in-context learning (ICL) as a primary paradigm for few-shot task adaptation without parameter updates. However, foundation models trained on massive, uncurated web crawls inevitably inherit and encode pervasive societal stereotypes and demographic disparities. Consequently, these models exhibit systematic gender skews across downstream tasks including image captioning, pronoun resolution, and visual question answering. While societal bias has been studied in purely textual contexts and discriminative vision-language models like CLIP, the vulnerability of generative, autoregressive LVLMs under multimodal few-shot demonstration sequences has remained largely unquantified.
The core tension stems from the double-edged nature of demonstration prompts and the blind spots of conventional evaluation benchmarks. Providing multimodal context demonstrations is standard practice to boost downstream accuracy, but whether demographic distributions within demonstration sequences amplify or trigger latent model biases was unknown. Moreover, widely adopted similarity-based retrieval strategies (e.g., CLIP-based image or text similarity) inherently replicate the demographic imbalances of the demonstration pool, while standard captioning quality metrics like BLEU-4 and CLIPScore remain entirely blind to substantial shifts in gender bias. This creates a dangerous illusion of stable generation quality while equity quietly deteriorates.
To address these challenges, this paper presents a systematic decomposition of bias dynamics across multimodal ICL settings, isolating the roles of demographic composition, retrieval mechanics, and task output spaces. Core idea: build a comprehensive evaluation suite, VL-BICLE, to uncover how gendered contexts systematically pull model behavior toward demonstrated genders via a "cross-gender degradation mechanism," and introduce an offline synthetic image replacement method using Stable Diffusion models that mitigates gender bias without degrading generation quality.
Method¶
Overall Architecture¶
The evaluation and debiasing pipeline centers on VL-BICLE (Vision-Language Gender Bias in ICL Evaluation). The framework incorporates 4 gender-composition selection strategies (Random Sample RS, Male-only Sample MS, Female-only Sample FS, Balanced Sample BS) alongside 2 similarity-based retrieval baselines (Similarity-based Image-Image Retrieval SIIR, Similarity-based Image-Text Retrieval SITR). These strategies are evaluated across 3 representative tasks (image captioning, pronoun prediction, visual question answering) on 4 benchmark datasets (COCOBias, DCI, VisoGender, VisualCoT). Empirical evaluations across 6 LVLMs reveal that demographic bias acts as a directional force, directly motivating an offline image replacement pipeline driven by stable diffusion models (SDMs).
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
IN["Multimodal Query Input<br/>(Query Image + Task Instruction)"] --> D1["VL-BICLE Multi-Dimensional Evaluation Matrix<br/>6 ICL Settings × 3 Downstream Tasks"]
D1 --> D2["Cross-Gender Degradation Analysis<br/>Opposite-Gender Error Spikes & Directional Shifts"]
D1 --> D3["Similarity Retrieval Imbalance Diagnosis<br/>SIIR / SITR Demonstration Pool Bias Inheritance"]
D2 --> D4["Diffusion-Based Context Synthetic Replacement<br/>SDM Offline Controllable Image Generation"]
D3 --> D4
D4 --> OUT["Equitable Multimodal Generation Output<br/>(Preserved Quality & Mitigated Disparity)"]
Key Designs¶
1. VL-BICLE Multi-Dimensional Evaluation Matrix: Disentangling Multimodal Context Biases
Addressing the absence of controlled benchmarks for multimodal ICL bias, this design establishes a formal subgroup performance gap formulation across 6 context construction setups. Given a dataset \(\mathcal{D}_t = \{(I_i, y_i, g_i)\}_{i=1}^N\) (with binary perceived gender annotations \(g_i \in \{\text{male}, \text{female}\}\) for single-person images), the subgroup error rates \(\text{ER}_m\) and \(\text{ER}_f\) are formulated, with their difference \(\text{ER}_{m-f}\) serving as the primary metric:
where \(\text{ER}_{m-f} > 0\) indicates a higher error rate on male-presenting queries (female bias), \(\text{ER}_{m-f} < 0\) denotes male bias, and values near zero represent equitable performance. The 6 context configurations systematically probe bias dynamics: RS maintains the pool distribution, MS and FS construct single-gender extremes, BS enforces alternating gender balance, and SIIR/SITR rank in-context examples by frozen CLIP cosine similarity to the query, placing top candidates closest to the query prompt.
2. Cross-Gender Degradation Analysis: Uncovering Asymmetric Negative Transfer
Addressing how single-gender demonstrations shift model decisions, this analysis tracks the fine-grained error rate trajectories of male and female subgroups. It demonstrates that single-gender contexts do not enhance in-group recognition but rather exert a "directional force" that triggers asymmetric cross-gender degradation: injecting male-only demonstrations (MS) leaves male error rates \(\text{ER}_m\) relatively flat while female error rates \(\text{ER}_f\) surge significantly; conversely, female-only contexts (FS) disproportionately elevate \(\text{ER}_m\). This mechanism is magnified in pronoun prediction (e.g., on VisoGender-OP at 8-shot, QwenVL's \(\text{ER}_f\) reaches 85.5% under MS compared to 9.7% under FS), verifying that single-gender context acts by destabilizing opposite-gender representations rather than debiasing the model.
3. Similarity Retrieval Imbalance Diagnosis: Auditing Demographic Skew in Semantic Retrieval
Targeting the widespread belief that similarity-based retrieval (SBR) yields superior in-context performance without drawbacks, this design audits its equity impact. Because standard demonstration pools (like MSCOCO) are intrinsically skewed toward male instances, SIIR (image-to-image) and SITR (image-to-text) disproportionately retrieve male-skewed demonstrations. Consequently, SBR methods inherit this demographic disparity and produce bias levels comparable to or exceeding pure MS settings (for instance, on 8-shot Phi-3.5-V, SIIR and SITR produce \(\text{ER}_{m-f}\) of -2.25% and -2.35%, exhibiting even stronger male bias than MS at -1.88%). Down-sampling the retrieval pool to perfect 1:1 gender balance only partially recovers this gap because top-\(k\) nearest-neighbor matching enforces no constraint on retrieved context gender composition.
4. Diffusion-Based Context Synthetic Replacement: Offline Disentangled Debiasing Without Quality Loss
Addressing the high annotation cost and risk of linguistic drift associated with text-side prompt rewriting, this design introduces an image-level replacement strategy powered by text-to-image diffusion models. Retaining the ground-truth demonstration captions \(y_i\) intact, those texts serve as prompts to offline stable diffusion models (FLUX or Stable Diffusion 3.5 Large) to synthesize corresponding replacement images \(I_i^{synth}\), forming a synthetic demonstration pool \(\mathcal{D}_s\). In-context examples during inference are drawn from \(\mathcal{D}_s\) instead of raw photographs. By synthesizing human subjects, diffusion models reconstruct visual scene contexts and dilute spurious real-world visual co-occurrences. On revealed samples (\(r_i = 1\), where explicit gender terms are generated), the absolute bias magnitude \(|\text{ER}_{m-f}|\) decreases substantially across models under both MS and FS settings, while reference-based BLEU-4 and reference-free CLIPScore remain fully preserved with zero inference-time computational overhead.
Loss & Training¶
This work represents a training-free inference-time evaluation and mitigation study. All evaluated LVLMs and vision feature extractors (CLIP) are frozen during inference, and token generation employs deterministic greedy search. Few-shot evaluations cover \(k \in \{0, 2, 4, 6, 8\}\). For random and balanced sampling setups, experiments are repeated across 5 independent sampling runs to compute mean and standard deviation, whereas similarity-based retrieval methods are deterministic. For VisualCoT evaluation, GPT-OSS-20B acts as an LLM-as-judge to score semantic correctness \(s_i \in [0, 1]\) against ground-truth answers, yielding a continuous sample error \(e_i = 1 - s_i\).
Key Experimental Results¶
Main Results¶
Zero-shot baseline error gaps (\(\text{ER}_{m-f}\)) across models and datasets are summarized in Table 1; Table 2 presents generation quality metrics (BLEU-4, CLIPScore, AvgL) and bias metrics (\(\text{ER}_{m-f}\)) for Phi-3.5-V on COCOBias under varying \(k\)-shot settings.
Table 1: Zero-shot baseline error rate gap \(\text{ER}_{m-f}\) across six LVLMs (negative indicates male bias, positive indicates female bias):
| Dataset | Task | QwenVL | MiniCPM | Idefics3 | Phi35V | InternVL35 | Qwen3VL |
|---|---|---|---|---|---|---|---|
| VisoGender-OO | Pronoun prediction | -2.90 | -27.54 | -1.45 | +13.04 | -4.35 | +2.90 |
| VisoGender-OP | Pronoun prediction | -8.89 | -45.36 | -8.10 | +35.91 | -3.36 | +3.64 |
| COCOBias | Image captioning | -0.94 | -0.09 | -1.69 | -0.09 | 0.00 | +0.19 |
| DCI | Image captioning | +0.84 | +0.84 | -1.68 | 0.00 | -0.84 | -0.84 |
| VisualCoT | Visual Question Answering | – | -0.31 | – | +6.42 | – | +1.61 |
Table 2: Caption quality metrics and \(\text{ER}_{m-f}\) across ICL configurations on COCOBias for Phi35V:
| ICL Setting | \(k\)-shot | BLEU-4 ↑ | CLIPScore ↑ | Avg Length (AvgL) | \(\text{ER}_{m-f}\) (Bias Gap) |
|---|---|---|---|---|---|
| Zero-shot | 0 | 18.86 | 33.13 | 20.71 | -0.09 |
| RS (Random Sample) | 2 | 35.84 | 32.08 | 10.70 | -1.30 |
| RS (Random Sample) | 8 | 35.89 | 31.84 | 10.20 | -1.01 |
| MS (Male-only) | 2 | 34.15 | 32.19 | 11.19 | -0.92 |
| MS (Male-only) | 8 | 35.77 | 32.00 | 10.50 | -1.88 |
| FS (Female-only) | 2 | 34.60 | 32.21 | 11.30 | +0.28 |
| FS (Female-only) | 8 | 36.41 | 31.73 | 10.05 | +2.14 |
| BS (Balanced Sample) | 8 | 36.50 | 32.02 | 10.59 | -0.75 |
| SIIR (Image-to-Image) | 8 | 36.65 | 32.07 | 10.81 | -2.25 |
| SITR (Image-to-Text) | 8 | 37.16 | 32.39 | 10.96 | -2.35 |
Ablation Study¶
Holding demonstration texts identical, in-context real images are replaced with SDM-generated synthetic images (FLUX, SD3.5-Large) alongside pure black control images and negative prompt variations. Table 3 shows the relative shift in absolute bias \(\Delta |\text{ER}_{m-f}|\) on revealed samples (\(r_i=1\)) compared against the raw COCOBias baseline (negative numbers denote reduced bias / improved equity):
Table 3: Bias mitigation on revealed samples under synthetic image replacement and ablations (delta relative to COCOBias baseline):
| Model | ICL Setting | \(k\) | COCOBias Baseline | FLUX (Synthetic) | SD35L (Synthetic) | Black Image | S-NP (No-Prompt) | F-NP (No-Prompt) |
|---|---|---|---|---|---|---|---|---|
| Idefics3 | MS (Male-only) | 4 | 2.82 | -0.42 | -0.27 | -0.32 | -0.12 | -0.05 |
| Idefics3 | MS (Male-only) | 8 | 2.87 | -0.12 | -0.30 | -0.33 | -0.02 | -0.25 |
| Phi35V | MS (Male-only) | 4 | 2.24 | -0.36 | -0.23 | +2.48 | +4.92 | +3.70 |
| Phi35V | MS (Male-only) | 8 | 2.04 | -0.10 | -0.20 | +3.84 | +9.41 | +5.79 |
| MiniCPM | MS (Male-only) | 8 | 0.32 | -0.24 | -0.08 | +0.68 | +1.63 | +1.10 |
| Qwen3VL | MS (Male-only) | 4 | 0.44 | -0.15 | -0.07 | -0.27 | +2.35 | +1.72 |
| Idefics3 | FS (Female-only) | 4 | 1.00 | -0.02 | -0.16 | +0.95 | +0.74 | +1.01 |
| Phi35V | FS (Female-only) | 8 | 2.53 | -0.90 | -0.59 | +3.36 | +27.91 | +19.52 |
| MiniCPM | FS (Female-only) | 8 | 0.14 | -0.12 | -0.05 | +1.47 | +1.45 | +0.94 |
| Qwen3VL | FS (Female-only) | 8 | 0.27 | -0.08 | -0.02 | +0.47 | +6.15 | +5.45 |
Key Findings¶
- Gendered demonstrations act as a directional force: Regardless of a model's zero-shot bias predisposition, MS prompts pull model outputs toward male bias while FS prompts pull them toward female bias, frequently flipping polarity entirely (e.g., MiniCPM on VisoGender-OO shifts from -27.54% at zero-shot to +27.20% under 8-shot FS).
- Decoupling of standard quality metrics from societal bias: Across all 6 models and shot configurations, CLIPScore fluctuates by less than 0.67, and BLEU-4 variations show no correlation with the direction or magnitude of bias shifts, confirming that conventional performance metrics are completely insensitive to equity drift.
- Task output space governs bias propagation: Across all tested models on VQA, error rates across RS, MS, FS, and BS remain virtually indistinguishable (shifts \(\le 1.82\%\)), demonstrating that ICL gender composition alters bias only when task outputs require gendered vocabulary.
Highlights & Insights¶
- Formulation of cross-gender degradation: Moves beyond empirical bias observation by demonstrating that single-gender context shifts bias primarily by degrading performance on opposite-gender queries, identifying the micro-mechanism of ICL bias induction.
- Refutation of similarity retrieval as a debiasing tool: Provides direct empirical evidence that CLIP-based demonstration retrieval inherits and magnifies demonstration pool demographic skews, disproving the assumption that semantic similarity naturally enhances fairness.
- Plug-and-play diffusion-based synthetic debiasing: Demonstrates that swapping raw context images with offline text-to-image synthetic renderings mitigates bias on revealed predictions while preserving downstream captioning fidelity without latency overhead.
Limitations & Future Work¶
- Author-admitted limitations: The benchmark relies on binary gender definitions and single-person images, omitting non-binary identities and complex multi-person demographic scenarios; synthetic image debiasing was primarily evaluated on image captioning.
- Self-identified limitations: The mechanistic link explaining why synthetic images diminish cross-gender degradation requires deeper attention-map analysis into multimodal token cross-attention layers.
- Future directions: Developing adaptive attention-reweighting schemes during multimodal ICL inference and expanding mitigation pipelines to intersectional social attributes such as race, age, and occupational stereotypes.
Related Work & Insights¶
- vs Zhao et al. (ICCV 2021) / Hendricks et al. (ECCV 2018): Earlier works examined static bias in fine-tuned captioning models via corpus constraints or regularized loss functions; this paper pioneers the analysis of dynamic bias transfer in frozen, autoregressive LVLMs under in-context learning.
- vs Yang et al. (NeurIPS 2024 LeverLM) / Doveh et al. (2024): Existing multimodal ICL literature focuses almost exclusively on maximizing accuracy metrics via sequence configuration and retrieval; this work reveals the severe fairness pitfalls embedded in retrieval-driven demonstration selection.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ First systematic formalization of gender bias dynamics and cross-gender degradation in multimodal ICL.
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorously tested across 6 popular LVLMs, 4 benchmarks, 3 tasks, and 6 demonstration strategies.
- Writing Quality: ⭐⭐⭐⭐⭐ Clear logical progression, insightful mechanistic formulations, and comprehensive experimental presentation.
- Value: ⭐⭐⭐⭐⭐ Establishes a foundational benchmark and actionable guidelines for equitable deployment of vision-language models.