Skip to content

Learning with Bilevel-Minimax Optimization for Efficient and Reliable Transfer Attacks

Conference: ECCV 2026
Paper: ECCV 2026 Poster
Code: https://github.com/callous-youth/BMAT
Area: Segmentation
Keywords: transfer attacks, bilevel-minimax optimization, initialization perturbation, semantic segmentation black-box attack, surrogate adaptation

TL;DR

Addressing the fragmentation where conventional surrogate-based black-box transfer attacks optimize initialization, perturbation crafting, and model adaptation in isolation, this paper presents BMAT (Bilevel-Minimax Adversarial Transfer), unifying all three into a principled bilevel-minimax optimization problem solved via a Soft Weight Modulator and an Implicit Gradient Approximator, substantially boosting cross-architecture transferability on classification and semantic segmentation benchmarks.

Background & Motivation

Transfer-based adversarial attacks craft imperceptible adversarial examples on white-box surrogate models and transfer them directly to unknown black-box victim models without requiring query feedback. Because attackers operate without active interaction or detection risk against target systems, transfer attacks present a critical threat to real-world vision and safety-critical applications. Existing research to enhance transferability has largely explored localized avenues: momentum-based techniques accumulate historical gradients to stabilize trajectory direction, input transformations apply scaling, padding, or patch mixing to prevent spatial overfitting, and surrogate ensemble approaches aggregate gradients across multiple architectures to mitigate single-model bias.

However, these approaches adhere to a fragmented and heuristic single-variable optimization paradigm. Fundamentally, transferability is governed by the ternary coupling interaction among three core components: the initialization perturbation (IP), which seeds the search trajectory and dictates the explored landscape regions; the adversarial perturbation, which captures and magnifies model vulnerabilities; and the surrogate parameters, which directly shape the back-propagated gradient surface. Prevailing works treat the initialization as zero or random noise and freeze surrogate weights at static pretrained checkpoints, focusing optimization strictly on the perturbation vector alone. This decoupled design leads to misaligned optimization dynamics, causing adversarial trajectories to overfit surrogate-specific artifacts and sharp local minima, resulting in drastic performance degradation when transferred across heterogeneous model families (such as from CNNs to Vision Transformers).

Overcoming these limitations requires a unified optimization framework capable of explicitly modeling and jointly coordinating the interdependent dynamics of initialization, perturbation, and surrogate adaptation. The core idea is to cast transfer attacks into a unified bilevel-minimax optimization formulation, where the outer level optimizes the initialization perturbation via implicit gradient approximation over a pseudo-surrogate to guide trajectory seeding, while the inner level leverages a minimax adversarial game to co-adapt perturbations and soft surrogate weights for flattening the loss landscape and extracting universal cross-architecture gradients.

Method

Overall Architecture

BMAT models transfer attacks through a hierarchical, bottom-up dynamic coordination process. Taking clean images and original ground-truth labels as input, the framework first runs a closed-loop bilevel-minimax optimization phase using the Soft Weight Modulator (SWM) and Implicit Gradient Approximator (IGA) to learn an optimal initialization seed. Subsequently, it transitions to a fast transfer stage that uses this learned seed as a warm start for standard projected gradient iterations on the original surrogate, yielding final adversarial examples endowed with strong cross-model transferability.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Clean Images and Base Surrogate Model"] --> B["Soft Weight Modulator (SWM)<br/>Single-Pass Co-Update of Perturbation and Soft Weights"]
    B --> C["Implicit Gradient Approximator (IGA)<br/>Conjugate Gradient Hypergradient Solver for Initialization Perturbation"]
    C --> D["Two-Stage Fast Transfer Mechanism<br/>Warm-Start Attack Generation from Learned Trajectory Seed"]
    D --> E["Black-Box Victim Model Evaluation<br/>Cross-CNN and Cross-Transformer Transfer Benchmarking"]

During Phase I (learning initialization perturbation), the outer loop optimizes the initialization perturbation \(\boldsymbol{\delta}\) under the supervision of a pseudo-surrogate \(\mathcal{P}\), treating it as the trajectory seed for the inner problem. Conditioned on \(\boldsymbol{\phi}_0 = \boldsymbol{\delta}_t\), the inner loop executes the Soft Weight Modulator (SWM) to jointly update the perturbation \(\boldsymbol{\phi}\) and the surrogate's soft weights \(\boldsymbol{\omega}\), encouraging the surrogate to adapt towards adversarial robustness while preserving clean accuracy, thereby generating a smoother loss landscape. The outer level employs the Implicit Gradient Approximator (IGA) to avoid costly unrolled back-propagation, utilizing implicit feedback to refine \(\boldsymbol{\delta}\). In Phase II (fast transfer), the method takes the optimized \(\boldsymbol{\delta}_T\) as a warm start and executes standard sign-gradient projected steps on the fixed surrogate to efficiently compute the final transferable attack.

Key Designs

1. Soft Weight Modulator (SWM): Single-pass co-adaptation of perturbation and soft surrogate weights

Standard transfer attacks evaluate gradients against static pretrained surrogates, locking parameters into fixed configurations that easily guide perturbations into model-specific sharp valleys. The Soft Weight Modulator resolves this by formulating an inner minimax adversarial objective:

\[ \min_{\boldsymbol{\phi} \in \mathcal{C}} \max_{\boldsymbol{\omega} \in \Omega} f(\boldsymbol{\phi}, \boldsymbol{\omega}) := -\mathcal{L}_{\text{s}}(\boldsymbol{\phi}, \mathcal{S}_{\boldsymbol{\omega}}; \mathcal{D}_i) - \tau \mathcal{R}(\mathcal{S}_{\boldsymbol{\omega}}; \mathcal{D}_i) \]

where \(\mathcal{L}_{\text{s}}\) is the task loss on perturbed inputs, \(\mathcal{R}(\mathcal{S}_{\boldsymbol{\omega}}; \mathcal{D}_i) := \mathcal{L}_{\text{s}}(\mathcal{S}_{\boldsymbol{\omega}}(u_i), v_i)\) acts as a natural-accuracy regularizer, and \(\tau > 0\) balances robustness against clean accuracy. Because minimizing \(-\mathcal{L}_{\text{s}}\) corresponds to maximizing adversarial disruption, this objective drives the surrogate weights \(\boldsymbol{\omega}\) to dynamically adapt against current perturbations, uncovering resilient feature representations.

In terms of execution, SWM extracts gradients for both the perturbation \(\nabla_{\boldsymbol{\phi}} f_k\) and model weights \(\nabla_{\boldsymbol{\omega}} f_k\) within a single backward pass. In each inner step, the perturbation performs constrained gradient ascent while the soft weights adapt along their gradient direction. Crucially, SWM treats pretrained weights \(\boldsymbol{\omega}_0\) as hard anchors, restoring \(\boldsymbol{\omega}_0\) at the start of each batch so that weight adaptation remains strictly local to the attack trajectory without inducing long-term parameter drift. This design flattens the surrogate loss landscape within 10 steps at negligible computational overhead compared to vanilla back-propagation.

2. Implicit Gradient Approximator (IGA): Unrolling-free initialization learning via conjugate gradient

Existing initialization schemes rely on Gaussian noise or heuristic ensemble tuning, lacking direct optimization toward black-box transferability. Directly computing the hypergradient of \(\boldsymbol{\delta}\) across an unrolled inner trajectory of \(\tilde{K}\) steps would incur prohibitive memory and computation costs due to high-order computational graphs. The Implicit Gradient Approximator overcomes this barrier by introducing an outer pseudo-surrogate objective:

\[ \min_{\boldsymbol{\delta} \in \mathcal{C}} F(\boldsymbol{\delta}, \boldsymbol{\phi}^*(\boldsymbol{\delta})) := -\mathcal{L}_{\text{p}}(\boldsymbol{\phi}^*(\boldsymbol{\delta}); \mathcal{P}, \mathcal{D}_i) \]

Applying the Implicit Function Theorem at the inner approximate optimum \(\boldsymbol{\phi}^*(\boldsymbol{\delta})\), the hypergradient is derived as:

\[ \nabla_{\boldsymbol{\delta}} F(\boldsymbol{\delta}, \boldsymbol{\phi}^*(\boldsymbol{\delta})) \approx -\left(\nabla^2_{\boldsymbol{\delta}\boldsymbol{\phi}} f\right)^\top \left(\nabla^2_{\boldsymbol{\phi}\boldsymbol{\phi}} f + \rho \mathbf{I}\right)^{-1} \nabla_{\boldsymbol{\phi}} F = -\left(\nabla^2_{\boldsymbol{\delta}\boldsymbol{\phi}} f\right)^\top \mathbf{h} \]

where \(\rho \mathbf{I}\) is a damping regularizer ensuring local invertibility and numerical stability. Instead of inverting the Hessian matrix explicitly, IGA solves the linear system \(\left(\nabla^2_{\boldsymbol{\phi}\boldsymbol{\phi}} f + \rho \mathbf{I}\right) \mathbf{h} = \nabla_{\boldsymbol{\phi}} F\) using the Fletcher-Reeves conjugate gradient (FR-CG) algorithm. Each iteration computes only Hessian-vector products, converging to an accurate approximation \(\mathbf{h}\) in few iterations. The initialization is then updated via:

\[ \boldsymbol{\delta}_{t+1} \leftarrow \Pi_{\mathcal{C}}\left(\boldsymbol{\delta}_t - \alpha \cdot \text{sgn}\left((\nabla^2_{\boldsymbol{\delta}\boldsymbol{\phi}} f)^\top \mathbf{h}\right)\right) \]

This unrolling-free formulation eliminates trajectory memory storage, equipping the initialization with the ability to steer the attack trajectory away from poor local extrema.

3. Two-Stage Fast Transfer Mechanism: Low-overhead warm-started attack deployment

Deploying full bilevel-minimax co-adaptation during test-time evaluation across large datasets would introduce noticeable latency. The Two-Stage Fast Transfer Mechanism decouples the attack procedure into "Phase I: Trajectory Seed Learning" and "Phase II: Fast Transfer Generation".

In Phase I, the solver runs \(T\) outer iterations with small step counts, steering the initialization perturbation from arbitrary noise into a structured seed \(\boldsymbol{\delta}_T\) through SWM and IGA interactions. Empirical analysis confirms that \(\boldsymbol{\delta}_T\) encodes cross-architecture feature shift directions before Phase II even begins. In Phase II, the algorithm warm-starts the attack at \(\boldsymbol{\phi}_0 = \boldsymbol{\delta}_T\) and executes standard projected gradient ascent steps on the clean surrogate model:

\[ \boldsymbol{\phi}_{k+1} \leftarrow \Pi_{\mathcal{C}}\left(\boldsymbol{\phi}_k + \alpha \cdot \text{sgn}\left(\nabla_{\boldsymbol{\phi}}\mathcal{L}_{\text{s}}(\boldsymbol{\phi}_k; \mathcal{S}_{\boldsymbol{\omega}}, \mathcal{D}_i)\right)\right) \]

By restricting Phase II to \(K\) standard forward/backward passes, BMAT combines the superior generalization of bilevel-minimax joint adaptation with the execution speed and low resource footprint of conventional transfer attacks.

Loss & Training

The complete BMAT optimization comprises the inner adversarial game objective and outer hypergradient guidance. For each input sample \((u_i, v_i)\), the inner problem jointly minimizes perturbation loss and maximizes soft-weight resilience, with clean accuracy weight \(\tau \in [0.1, 0.5]\) and damping parameter \(\rho = 10^{-3}\). Outer loop iterations are set to \(T \in [1, 3]\) and inner steps to \(\tilde{K} \in [5, 10]\). The pseudo-surrogate \(\mathcal{P}\) can be configured either as an auxiliary model (e.g., Inception-v3 or GCNet) or as a Bayesian weight-sampled version of the primary white-box surrogate, achieving zero-prior black-box transfer without violating strict threat model constraints.

Key Experimental Results

Main Results

The effectiveness of BMAT was comprehensively evaluated on ImageNet classification and Cityscapes / ADE20K semantic segmentation benchmarks. Table 1 summarizes black-box transfer mIoU results on Cityscapes using FCN, DeepLabV3-Res50 (DLV3-R50), and Segformer as white-box surrogates against 10 unknown victim models (8 CNN variants and 2 Transformer variants).

Surrogate Model Basic Attacker CNN Victim: FCN CNN Victim: UPerNet CNN Victim: DLV3-R101 CNN Victim: PSP-R101 Transformer: Segformer Transformer: Setr Transfer Behavior
Clean Data N/A 72.25 77.10 80.20 78.34 76.54 78.10 Baseline clean accuracy
FCN PGD 1.97 3.74 8.09 6.83 33.96 42.09 Standard single-level baseline
FCN SegPGD 2.02 3.60 10.04 7.98 36.65 44.73 Dense prediction baseline
FCN MI 2.40 3.57 6.52 5.39 27.81 38.75 Momentum-based baseline
FCN EBAD 2.00 3.83 8.40 7.05 33.91 42.10 Ensemble-based baseline
FCN BMAT (Ours) 1.75 2.74 5.30 4.44 26.58 38.35 Consistent cross-model gains
DLV3-R50 PGD 7.55 4.85 15.02 11.43 41.04 48.37 Limited CNN & ViT transfer
DLV3-R50 MI 5.57 4.29 9.55 6.52 32.99 43.38 Momentum improves CNN transfer
DLV3-R50 BMAT (Ours) 2.60 1.96 4.54 3.57 28.53 42.61 Broad and deep mIoU drops
Segformer PGD 31.43 32.61 35.22 33.56 1.85 43.16 ViT surrogate fails on CNNs
Segformer MI 24.09 24.37 26.21 23.03 2.09 38.95 Heavy cross-architecture drop
Segformer BMAT (Ours) 10.48 13.08 14.10 11.11 2.70 37.52 Nearly 2× mIoU degradation

Under a normalized computational budget (Backward Passes = 40), BMAT demonstrates superior attack success rate (ASR ↑) and efficiency on ImageNet:

Method Budget (BP) CNN Avg. ASR (%) CNN Ensemble ASR (%) Transformer Avg. ASR (%) Overall Avg. ASR (%) Memory (GB) Runtime (s)
PGD BP=10 10.93 5.73 7.16 8.24 3.21 2.68
PGD BP=40 11.05 5.25 6.70 8.00 3.22 3.15
RAP BP=40 7.31 4.67 3.71 5.44 3.75 3.07
RAP BP=400 13.74 6.46 7.47 9.68 5.46 6.91
BETAK BP=40 17.16 8.07 9.25 12.06 22.69 6.37
BMAT (Ours) BP=40 22.52 8.75 11.34 15.03 7.89 4.64

Ablation Study

To isolate the individual and synergistic contributions of IGA and SWM, the table below details the ablation results (mIoU ↓) on Cityscapes using MI as the base attacker:

Surrogate Model Attack Configuration FCN UPerNet DLV3-R101 PSP-R101 Segformer (ViT) Setr (ViT) Analysis
FCN MI (Baseline) 2.40 3.57 6.52 5.39 27.81 38.75 Standard momentum updates
FCN MI + IGA 1.56 2.40 4.64 4.29 27.16 40.07 Strong intra-CNN gains, stagnant on Transformers
FCN MI + IGA + SWM 1.75 2.74 5.30 4.44 26.58 38.35 SWM resolves cross-architecture bottleneck
DLV3-R50 MI (Baseline) 5.57 4.29 9.55 6.52 32.99 43.38 Base CNN surrogate
DLV3-R50 MI + IGA 1.92 1.56 3.96 3.11 29.78 44.42 CNN mIoU halved, but minor regression on Setr
DLV3-R50 MI + IGA + SWM 2.60 1.96 4.54 3.57 28.53 42.61 Restores robust cross-architecture transfer

Key Findings

  • Complementary synergy between SWM and IGA: The ablation experiments confirm that IGA predominantly enhances intra-architecture transfer (CNN to CNN) by steering trajectory seeding, but can experience plateauing on heterogeneous targets. Incorporating SWM flattens the surrogate landscape and extracts universal features, effectively removing cross-architecture bottlenecks on Vision Transformers.
  • Overcoming single-level over-optimization: Under normalized computation budgets, scaling PGD iterations from 10 to 40 steps fails to improve black-box transfer (average ASR declines from 8.24% to 8.00%), confirming that aggressive single-level white-box optimization leads to severe surrogate overfitting. In contrast, BMAT reaches 15.03% ASR under the identical budget of 40 backward passes.
  • Zero-prior single-surrogate viability: In the zero-prior setting where no auxiliary models are available and \(\mathcal{P}\) is configured via Bayesian weight perturbation of the single white-box surrogate, BMAT still yields a 30.17% relative ASR improvement over PGD, proving that its performance advantages stem from the bilevel optimization dynamics rather than ensemble knowledge leakage.

Highlights & Insights

  • Elevating heuristic attack techniques to a principled bilevel-minimax optimization formulation: Rather than stacking ad-hoc data augmentations or gradient smoothing tricks, BMAT formally unifies initialization perturbation, adversarial perturbation, and surrogate adaptation into a mathematically grounded Bilevel-Minimax problem.
  • Efficient single-pass co-updates paired with implicit conjugate gradient solving: By computing both perturbation and soft-weight gradients within a single backward pass in SWM, and solving hypergradients via Fletcher-Reeves conjugate gradients in IGA, BMAT circumvents expensive unrolled trajectory storage, making bilevel optimization practical for high-resolution visual attacks.
  • Dramatic transferability improvements on dense prediction tasks: In semantic segmentation, transferring attacks from Vision Transformers (Segformer) to CNN backbones is notoriously difficult. BMAT slashes victim mIoU by nearly 2× (dropping from 20%~35% down to 10%~14%), establishing a new benchmark for cross-architecture dense attacks.

Limitations & Future Work

  • Computational overhead of second-order gradient approximations: Although IGA avoids unrolling memory spikes and explicit Hessian matrix construction, the conjugate gradient solver still requires several Hessian-vector products per iteration, resulting in runtime approximately 1.5× to 2× that of standard single-level attacks. Developing first-order approximations could further streamline execution.
  • Broader task validation: The current empirical study centers on image classification and semantic segmentation. Extending the bilevel-minimax formulation to object detection, 3D point cloud understanding, and vision-language foundation models (VLMs) remains an open direction.
  • vs RAP (NeurIPS 2022): RAP operates within a single-level minimax framework through explicit reverse adversarial iterations to seek flat minima, which scales poorly in computational overhead and ignores initialization dynamics; BMAT introduces a bilevel hierarchy that co-optimizes initialization and model adaptation, outperforming RAP even when RAP is granted 10× more backward passes.
  • vs BETAK (IJCAI 2024): BETAK derives initialization perturbations via bilevel optimization but heavily depends on surrogate ensembles and unrolled trajectory truncation, incurring a 22.69GB memory footprint; BMAT achieves higher transferability with only 7.89GB memory by coupling inner minimax adaptation with unrolling-free implicit gradients.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ [Pioneering application of bilevel-minimax optimization to transfer attacks, systematically formalizing the ternary coupling of initialization, perturbation, and surrogate adaptation]
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Comprehensive evaluations across classification and segmentation tasks, 30+ victim models, 12 competitive baselines, normalized budget analyses, and detailed ablations]
  • Writing Quality: ⭐⭐⭐⭐⭐ [Clear mathematical derivations, tightly argued motivation, well-structured algorithmic solutions, and thorough empirical analyses]
  • Value: ⭐⭐⭐⭐⭐ [Sets a strong, reproducible benchmark for black-box physical and digital robustness testing while demonstrating efficient bilevel optimization in deep neural networks]