Skip to content

CRISP: Calibration-Aware Visual State Space Duality for Remote Sensing Image Segmentation

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/crazylifeha/crisp
Area: Segmentation
Keywords: remote sensing semantic segmentation, state space models, visual state space duality, frequency calibration, prototype learning

TL;DR

Targeting high-frequency energy decay and excessive boundary smoothing caused by global normalized aggregation in Visual State Space Duality (VSSD), CRISP introduces a linear-complexity Duality Calibration Operator (DCO) that reuses aggregation statistics alongside an end-to-end Orthogonal Multi-Prototype (OMP) decoder, effectively preserving fine-grained multi-modal features and crisp boundaries with only ~32M parameters.

Background & Motivation

High-resolution remote sensing image semantic segmentation serves as a fundamental cornerstone for Earth observation tasks such as urban planning, disaster monitoring, and ecological assessment. Overhead imagery typically presents exceptionally high spatial resolutions, an abundance of dense, miniature structures (such as compact vehicles and isolated dwellings), and intricate category boundaries. These distinctive characteristics impose stringent demands on a model's ability to preserve delicate high-frequency textures and sharp boundary transitions. While early convolutional neural networks (CNNs) possessed strong local inductive biases, they lacked sufficient capability for modeling long-range contextual dependencies. Subsequent Vision Transformers (ViTs) successfully captured broad global contexts via self-attention; however, the quadratic computational and memory scaling of self-attention with respect to sequence length becomes practically prohibitive when processing ultra-high-resolution aerial tiles.

State Space Models (SSMs) have emerged as an appealing linear-time alternative to Transformers, progressing from selective state spaces in Mamba to State Space Duality (SSD) in Mamba-2, which unifies state recurrence with masked linear attention. To adapt to 2D vision tasks lacking a canonical 1D temporal ordering, Visual State Space Duality (VSSD) performs multi-directional sweeping and fuses the results into a global, position-agnostic aggregation, roughly halving scanning redundancy while maintaining an \(O(L)\) computational bound. Nonetheless, this design harbors a long-overlooked intrinsic dilemma: collapsing multi-directional scans into a global normalized spatial aggregation mathematically acts as a broad-window, low-pass filter that systematically averages out oscillatory, zero-mean high-frequency variations. Feature spectral energy analysis (\(E_{\text{high}}/E_{\text{low}}\) via 2D FFT) demonstrates that high-frequency energy decays monotonically across backbone layers, inevitably smoothing out narrow roads, building footprints, and small object silhouettes.

Existing remedies for spatial detail recovery typically append parallel frequency-domain branches, convolutional side paths, or complex decoder-side feature fusion modules. However, these external add-ons break the hardware-efficient continuity of native state-space scans while introducing substantial parameter and FLOP overheads. The core idea of this work is to construct an in-place Duality Calibration Operator (DCO) that reuses existing backbone aggregation statistics to compensate for value-rescaling attenuation and high-frequency decay within linear complexity, paired with an end-to-end Orthogonal Multi-Prototype (OMP) head in the decoder to prevent the recovered intra-class multi-modal representations from collapsing.

Method

Overall Architecture

The CRISP framework comprises two coordinated components: the backbone-level Duality Calibration Operator (DCO) and the decoder-level Orthogonal Multi-Prototype (OMP) head. In the backbone, input image patches pass through stacked VSSD stages; inside each VSSD block, following multi-directional scanning and global normalized fusion, DCO is applied in place to perform residual injection, de-meaned high-pass modulation, and channel-wise DC/HC rebalancing using precomputed key, value, and query statistics. In the decoding stage, multi-scale features enter the lightweight MultiProtoDecoder, which derives content-adaptive class tokens via dual-path pooling and refines them through bidirectional cross-attention. A hyper-network then generates \(K\) orthogonally constrained sub-prototypes per semantic category, which are matched against normalized pixel descriptors via max-response matching to yield the final segmentation prediction.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Remote Sensing Input<br/>H x W x 3"] --> Emb["Patch Embedding & Encoding"]
    Emb --> VSSD["VSSD Backbone Multi-Scan<br/>State Aggregation & Global Normalization"]
    VSSD --> DCO["Duality Calibration Operator (DCO)<br/>Residual Injection + Hi-Pass + DC/HC Rebalance"]
    DCO --> Dec["Multi-Scale Fusion & Dynamic Token Generation"]
    Dec --> OMP["Orthogonal Multi-Prototype Matching (OMP)<br/>Hyper-Network Prototypes + Max-Response"]
    OMP --> Reg["Orthogonal & Margin Regularization<br/>Intra-Class Orthogonality + Inter-Class Margin"]
    Reg --> Out["Semantic Segmentation Output"]

Key Designs

1. Duality Calibration Operator (DCO): In-place statistic reuse for linear-time high-frequency recovery

To counter the low-pass filtering effect of VSSD global normalized aggregation without incurring quadratic attention costs or external spectral branches, DCO extracts the residual signal stemming from value scaling. In VSSD, input features are scaled by input-dependent gain factors to produce values \(v'_j = \rho_j \odot v_j\). DCO isolates the attenuated variation as \(\Delta v_j = v_j - v'_j\) and reuses the existing kernelized key statistics \(k_j = \psi_v(b_j)\) (where \(\psi_v(x) = \text{ELU}(x) + 1\)), accumulating the global residual state matrix across all tokens in linear time: $$ \bm{S}{\Delta} = \sum k_j \Delta v_j^\top $$ For any query location }^{L\(i\), the residual compensation is retrieved using its corresponding query representation \(q_i = \psi_v(c_i)\) via \(r_i \approx q_i^\top \bm{S}_{\Delta}\). Because the summary matrix \(\bm{S}_{\Delta}\) is shared globally across all positions and computed only once, the calibration strictly preserves \(O(L)\) linear time complexity. The extracted residual is injected into the backbone stream via a small-gain near-identity gate: \(\tilde{y}_i = \hat{y}_i + (\varepsilon + \lambda_h) \odot r_i\), initialized with \(\lambda_h = 0.01\). This ensures stable optimization from an identity initialization while restoring local contrast attenuated by discretization and decay.

2. Token-level de-meaned high-pass modulation and DC/HC rebalancing

To further amplify fine boundary details suppressed across deep backbone layers, DCO introduces a token-level high-pass modulation step based on spatial deviation from the global average: $$ y'i = \tilde{y}_i + \omega_h \odot (\tilde{y}_i - \bar{y}), \quad \text{where} \quad \bar{y} = \frac{1}{L}\sum_j $$ Subtracting the spatial mean }^L \tilde{y\(\bar{y}\) isolates high-frequency spatial perturbations, which are then sharpened by the learnable scaling factor \(\omega_h\). Subsequently, the calibrated feature is decomposed into low-frequency semantics and high-frequency structures \(Y' = Y_{\text{dc}} + Y_{\text{hc}}\), and channel energies are adaptively reallocated via \(Y_{\text{out}} = (1+s) \odot Y_{\text{dc}} + (1+t) \odot Y_{\text{hc}}\). The parameters \(\omega_h\), \(s\), and \(t\) are bounded and initialized near zero (0.01), allowing the network to adaptively rebalance macro-contextual semantics and micro-boundary geometry throughout training.

3. Orthogonal Multi-Prototype (OMP) decoder: Modeling intra-class spectral multimodality

Even when high-frequency spatial details are restored by the backbone, conventional single-prototype classification heads (which dedicate one weight vector per category) collapse rich intra-class modalitiesโ€”such as identical building categories displaying disparate roof textures or vegetation exhibiting varied illuminationโ€”onto a single direction. OMP tackles this by learning content-adaptive dynamic tokens via average and max pooling combined with an embedding bank. A hyper-network then projects these tokens into \(K\) normalized sub-prototypes \(\hat{P}_{c,k}\) for each class \(c\). During inference, normalized pixel descriptors \(\hat{U}_i\) are classified via max-response matching: $$ z_{c,i} = \max_{1 \le k \le K} \langle \hat{U}i, \hat{P} \rangle $$ Each pixel dynamically selects its closest sub-prototype, granting the model sufficient representational capacity to capture complex multi-modal distributions without forcing them into a compromise mean vector.

4. Dual geometric prototype regularization: Intra-class orthogonality and inter-class angular margin

Without explicit geometric constraints, unregularized sub-prototypes often collapse toward redundant, clustered vectors. OMP resolves this issue by introducing two complementary regularizers: $$ \mathcal{L}{\text{orth}} = \frac{1}{B |\mathcal{C}| K (K-1)} \sum} \sum_{k \neq l} \big| \langle \hat{P{b,c,k}, \hat{P} \rangle \big| $$ $$ \mathcal{L}{\text{margin}} = \frac{1}{B} \sum} \sum_{c \neq c'} \text{ReLU}\big( \langle \hat{\bar{P}{b,c}, \hat{\bar{P}} \rangle - \delta \big) $$ The intra-class orthogonality penalty \(\mathcal{L}_{\text{orth}}\) forces sub-prototypes within the same class to be mutually orthogonal, maximizing their span over diverse intra-class modes. Concurrently, the inter-class angular margin penalty \(\mathcal{L}_{\text{margin}}\) enforces a minimum angular separation \(\delta\) between mean class centers \(\hat{\bar{P}}_{b,c}\), preventing intra-class expansion from encroaching upon foreign class decision boundaries. Both terms optimize end-to-end with the segmentation objective without requiring offline clustering.

Loss & Training

The overall training objective combines the segmentation task loss with the dual prototype regularizations: $$ \mathcal{L}{\text{total}} = \mathcal{L}}} + \lambda_{\text{orth}} \mathcal{L{\text{orth}} + \lambda $$ where }} \mathcal{L}_{\text{margin}\(\mathcal{L}_{\text{seg}}\) denotes standard cross-entropy loss, and \(\lambda_{\text{orth}}\) and \(\lambda_{\text{margin}}\) balance the respective penalties. Models are trained under a unified MMSegmentation protocol using AdamW (initial learning rate \(6 \times 10^{-5}\), weight decay 0.05), linear warmup, and polynomial learning rate decay (\(p=0.9\)) on \(512 \times 512\) random crops.

Key Experimental Results

Main Results

CRISP was extensively benchmarked against representative CNN, Transformer, and state-of-the-art SSM architectures across LoveDA, ISPRS Potsdam, and ISPRS Vaihingen datasets.

Dataset Model Params (M) mFscore (%) mIoU (%) Representative Class Performance (Building / Tree / Forest)
LoveDA VSSD (Baseline) 56.1 - 47.15 Bu: 50.96 / Fo: 40.57
LoveDA UNetFormer 11.7 - 49.88 Bu: 58.89 / Fo: 41.49
LoveDA AerialFormer 43.2 - 50.41 Bu: 64.05 / Fo: 41.78
LoveDA CRISP (Ours) 32.32 - 51.56 Bu: 65.16 / Fo: 43.92
LoveDA D2LS (Large) 94.0 - 53.53 Bu: 63.27 / Fo: 41.10
Potsdam VSSD (Baseline) 56.1 92.07 85.57 Build: 92.81 / Tree: 77.30
Potsdam UNetFormer 11.7 92.68 86.59 Build: 96.83 / Tree: 88.07
Potsdam AerialFormer 43.2 93.02 87.17 Build: 96.74 / Tree: 88.77
Potsdam CRISP (Ours) 32.32 93.93 88.77 Build: 97.78 / Tree: 89.47
Potsdam D2LS (Large) 94.0 94.15 89.17 Build: 97.99 / Tree: 89.71
Vaihingen VSSD (Baseline) 56.1 88.54 79.75 Build: 90.37 / Tree: 79.43
Vaihingen UNetFormer 11.7 87.77 78.64 Build: 94.65 / Tree: 89.04
Vaihingen AerialFormer 43.2 89.98 82.03 Build: 95.75 / Tree: 89.00
Vaihingen CRISP (Ours) 32.32 90.55 83.00 Build: 96.62 / Tree: 89.46
Vaihingen D2LS (Large) 94.0 91.18 84.06 Build: 96.95 / Tree: 89.71

Ablation Study

1. DCO sub-component ablation on Potsdam (evaluated with UPerHead, single-scale)

Config Residual Injection (Res. Inj.) High-Pass Modulation (Hi-Pass) DC/HC Rebalance mIoU (%) Gain over Baseline
Baseline - - - 87.32 -
Config A โœ“ - - 87.65 +0.33
Config B - โœ“ - 87.68 +0.36
Config C - - โœ“ 87.64 +0.32
Config D - โœ“ โœ“ 87.83 +0.51
Full DCO โœ“ โœ“ โœ“ 88.10 +0.78

2. End-to-end component synergy ablation on Potsdam (VSSD backbone)

Decoder DCO Included Intra-Class Orthogonality (Orth) Inter-Class Margin (Margin) mIoU (%) Note & Insight
UPerHead - - - 87.32 Standard VSSD baseline
UPerHead โœ“ - - 88.10 Backbone DCO alone yields +0.78%
OMP - โœ“ โœ“ 85.63 Multi-prototypes degrade without DCO (-1.69%)
OMP โœ“ - - 86.85 DCO with unregularized OMP
OMP โœ“ โœ“ - 88.06 Orthogonality prevents collapse (+1.21%)
OMP โœ“ - โœ“ 88.02 Margin enforces class separability (+1.17%)
OMP (Full) โœ“ โœ“ โœ“ 88.25 Dual regularizations achieve peak synergy

Key Findings

  • Complementary and lightweight calibration: Each DCO sub-component provides an isolated gain of roughly +0.33%, culminating in a +0.78% boost when combined. This confirms that state-level residual injection, spatial deviation modulation, and channel rebalancing address distinct aspects of spectral degradation while incurring negligible overhead (<0.001M params, 3.19G FLOPs).
  • Sequential dependency between DCO and OMP: Ablation reveals that introducing OMP without DCO causes mIoU to plunge to 85.63% (substantially below the 87.32% baseline). Without high-frequency feature separability restored by DCO, multi-prototype learning disperses along uninformative directions. Once DCO is active, OMP leverages the refined multi-modal features to reach 88.25% mIoU.
  • Clear saturation behavior across sub-prototypes: Increasing the number of prototypes from \(K=2\) to \(K=3\) produces notable gains, whereas performance plateaus at \(K=4, 5\). This confirms that performance gains stem from genuine multi-modal mode coverage rather than extraneous parameter capacity.

Highlights & Insights

  • Turning duality approximation error into calibration signal: Rather than treating global normalized aggregation as an unavoidable low-pass filter, CRISP creatively leverages the difference between raw and scaled values as a rich residual source, executing linear-time calibration by reusing existing attention-like statistics.
  • Cross-hierarchy synergy: The architecture establishes a tight methodological loop: low-level DCO restores fine-grained geometric separability in the feature space, while high-level OMP accommodates the resulting complex distributions via orthogonally regularized sub-prototypes.
  • Transferable SSM frequency adaptation: The mechanism of reusing kernelized key/query statistics to inject \(O(L)\) residual corrections is highly generalizable and can be applied to medical imaging, industrial inspection, and edge-sensitive dense prediction tasks.

Limitations & Future Work

  • Lack of spatially adaptive modulation: Current DCO gating parameters rely primarily on channel-wise vectors and global spatial pooling, lacking dynamic spatial adaptivity based on localized texture roughness.
  • Prototype vulnerability in ultra-sparse classes: In remote sensing datasets with extreme class imbalance, rare categories may lack sufficient diverse samples to support multiple orthogonal prototypes effectively.
  • Future directions: The authors outline plans to develop spatially adaptive calibration conditioned on local frequency content and explore deployment on ultra-high-resolution hyperspectral imagery.
  • vs VSSD [ICCV 2025]: VSSD achieves linear scanning speed via non-causal duality and global aggregation but suffers from low-pass over-smoothing; CRISP resolves this with in-place DCO calibration and replaces the cumbersome UPerHead with a streamlined OMP decoder.
  • vs FreqFusion [TPAMI 2024] / LocalMamba [ECCVW 2024]: FreqFusion relies on expensive decoder-side spectral operations, and LocalMamba introduces windowed scans that interrupt hardware continuity; CRISP preserves full linear-scan efficiency with negligible computational cost.
  • vs NAPG [TGRS 2025] / Prototypical Contrastive Learning [TMM 2025]: Previous multi-prototype methods in remote sensing rely on offline clustering or complex contrastive pretraining; OMP achieves fully end-to-end prototype learning through dual intra-class and inter-class geometric penalties.

Rating

  • Novelty: โญโญโญโญโ˜† First to uncover the spectral low-pass bias of VSSD aggregation and introduce an in-place linear-time duality calibration operator.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across three challenging benchmarks, accompanied by rigorous component-level ablations, efficiency audits, and spectral visualizations.
  • Writing Quality: โญโญโญโญโญ Rigorous mathematical derivation combined with intuitive frequency-domain insights and clean design structure.
  • Value: โญโญโญโญโญ Provides a practical and lightweight blueprint for deploying modern State Space Models to dense high-resolution Earth observation tasks.