LoCA: Spatially-Aware Low-Rank Convolutional Adaptation of Vision Foundation Models¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/ssojungan/loca
Area: Segmentation
Keywords: parameter-efficient fine-tuning, low-rank adaptation, convolutional networks, spatial inductive bias, domain-generalized semantic segmentation
TL;DR¶
To tackle topological collapse and effective receptive field reconvergence caused by naive 4D kernel flattening in LoRA, LoCA introduces a spatial-channel decoupled convolutional adaptation framework combining low-rank channel mixing, covariance SVD spatial basis refinement, and hierarchical rank scheduling, significantly outperforming full fine-tuning and prior PEFT baselines on domain-generalized segmentation and visual adaptation.
Background & Motivation¶
Vision Foundation Models (VFMs) have demonstrated exceptional visual representation capabilities across diverse downstream perception tasks. However, full fine-tuning (FFT) of large-scale models incurs prohibitive memory and computational overhead while causing catastrophic forgetting of pre-trained generic priors. Parameter-Efficient Fine-Tuning (PEFT), typified by Low-Rank Adaptation (LoRA), has achieved remarkable success by freezing pre-trained backbones and learning low-rank incremental matrices for linear self-attention layers in Transformers. Nonetheless, modern vision backbonesโsuch as ConvNeXt, hybrid MambaVision, and the U-Net architecture in Stable Diffusionโstill rely fundamentally on convolutional operators to provide weight sharing, sliding-window operations, and crucial spatial inductive biases.
When extending LoRA to convolutional layers, prior methods typically flatten the 4D convolutional kernel tensor into a 2D matrix before applying low-rank decomposition. This naive flattening forces spatial geometric topology (locality, directionality) into a monolithic cross-channel parameterization, leading to severe spatial-channel entanglement. Consequently, the Effective Receptive Field (ERF) undergoes a detrimental "spatial reconvergence" phenomenon during fine-tuning: the receptive field expands only transiently before rapidly reconverging to localized patterns, accompanied by high-frequency noise amplification and loss of informative low-frequency structures. Conversely, filter subspace approaches like FSF decompose kernels into spatial atoms and mixing coefficients, but they rely on lossy sparse coding approximations and freeze cross-channel mixing coefficients, which introduces accumulation error and strips the model of the flexibility to adapt inter-channel correlations in new domains.
Caught between the structural distortion of kernel flattening and the inflexibility of approximate subspace decomposition, adapting convolution layers demands a structure-preserving mechanism that ensures zero reconstruction error while orthogonally decoupling spatial basis evolution from cross-channel mixing. Core idea: propose Low-Rank Convolutional Adaptation (LoCA), a convolution-aware PEFT framework that decouples the kernel update into a low-rank channel adaptation branch for dense cross-channel mixing and an SVD-derived spatial basis refinement branch on depthwise diagonals, coupled with hierarchical rank scheduling to preserve pre-trained spatial priors without reconstruction loss.
Method¶
Overall Architecture¶
The core mechanism of LoCA parameterizes incremental convolutional updates into two decoupled yet collaborative pathways: a low-rank channel adaptation branch and a singular value decomposition (SVD) based spatial basis refinement branch. The input feature map passes through the frozen pre-trained convolutional kernel while concurrently feeding into the decoupled adaptation branch, after which their outputs are element-wise combined.
Specifically, for a convolutional layer parameterized by \(W_0 \in \mathbb{R}^{C_{\text{out}} \times C_{\text{in}} \times k_h \times k_w}\), LoCA leaves the original pre-trained weight \(W_0\) frozen without altering its values. The low-rank channel adaptation path models dense inter-channel transformations, whereas the spatial basis path extracts orthogonal spatial geometric bases via the spatial covariance matrix of pre-trained weights and modulates them along depthwise diagonals with learnable channel-specific coefficients. These components are composed into a structured incremental tensor \(\Delta W_{cs}\), strictly satisfying zero-initialization conditions to ensure functional equivalence at the start of training. Furthermore, channel ranks across stages are dynamically scheduled according to layer width and resolution.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Feature Map X"] --> PreW["Frozen Pre-trained Conv Layer<br/>W0 โ R^{Cout ร Cin ร kh ร kw}"]
In --> LoCA_Branch["LoCA Decoupled Adaptation Branch"]
subgraph LoCA_Branch["LoCA Core Module"]
direction TB
ChAdapt["Low-Rank Channel Adaptation<br/>BcAc Dense Cross-Channel Mixing"]
SpAdapt["SVD Spatial Basis Refinement<br/>Covariance SVD Basis S & Diagonal Udw"]
ChAdapt --> Comp["Channel-Spatial Composition<br/>ฮWcs = ฮWch + diag(ฮWsp)"]
SpAdapt --> Comp
Sched["Hierarchical Rank Scheduling<br/>Dynamic r_ch allocation by stage width"] -.-> ChAdapt
end
PreW --> OutAdd["Element-wise Addition โ<br/>W' = W0 + ฮWcs"]
Comp --> OutAdd
OutAdd --> Out["Output Feature Map Y"]
Key Designs¶
1. Low-Rank Channel Adaptation: Isolating Dense Cross-Channel Dependencies
Although flattening a 4D convolutional kernel enables standard matrix multiplication, burdening a single factorization with both channel transformations and spatial topology leads to feature entanglement. LoCA introduces a dedicated channel adaptation update \(\Delta W_{\text{ch}}\) to model cross-channel interactions in downstream tasks, parameterized by channel rank \(r_{\text{ch}}\): $$ \Delta W_{\text{ch}}^{\flat} = \frac{\alpha}{r_{\text{ch}}} B_c A_c $$ where \(A_c \in \mathbb{R}^{r_{\text{ch}} \times (C_{\text{in}} k_h k_w)}\) is initialized via Kaiming uniform distribution and \(B_c \in \mathbb{R}^{C_{\text{out}} \times r_{\text{ch}}}\) is initialized to zero before reshaping back to \(\Delta W_{\text{ch}} \in \mathbb{R}^{C_{\text{out}} \times C_{\text{in}} \times k_h \times k_w}\). Zero-initializing \(B_c\) guarantees \(\Delta W_{\text{ch}} = 0\) at training onset, perfectly preserving pre-trained representations without reconstruction distortion.
2. SVD Spatial Basis Refinement: Preserving Spatial Priors via Covariance Bases
To refine spatial geometric features (e.g., edges, textures, and orientations) without cross-channel interference, LoCA extracts deterministic bases from the spatial statistics of pre-trained kernels. Each kernel slice is reshaped into a length-\((k_h k_w)\) vector and standardized to zero-mean and unit-variance to construct \(W_{\text{norm}} \in \mathbb{R}^{C_{\text{out}}C_{\text{in}} \times k_h k_w}\), forming a spatial covariance matrix: $$ C_{\text{sp}} = W_{\text{norm}}^\top W_{\text{norm}} \in \mathbb{R}^{(k_h k_w) \times (k_h k_w)} $$ Performing SVD on \(C_{\text{sp}}\) with spatial rank \(r_{\text{sp}} = k_h k_w\) yields an initial basis tensor \(\mathcal{S} \in \mathbb{R}^{r_{\text{sp}} \times k_h \times k_w}\), which remains learnable for fine-tuning. To avoid channel crosstalk and focus exclusively on spatial modulation, spatial updates are restricted to depthwise diagonals. For channel \(i\) with \(i \le D = \min(C_{\text{out}}, C_{\text{in}})\), the update is formed by combining learnable coefficients \(U_{\text{dw}} \in \mathbb{R}^{D \times r_{\text{sp}}}\) with spatial bases: $$ \Delta W_{\text{sp}}^{\text{diag}}[i] = \sum_{m=1}^{r_{\text{sp}}} U_{\text{dw}}[i, m] \cdot \mathcal{S}_m $$ Using the Kronecker delta \(\delta_{ij}\), the full spatial tensor is \(\Delta W_{\text{sp}}[i, j, :, :] = \delta_{ij} \Delta W_{\text{sp}}^{\text{diag}}[i]\). Zero-initializing \(U_{\text{dw}}\) smoothly introduces spatial adaptation and effectively prevents the spatial reconvergence of receptive fields during later training phases.
3. Channel-Spatial Composition: Unifying Decoupled Weights for Zero-Cost Inference
To synthesize the two independent pathways without memory or latency fragmentation, LoCA combines them directly in weight space: $$ \Delta W_{cs}[i, j, :, :] = \Delta W_{\text{ch}}[i, j, :, :] + \delta_{ij} \Delta W_{\text{sp}}^{\text{diag}}[i] $$ The effective adapted convolutional weight becomes \(W' = W_0 + \Delta W_{cs}\). During inference deployment, \(\Delta W_{cs}\) is seamlessly merged into the frozen pre-trained kernel \(W_0\), adding zero inference latency or extra runtime compute.
4. Hierarchical Rank Scheduling: Aligning Adaptation Capacity with Backbone Hierarchy
Convolutional VFMs feature hierarchical architectures with channel dimensions increasing progressively across stages. Applying a uniform rank across all layers restricts model expressivity in semantic stages while wasting parameters in shallow stages. LoCA keeps the spatial rank \(r_{\text{sp}} = k_h k_w\) fixed while dynamically scaling the channel rank according to stage-specific width \(C_{\text{out}}^{(s)}\) under global budget \(R\): $$ r_{\text{ch}}^{(s)} = \left\lfloor R \cdot \frac{C_{\text{out}}^{(s)}}{\sum_i C_{\text{out}}^{(i)}} \right\rfloor $$ This allocates compact ranks to shallow stem layers while provisioning broader capacity to deeper semantic blocks, maximizing parameter efficiency for dense visual downstream tasks.
Key Experimental Results¶
Main Results¶
On out-of-distribution domain-generalized semantic segmentation (DGSS, GTAV \(\rightarrow\) Cityscapes, BDD100K, Mapillary), LoCA outperforms Full Fine-Tuning (FFT) and leading PEFT approaches across diverse backbones:
| Method | Backbone | Trainable Param. (M) | GFLOPs | Cityscapes (mIoU) | BDD100K (mIoU) | Mapillary (mIoU) | Avg. mIoU (%) |
|---|---|---|---|---|---|---|---|
| FFT | ConvNeXt-B | 87.56 | 81 | 62.18 | 57.01 | 65.00 | 61.40 |
| LoRA (Linear) | ConvNeXt-B | 2.90 | 81 | 63.90 | 57.87 | 65.53 | 62.43 |
| LoRA (4D Flatten) | ConvNeXt-B | 17.20 | 81 | 64.17 | 56.98 | 65.74 | 62.30 |
| Conv-Adapter | ConvNeXt-B | 2.30 | 81 | 59.94 | 56.47 | 63.48 | 59.96 |
| CoLoRA | ConvNeXt-B | 2.20 | 81 | 61.87 | 55.70 | 64.04 | 60.53 |
| FSF | ConvNeXt-B | 0.60 | 81 | 60.15 | 56.98 | 63.70 | 60.27 |
| LoCA (Ours) | ConvNeXt-B | 3.40 | 81 | 66.46 | 58.53 | 66.29 | 63.76 |
| FFT | ConvNeXt-L | 196.20 | 152 | 65.50 | 59.10 | 67.01 | 63.87 |
| LoRA (Linear) | ConvNeXt-L | 4.30 | 152 | 66.95 | 60.46 | 68.45 | 65.29 |
| LoCAโก (Ours+SVD) | ConvNeXt-L | 5.00 | 152 | 69.73 | 62.03 | 70.62 | 67.46 |
| FFT | MambaVision-B | 96.70 | 211 | 36.05 | 30.13 | 31.39 | 32.52 |
| LoCA (Ours) | MambaVision-B | 2.50 | 211 | 45.21 | 41.68 | 45.70 | 44.20 |
In generative subject-driven personalization using Stable Diffusion v1.4 on DreamBooth, LoCA demonstrates superior subject fidelity while maintaining competitive prompt responsiveness:
| Method | DINO Score โ | CLIP-I Score โ | CLIP-T Score โ | Note |
|---|---|---|---|---|
| Pretrained | 0.320 | 0.643 | 0.267 | zero adaptation |
| Real Images | 0.711 | 0.857 | โ | upper bound reference |
| Textual Inversion | 0.564 | 0.739 | 0.213 | prompt embedding tuning |
| DreamBooth (FFT) | 0.642 | 0.794 | 0.236 | full U-Net tuning |
| LoRA | 0.637 | 0.792 | 0.239 | standard linear LoRA |
| FSF | 0.572 | 0.715 | 0.313 | high text score, severe subject collapse |
| LoCA (r16, Ours) | 0.709 | 0.801 | 0.280 | highest DINO/CLIP-I, strong text fidelity |
Ablation Study¶
Incremental ablation on domain-generalized semantic segmentation (DGSS with ConvNeXt-L) highlights the cumulative contributions of each core module:
| Configuration | Cityscapes (mIoU) | BDD100K (mIoU) | Mapillary (mIoU) | Avg. mIoU (%) | Gain vs. Baseline |
|---|---|---|---|---|---|
| LoRA Linear Baseline | 66.95 | 60.46 | 68.45 | 65.29 | baseline |
| ๏ผ Naive Conv LoRA (4D Flatten) | 65.74 (-1.21) | 60.50 (+0.04) | 68.62 (+0.17) | 64.95 | -0.34 |
| ๏ผ Channel Mixing | 68.22 (+1.27) | 61.12 (+0.66) | 69.13 (+0.68) | 66.16 | +0.87 |
| ๏ผ Spatial Basis Refinement | 69.44 (+2.49) | 61.39 (+0.93) | 70.26 (+1.81) | 67.03 | +1.74 |
| ๏ผ Hierarchical Rank Scheduling | 69.73 (+2.78) | 62.03 (+1.57) | 70.62 (+2.17) | 67.46 | +2.17 |
Comparison of spatial basis initialization strategies on VTAB-1k and DGSS:
| Initialization Strategy | VTAB-1k Avg. Acc (%) | DGSS Avg. mIoU (%) | Note |
|---|---|---|---|
| Zero | 75.7 | 65.6 | no pre-trained spatial prior at onset |
| Flatten SVD | 73.0 | 66.2 | collapses spatial geometry, severe classification drop |
| Uniform Random | 75.7 | 66.3 | lacks directional and texture statistics |
| Covariance SVD (Ours) | 75.9 | 66.5 | extracts principal spatial patterns effectively |
Key Findings¶
- Naive flattening causes negative transfer: Directly flattening 4D convolution kernels degrades DGSS performance by 0.34% on average (and by 1.21% on Cityscapes), proving that destroying 2D spatial locality harms dense visual prediction.
- Decoupled spatial-channel adaptation delivers highest gains: Introducing channel mixing and covariance-based spatial refinement provides consecutive boosts of +0.87% and +0.87% (cumulative +1.74%), verifying that decoupled parameterization restores broad effective receptive fields.
- Substantial benefits for emerging hybrid architectures: On MambaVision-B, LoCA improves mIoU by +11.68 points (+35.9% relative gain) over FFT while utilizing only ~2.5% of parameters, confirming that preserving convolutional inductive biases is critical for dense prediction with hybrid backbones.
Highlights & Insights¶
- Spatial-Channel Decoupled Adaptation: Elegantly addresses spatial-channel entanglement in convolution by separating full cross-channel transformations from depthwise spatial basis refinement.
- Covariance SVD Spatial Basis Extraction: Extracts geometrically meaningful principal patterns directly from pre-trained spatial covariances, circumventing the lossy reconstruction of sparse coding and the topological collapse of tensor flattening.
- Broad Versatility & Zero Inference Overhead: Seamlessly integrates into pure ConvNets (ConvNeXt, ResNet), hybrid vision backbones (MambaVision), and generative diffusion models (Stable Diffusion U-Net), collapsing into base weights during deployment.
Limitations & Future Work¶
- Computational Scaling for Very Large Kernels: The spatial covariance matrix has dimension \((k_h k_w) \times (k_h k_w)\), which is negligible for \(3 \times 3\) or \(7 \times 7\) kernels but may incur non-trivial SVD initialization overhead if applied to extreme large-kernel networks (e.g., \(31 \times 31\)).
- Fixed Spatial Rank: Currently, spatial rank \(r_{\text{sp}}\) is set to \(k_h k_w\) throughout the network; joint dynamic rank allocation across both spatial and channel dimensions across different frequencies represents an appealing future direction.
Related Work & Insights¶
- vs LoRA / LoRA-C: Conventional LoRA flattens 4D kernels into 2D matrices, leading to spatial reconvergence and degraded receptive fields; LoCA maintains independent spatial and channel adaptation to preserve spatial inductive biases.
- vs Filter Subspace Fine-Tuning (FSF): FSF relies on approximate sparse coding and freezes cross-channel coefficients, introducing reconstruction error and limiting inter-channel adaptation; LoCA ensures zero reconstruction error and jointly updates channel and spatial components.
- vs SoMA / PiSSA: While SoMA and PiSSA apply SVD to linear Transformer layers, LoCA innovatively applies SVD to spatial covariance matrices, establishing a principled orthogonal adaptation paradigm for convolution.
Rating¶
- Novelty: โญโญโญโญโญ Decouples channel and spatial adaptation for 4D convolutions with covariance SVD spatial bases.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation across classification (VTAB-1k, FGVC), domain-generalized segmentation (DGSS), DreamBooth generation, and multiple backbones.
- Writing Quality: โญโญโญโญโญ Clear progression from ERF dynamics to mathematical formulations and spectral analysis.
- Value: โญโญโญโญโญ Sets a robust, parameter-efficient fine-tuning benchmark for convolutional and hybrid vision foundation models.