COLA: Continual Orthogonal Low-Rank Adaptation for Class-Incremental Learning¶
Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/autovisionproject/COLA
Area: AI Safety / Continual Learning
Keywords: continual learning, class-incremental learning, low-rank adaptation, orthogonal projection, catastrophic forgetting
TL;DR¶
To eliminate unbounded parameter growth and heavy replay storage in class-incremental learning, COLA introduces a rehearsal-free continual learning framework that tracks principal feature covariance structures online via an Oja-inspired learning rule and dynamically projects novel task representations onto complementary orthogonal subspaces within a shared, fixed-capacity LoRA adapter.
Background & Motivation¶
Class-Incremental Learning (CIL) aims to enable deep neural models to continually incorporate emerging object classes without suffering from catastrophic forgetting or destabilizing the stability-plasticity trade-off. In recent years, adapting large-scale Pre-trained Models (PTMs) has emerged as the dominant paradigm for CIL, as rich pre-trained representations provide strong generalization priors against distribution shifts across sequential tasks. Nevertheless, prominent parameter-efficient adaptation strategies remain constrained by severe scalability bottlenecks: prompt-based methods such as L2P, DualPrompt, and CODA-Prompt maintain dynamically expanding prompt pools or require test-time prompt selection, incurring inference overhead that scales with task count; meanwhile, rehearsal-based methods like HiDe-Prompt and InfLoRA rely on caching raw historical exemplars or intermediate representations, introducing substantial memory overhead and acute data privacy concerns.
To circumvent replay dependencies, gradient projection techniques such as OGD, GPM, InfLoRA, and CoSO seek to project model updates onto orthogonal subspaces. However, these methods depend on computing explicit Singular Value Decompositions (SVD) at task boundaries to construct hard null-space constraints. As the task sequence grows longer, accumulated historical bases become vulnerable to estimation noise and distribution drift, severely suppressing model plasticity and impairing adaptation to new tasks.
This paper tackles this dilemma by uniting low-rank parameter-efficient adaptation with streaming covariance-guided subspace orthogonalization. Core idea: insert fixed-capacity shared LoRA adapters into a frozen Vision Transformer backbone, incrementally track dominant feature eigenstructures across tasks using an Oja-inspired online update rule, and steer new task updates onto orthogonal complementary subspaces via a soft projection operator, achieving rehearsal-free, parameter-constant, and fully end-to-end class-incremental learning.
Method¶
Overall Architecture¶
COLA builds upon a frozen Vision Transformer (ViT-B/16) backbone, keeping base parameters \(\theta_0\) entirely frozen while inserting shared low-rank adaptation modules (rank \(r=10\)) into the Query and Value attention projection layers across all transformer blocks. Rather than caching historical exemplars or expanding task-specific adapters, COLA tracks the geometric evolution of learned representations entirely through streaming second-order statistics.
During forward propagation, input representations mapped into the low-rank subspace are first modulated by a soft orthogonal projection operator derived from the normalized global feature covariance matrix, thereby diverting feature activations away from subspaces dominated by earlier tasks. In parallel, mini-batch feature statistics are tracked online without backpropagation graphs via an Oja-inspired streaming rule, periodically consolidated into a global covariance matrix. The projected representations are dynamically combined with original low-rank features through a learnable residual scaling parameter before entering the attention computation, and the model is optimized end-to-end with standard cross-entropy loss over all encountered classes. At inference time, the model executes through the consolidated shared adapters with zero task-identity routing or sample retrieval overhead.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Sequential input x<br/>Frozen ViT backbone θ0"] --> B["Low-rank feature extraction<br/>Shared LoRA projection z = A x"]
B --> C["Soft orthogonal projection operator<br/>P matrix derived from global covariance Σ_global"]
C --> D["Oja-inspired online covariance accumulation & consolidation<br/>Streaming update of Σ_task and integration into Σ_global"]
C --> E["Adaptive residual scaling & feature fusion<br/>Balance stability and plasticity with learnable γ"]
D -.->|Incremental statistical guidance| C
E --> F["Integrated attention QKV & incremental classification<br/>Unified end-to-end rehearsal-free inference"]
Key Designs¶
1. Covariance-Guided Soft Orthogonal Projection Operator: Alleviating Cross-Task Representational Interference
To prevent parameter and representational collisions across sequential tasks in a shared low-rank space, prior projection methods enforce hard null-space constraints via explicit SVD, which is computationally expensive and overly sensitive to estimation bias. COLA maintains a compact \(r \times r\) global covariance matrix \(\boldsymbol{\Sigma}_{\text{global}}^{(i)}\) at each layer \(i\), summarizing dominant second-order statistics of all previously completed tasks. The soft orthogonal projection matrix is analytically defined as:
where \(\alpha > 0\) controls the projection intensity. Multiplying low-rank features to obtain \(\tilde{\mathbf{z}}^{(i,t)} = \mathbf{P}^{(i,t)} \mathbf{z}^{(i,t)}\) attenuates feature components aligned with historical dominant eigenvectors while steering new representations into complementary orthogonal directions. Unlike rigid orthogonal truncation, this soft alignment operator accommodates minor statistical estimation bias and preserves sufficient plasticity for novel class integration.
2. Oja-Inspired Online Covariance Accumulation and Consolidation: Rehearsal-Free Continual Knowledge Preservation
To eliminate sample buffers without suffering from catastrophic forgetting, COLA tracks the geometric structure of task representations online. Within each mini-batch of task \(t\), the batch covariance of projected features is computed:
The task-specific covariance matrix is then updated online following an Oja-inspired learning rule:
where \(\eta > 0\) is the learning rate and \(\beta > 0\) denotes a decay factor preventing unbounded scale expansion. This formulation enables \(\boldsymbol{\Sigma}_{\text{task}}\) to incrementally capture the principal eigenstructure of the task subspace without retaining computation graphs. Every \(K\) batches, task statistics are consolidated into the global covariance matrix via \(\boldsymbol{\Sigma}_{\text{global}}^{(i)} \leftarrow (1-\lambda)\boldsymbol{\Sigma}_{\text{global}}^{(i)} + \lambda \boldsymbol{\Sigma}_{\text{task}}^{(i,t)}\) with rate \(\lambda \in (0, 1)\). At task transitions, remaining task statistics are accumulated and \(\boldsymbol{\Sigma}_{\text{task}}\) is reset to \(\epsilon \mathbf{I}_r\). This decouples memory complexity entirely from sequence length.
3. Adaptive Residual Scaling and Shared LoRA Reuse: Balancing Stability and Plasticity
While orthogonal projections effectively shield historical knowledge from degradation, overly rigid projection can restrict the model's expressiveness on newly introduced tasks. COLA dynamically recombines projected and original features:
where \(\gamma\) is a learnable residual scaling parameter. The projected branch \(\tilde{\mathbf{z}}\) provides structural geometric regularization against catastrophic forgetting, while the residual path \(\gamma \mathbf{z}\) preserves task-specific feature expressiveness. The fused vectors modulate the attention Query and Value projections, allowing a single set of LoRA parameters and a unified linear head to be continually trained end-to-end across arbitrary task sequences without expanding parameter counts.
Loss & Training¶
During the training of task \(t\), the backbone remains frozen, and the network optimizes the learnable LoRA parameters \(\{\mathbf{A}^{(i,t)}, \mathbf{B}^{(i,t)}\}_{i=1}^L\) and \(\gamma\) using standard cross-entropy loss over all classes observed up to task \(t\):
The framework is trained using the Adam optimizer with a learning rate of 0.008 and a batch size of 128 (30 epochs on ImageNet-R and 20 epochs on other datasets). Hyperparameters are configured to \(\alpha=0.02\), \(\eta=10^{-4}\), \(\beta=0.01\), \(\lambda=0.3\), and consolidation interval \(K=10\). Covariance updates operate outside backpropagation, requiring neither replay buffers nor multi-stage distillation steps.
Key Experimental Results¶
Main Results¶
On the artistic domain variation benchmark ImageNet-R across 10-task, 20-task, and 25-task sequences, COLA consistently outperforms state-of-the-art prompt-based and LoRA-based continual learning methods. Results below report mean ± standard deviation across four independent runs using a ViT-B/16 backbone with LoRA rank fixed at 10:
| Method | ImageNet-R (10 Task) Acc (%) | ImageNet-R (10 Task) AAA (%) | ImageNet-R (20 Task) Acc (%) | ImageNet-R (20 Task) AAA (%) | ImageNet-R (25 Task) Acc (%) | ImageNet-R (25 Task) AAA (%) |
|---|---|---|---|---|---|---|
| Full Fine-Tuning | 60.57 ± 1.30 | 72.31 ± 1.23 | 49.95 ± 1.31 | 65.32 ± 0.92 | 40.12 ± 1.43 | 46.11 ± 0.98 |
| L2P (CVPR 2022) | 71.26 ± 0.65 | 76.13 ± 0.86 | 68.97 ± 0.57 | 74.16 ± 0.16 | 60.11 ± 1.09 | 65.12 ± 0.94 |
| DualPrompt (ECCV 2022) | 68.22 ± 0.30 | 73.81 ± 0.39 | 65.23 ± 0.45 | 71.30 ± 0.15 | 58.00 ± 0.66 | 64.32 ± 0.76 |
| CODA-Prompt (CVPR 2023) | 74.05 ± 0.41 | 78.14 ± 0.39 | 69.38 ± 0.34 | 73.95 ± 0.76 | 62.92 ± 0.34 | 67.11 ± 0.49 |
| HiDe-Prompt (WACV 2024) | 74.65 ± 0.14 | 78.46 ± 0.28 | 73.59 ± 0.21 | 77.93 ± 0.38 | 67.11 ± 0.43 | 70.56 ± 0.56 |
| RanPAC (NeurIPS 2023) | 74.58 ± 0.77 | 81.67 ± 0.48 | 72.40 ± 0.60 | 79.27 ± 0.40 | 71.17 ± 1.56 | 74.11 ± 0.32 |
| InfLoRA (CVPR 2024) | 74.75 ± 0.65 | 80.67 ± 0.56 | 69.89 ± 0.66 | 76.68 ± 0.76 | 68.01 ± 0.44 | 75.12 ± 0.16 |
| VPT-NSP2 (2024) | 78.25 ± 0.56 | 79.45 ± 0.76 | 74.68 ± 0.82 | 79.85 ± 0.61 | 67.12 ± 0.34 | 75.31 ± 0.46 |
| PLAN (ICCV 2025) | 75.25 ± 0.56 | 80.41 ± 0.33 | 73.25 ± 0.42 | 78.41 ± 0.26 | 71.43 ± 0.54 | 75.23 ± 0.45 |
| CoSO (NeurIPS 2025) | 76.52 ± 0.87 | 83.12 ± 0.76 | 72.52 ± 0.34 | 79.67 ± 0.65 | 67.35 ± 0.45 | 77.24 ± 0.43 |
| SD-LoRA (ICLR 2025) | 77.04 ± 0.51 | 82.55 ± 1.50 | 74.96 ± 0.40 | 80.02 ± 0.69 | 70.78 ± 0.43 | 74.71 ± 0.54 |
| COLA (Ours) | 77.56 ± 0.35 | 83.40 ± 0.38 | 76.12 ± 0.44 | 82.27 ± 0.28 | 76.01 ± 0.38 | 82.11 ± 0.47 |
On the adversarial benchmark ImageNet-A and fine-grained bird classification dataset CUB200 (10 tasks), COLA similarly outperforms prior techniques:
| Method | ImageNet-A (10 Task) Acc (%) | ImageNet-A (10 Task) AAA (%) | CUB200 (10 Task) Acc (%) | CUB200 (10 Task) AAA (%) |
|---|---|---|---|---|
| L2P | 42.94 ± 1.22 | 51.40 ± 1.96 | 65.18 ± 2.46 | 76.12 ± 1.29 |
| DualPrompt | 45.49 ± 0.90 | 54.68 ± 1.29 | 68.00 ± 1.20 | 79.40 ± 0.96 |
| CODA-Prompt | 45.36 ± 0.78 | 57.06 ± 0.96 | 71.92 ± 0.45 | 78.76 ± 0.77 |
| HiDe-Prompt | 42.70 ± 0.67 | 56.32 ± 0.48 | 69.12 ± 1.14 | 77.46 ± 1.01 |
| InfLoRA | 49.20 ± 1.12 | 60.92 ± 0.76 | 70.82 ± 0.44 | 81.39 ± 0.16 |
| PLAN | 54.12 ± 0.15 | 61.45 ± 0.76 | 67.14 ± 0.54 | 76.26 ± 1.17 |
| CoSO | 56.43 ± 0.16 | 65.54 ± 1.20 | 72.23 ± 0.87 | 81.62 ± 1.67 |
| SD-LoRA | 57.76 ± 0.51 | 66.56 ± 1.50 | 72.90 ± 0.79 | 82.07 ± 1.12 |
| COLA (Ours) | 59.20 ± 0.87 | 67.46 ± 0.58 | 74.51 ± 0.88 | 84.38 ± 0.97 |
In terms of computational efficiency and storage, COLA requires only 32.22 GFLOPs per forward pass (versus 35.12 GFLOPs for InfLoRA/SD-LoRA and ~70 GFLOPs for prompt baselines), maintains 0.37M learnable parameters, and requires zero stored features (0.00M stored features).
Ablation Study¶
Sensitivity analysis regarding the initialization value of residual scaling factor \(\gamma\) (mean over four independent runs):
| Initial \(\gamma\) | ImageNet-A Acc (%) | ImageNet-A AAA (%) | ImageNet-R Acc (%) | ImageNet-R AAA (%) | CUB200 Acc (%) | CUB200 AAA (%) | Note |
|---|---|---|---|---|---|---|---|
| 0.08 | 56.55 ± 2.15 | 65.95 ± 0.65 | 76.87 ± 0.24 | 81.44 ± 0.27 | 73.75 ± 0.86 | 82.10 ± 0.57 | Excessive projection constraint restricts plasticity |
| 0.10 | 57.93 ± 1.35 | 66.06 ± 0.78 | 76.98 ± 0.44 | 82.12 ± 0.32 | 74.02 ± 0.84 | 83.10 ± 0.97 | Steady performance gains |
| 0.20 (Default) | 59.20 ± 0.87 | 67.46 ± 0.58 | 77.56 ± 0.35 | 83.40 ± 0.38 | 74.51 ± 0.88 | 84.38 ± 0.97 | Optimal balance between stability and plasticity |
| 0.40 | 58.45 ± 1.87 | 65.43 ± 0.55 | 77.04 ± 0.52 | 82.10 ± 0.26 | 74.10 ± 0.81 | 83.03 ± 0.95 | Weakened orthogonal guidance |
| 0.60 | 55.43 ± 1.34 | 62.88 ± 0.68 | 73.23 ± 0.54 | 78.87 ± 0.98 | 69.45 ± 1.08 | 80.12 ± 1.02 | Constraint collapse; reverts to vanilla LoRA with severe forgetting |
Regarding LoRA projection placement, Q/V adaptation consistently outperforms K/V across 10-task and 20-task sequences on all three benchmarks (e.g., ImageNet-R 10-task Q/V achieves 77.56% Acc / 83.40% AAA versus 77.13% Acc / 83.09% AAA for K/V). Across architectures, ViT-B/16 incurs only 9.02% forgetting on ImageNet-R, while lightweight DeiT-S/16 achieves 53.92% Acc and 62.93% AAA with a modest peak memory footprint of 5.28GB.
Key Findings¶
- Soft orthogonal projection is the decisive anti-forgetting driver: Removing the projection matrix \(\mathbf{P}\) causes final accuracy to plunge by 2.6% on ImageNet-A, 2.7% on ImageNet-R, and 3.2% on CUB200, verifying that low-rank adaptation alone cannot prevent representational overwriting without geometric subspace decoupling.
- Superior scalability on extended task sequences: On the 25-task ImageNet-R benchmark, existing baselines degrade markedly (CoSO drops to 67.35% Acc / 77.24% AAA, SD-LoRA to 70.78% Acc / 74.71% AAA), whereas COLA preserves 76.01% Acc and 82.11% AAA, outperforming SD-LoRA by 5.23% Acc / 7.40% AAA and CoSO by 4.87% AAA.
- Residual scaling controls stability-plasticity polarization: Extreme choices of \(\gamma\) either impede learning of novel classes (small \(\gamma\)) or render orthogonal regularization ineffective (large \(\gamma\)), confirming that \(\gamma=0.20\) provides an optimal operating point.
Highlights & Insights¶
- Oja-inspired rule for parameter-efficient continual learning: By adapting classical streaming principal component estimation to low-rank covariance accumulation, COLA eliminates offline SVD operations at task boundaries, yielding an elegant, gradient-free method for knowledge consolidation.
- Soft projection operator replaces rigid null-space truncation: Normalizing global covariance into a soft decay operator suppresses historical interference while gracefully accommodating noise in streaming statistics and preserving gradient flow.
- Constant inference compute with zero parameter expansion: Because adapters are shared across all tasks and require no exemplar retrieval or routing, deployment complexity remains strictly invariant to the number of incremental tasks.
Limitations & Future Work¶
- Linear classification head imbalance: While the feature extractor is shielded by orthogonal projection, the final linear head remains susceptible to class imbalance between novel and past classes, where logit scale disparities could be further addressed via cosine classifiers.
- Capacity ceiling under extreme sequence lengths: Sharing a fixed rank (\(r=10\)) across hundreds of highly heterogeneous domains could eventually saturate the orthogonal subspace, motivating future research into dynamically expanding or sparse orthogonal bases.
Related Work & Insights¶
- vs InfLoRA (CVPR 2024): InfLoRA relies on hard gradient projections alongside exemplar replay to correct feature drift; COLA is completely rehearsal-free and uses soft covariance projection to reduce computation and storage overhead.
- vs CoSO (NeurIPS 2025): CoSO executes explicit SVD at task transitions to construct sequential orthogonal subspaces, which over-constrains optimization on long sequences (25-task ImageNet-R); COLA consolidates online covariance summaries into a soft constraint, sustaining robust accuracy over extended task sequences.
- vs SD-LoRA (ICLR 2025): SD-LoRA separates magnitude and directional learning but accumulates task-specific adapters, causing memory growth proportional to task count; COLA reuses a single low-rank module, maintaining a constant 0.37M parameter budget.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Unifies Oja's streaming rule with soft orthogonal subspace projection in low-rank adaptation, eliminating SVD and memory buffers]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Evaluated across ImageNet-R 10/20/25 tasks, ImageNet-A, and CUB200, including FLOPs, memory footprint, multi-backbone scalability, and extensive ablations]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, cohesive mathematical formulation, and well-structured empirical analysis]
- Value: ⭐⭐⭐⭐⭐ [Offers a highly practical, parameter-constant, and privacy-preserving paradigm for continual learning with foundation vision models]