Skip to content

VD-LoRA: Adaptive Reuse of Low-Rank Directions for Continual Learning

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/LuQiongD/VD-LoRA
Area: Model Compression
Keywords: Continual Learning / Class-Incremental Learning / Low-Rank Adaptation / Bayesian Inference / Uncertainty Estimation

TL;DR

Addressing capacity collapse and plasticity degradation in rehearsal-free class-incremental learning caused by rigid orthogonality in LoRA, VD-LoRA introduces recursive direction-wise variational Bayesian modeling that leverages posterior variance to adaptively calibrate protection strength and guide low-risk direction reuse under strict rank constraints.

Background & Motivation

In class-incremental learning (CIL), models must sequentially acquire new classes without suffering catastrophic forgetting of previously learned ones. In the era of vision foundation models, parameter-efficient fine-tuning (PEFT) based on Low-Rank Adaptation (LoRA) has emerged as the dominant paradigm for rehearsal-free CIL because it freezes the pre-trained backbone and adapts only compact low-rank residual matrices. To suppress cross-task interference, recent representative methods such as InfLoRA, BiLoRA, and PLAN enforce strict orthogonality constraints, decomposing weight updates over a fixed dictionary of column basis vectors and requiring each new task to inhabit an isolated, non-overlapping coordinate subspace.

However, strict rank budgets and long task sequences render deterministic hard orthogonality overly conservative. In conventional orthogonal LoRA architectures, once a low-rank direction is allocated, it is permanently locked and protected. As tasks accumulate, available update directions are rapidly exhausted, forcing subsequent tasks to adapt within an increasingly restricted residual subspace. This induces severe gradient–subspace mismatch and under-adaptation, severely crippling the model's plasticity. Crucially, uniform hard protection treats all historical directions as equally critical, ignoring the inherent heterogeneity in direction sensitivity: while certain directions preserve highly sensitive decision boundaries, others carry low interference risk and could be safely reused.

Overcoming this trade-off requires abandoning binary "once-used, always-frozen" heuristics in favor of fine-grained directional uncertainty estimation to discern truly sensitive components from adaptable ones. The core idea is to establish a recursive variational posterior over LoRA basis coefficients, using posterior precision to adaptively scale protection strength and formulating an uncertainty-aware utility score to selectively reuse low-risk directions while preserving task-relevant gradient alignment under fixed rank constraints.

Method

Overall Architecture

VD-LoRA operates on a frozen pre-trained backbone paired with a shared LoRA dictionary bank, decomposing the low-rank update into column-basis directions \(e_i\) scaled by coefficient vectors \(b_i\). When a new task arrives, the pipeline executes three sequential stages: first, lightweight online gradient probing interacts with historical uncertainty variances to compute utility scores and dynamically activate the top-\(r\) directions best suited for adaptation; second, variational updates optimize the selected directions against task loss while a recursive prior penalizes deviation from historical posteriors; concurrently, unselected inactive directions receive uncertainty-scaled perturbations for stability consistency regularization, preventing future activation shifts from disrupting current boundaries; finally, inference uses posterior mean representations or Monte Carlo predictive averaging to deliver well-calibrated predictions.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Current task data and bank state<br/>Input batches and recursive prior from previous task"] --> B["Uncertainty-guided direction selection<br/>Probe gradients combined with prior variance for Top-r selection"]
    B --> C["Variational modeling of direction uncertainty<br/>Optimize task loss and recursive Gaussian KL divergence"]
    B --> D["Stability regularization for unselected directions<br/>Inject uncertainty-scaled noise and enforce prediction consistency"]
    C --> E["Write back updated statistics<br/>Store posterior mean and variance in shared dictionary bank"]
    D --> E
    E --> F["Posterior predictive averaging<br/>Perform robust calibrated inference across task sequences"]

Key Designs

1. Variational Modeling of Direction Uncertainty: Tracking direction consolidation via recursive variance To overcome the blind uniform freezing of traditional orthogonal LoRA, this design reformulates deterministic basis coefficients into probabilistic variables. Given a fixed dictionary of column basis vectors \(\{e_i\}_{i=1}^{d_{\mathrm{in}}}\), the weight update for task \(t\) is formulated as \(\Delta W_t = \sum_{i \in S_t} b_{t,i} e_i^\top\). The method places a diagonal Gaussian variational posterior on the coefficient of each active direction, \(q_t(b_i) = \mathcal{N}(\mu_{t,i}, \operatorname{diag}(\sigma_{t,i}^2))\). Here, the posterior variance \(\sigma_{t,i}^2\) directly quantifies directional uncertainty: a smaller variance signals strong consolidation by previous tasks demanding strict preservation, whereas a larger variance indicates residual flexibility available for adaptation. Across task boundaries, the posterior from the most recent task acts recursively as the prior for the next, \(p_t(b_i) = q_{t-1}(b_i)\) (with \(p_1(b_i) = \mathcal{N}(0, I)\)). The variational objective balances task empirical loss with directional KL divergence: $\(\mathcal{L}_v = \mathbb{E}_{q_t}\left[\mathcal{L}_{\mathrm{task}}(W_t; \mathcal{D}_t)\right] + \beta \sum_{i \in S_t} \mathrm{KL}\left(q_t(b_i) \,\|\, p_t(b_i)\right)\)$ Because prior and posterior are diagonal Gaussians, the KL divergence admits a closed-form penalty that penalizes movements on low-variance directions, implementing smooth adaptive soft protection instead of rigid geometric projections.

2. Uncertainty-Guided Direction Selection: Utility scoring balancing adaptation gain and interference risk Under a strict rank budget \(r\), determining which subset \(S_t\) of directions to activate governs both plastic learning and stability. Relying solely on task gradients greedily overwrites consolidated knowledge, whereas strict orthogonal exclusion rapidly exhausts capacity. This design probes the gradient \(g_{t,i} = \nabla_{b_i} \mathcal{L}_{\mathrm{task}}\) using a small number of initial mini-batches and blends it with the inherited posterior uncertainty \(\sigma_{t-1,i}^2\) into a principled directional utility score: $\(u_{t,i} = g_{t,i}^\top \operatorname{diag}\left((\sigma_{t-1,i}^2 + \varepsilon)^\gamma\right) g_{t,i}\)$ where \(\varepsilon\) guarantees numerical stability and \(\gamma\) regulates the uncertainty weighting. A candidate direction achieves a high utility score only when it satisfies two criteria simultaneously: it aligns strongly with current gradient demand (large \(\|g_{t,i}\|\)) and carries sufficient historical uncertainty (large \(\sigma_{t-1,i}^2\), denoting low interference risk). Highly sensitive directions with tiny variances are automatically penalized even if their gradients are substantial. The model selects the top-\(r\) directions to form \(S_t\), realizing safe, selective reuse of low-risk directions alongside unallocated coordinates.

3. Stability Regularization for Unselected Directions: Injecting uncertainty perturbations against future drift In long incremental sequences, directions unselected for the current task (\(i \notin S_t\)) remain frozen in the dictionary, but could be repurposed by subsequent tasks. If the model's current representations are fragile to fluctuations along these dormant coordinates, future updates will induce severe feature drift and undermine the current task's decision boundaries. To prevent this, the method injects forward-looking perturbations into layer activations \(h\) using the uncertainty of unselected directions: $\(\tilde{h} = h + \rho \sum_{i \notin S_t} \sigma_{t-1,i} \odot \epsilon_i, \quad \epsilon_i \sim \mathcal{N}(0, I)\)$ The injected perturbation magnitude is directly proportional to each direction's uncertainty \(\sigma_{t-1,i}\), subjecting flexible directions that are prone to future reuse to stronger synthetic noise. The network is then constrained to output invariant predictive distributions through stop-gradient consistency regularization: $\(\mathcal{L}_{\mathrm{stab}} = \mathrm{KL}\left(\mathrm{sg}(p) \,\|\, \tilde{p}\right)\)$ The combined training objective is \(\mathcal{L} = \mathcal{L}_v + \lambda \mathcal{L}_{\mathrm{stab}}\), conferring structural robustness against subsequent parameter evolution.

Loss & Training

The overall optimization objective is \(\mathcal{L} = \mathcal{L}_v + \lambda \mathcal{L}_{\mathrm{stab}}\). LoRA modules are inserted into the key and value projection matrices of the pre-trained ViT-B/16 architecture. Training utilizes the Adam optimizer (\(\beta_1=0.9, \beta_2=0.999\)) with a batch size of 128. Online gradient probing requires only 1–2 initial mini-batch forward-backward passes to rank direction utilities before freezing \(S_t\) for the rest of the task. Reparameterization gradients update variational parameters smoothly. Hyperparameters remain robust across configurations: \(\beta_{\mathrm{max}} \in [10^{-4}, 10^{-3}]\), \(\lambda \in [0.05, 0.2]\), and \(\gamma \in [0.25, 0.75]\).

Key Experimental Results

Main Results

Evaluations span ImageNet-R (partitioned into 5, 10, and 20 tasks), CIFAR-100, ImageNet-A, and VTAB on a pre-trained ViT-B/16 backbone. Metrics report final accuracy (Acc) and Average Anytime Accuracy (AAA).

Dataset / Setting Metric VD-LoRA (Ours) Best Baseline (CL-LoRA / LoRA-DRS) Gain
ImageNet-R (N = 5) Acc / AAA (%) 80.22 (±0.52) / 85.77 (±0.38) 79.88 (±0.52) / 84.98 (±0.49) +0.34 / +0.79
ImageNet-R (N = 10) Acc / AAA (%) 79.55 (±0.39) / 84.97 (±1.23) 78.94 (±0.54) / 84.49 (±0.61) +0.61 / +0.48
ImageNet-R (N = 20) Acc / AAA (%) 77.53 (±0.35) / 83.90 (±0.77) 76.11 (±0.24) / 83.23 (±0.92) +1.42 / +0.67
CIFAR-100 (N = 10) Acc / AAA (%) 89.69 (±0.31) / 92.72 (±0.11) 88.52 (±0.37) / 92.01 (±0.16) +1.17 / +0.71
ImageNet-A (N = 10) Acc / AAA (%) 60.27 (±0.88) / 70.18 (±1.24) 59.14 (±0.39) / 69.49 (±1.73) +1.13 / +0.69
VTAB (19 tasks) Acc / AAA (%) 94.11 (±0.22) / 94.45 (±0.63) 93.66 (±0.29) / 94.13 (±0.48) +0.45 / +0.32

Ablation Study

Progressive ablations on ImageNet-R isolate the impact of recursive KL regularization (RKL), uncertainty-guided direction selection (SEL, replaced with random selection when disabled), and stability regularization (STR). Calibration quality is examined via Expected Calibration Error (ECE) and Negative Log-Likelihood (NLL).

Configuration & Modules RKL (Recursive KL) SEL (Adaptive Selection) STR (Stability Noise) N = 5 AAA (%) N = 10 AAA (%) N = 20 AAA (%)
Baseline (Random Selection, No KL) ✗ ✗ ✗ 80.22 76.54 68.95
+ Recursive Variational KL ✓ ✗ ✗ 83.30 80.20 73.80
+ Uncertainty-Guided Selection ✓ ✓ ✗ 85.10 83.60 82.60
Full VD-LoRA Model ✓ ✓ ✓ 85.77 84.97 83.90

Calibration evaluations on ImageNet-R (\(N=10\)) confirm substantial uncertainty benefits: VD-LoRA drops ECE to 0.0120 (versus 0.0204 for CL-LoRA and 0.0256 for InfLoRA) and reduces NLL to 0.7001 (compared to 0.8457 for CL-LoRA and 1.1271 for InfLoRA).

Key Findings

  • Widening advantage over longer task horizons: When task sequences are short (\(N=5\)), subspace capacity is plentiful and VD-LoRA maintains a modest +0.34% Acc margin. But as sequences expand to \(N=20\), hard orthogonality methods suffer severe capacity starvation, causing VD-LoRA's margin to surge to +1.42% (77.53% vs. 76.11%), demonstrating that adaptive reuse breaks the capacity ceiling.
  • Superior plasticity under extreme rank constraints: In rank ablation benchmarks (\(N=20\)), setting \(r=1\) yields 75.92% Acc for VD-LoRA, outperforming InfLoRA's 69.52% by +6.40%. Tracking the directional plasticity ratio \(\eta^2(t)\) indicates that VD-LoRA preserves dramatically higher task-relevant gradient energy throughout the entire continuum.
  • Posterior variance faithfully reflects empirical sensitivity: Directly perturbing posterior means by \(\pm\delta\) on held-out old classes shows a strong inverse correlation between posterior standard deviation and empirical loss degradation (Pearson \(r = -0.79\), Spearman \(r = -0.71\)), validating the theoretical premise that posterior variance accurately gauges directional sensitivity.

Highlights & Insights

  • Transitioning from geometric exclusion to probabilistic soft consolidation: While prior literature viewed interference elimination as a rigid orthogonal projection problem, VD-LoRA reframes it as directional uncertainty management, replacing binary subspace freezing with smooth precision-based regularization.
  • Utility formulation resolves dual exploration-exploitation tension: By multiplying online gradient energy with prior variance, the utility score elegantly identifies directions providing maximal task adaptation with minimal retroactive interference in minimal compute time.
  • Preemptive stability regularization against dormant directions: Injecting uncertainty-scaled noise along unselected directions forces representations to be invariant to future subspace shifts, offering an insightful design principle for modular lifelong learning.

Limitations & Future Work

  • Authors' admitted limitations: The architecture enforces a uniform, static rank budget \(r\) across all tasks and layers, lacking dynamic rank adjustment tailored to task complexity. Furthermore, the training paradigm assumes explicit, discrete task boundaries rather than task-free continuous data streams.
  • Independent observations: Gradient probing relies on initial mini-batches; in extreme few-shot or heavily class-imbalanced incremental regimes, initial gradient estimates could exhibit sampling noise. Additionally, total direction capacity is bound by the input embedding dimension \(d_{\mathrm{in}}\).
  • Promising directions: Future exploration could combine directional Bayesian tracking with differentiable rank allocation (\(r_t\)) and expand direction-level continual adaptation to test-time adaptation (TTA) and autoregressive multimodal generation.
  • vs InfLoRA / BiLoRA: InfLoRA and BiLoRA enforce strict or near-orthogonal subspace isolation that permanently locks directions upon allocation, causing capacity exhaustion in long sequences. VD-LoRA replaces hard exclusion with variance-calibrated soft penalties, permitting safe reuse of low-risk directions.
  • vs CL-LoRA / SD-LoRA: CL-LoRA and SD-LoRA rely on heuristic adapter composition and task routing that can complicate inference. VD-LoRA merges directional statistics directly into a shared dictionary bank, achieving clean single-pass inference without auxiliary routing overhead.
  • vs BLoB / Bayesian-LoRA: Existing Bayesian LoRA approaches primarily address uncertainty calibration or parameter pruning within single-task fine-tuning. VD-LoRA pioneers directional recursive posterior tracking specifically formulated to balance stability and plasticity in lifelong continual learning.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Formulates direction-wise variational Bayesian inference to replace rigid orthogonal freezing with principled direction reuse.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Comprehensive evaluations across 4 benchmarks, multiple task horizons, plasticity ratios, and empirical sensitivity correlations.
  • Writing Quality: ⭐⭐⭐⭐⭐ Flawless narrative structure, natural integration of theoretical Bayesian formulation with empirical challenges.
  • Value: ⭐⭐⭐⭐⭐ Provides an elegant and highly effective solution to the stability-plasticity dilemma in parameter-efficient continual learning.