Skip to content

Stabilizing the Dynamic Low-Rank Training

Conference: NeurIPS2026
arXiv: 2609.32615
Area: Model Compression
Keywords: dynamic low-rank training, curvature coupling, compensation buffer, rank adaptation, parameter-efficient fine-tuning

TL;DR

SDLRT retains the leading singular directions discarded in the previous iteration and combines compensated basis augmentation with negative feedback on truncation tolerance to mitigate rank collapse under aggressive compression, achieving 85.15% average accuracy on six SuperGLUE validation tasks.

Background & Motivation

Training low-rank weights directly differs from training a full model and then pruning or factorizing it: each layer starts with two narrow basis matrices and a small core, and both training and inference can use this compact representation. Dynamic low-rank training (DLRT) updates the factors through manifold projection and adapts layer ranks during training, avoiding repeated dense training and post-hoc compression. However, how much information a low-rank representation stores and whether optimization can discover suitable directions are separate questions. The paper's VGG-19/CIFAR-100 diagnostic shows that, at approximately 69.6% initial compression, DLRT trails the full model by 25.9 percentage points in test accuracy and also has low training accuracy. This is not simply overfitting: the compact network has lost sufficient trainable expressive capacity.

The difficulty is that singular-value truncation removes not only small components of the current weights but also information affecting subsequent subspace motion. A direction that is weak now can still influence how retained directions rotate next. Comparing the best low-rank approximation at each point of a full-parameter training trajectory with training confined to the low-rank manifold, the authors identify two discrepancies: gradients are evaluated at different points, and curvature coupling is omitted. Under aggressive compression, the spectral gap between retained and discarded singular values can be small, making the coupling non-negligible. Rank reduction weakens expressivity and encourages further truncation, creating a reinforcing loop.

Permanently increasing layer ranks could restore capacity but would weaken compression benefits. Instead, the paper gives training a short-term spectral memory and reduces truncation when ranks fall, allowing discarded yet potentially useful directions to re-enter the next update. Core idea: keep the final weights low-rank without treating each truncation as permanent forgetting; reuse recently discarded directions in a temporarily augmented space and control tolerance to suppress continuing rank collapse.

Method

Overall Architecture

SDLRT maintains each layer's low-rank factors and the previous iteration's compensation buffer as its training state, producing updated factors and a new buffer. Each iteration proceeds through compensated basis augmentation, a Galerkin core update, and spectral truncation with rank feedback. The buffer assists training only; it is not an additional branch of the final inference weights.

For a matrix layer, weights are represented as \(W=USV^{\top}\), with \(r\) columns in each basis and an \(r\times r\) core. The algorithm first updates auxiliary matrices for the left and right factors along the gradient flow, then orthogonalizes a combination of these updates, the old bases, and buffered directions. It updates the core in this augmented space, performs an SVD of the core, retains the directions corresponding to the new rank, and buffers the leading directions immediately beyond the cutoff.

Thus, “not increasing the final rank” does not mean that the training space never expands. Temporary QR bases and the core can be larger, while the final retained rank remains tolerance-adaptive. Inference uses only the final three retained factors. The full-parameter trajectory in the derivation is a theoretical reference, and the full model in the stability experiment is a measurement reference; neither is a full-parameter shadow model maintained at every SDLRT update.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Training batch and<br/>low-rank factors"] --> B["Compensated basis<br/>augmentation"]
    P["Previous training buffer"] --> B
    B --> C["Galerkin core update"]
    A -.->|Task-loss gradient supervision| C
    C --> D["Spectral truncation<br/>and rank feedback"]
    D -->|New discarded directions, training only| P
    D -->|Factors and tolerance, next iteration| A
    D -->|Training ends, retained factors only| E["Final factorized inference"]

Key Designs

1. Compensated basis augmentation: let recently discarded directions participate in subspace updates again

The paper first explains why retaining the largest singular values does not automatically make low-rank training follow the full model. Consider full weights evolving under their gradient flow, and take their best rank \(r\) approximation at a given time. Differentiating its left and right singular subspaces introduces coupling between retained and discarded directions. The two coupling coefficients given in the paper are:

\[ \mathcal{P}_{ij}=\frac{\sigma_{r+j}^{2}}{\sigma_i^{2}-\sigma_{r+j}^{2}},\qquad \mathcal{Q}_{ij}=\frac{\sigma_i\sigma_{r+j}}{\sigma_i^{2}-\sigma_{r+j}^{2}}. \]

Here \(\sigma_i\) belongs to the retained part and \(\sigma_{r+j}\) to the discarded part. These coefficients multiply cross-directional projections of the gradient along the full trajectory; they are not simply penalties on tail energy. The denominator shows that risk depends not only on the magnitude of discarded singular values but also on the spectral gap across the cutoff. The derivation assumes no repeated singular values crossing the truncation boundary, so the formulas cannot be applied directly when their denominator vanishes.

Standard dynamic low-rank approximation (DLRA) projects the gradient evaluated at the low-rank point: it neither evaluates it at the full-parameter point nor retains the coupling correction. SDLRT does not explicitly compute every coupling coefficient. Instead, it uses a practical approximation: retain the leading left and right singular directions discarded in the previous iteration so that the augmented space permits these directions to influence weight changes again. This recovers some directional information, not exact full-gradient tracking or all curvature corrections.

Specifically, the algorithm initializes \(K_0=US\) and \(L_0=VS^{\top}\) from the old factors and updates them along the gradient flow to obtain \(K_1\) and \(L_1\). On the left, it concatenates the updated matrix, the old left basis, and the left compensation buffer; the right side is symmetric. QR then supplies orthogonal augmented bases. The old bases ensure that the new space still represents the previous weights, the updates supply current descent directions, and the buffer prevents recent spectral-tail directions from being forgotten at every step. These columns serve different roles; the buffer is not an independently trained copy of the full weights.

2. Galerkin core update: enlarge the representation space without first changing the old weights

Once augmented bases are available, arbitrarily initializing a larger core would introduce an unjustified weight perturbation into the claimed stability improvement. SDLRT instead projects the previous weights into the augmented coordinates for initialization and updates only the core within that coordinate space. Denoting the augmented bases in this section by \(\bar U,\bar V\), the core operation is:

\[ \bar S(0)=\bar U^{\top}W_{\mathrm{prev}}\bar V,\qquad \dot{\bar S}=-\bar U^{\top}\nabla_W\mathcal{L}(\bar U\bar S\bar V^{\top})\bar V. \]

Because the augmented bases contain the old left and right bases, the initial augmented representation reconstructs the same previous weights. Buffer columns provide freedom for subsequent motion rather than immediately adding discarded weights to the network. Compensation therefore changes where the next update can move, not whether every removed value is directly restored. The gradient is still evaluated at the current augmented low-rank representation, without separately training a full model.

After the core update, SVD acts on the small matrix rather than refactorizing a full-sized dense weight matrix at every step. Under \(r\ll\min(m,n)\), the authors give a per-layer update complexity of \(O(r^2(m+n))\). This preserves the low-rank computational route, but does not mean practical overhead depends only on the core: buffers and augmented bases occupy memory, and QR and basis rotation require computation. Hardware measurements remain necessary to establish actual incremental cost.

3. Spectral truncation and rank feedback: retain a limited tail and reduce deletion when ranks fall

After the core SVD, the algorithm retains leading singular directions according to a relative tolerance. The appendix analyzes truncation error by bounding the discarded tail's Frobenius norm, while the main text sets the absolute tolerance to \(\vartheta=\tau\|\hat S\|\). Retained left and right singular vectors are mapped through the augmented bases to form the new network factors. Starting immediately after the new rank, the algorithm stores at most another new-rank-sized group of leading discarded directions in the next compensation buffer. The more distant tail is still removed, so this is limited short-term directional memory, not storage of the full spectrum.

Rank collapse also requires an active brake. If the new rank falls below the layer's initial rank, the algorithm multiplies the relative tolerance factor by a feedback coefficient smaller than one:

\[ \tau\leftarrow\omega\tau\quad\text{if }r_1<r_0,\qquad 0<\omega<1. \]

A smaller tolerance permits less tail energy to be discarded at the next truncation, suppressing further rank reduction. This does not hard-lock ranks at their initial values or restore every deleted column at each iteration; it changes subsequent truncation rules. Comparisons must therefore distinguish initial rank, temporary training dimension, and final retained rank rather than treating them as an entirely fixed identical rank budget.

The paper reuses \(U_1,V_1\) in its truncation update: the same displayed system first reassigns them to retained bases and then uses them to map tail singular vectors. Dimensional consistency requires the pre-truncation augmented bases for the tail mapping. This note distinguishes augmented and retained bases when explaining the procedure rather than copying the reused notation as directly executable assignments.

A Worked Example

Consider a layer whose retained directions explain most of its current weights, while a direction near the cutoff was discarded in the previous iteration. Ordinary DLRT must seek new directions using the old retained bases and the new factor updates. SDLRT also brings that tail direction into the temporary QR space, allowing the core update to determine whether it should regain a retained position.

If its importance increases after the core update, it may enter the new leading singular-direction set. Otherwise, it remains buffered or is discarded farther into the tail. If the retained rank is still below the initial rank, the tolerance factor shrinks and the next truncation becomes more conservative. At the end of training, regardless of which directions the buffer has held, prediction uses only the final retained bases and core. This is an illustration of the mechanism, not a numerical training curve reported in the paper.

Loss & Training

SDLRT changes optimization and truncation rather than introducing a new task loss. The vision experiments train for 120 epochs with batch size 128, initial learning rate 0.05, a factor-of-0.1 decay at epochs 60 and 100, and momentum 0.1. Relative tolerances are \(\tau\in\{0.1,0.2,0.25,0.3,0.4\}\), and the feedback coefficient is \(\omega=0.8\).

The parameter-efficient fine-tuning (PEFT) experiments apply low-rank adaptation to query, key, and value projections in every self-attention layer. All methods start at rank 10; DLRT and SDLRT use \(\tau=0.02\). In the order listed in the setup—BoolQ, CB, COPA, RTE, WiC, WSC—the six tasks train for 30, 20, 20, 20, 30, and 30 epochs. Batch size is 16, the learning rate is \(6\times10^{-4}\), linear warm-up covers the first 6% of steps, and weight decay is 0.01. LoRA+ uses a learning rate of \(2\times10^{-4}\) and a factor learning-rate ratio of 8.

The theoretical guarantees should be read separately. Appendix B proves exactness when the trajectory remains exactly rank \(r\) throughout the interval, relevant subspace-overlap matrices are invertible, substeps are integrated exactly, and truncation retains every nonzero component. It shows that compensation does not undermine reconstruction in this special setting, not that practical discrete deep-network training reproduces any full trajectory.

Appendix C gives the global error bound:

\[ \|\hat U_t\hat S_t\hat V_t^{\top}-W_f(n\eta)\|_F \le c_0\delta+c_1\gamma\varepsilon+c_2\eta+c_3\vartheta/\eta. \]

Here \(\delta\) is the initial projection error, \(\varepsilon\) bounds the component of the gradient flow outside the low-rank tangent space, \(\eta\) is the integration step size, and \(\gamma\) represents projection contraction from compensation. Conditions include a bounded Lipschitz gradient vector field, proximity of the trajectory neighborhood to the low-rank manifold, and compensation actually producing contraction. Enlarging a subspace directly guarantees only non-increasing projection error; inclusion alone does not establish a uniform strict \(\gamma<1\) at every step. Strict improvement is therefore conditional, not proof of global convergence for all deep networks. The \(\vartheta/\eta\) term also warns that reducing step size without controlling truncation error need not improve total error.

Key Experimental Results

Main Results

The following excerpt from Table 1 reports mean validation accuracy on six SuperGLUE tasks, in percent. The original table also reports dispersion across six random seeds for individual tasks; only means are shown here for readability. Avg. is the average over these six tasks, not the official aggregate over all SuperGLUE tasks.

Method CB COPA WSC RTE WiC BoolQ Avg.
SDLRT 83.63 91.00 91.51 86.52 73.93 84.29 85.15
DLRT 81.85 91.17 90.06 86.28 73.41 84.33 84.52
LoRA 85.72 90.50 84.93 87.30 73.07 84.14 84.28
AdaLoRA 82.74 90.67 63.46 85.92 73.49 83.92 80.03
LoRA+ 85.12 89.83 89.58 87.24 73.27 83.80 84.81

SDLRT's average exceeds LoRA+ by 0.34 percentage points, DLRT by 0.63 points, and LoRA by 0.87 points, but it has the highest mean only on WSC and WiC. LoRA leads on CB and RTE, while DLRT leads on COPA and BoolQ; the best average does not imply winning every task. On WSC, SDLRT reports \(91.51\pm1.13\) versus LoRA's \(84.93\pm10.53\), suggesting that stability differences deserve attention beyond average scores.

Trainable parameter counts are approximately 1.145M each for LoRA, LoRA+, and AdaLoRA, 1.159M for DLRT, and 1.177M for SDLRT. The approximately 2.8% increment is relative to LoRA's adapter parameter count, not a 2.8% increase in the entire DeBERTa model, and the comparison is not strictly parameter-matched. The table caption says DeBERTa-base, whereas the setup and Appendix Figure 5 say DeBERTa-v3-base; this note preserves the source's naming discrepancy.

Ablation Study

The direction-compensation comparison uses \(\mathrm{SDLRT}_{\mathrm{2dim}}\), which applies compensation directly to the original \(K,L\) factors by combining each updated factor with its compensation basis. Comparing this variant with DLRT under corresponding dimensional settings supports an improvement beyond merely enlarging the temporary space. However, the local text does not provide pointwise curve values or a numerical table removing negative feedback alone, so exact ablation drops cannot be reconstructed.

The second table uses the measured overhead analysis in Appendix D.2: VGG-19/CIFAR-100 on RTX3080, seed=30, \(\tau=0.45\), and batch size 128. This is one specific single-seed configuration, not an average over the five main-experiment runs.

Method GPU memory Average training time per epoch Test accuracy
DLRT 2680 MB 56.88 s 53.87%
SDLRT 2710 MB 57.53 s 61.86%

In this configuration, SDLRT uses 30 MB more memory and 0.65 s more per epoch while improving accuracy by 7.99 percentage points. This supports modest measured overhead here, but does not establish identical overhead across architectures, ranks, or devices.

Key Findings

  • Each CIFAR-10 configuration has five independent runs. The text reports that above approximately 80% compression, SDLRT maintains accuracy above 90% on VGG-16 and above 84% on AlexNet. Differences are smaller at low compression, where compensation can cause mild overfitting. These are textual thresholds, not inferred exact values for other plotted points.
  • On VGG-19/CIFAR-100, compression above 76.1% can cause DLRT accuracy collapse and substantial seed variation, while SDLRT is more stable. This is a setting-specific observation, not a universal safe-compression threshold.
  • The stability metric is the Frobenius distance between the best same-rank SVD approximation of the current full-model weights and the low-rank model weights under the same initialization: \(d_t=\|\mathrm{SVD}(W_t,r_t)-U_tS_tV_t^{\top}\|_F\). At \(\tau=0.4\), the text reports an approximately 10% smaller terminal distance for SDLRT than DLRT. This measures proximity of weight trajectories, not generalization error directly.
  • The captions of Figures 3 and 4 define error bars as minima and maxima across five runs, whereas Section 4.2 says it reports means and standard deviations. The source contains a statistical-description conflict; the caption-defined ranges are not relabeled as standard deviations here.

Highlights & Insights

  • Importance depends not only on present singular-value magnitude but also on how a direction influences future subspace rotation. Retaining a few recently discarded directions separates “weak now” from “permanently useless,” making the change more targeted than mechanically increasing rank.
  • The buffer expands admissible update directions without changing the initial weights in the augmented representation. This provides a clear mechanism for the stability improvement and avoids confusing compensation with directly adding tail weights back.
  • Truncation itself can be a feedback-controlled operation. Reducing tolerance after rank drops is transferable to resource-constrained training, provided actual parameter-budget changes are also examined.

Limitations & Future Work

  • The authors acknowledge heuristic feedback coefficients and leave scaling to large architectures for future work. Current evidence primarily covers classical convolutional networks on CIFAR and small DeBERTa adapters; it does not directly establish large language model pretraining performance.
  • Strict projection contraction does not follow automatically from subspace expansion; the theory also depends on boundedness, Lipschitz continuity, and manifold proximity. Future work could characterize when buffers cover missing descent directions rather than merely retaining more columns.
  • There is no fully separated numerical ablation for removing the buffer versus removing feedback. Their independent contributions remain insufficiently isolated; fixed final parameter budgets and separate sweeps of buffer length and feedback strength would help.
  • The overhead table contains only one configuration and seed. More ranks, networks, and hardware should be evaluated for training throughput, peak memory, and final factorized inference speed before claiming consistent deployment gains from algorithmic complexity.
  • Reused basis notation, model naming, and error-bar descriptions reduce reproducibility clarity. The text says code is in supplementary material, but the cache provides no verifiable repository link, so no code URL is added.
  • vs DLRT / DLRA: Both optimize low-rank representations; SDLRT adds recently discarded directions to the existing basis augmentation and adjusts truncation tolerance. The distinction is training-space memory and feedback, not separate full-model training followed by projection.
  • vs CondLR: CondLR stabilizes factor training through approximate orthonormal constraints; SDLRT targets directional information lost through truncation. High-compression comparisons support the latter mechanism, but do not show that the two approaches cannot be combined.
  • vs LoRA / AdaLoRA / LoRA+: LoRA uses fixed-rank adaptation, AdaLoRA allocates rank budgets through importance scores, and LoRA+ adjusts factor learning rates. SDLRT brings dynamic low-rank updates and short-term spectral compensation to adapters; its six-task average advantage should be read alongside the slightly larger trainable parameter count.
  • vs post-hoc SVD and the lottery ticket hypothesis: Post-hoc factorization compresses already trained weights, whereas SDLRT seeks trainable low-rank subnetworks during training. Its winning tickets are subnetworks in low-rank spectral space, not the entire set of claims associated with classical sparse-mask lotteries.

Rating

  • Novelty: 4/5. Retained/discarded curvature coupling motivates a targeted short-term spectral buffer.
  • Experimental Thoroughness: 3/5. Vision compression, trajectory distances, six-task PEFT, and overhead are covered, but isolated module ablations and large-scale validation remain limited.
  • Writing Quality: 3/5. Motivation connects clearly to the algorithm, while strict contraction claims and some notation and statistical descriptions need clarification.
  • Value: 4/5. Useful for studying trainability under aggressive compression and low-rank optimization, but not a substitute for systematic evaluation on large architectures.