Going Deep: Deep Visual Prompting with LoTeP¶
Conference: ECCV 2026
Paper: ECCV Official Link
Area: Multimodal VLM
Keywords: Visual Prompting, Low-Rank Tensor Decomposition, Deep Feature Adaptation, Capacity-Integrity Paradox, Parameter-Efficient Fine-Tuning
TL;DR¶
To tackle the severe parameter explosion and semantic disruption when visual prompting is extended into deep activations, this paper presents LoTeP, a low-rank tensor visual prompting method that factorizes prompts via Tucker and CP decompositions coupled with a layer-wise cosine rank-decaying schedule to achieve seamless adaptation.
Background & Motivation¶
Parameter-efficient fine-tuning (PEFT) has become foundational for adapting visual foundation models to downstream tasks. Visual Prompting (VP), as a prominent non-invasive adaptation paradigm, traditionally prepends or adds trainable perturbation parameters purely in the input pixel space (e.g., through border padding, patch modification, or whole-image low-rank additions). While this input-only paradigm provides model-agnostic flexibility without touching internal architectures, it suffers from severe signal attenuation as information propagates through deep layers. Centered Kernel Alignment (CKA) feature similarity analysis reveals that representations adapted solely by input-level VP remain tightly aligned with the initial, unadapted pre-trained representations, failing to extract the downstream task-discriminative features captured by full fine-tuning (FT).
However, directly injecting learnable prompts into deep intermediate activations triggers a fundamental Capacity-Integrity Paradox. On the one hand, feature channel dimensions expand dramatically in deeper layers. Assigning independent prompt parameters per channel triggers a prohibitive parameter surge, contradicting the core premise of PEFT and inviting severe overfitting. On the other hand, unconstrained deep prompts offer excessive degrees of freedom, overwriting the stable semantic manifolds established during pre-training and causing catastrophic disruption to intermediate representations.
Preliminary investigations on channel grouping indicate strong structural redundancy and inter-channel correlations among feature maps; concurrently, theoretical derivation from backpropagation gradient flows demonstrates that deep visual prompts inherently exhibit a low-rank property that contracts with network depth. The core idea is to model deep activation prompts as low-rank tensors (LoTeP) that simultaneously decouple and compress channel and spatial dimensions, while enforcing a layer-wise cosine rank-decaying schedule to provide high adaptation capacity in shallow layers and preserve semantic integrity in deep layers.
Method¶
Overall Architecture¶
LoTeP extends visual prompting from input pixels into deep intermediate activations while leveraging tensor decomposition and dynamic rank decay to control layer-wise prompting capacity. Given the first \(d\) feature blocks of a pre-trained visual backbone, for the \(l\)-th intermediate feature map \(\mathcal{X}_l \in \mathbb{R}^{C \times H \times W}\), LoTeP introduces a learnable prompt tensor \(\mathcal{P}_l \in \mathbb{R}^{C \times H \times W}\) that is added element-wise to the feature map before being processed by the frozen block \(f_l\):
To prevent parameter explosion while modeling multi-dimensional correlations, LoTeP introduces two low-rank tensor factorizations: a Tucker-based variant (LoTeP-TK) and a CP-based variant (LoTeP-CP). In addition, considering that higher-level representations become increasingly vulnerable to distortion deeper in the network, LoTeP applies a depth ratio \(\rho\) and a smooth cosine rank-decaying schedule, gracefully tapering the prompt rank down to zero to balance task capacity and semantic integrity.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
In["Input Image / Early Features"] --> P1["Multi-Dimensional Low-Rank Tensor Factorization<br/>Tucker / CP decomposition decouples channels and space"]
P1 --> P2["Zero-Tensor Initialization<br/>Zeroed channel factors ensure transparent start"]
P2 --> P3["Layer-Wise Cosine Rank-Decaying Schedule<br/>Smoothly reduces rank with depth to protect deep semantics"]
P3 --> P4["Deep Feature Residual Injection<br/>Element-wise addition across the first d blocks"]
P4 --> Out["Frozen Deep Blocks & Final Classifier"]
Key Designs¶
1. Multi-Dimensional Low-Rank Tensor Factorization: Decoupling Channel Expansion and Spatial Redundancy
Naively adding deep prompts requires \(\mathcal{O}(C \cdot H \cdot W)\) parameters, while straightforwardly extending spatial low-rank matrices across channels (as in Deep LoR-VP) demands independent factorizations per channel, leading to \(\mathcal{O}(C \cdot r \cdot (H + W))\) parameters that scale linearly with channel count \(C\). LoTeP overcomes this bottleneck by conceptualizing the prompt \(\mathcal{P}_l\) as a 3rd-order tensor, uniting channel and spatial dimensions into a cohesive low-rank factorization framework.
In LoTeP-TK, the prompt tensor is formulated using Tucker decomposition with a core tensor \(\mathcal{C}^{(l)} \in \mathbb{R}^{r_l \times r_l \times r_l}\) and factor matrices \(U^{(l)} \in \mathbb{R}^{r_l \times C}\), \(V^{(l)} \in \mathbb{R}^{r_l \times H}\), and \(W^{(l)} \in \mathbb{R}^{r_l \times W}\) via mode products:
In LoTeP-CP, the prompt tensor is decomposed as the sum of \(r_l\) rank-one outer-product components parameterized by channel vectors \(\mathbf{a}_t^{(l)} \in \mathbb{R}^C\), height vectors \(\mathbf{b}_t^{(l)} \in \mathbb{R}^H\), and width vectors \(\mathbf{c}_t^{(l)} \in \mathbb{R}^W\):
This compresses the parameter complexity to \(\mathcal{O}(r_l(C + H + W))\). Because \(r_l \ll C\), channel and spatial dimensions are effectively decoupled, eliminating inter-group isolation while maintaining rich cross-channel interactions at minimal parameter cost.
2. Zero-Tensor Initialization: Eliminating Early Representational Noise
At the onset of adaptation, randomly initialized deep prompts introduce violent perturbations into pre-trained activations, disrupting intermediate representations and destabilizing early gradient descent. To guarantee that the model exactly matches the pre-trained backbone at step zero, LoTeP incorporates an asymmetric zero-initialization strategy: factor matrices or vectors along the channel dimension (i.e., \(U^{(l)}\) in LoTeP-TK or \(\mathbf{a}_t^{(l)}\) in LoTeP-CP) are initialized strictly to zero, while spatial factors are initialized using standard Gaussian random variables. Because tensor multiplication by a zero-valued factor yields a zero tensor, \(\mathcal{P}_l\) evaluates to an exact zero tensor initially, ensuring a clean identity mapping and smooth gradient flow.
3. Layer-Wise Cosine Rank-Decaying Schedule: Reconciling Capacity and Semantic Preservation
While tensor decomposition constrains intra-layer parameter freedom, maintaining a constant rank throughout all depths still risks overwriting critical high-level semantics accumulated in later layers. To resolve this dilemma, LoTeP limits prompt injection to the first \(d = \lfloor L \cdot \rho \rceil\) layers and introduces a dynamic cosine rank-decaying schedule where layer-wise rank \(r_l\) decreases smoothly from an initial rank \(r_0\):
This cosine trajectory decays slowly in shallow layers to provide ample capacity for domain-specific visual features, and then accelerates its decay towards zero in deeper layers to protect semantic integrity. Theoretical derivation under CNN backpropagation dynamics corroborates that the intrinsic rank upper bound of gradients naturally contracts monotonically with depth, making this decaying schedule mathematically rigorous.
Loss & Training¶
During adaptation, all convolutional and self-attention weights of the backbone remain frozen; only the tensor factor matrices/vectors and the linear classification head are optimized. Training is guided by the standard cross-entropy objective:
where \(K\) denotes the number of target classes and \(\hat{y} = \text{Softmax}(\text{Head}(\mathcal{X}_L))\). The initial input rank is set to \(r_0 = 6\), and depth ratio \(\rho\) ranges between 0.5 and 1.0 depending on backbone scale. No specialized data augmentation beyond standard image resizing is required.
Key Experimental Results¶
Main Results¶
The method is evaluated across 9 diverse vision architectures (covering ResNet, ViT, Swin, and ConvNeXt) on 10 downstream benchmarks including CIFAR-100 and Tiny-ImageNet. As reported in the table below, LoTeP substantially surpasses the previous state-of-the-art input-level VP baseline (LoR-VP) and even outperforms full fine-tuning (FT) and LoRA on large vision models.
| Architecture | Method | CIFAR-100 Acc (%) | Tiny-ImageNet Acc (%) | Tunable Params (M) |
|---|---|---|---|---|
| ViT-B/16 | Full Fine-Tuning (FT) | 92.15 | 91.24 | 575.39 |
| ViT-B/16 | Linear Probing (LP) | 87.12 | 89.25 | 1.00 |
| ViT-B/16 | LoRA (rank=8) | 91.38 | 91.19 | 11.05 |
| ViT-B/16 | AutoVP | 88.84 | 89.53 | 9.97 / 18.39 |
| ViT-B/16 | LoR-VP | 89.01 | 90.53 | 1.05 / 2.05 |
| ViT-B/16 | LoTeP-CP (Ours) | 91.59 | 91.90 | 1.13 / 2.13 |
| ViT-B/16 | LoTeP-TK (Ours) | 91.82 | 91.93 | 1.13 / 2.14 |
| ResNet-50 | Full Fine-Tuning (FT) | 83.46 | 81.56 | 575.39 |
| ResNet-50 | LoR-VP | 74.86 | 77.55 | 1.05 / 2.05 |
| ResNet-50 | LoTeP-CP (Ours) | 80.39 | 80.41 | 1.13 / 2.13 |
| ResNet-50 | LoTeP-TK (Ours) | 80.82 | 80.63 | 1.13 / 2.14 |
Across the 10-dataset benchmark suite (EuroSAT, Pets, Food, DTD, Flowers, CIFAR-10/100, SVHN, GTSRB), LoTeP-TK achieves an average accuracy of 91.44% on ViT-B/32, surpassing LoR-VP (88.94%) by +2.50% and outperforming FT (91.11%) and LoRA (91.22%).
Ablation Study¶
The ablation investigations on CoAtNet-2 and ViT-B/16 examine the decay schedule formulation and computational training/inference efficiency:
| Evaluation Dimension | Configuration / Schedule | CIFAR-100 Acc (%) / GPU Usage | Training Time (h) / Latency (ms) | Note |
|---|---|---|---|---|
| Schedule (CoAtNet-2) | Constant Schedule (Fixed Rank) | 89.65 | - | Lacks deep semantic regularization |
| Schedule (CoAtNet-2) | Linear Schedule | 90.15 | - | Tapers shallow capacity too abruptly |
| Schedule (CoAtNet-2) | Cosine Schedule | 90.45 | - | Optimal capacity-integrity balance |
| Efficiency (ViT-B/16) | LoRA | 91.19 (22.08 GB) | 2.89h / 5.26ms | Gradients propagate to internal weights |
| Efficiency (ViT-B/16) | VPT-Deep | 91.10 (12.60 GB) | 3.48h / 6.33ms | Adds prefix tokens and attention overhead |
| Efficiency (ViT-B/16) | LoR-VP | 90.53 (14.34 GB) | 1.65h / 5.74ms | Optimizes low-rank input-pixel matrix |
| Efficiency (ViT-B/16) | LoTeP-TK (Ours) | 91.93 (14.34 GB) | 1.59h / 5.77ms | Best accuracy with lowest training cost |
Key Findings¶
- Rank decay is essential: Maintaining a constant rank across all prompted layers leads to noticeable performance drops, demonstrating that unconstrained prompt capacity in deeper activations actively erodes pre-trained representations.
- Model capacity dictates depth sensitivity: For smaller backbones like ResNet-50, increasing the depth ratio \(\rho\) toward 1.0 continuously improves accuracy because the network benefits from additional representational capacity. Conversely, for large models like Swin-B, higher \(\rho\) harms accuracy, verifying that deep features in large pre-trained models are highly sensitive to invasive intervention.
- Superior OOD robustness: On the challenging ImageNet-A benchmark using ConvNeXt-B, LoTeP-TK achieves 44.41% accuracy, outperforming LoR-VP (39.36%) by +5.05%, which suggests that low-rank tensor prompting reduces overfitting to spurious source-domain artifacts.
Highlights & Insights¶
- From discrete grouping to continuous tensorization: The authors effectively generalize the observation of inter-channel redundancy into an elegant multilinear tensor decomposition, resolving the exponential channel parameter growth with an \(\mathcal{O}(r(C+H+W))\) formulation.
- Theory-driven schedule design: Rather than relying on trial-and-error heuristics, the cosine rank-decaying schedule is grounded in gradient backpropagation analysis showing that prompt rank naturally contracts with depth.
- Challenging the performance ceiling of visual prompting: Traditionally, visual prompting has been considered an inferior alternative to LoRA and full fine-tuning. LoTeP demonstrates that deep activation prompting with proper rank decay can match or exceed LoRA and full fine-tuning on large foundation models while maintaining minimal overhead.
Limitations & Future Work¶
- Performance lag on small architectures: On compact models like ResNet-18 where backbone capacity is limited, LoTeP significantly outperforms prior VP methods but still trails full fine-tuning (75.31% vs. 82.51%), indicating that prompting frozen features cannot substitute for weight updates when pre-trained representations are weak.
- Fixed hyperparameter heuristics: The starting rank \(r_0\) and depth ratio \(\rho\) are currently selected heuristically based on architecture size; integrating an adaptive singular value or gradient-aware rank allocation mechanism could further automate optimization.
Related Work & Insights¶
- vs LoR-VP: LoR-VP introduces low-rank matrix approximation in input pixel space, offering high parameter efficiency but suffering from deep signal attenuation. Directly extending LoR-VP to deep layers leads to linear parameter scaling across channels and severe semantic disruption. LoTeP solves this via 3rd-order tensor modeling and cosine rank decay.
- vs VPT (Visual Prompt Tuning): VPT-Deep injects prefix tokens into each self-attention layer of Transformers, which is inapplicable to CNNs and extends sequence length, causing higher memory usage and inference latency. LoTeP uses element-wise residual addition, achieving architectural generality across CNNs, Transformers, and hybrid models with zero latency overhead.
Rating¶
- Novelty: โญโญโญโญโญ Formulates deep visual prompting as a low-rank tensor problem and tackles the capacity-integrity paradox with theoretical and empirical rigor.
- Experimental Thoroughness: โญโญโญโญโญ Comprehensive evaluation covering 9 backbones, 10 datasets, 4 OOD benchmarks, and multi-faceted ablations.
- Writing Quality: โญโญโญโญโญ Clearly motivated, logically cohesive, and supported by solid mathematical and empirical evidence.
- Value: โญโญโญโญโญ Provides a practical and principled PEFT paradigm for adapting visual foundation models.