COVERT: Privacy-Preserving Covariant Obfuscation for VLMaaS via Exact Reparameterization and Tailored Tuning¶
Conference: ECCV 2026
Paper: ECCV Official
PDF: ECCV PDF
Area: LLM Safety
Keywords: Vision-Language Models, Covariant Obfuscation, Privacy-Preserving Inference, Hidden-State Inversion, Model-as-a-Service
TL;DR¶
To defend against visual prompt privacy leakages in Vision-Language Model-as-a-Service (VLMaaS), COVERT introduces the first multimodal covariant obfuscation framework, combining null-space perturbed convolutional patchification, dimensional-augmented bias absorption, and tailored tuning of shallow ViT attention to fully neutralize inversion attacks while keeping accuracy degradation below 3% and throughput overhead below 6%.
Background & Motivation¶
As vision-language models (VLMs) become deeply integrated into privacy-critical sectors such as healthcare diagnostics, financial analytics, and personal assistants, resource-constrained clients increasingly rely on the Vision-Language Model-as-a-Service (VLMaaS) paradigm by offloading multimodal prompts to cloud servers. However, multimodal visual prompts inherently encompass sensitive semantics including biometric signatures, confidential identities, and geolocation metadata. An untrusted or honest-but-curious cloud server can reconstruct raw high-resolution images via feature inversion or infer private attributes from intermediate activations. Conventional cryptographic solutions (e.g., three-party MPC in PrivMLLM) incur prohibitive communication volume and latency bottlenecks when processing continuous high-dimensional visual tensorsโoften requiring gigabytes of transmission and several minutes of WAN latency per single requestโrendering interactive real-time serving infeasible. Conversely, perturbation approaches based on differential privacy (DP) inject uncalibrated noise that fundamentally destroys delicate cross-modal semantic alignments, triggering catastrophic utility collapse.
In text-only language models, covariant obfuscation (such as AloePri) achieves an optimal balance between algebraic zero-error cancellation and robust privacy by jointly transforming inputs and model weights. However, extending this paradigm to VLMs faces critical barriers stemming from architectural and representational heterogeneity across modalities. On the one hand, the vision pipeline incorporates convolutional patchification, non-linear activations, and pervasive additive biases. Under strict algebraic constraints like Rotary Position Embeddings (RoPE), the admissible degrees of freedom for multiplicative masking matrices shrink drastically to \(O(n)\), making naive bias masks trivial to invert algebraically. On the other hand, continuous visual signals exhibit strong local smoothness and dense spatial priors. Attackers can bypass global token permutations to launch adaptive hidden-state inversion attacks from shallow Vision Transformer (ViT) attention scores, reconstructing legible visual semantics such as license plates and faces.
To address both structural cross-modal heterogeneity and spatial inversion risks, the core angle of attack is to decouple text-centric permutation obfuscation from exact algebraic visual reparameterizationโachieving deterministic error cancellation in convolutions and projectors while neutralizing shallow spatial vulnerabilities through lightweight, utility-aware parameter tuning. Core idea: develop COVERT, a modality-adaptive covariant obfuscation framework that combines null-space noise injection with exact architectural reparameterization to eliminate visual feature inversion, absorbs additive biases via dimensional augmentation, and selectively tunes the first ViT attention layer using a private key seed to sever spatial hidden-state inversion pathways.
Method¶
Overall Architecture¶
COVERT is tailored for the dominant VLM architecture consisting of a vision encoder, a cross-modal projection layer, and an autoregressive LLM decoder. In the offline initialization phase, the client samples stochastic transformation matrices and a cryptographic noise seed using a private key, applies collaborative weight transformations to the VLM parameters, and deploys the obfuscated model to the cloud server. During online inference, the client performs lightweight flattening and injects orthogonal null-space noise into the image representation; the server executes forward inference entirely over obfuscated representations, and the client decodes the returned response using its private label de-obfuscator.
The visual defense pipeline encompasses four tightly coupled stages: first, convolutional patchification is reparameterized with volume expansion and row-wise null-space noise to cancel patch-level feature visibility; second, dimensional augmentation and activation scaling absorb additive biases within ViT linear layers and FFNs; third, utility-aware parameter tuning on the first ViT attention block disrupts shallow spatial attention topology using a private seed; and fourth, a block-diagonal inverse mask coordinates multi-patch spatial aggregation within the cross-modal projection layer, seamlessly feeding obfuscated representations into the backbone LLM.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Raw Visual Input X_v"] --> B["Covariant Convolutional Patchification & Null-Space Perturbation<br/>Dimension inflation and orthogonal null-space noise injection"]
B --> C["Dimensional Augmentation & Bias Absorption<br/>Append unit channel to absorb ViT linear layer biases"]
C --> D["Tailored Parameter Tuning Against Shallow Inversion<br/>Private-seeded noise and tuning on Layer 1 attention weights"]
D --> E["Block-Diagonal Inverse Masking & Projection Alignment<br/>Kronecker-product inverse cancels spatial aggregation masks"]
E --> F["Obfuscated Multimodal Tokens to Backbone LLM"]
Key Designs¶
1. Covariant Convolutional Patchification and Null-Space Perturbation: Eliminating Patch-Level Observability of Continuous Visual Inputs
Transmitting unshielded continuous image patches directly exposes visual inputs to the server. COVERT formulates 2D/3D convolution as a vectorized matrix multiplication \(f_{\text{Conv}}(X_v) = X_v W_{\text{conv}}\), where the unrolled kernel dimension is \(D = C_{in} \cdot T_c \cdot H_c \cdot W_c\) across \(N\) patches. The client generates an expanded stochastic transformation pair \((\hat{P}_v, \hat{Q}_v)\) satisfying \(\hat{P}_v \hat{Q}_v = I_D\), where \(\hat{P}_v \in \mathbb{R}^{D \times (D + d_{extra})}\) inflates the input volume. In the offline phase, convolutional weights are transformed into \(\tilde{W}_{\text{conv}} = \hat{Q}_v W_{\text{conv}} \hat{P}_{\text{Pat}}\). During online inference, the client injects a stochastic perturbation vector \(v_i\) into each row of the projected input \(X_v \hat{P}_v\), strictly constrained by \(v_i \hat{Q}_v = \mathbf{0}\). The transmitted obfuscated representation is \(\tilde{r}_i = r_i + v_i\). When the server computes the convolution \(\tilde{X}_v \tilde{W}_{\text{conv}}\), the injected perturbation \(v_i \hat{Q}_v\) cancels out identically with zero algebraic divergence (\(e_C = 0\)), preventing deterministic feature reconstruction without impacting multimodal task utility.
2. Dimensional Augmentation and Bias Absorption: Resolving Algebraic Conflicts between Additive Biases and RoPE Constraints
Standard covariant obfuscation requires multiplicative associativity, which is broken by additive biases \(W X + b\) throughout ViT blocks. Applying naive multiplicative masks \(b \hat{P}\) fails under Rotary Position Embeddings (RoPE), because RoPE forces \(\hat{P}\) into block-diagonal structures, collapsing the system's unknown degrees of freedom to \(O(n)\) and making the private mask algebraically solvable by an adversary. COVERT resolves this by dimensional augmentation: appending a constant unit channel to the input \(X^+ = (X, 1)\) and concatenating biases into the weight matrix \(W_{QKV}^+ = [W_{QKV}^\top; b_{QKV}^\top]^\top\), converting the affine mapping into a homogeneous linear operation \(\tilde{W} = \hat{Q} W^+\). For the output projection \(W_O\), where the global non-linearity of Softmax cannot propagate the augmented dimension, the extra coordinate is discarded prior to \(W_O\). Because \(W_O\) is free from RoPE constraints, it is safely protected via dense bias masking \(\tilde{b}_O = b_O \hat{P}_O\), presenting \(O(n^2)\) unknown variables against \(O(n)\) equations to ensure mathematical intractability. For non-linear activations like SiLU in FFNs that map the unit dimension to \(\text{SiLU}(1)\), an exact scaling factor \(\text{SiLU}(1)^{-1}\) is applied to the augmented down-projection weights to restore exact functional equivalence.
3. Tailored Parameter Tuning Against Shallow Inversion: Disrupting Spatial Geometric Leakage in ViT Layer 1
While exact reparameterization ensures mathematical isomorphism, empirical findings reveal that adversaries can exploit spatial locality in shallow layers to launch high-fidelity hidden-state inversion attacks using extracted attention scores, reconstructing confidential details like vehicle license plates. Empirical audits demonstrate that this vulnerability is strictly localized to the shallowest block (Layer 1), whereas deeper layers naturally suppress spatial details through high-level semantic abstraction. COVERT introduces a utility-aware tailored tuning defense: Gaussian noise \(\mathcal{N}(0, \alpha \cdot \text{std}(W))\) seeded by a client-held private key is added to the attention parameters \(\{W_Q, W_K, W_V, W_O\}\) of Layer 1. Then, using 5,000 public images, the client fine-tunes these attention matrices and the preceding normalization layer \(W_{\text{norm1}}\) while freezing all FFNs and subsequent modules. The tuning minimizes the cosine distance between the tuned block and the original block outputs:
$\(L = 1 - \cos\left(B_{\text{tuned}}^{(1)}(x_i), B_{\text{orig}}^{(1)}(x_i)\right)\)$
This aligns output representations with the original semantic space to preserve downstream task accuracy, while thoroughly disrupting the internal spatial geometry of attention scores. Because the noise trajectory depends on the private cryptographic seed, gradient-based inversion becomes mathematically irreproducible.
4. Block-Diagonal Inverse Masking and Cross-Modal Projection Alignment: Unifying Visual and Text Semantic Spaces
Visual features from the ViT must be projected into the LLM embedding space via an MLP projector. Modern architectures (e.g., Qwen2.5-VL) merge adjacent \(2 \times 2\) spatial tokens prior to projection, expanding the input channel dimension to \(D_{in} = 4 D_{\text{Feat}}\). To cancel the visual feature mask \(\hat{P}_{\text{Feat}}\), the client constructs a block-diagonal inverse transformation \(\hat{M} = I_4 \otimes \hat{Q}_{\text{Feat}}^{-1}\) using the Kronecker product. It also generates a hidden permutation matrix \(Z\) and an alignment mask \(\hat{P}_{\text{Emb}}\) shared across vision and text modalities. The projection weights are obfuscated as:
$\(\tilde{W}_1 = \hat{M} W_1 Z, \quad \tilde{W}_2 = Z^{-1} W_2 \hat{P}_{\text{Emb}}\)$
During inference, the server computes \(\tilde{f}_{\text{Proj}}(\tilde{F}_{\text{concat}}) = \text{GeLU}(\tilde{F}_{\text{concat}} \tilde{W}_1) \tilde{W}_2 = f_{\text{Proj}}(F_{\text{concat}}) \hat{P}_{\text{Emb}}\), which cancels the internal permutation and visual masks exactly, producing tokens aligned with \(\hat{P}_{\text{Emb}}\) ready for the obfuscated LLM decoder.
Key Experimental Results¶
Main Results¶
Evaluations are conducted under the VLMEvalKit benchmark suite across five datasets covering general VQA (Dataset 1), multidisciplinary reasoning (Dataset 2), vision-centric VQA (Dataset 3), scientific diagram understanding (Dataset 4), and mathematical reasoning (Dataset 5). Models include Qwen2.5-VL (7B, 32B, 72B), MiMO-7B, and Fara-7B under standard and 8-bit quantized (Quant) configurations.
| Model & Config | General VQA (D1) | Multidisciplinary (D2) | Vision-Centric (D3) | Scientific Diagram (D4) | Math Reasoning (D5) | Average Accuracy | Average Drop |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL-7B (Plaintext) | 82.97% | 55.11% | 64.87% | 85.01% | 68.10% | 71.21% | - |
| COVERT | 79.79% | 51.89% | 61.80% | 83.87% | 65.80% | 68.63% | โ2.58% |
| COVERT (Quant-8bit) | 78.41% | 51.56% | 61.67% | 83.94% | 66.40% | 68.40% | โ2.81% |
| Qwen2.5-VL-32B (Plaintext) | 86.38% | 68.67% | 69.07% | 84.84% | 73.70% | 76.53% | - |
| COVERT | 83.59% | 65.33% | 66.93% | 83.65% | 74.00% | 74.70% | โ1.83% |
| COVERT (Quant-8bit) | 83.90% | 65.44% | 67.23% | 84.49% | 73.20% | 74.85% | โ1.68% |
| Qwen2.5-VL-72B (Plaintext) | 88.08% | 67.55% | 70.80% | 89.11% | 75.20% | 78.15% | - |
| COVERT | 87.00% | 64.56% | 67.33% | 87.18% | 73.30% | 75.87% | โ2.28% |
| COVERT (Quant-8bit) | 86.53% | 64.22% | 67.47% | 87.05% | 73.40% | 75.73% | โ2.42% |
| MiMO-7B (Plaintext) | 84.13% | 65.55% | 69.80% | 86.85% | 70.30% | 75.32% | - |
| COVERT | 82.66% | 63.44% | 68.20% | 85.36% | 69.30% | 73.79% | โ1.53% |
| Fara-7B (Plaintext) | 65.09% | 52.78% | 55.93% | 80.47% | 58.30% | 62.51% | - |
| COVERT | 61.07% | 47.22% | 53.20% | 79.83% | 57.20% | 59.70% | โ2.81% |
Ablation Study¶
The defense against adaptive hidden-state inversion is quantified via SSIM (lower is more private) and MSE (higher is more private) on reconstructed images from ViT layers, paired with real-world serving efficiency measured on Qwen2.5-VL-32B under vLLM (30 concurrent streams, 1,064 visual tokens per prompt).
| Target & Configuration | SSIM (โ) | MSE (โ) | Privacy Protection Note |
|---|---|---|---|
| Untuned Baseline (Layer 1 - Vulnerable) | 0.3404 | 3275.28 | Catastrophic spatial leakage; license plates clearly readable |
| Untuned Baseline (Layer 2) | 0.2205 | 5737.62 | Natural semantic abstraction provides moderate safety |
| Untuned Baseline (Layer 3) | 0.1575 | 5541.46 | Deep feature defocusing thwarts contour recovery |
| Untuned Baseline (Layer 4) | 0.1078 | 5855.34 | Spatial information largely extinguished |
| COVERT Protected (Layer 1) | 0.1550 | 4577.28 | SSIM drops by 54.46%, MSE increases by 39.75%, details unrecoverable |
System latency and communication performance under concurrent vLLM serving:
| Serving Configuration | TTFT (ms) โ | TPOT (ms) โ | Throughput (tokens/s) โ | Expansion Ratio |
|---|---|---|---|---|
| Plaintext Serving | 3864.89 | 26.73 | 25.17 | 1.00ร (Baseline) |
| COVERT Full (Lossless PNG) | 3898.38 (+0.86%) | 29.26 (+8.65%) | 23.76 (-5.93%) | 7.68ร (PNG) |
| COVERT Optimized (64-bin Quant.) | 3898.38 (+0.86%) | 29.26 (+8.65%) | 23.76 (-5.93%) | 1.30ร (Only 2.39% accuracy penalty) |
Key Findings¶
- Distinct scaling resilience: Larger models display significantly tighter accuracy bounds under obfuscation; the average degradation narrows from 2.58% in 7B to 1.83% in 32B and 2.28% in 72B, verifying that increased parameter capacity better tolerates normalization heuristics and quantization noise.
- Spatial inversion vulnerability is strictly concentrated at Layer 1: Layers 2 through 4 exhibit natural defense with MSE values exceeding 5500. Tailored tuning on just ~4M parameters of Layer 1 (taking only 1.5 hours on a single GPU) reduces SSIM from 0.3404 to 0.1550, aligning Layer 1 security with deeper layers.
- Overwhelming utility advantage over differential privacy: While DP-Forward completely destroys multimodal utility (e.g., accuracy collapsing from 82.97% to 0.07% on Dataset 1 even at \(\epsilon=128\)), COVERT preserves 79.79% accuracy under robust defense.
- Negligible serving throughput overhead (<6%): Because algebraic transformations are handled offline, online inference on vLLM incurs only a +0.86% TTFT penalty and a minor 5.93% throughput drop, avoiding the hundreds-of-seconds bottlenecks of MPC.
Highlights & Insights¶
- Null-space injection achieves algebraic zero-divergence feature obfuscation: By mathematically restricting dynamic client-side noise to the null space of the weight matrix inverse, random noise vanishes automatically during server-side matrix multiplication, preventing feature reconstruction while retaining functional equivalence (\(e_C = 0\)).
- Homogeneous dimensional augmentation cleanly circumvents RoPE constraints: Appending a unit coordinate and incorporating bias vectors into augmented weight matrices converts affine layers into pure linear operations, sidestepping block-diagonal rank collapse in QKV attention without altering token semantics.
- Surgical tuning delivers high defense efficiency at minimal cost: Rather than retraining the entire vision backbone, fine-tuning only 4M parameters in the first attention block using private noise destroys the spatial attention geometry needed by inversion attacks while preserving downstream representations.
Limitations & Future Work¶
- Payload expansion under lossless transmission: Because obfuscation breaks natural image spatial locality (driving differential entropy from ~2.0 to ~7.0), standard image compression fails, leading to a 7.68ร payload increase in lossless PNG mode; future work should explore frequency-domain transform operators that preserve compressibility.
- Client-side offline tuning requirements: Although tuning requires only 1.5 GPU hours or 6.5 CPU hours, extremely low-power edge microcontrollers may still require trusted external pre-computation to prepare obfuscated weights.
- Extension to dynamic non-uniform spatial token schemes: Current closed-form derivations focus on uniform grid aggregation (e.g., \(2 \times 2\) patch packing); adapting algebraic inverses to dynamic scale-varying visual token patchification warrants further investigation.
Related Work & Insights¶
- vs AloePri: While AloePri pioneered covariant obfuscation for text-only LLMs, it is fundamentally incompatible with continuous image inputs, convolutional operators, and shallow spatial inversion vulnerabilities. COVERT establishes the necessary visual patchification transforms, bias absorption techniques, and targeted attention tuning to bridge this gap.
- vs DP-Forward & Perturbation Baselines: Differential privacy mechanisms inject uncalibrated noise into activation vectors, destroying delicate multimodal alignments and leading to catastrophic accuracy drops. COVERT maintains strict mathematical cancellation and representation alignment, bounding utility drops within 3%.
- vs PrivMLLM & Cryptographic Protocols: Three-party MPC frameworks achieve privacy at the cost of massive network traffic (2.69 GB per request) and multi-minute latency. COVERT shifts transformations entirely offline, requiring zero multi-round online communication and reducing throughput by less than 6%.
Rating¶
- Novelty: โญโญโญโญโญ [Pioneers practical covariant obfuscation for multimodal VLMs, introducing elegant null-space projection and dimensional bias absorption]
- Experimental Thoroughness: โญโญโญโญโญ [Evaluated across 5 benchmark datasets, 3 parameter scales, multiple architectural families, adaptive inversion attacks, and concurrent vLLM deployments]
- Writing Quality: โญโญโญโญโญ [Clear structural formulation, rigorous algebraic derivation, and transparent threat modeling]
- Value: โญโญโญโญโญ [Offers a practical privacy solution for VLMaaS that preserves high accuracy (<3% drop) and real-time inference throughput (<6% overhead)]