Skip to content

ORFC: Orthogonal Reparameterization for Low-Bitrate ViT Feature Coding

Conference: ECCV 2026
Paper: ECCV Official
Code: https://github.com/zhangletian2/ORFC
Area: Model Compression / LLM Efficiency
Keywords: Feature Coding, Product Quantization, Orthogonal Reparameterization, Split Inference, Vision Transformer

TL;DR

Addressing the severe task degradation and low-bitrate collapse in edge-cloud split ViT inference caused by feature-space MSE misalignment and highly uneven channel subspace task sensitivity, ORFC incorporates a differentiable Cayley-parameterized orthogonal rotation prior to fixed-group product quantization and optimizes an entropy-constrained rate-distortion objective using frozen-tail output deviation as a proxy distortion, substantially eliminating task performance collapse under 0.1 BPFP.

Background & Motivation

Vision Transformers (ViTs) have emerged as foundational backbones across diverse visual representation learning and downstream perception tasks, yet their substantial model parameters and intensive self-attention computational overhead critically obstruct real-time deployment on resource-constrained edge devices. Edge-cloud split inference addresses this dilemma by partitioning the model into an edge front (executing early layers to leverage on-device compute) and a cloud tail (handling the heavier remaining workload). However, this architectural partition shifts the primary system bottleneck from on-device computation to bandwidth-constrained edge-cloud communication channels, rendering ultra-low-bitrate lossy compression of intermediate high-dimensional feature tensors the decisive factor for practical split deployment.

Conventional intermediate feature coding frameworks predominantly adapt standard video codecs (such as VVC/VTM) or neural feature compression networks, minimizing Euclidean reconstruction errors such as mean squared error (MSE) directly within the feature space. This paradigm implicitly hinges on a foundational premise: preserving Euclidean numerical fidelity in the feature space is sufficient to guarantee the performance of the remaining network on downstream tasks. Nevertheless, for intermediate ViT representations operating in the ultra-low-bitrate regime (e.g., below 0.1 BPFP), this assumption suffers from two structural breakdowns. First, minor perturbations in intermediate features are nonlinearly magnified through multi-layer global self-attention across the cloud tail, meaning that small feature-space MSE can still induce catastrophic deviation in downstream predictions. Second, the task sensitivity across channel subspaces exhibits severe anisotropy: perturbations along certain channel directions critically compromise task accuracy, whereas other directions contribute negligibly.

When employing fixed-group product quantization (Grouped PQ)โ€”which offers substantial hardware throughput and low memory footprint in productionโ€”such sensitivity disparity induces severe bit allocation inefficiencies. Classical transform coding rate-distortion theory establishes that optimal bit allocation must satisfy the equal-slope criterion: the marginal reduction in task distortion per consumed bit must remain balanced across all quantized components. Because fixed-group PQ assigns an identical codebook size and implicit uniform bit budget to every group, precious bits are squandered on task-insensitive subspaces while critical task-sensitive directions remain underrepresented, culminating in a steep performance collapse at low bitrates. Rather than introducing complex variable-rate codebooks or non-uniform quantization hardware, this paper tackles the dilemma directly from representation geometry: Core idea: learn a differentiable orthogonal rotation matrix via the Cayley transform to reparameterize intermediate features before grouped product quantization, redistributing heterogeneous task sensitivities evenly across PQ groups to satisfy the equal-slope criterion, optimized end-to-end against a task-aligned proxy distortion measured by frozen-tail output deviation.

Method

Overall Architecture

ORFC is specifically engineered for edge-cloud split inference, comprising an edge encoder, a differentiable rate-distortion training pipeline, and a cloud decoder. An input image is processed through the edge front Transformer blocks to extract intermediate features \(F \in \mathbb{R}^{T \times D}\) (where \(T\) denotes the sequence length and \(D\) the channel dimension). At the edge, features undergo magnitude normalization to decouple global activation scales, followed by a learnable orthogonal rotation matrix \(R \in \mathbb{R}^{D \times D}\) to project the representation into the rotated space \(Z = F R\). The rotated tensor \(Z\) is partitioned along channels into \(G\) disjoint subspaces, each quantized via entropy-constrained product quantization against dedicated codebooks to yield discrete indices. These indices are compressed via arithmetic entropy coding conditioned on learned group-specific prior probabilities and transmitted across the wireless link. At the cloud server, the entropy decoder recovers the discrete indices, reconstructs rotated subvectors \(\hat{Z}\) by codebook lookup, restores original feature geometry via transposed inverse rotation \(\hat{F} = \hat{Z} R^\top\) and scale recalibration, and feeds \(\hat{F}\) into the frozen cloud tail and task heads to execute downstream inference.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    In["Input Intermediate Feature F<br/>Magnitude Normalization"] --> R1["Cayley Orthogonal Rotation Transform<br/>Rotate Channels with Matrix R: Z=FR"]
    R1 --> R2["Entropy-Constrained Grouped Product Quantization<br/>Fixed Grouping & Group-Adaptive Priors"]
    R2 --> Comm["Arithmetic Entropy Coded Bitstream<br/>Edge-Cloud Wireless Transmission"]
    Comm --> Dec["Cloud Dequantization Lookup<br/>Inverse Transpose Transform F_hat=Z_hat R^T"]
    Dec --> R3["Frozen-Tail Output Deviation Proxy Distortion<br/>Compute Tail Output Deviation vs Ground Truth"]
    R3 --> R4["Straight-Through End-to-End RD Optimization<br/>Annealed Soft Assignment for Rotation & Codebooks"]

Key Designs

1. Cayley Orthogonal Rotation Transform: balancing cross-group task sensitivities To eliminate bit wastage caused by severely anisotropic task sensitivity across channel subspaces under uniform group budgets, ORFC inserts an orthogonal transformation matrix \(R \in \mathbb{R}^{D \times D}\) prior to feature grouping. The intrinsic norm-preserving property of orthogonal transforms (\(R^\top R = I\)) preserves global energy and Riemannian manifold structure without scale distortion, while yielding an exact inverse via trivial matrix transposition (\(R^{-1} = R^\top\)) with zero inversion overhead. Directly optimizing orthogonal constraints on the Stiefel manifold using projection steps or regularization penalties often destabilizes gradient dynamics. Instead, ORFC leverages the classic Cayley transform to achieve an unconstrained parameterization: optimizing a strictly skew-symmetric matrix \(A \in \mathbb{R}^{D \times D}\) (where \(A^\top = -A\)) and analytically generating a strictly orthogonal matrix in the forward pass via: $\(R(A) = (I - A)^{-1}(I + A)\)$ Driven by downstream task-aligned gradients, this rotation reorganizes coordinate axes within the feature space, dispersing concentrated high-sensitivity geometric components across all \(G\) fixed-dimensional groups. This ensures that every group exhibits comparable marginal task utility per bit without requiring irregular codebooks or hardware-unfriendly dynamic bit-allocation logic.

2. Entropy-Constrained Grouped Product Quantization: achieving unified rate-distortion assignment Following channel rotation, the feature tensor \(Z\) is sliced along the channel dimension into \(G\) distinct subvectors \(z_{t,g} \in \mathbb{R}^d\) of dimension \(d = D/G\). Each subspace is assigned an independent codebook \(C_g = \{c_{g,1}, \dots, c_{g,N}\}\). Conventional nearest-neighbor assignment relies solely on Euclidean distance \(\text{dist}_{t,g}(j) = \|z_{t,g} - c_{g,j}\|_2\), ignoring codeword empirical frequencies and their impact on entropy-coded bitstream length. ORFC introduces a learnable categorical prior distribution for each group, parameterized by logits \(\mathbf{p}_g \in \mathbb{R}^N\) and normalized via Softmax into codeword prior probabilities \(\boldsymbol{\pi}_g = \text{softmax}(\mathbf{p}_g)\). Codeword index selection is guided by a joint rate-distortion cost function governed by a shared Lagrange multiplier \(\lambda\): $\(\mathrm{cost}_{t,g}(j) = \mathrm{dist}_{t,g}(j) + \frac{1}{\lambda}\left(-\log_2 \pi_{g,j}\right)\)$ with hard assignment determined by \(k_{t,g} = \arg\min_j \mathrm{cost}_{t,g}(j)\). This formulation directly integrates rate constraints into edge-side local quantization decisions, biasing selections toward high-probability codewords with minimal distortion, while providing an exact and differentiable rate estimator \(-\log_2 \pi_{g,k_{t,g}}\) without requiring computationally demanding autoregressive context models.

3. Frozen-Tail Output Deviation: establishing a task-aligned proxy distortion Traditional feature codecs optimize reconstruction MSE directly in the intermediate feature space; however, empirical findings confirm that intermediate feature MSE correlates poorly with downstream degradation in classification, dense detection, and multimodal retrieval. Conversely, directly propagating supervision from diverse task-specific loss heads (e.g., cross-entropy or bounding box regression) causes the codec to overfit to a single task, destroying universal visual representations and demanding extensive labeled annotations. ORFC overcomes this dilemma by treating the cloud-side tail network \(M_{l+1\to L}(\cdot)\) as a frozen, label-free, task-agnostic geometric evaluator. Denoting uncompressed feature forward output as \(o = M_{l+1\to L}(F)\) and reconstructed feature forward output as \(\hat{o} = M_{l+1\to L}(\hat{F})\), the task proxy distortion is defined as their Euclidean deviation in the deep representation space: $\(\mathcal{L}_{\text{ref}} = \|o - \hat{o}\|_2\)$ Because \(o\) incorporates non-linear cross-token self-attention and deep contextual semantics, its deviation faithfully mirrors the tail's sensitivity to intermediate feature degradation, effectively steering the orthogonal rotation and codebooks toward preserving task-critical feature subspaces without requiring task labels.

4. Straight-Through End-to-End RD Optimization: stabilizing joint learning of codebooks and rotation The complete training objective is formalized as a rate-distortion Lagrangian: \(\mathcal{L} = R_{\text{bits}} + \lambda \mathcal{L}_{\text{ref}}\), where total bitrate is \(R_{\text{bits}} = \sum_{t=1}^T \sum_{g=1}^G (-\log_2 \pi_{g, k_{t,g}})\). Standard Straight-Through Estimators (STE) struggle in this setting: because the entropy penalty heavily penalizes low-frequency codewords, naive hard STE causes codebook collapse during early iterations. ORFC incorporates temperature-annealed soft assignment straight-through estimation. A temperature-controlled soft assignment probability over candidate codewords is defined as \(\tilde{\mathbf{w}}_{t,g} = \text{softmax}(-\mathrm{cost}_{t,g} / \tau)\), which is fused with the hard one-hot assignment \(\mathbf{h}_{t,g} = \mathbf{e}_{k_{t,g}}\) via: $\(\mathbf{w}_{t,g} = \mathbf{h}_{t,g} - \text{sg}[\tilde{\mathbf{w}}_{t,g}] + \tilde{\mathbf{w}}_{t,g}\)$ In the forward pass, subvectors are reconstructed using exact discrete codewords; in the backward pass, continuous gradients flow smoothly across all candidate codewords, prior distributions \(\boldsymbol{\pi}_g\), and the orthogonal matrix \(R\). By exponentially annealing temperature \(\tau\) from 0.5 down to 0.005, the optimization smoothly transitions from exploration to sharp discrete quantization, ensuring stable convergence of both discrete codebooks and continuous orthogonal rotations.

Loss & Training

The framework is trained on unannotated natural images (e.g., 5,000 images sampled from ImageNet) by extracting intermediate representations at target split layers. Optimization is carried out via Adam with a learning rate of \(3 \times 10^{-4}\) and Lagrange multiplier \(\lambda = 0.5\). Codebook capacity is explored across \(N \in \{4, 8, 16, 64, 256\}\), subvector dimensions \(d \in \{16, 32\}\), and group configurations \(G \in \{32, 64\}\). Soft assignment temperature \(\tau\) decays exponentially \(\tau_t = \tau_0 \cdot \gamma^t\) towards the terminal value of 0.005 to guarantee high codebook utilization and stable convergence.

Key Experimental Results

Main Results

On DINOv2 ViT-L/14 (layer 20 split point), VitDet ViT-B/16 (detection layer 3), and SigLIP2 (multimodal retrieval layer 8), ORFC is benchmarked against leading compression baselines: standard video codec VTM, neural hyperprior model LaMoFC, and Optimized Product Quantization (OPQ). BD-rates are calculated relative to VTM (negative values denote bitrate savings).

Model & Backbone Downstream Task / Dataset Bitrate Point (BPFP) Ours (ORFC) OPQ VTM LaMoFC
DINOv2 ViT-L/14 (blk10) ImageNet Classification (Acc@1) 0.06 BPFP 96.0% 34.2% 19.8% (0.11 BPFP) <1.0%
DINOv2 ViT-L/14 (blk20) ImageNet Classification (BD-rate) Full Range -77.5% Baseline Ref 0.0% (Anchor) +120.5%
DINOv2 ViT-G/14 (blk20) VOC2012 Segmentation (mIoU) 0.06 BPFP 83.1% 52.4% 65.5% (0.09 BPFP) 31.2%
VitDet ViT-B/16 (blk03) COCO2017 Detection (mAP) 0.05 BPFP 91.6% 36.7% 49.5% <5.0%
SigLIP2 So400m/14 (blk08) COCO2017 Retrieval (Recall@1) 0.08 BPFP 79.4% 58.1% 62.3% 2.8% (Collapsed)

Ablation Study

On DINOv2 ViT-L/14 block 20 features, an extensive ablation assesses the isolated contributions of proxy distortion \(\mathcal{L}_{\text{ref}}\), orthogonal rotation \(R\), and annealed soft assignment (Soft assign) across classification (BD-rateCls) and dense segmentation (BD-rateSeg) relative to VTM:

Config ID Tail Proxy Distortion \(\mathcal{L}_{\text{ref}}\) Orthogonal Rotation \(R\) Soft Assign Annealing BD-rateCls (vs VTM) BD-rateSeg (vs VTM) Note
(a) Vanilla Baseline โŒ (Feature MSE) โŒ (Identity) โŒ (Hard assign) -6.7% +190.7% Conventional MSE-guided grouped PQ suffers severe segmentation degradation
(b) Add Rotation โŒ (Feature MSE) โœ”๏ธ โŒ -5.2% +153.2% Pure MSE-driven rotation fails to align with downstream task sensitivity
(c) Add Proxy Distortion โœ”๏ธ โŒ (Identity) โœ”๏ธ -57.0% +14.4% Without rotation, severe subspace sensitivity imbalance limits dense prediction
(d) Remove Soft Annealing โœ”๏ธ โœ”๏ธ โŒ (Hard assign) -24.8% +181.2% Hard assignment causes codebook collapse and sub-optimal RD trade-offs
(e) Full Model (ORFC) โœ”๏ธ โœ”๏ธ โœ”๏ธ -77.5% -41.4% All modules collaborate synergistically to achieve massive rate savings on both tasks

Key Findings

  • Orthogonal rotation is decisive for dense prediction: Comparing config (c) with (e), under the same tail proxy distortion, introducing orthogonal rotation \(R\) slashes segmentation BD-rate from +14.4% to -41.4% (over 55% bitrate savings). Dense segmentation requires fine-grained spatial coherence, which benefits crucially from redistributing error across subspaces.
  • Feature MSE is fundamentally misaligned: Config (b) with MSE-driven rotation yields negligible gain over baseline (a) in classification (-5.2% vs -6.7%) and remains disastrous in segmentation (+153.2%), demonstrating that minimizing Euclidean feature reconstruction error does not safeguard machine task utility.
  • Robust cross-layer and cross-task generalization: When trained on ImageNet features and transferred directly to NYU Depth v2 depth estimation, ORFC achieves BD-rate reductions of -73.3%, -90.0%, -64.1%, and -50.6% across blocks 5, 10, 15, and 20, confirming that the learned orthogonal representation geometry is highly task-agnostic.
  • Negligible computational and memory overhead: On an NVIDIA RTX 4090 GPU with ViT-L/14 (\(D=1024\)), the orthogonal transform adds only 0.36 ms latency and 25.5 MB memory, bringing total edge encoding time to just 1.23 ms, fully satisfying real-time latency constraints in edge-cloud systems.

Highlights & Insights

  • Bridging transform coding theory and deep feature geometry: By transposing the equal-slope rate-distortion principle from classical signal processing to foundation model split inference, ORFC elegantly resolves subspace sensitivity imbalance with a lightweight, parameter-efficient orthogonal rotation.
  • Label-free surrogate evaluation via frozen-tail forward deviation: Using the frozen tail network as an implicit task evaluator captures multi-layer self-attention error propagation without requiring task labels, preventing overfitting while ensuring robust alignment with machine vision tasks.
  • Hardware-friendly uniform grouping meets mathematical differentiability: Preserving regular, parallelizable fixed-size codebooks while offloading complexity to continuous coordinate rotation via Cayley reparameterization delivers an optimal compromise between implementation efficiency and algorithmic efficacy.

Limitations & Future Work

  • Split-point specificity: The learned orthogonal rotation matrix \(R\) and tail evaluator are tied to a specific split layer \(l\). In dynamic split computing frameworks where split points adapt on-the-fly to fluctuating wireless bandwidth, separate codec checkpoints must be maintained.
  • Independence across spatial tokens: The current design quantizes each visual token independently, omitting spatial correlations and semantic redundancy among adjacent visual patches, leaving potential compression margins untapped.
  • Future directions: Integrating spatial token pruning or dynamic token merging within the rotated orthogonal subspace to jointly eliminate spatial and channel redundancies at extreme low bitrates.
  • vs OPQ (Optimized Product Quantization): OPQ optimizes orthogonal rotation to minimize feature-space MSE, which fails to balance downstream task sensitivities under low bitrates, leading to sharp accuracy collapse; ORFC optimizes rotation guided by tail output deviation, achieving task-level equal-slope bit allocation.
  • vs VTM (Versatile Video Coding Reference Software): VTM provides strong high-bitrate compression but suffers severe cliff-effect degradation below 0.1 BPFP, accompanied by hundreds of milliseconds of encoding latency; ORFC eliminates the low-bitrate collapse while maintaining ultra-low 1.2 ms encoding latency.
  • vs LaMoFC: LaMoFC relies on neural hyperpriors optimized for feature reconstruction, which struggle to converge under complex transformer feature distributions; ORFC pairs robust grouped PQ with orthogonal reparameterization, delivering substantially superior rate-task trade-offs.

Rating

  • Novelty: โญโญโญโญโญ First work to identify subspace sensitivity imbalance via the equal-slope criterion in ViT feature coding and address it via Cayley orthogonal reparameterization.
  • Experimental Thoroughness: โญโญโญโญโญ Comprehensive validation across 5 ViT architectures, 4 diverse downstream tasks, multiple split layers, with extensive ablation, generalization, and latency analyses.
  • Writing Quality: โญโญโญโญโญ Rigorous motivation, coherent mathematical derivations, and crystal-clear presentation connecting theory to empirical validation.
  • Value: โญโญโญโญโญ Provides an immediately deployable, ultra-lightweight, and high-performance communication compression paradigm for edge-cloud collaborative intelligence.