Low-Rank Ternary Adaptation for Fine-Tuning Transformers¶
Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/alexmanoo/ternary_adaptation
Area: Model Compression
Keywords: Ternary Transformers / Parameter-Efficient Fine-Tuning / Multiplicative Adaptation / Kronecker Factorization / 1.58-bit Models
TL;DR¶
To address the fundamental inability of existing low-bit LoRA methods to adapt ternary weights without breaking discrete domain constraints upon merging, this paper proposes a low-rank ternary multiplicative adapter via Kronecker factorization that performs discrete keep/zero/flip updates directly within \(\{-1, 0, 1\}\), achieving dequantization-free merging into pure 1.58-bit models.
Background & Motivation¶
Transformer models across vision and language continue to demonstrate remarkable capabilities, yet their steep memory and computational demands make edge and resource-constrained deployment prohibitively expensive. Weight quantization serves as an indispensable pillar of model compression, pushing past standard 8-bit and 4-bit boundaries into extreme low-bit regimes. In this landscape, ternary quantization—which restricts weights strictly to \(\{-1, 0, 1\}\) and demands a theoretical storage budget of \(\log_2 3 \approx 1.58\) bits per weight—represents a premier paradigm for sub-2-bit deployment. For example, it slashes the theoretical memory footprint of an 8B model from 16GB down to just 1.6GB, while replacing power-hungry floating-point multiply-accumulate operations with lightweight additions and sign manipulations.
However, adapting and fine-tuning these highly compressed ternary backbones for specialized downstream tasks exposes severe structural dilemmas. Conventional parameter-efficient fine-tuning (PEFT) frameworks such as LoRA inherently rely on additive continuous updates in real-valued floating-point space. When applied to quantized models, QLoRA must dequantize the low-bit weights back to full precision, add continuous updates, and then requantize the merged model for low-bit inference. In sub-2-bit regimes, this dequantize-requantize loop introduces devastating quantization noise that wipes out fine-tuning gains. While QA-LoRA mitigates this at 2 bits by folding additive updates into group-wise quantization parameters, it breaks down on ternary models: ternary networks offer no continuous zero-points or affine scaling parameters capable of absorbing additive offsets without abandoning the \(\{-1, 0, 1\}\) weight domain. The update space of ternary weights is inherently discrete and non-linear, with the core degrees of freedom being keeping, zeroing, or flipping signs.
Faced with this bottleneck, standard additive adaptations inevitably force the merged model to depart from the ternary domain, forfeiting ternary hardware kernel efficiency. The core idea is to model discrete ternary updates via element-wise multiplicative adaptation, and parameterize this high-rank discrete mask as the Kronecker product of two compact ternary factor matrices, achieving algebraic domain closure and enabling dequantization-free merging into a pure ternary model.
Method¶
Overall Architecture¶
The overall architecture of low-rank ternary multiplicative adaptation is established upon the principle of strict domain closure over \(\{-1, 0, 1\}\). During fine-tuning, the pre-trained ternary base weights remain frozen in packed low-bit format. Instead of calculating continuous additive residuals, the adapter uses two compact real-valued proxy matrices to generate two miniature ternary factor matrices via straight-through projection, which expand via a Kronecker product into a full-sized ternary modulation mask \(\Delta_{\text{tern}} \in \{-1, 0, 1\}^{d_{\text{out}} \times d_{\text{in}}}\). In the forward pass, this mask modifies base weights through an element-wise Hadamard product. Upon completing downstream tuning, the adapter is merged multiplicatively into the base model in a single step, yielding a standalone 1.58-bit model requiring zero inference overhead or runtime dequantization.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Input features and frozen ternary weights W_tern"] --> B["Ternary multiplicative adaptation<br/>Element-wise Hadamard product ensures domain closure"]
B --> C["Low-rank Kronecker factorization<br/>Outer product of two miniature ternary factors"]
C --> D["Real-valued proxies and STE optimization<br/>Backpropagating gradients to continuous proxy weights"]
D --> E["Zero-overhead discrete merge and deployment<br/>W_merged = W_tern ⊙ (A ⊗ B) strictly in 1.58-bit"]
Key Designs¶
1. Ternary multiplicative adaptation: preserving algebraic domain closure via element-wise Hadamard modulation
Standard additive updates inevitably violate ternary constraints: when computing \(W_{\text{merged}} = W_0 + \Delta W\), even a ternary \(\Delta W\) pushes entries into \(\{-2, -1, 0, 1, 2\}\), necessitating lossy requantization. This method departs from additive schemes by formulating the adapted ternary weight matrix \(W'_{\text{tern}}\) as: $\(W'_{\text{tern}} = W_{\text{tern}} \odot \Delta_{\text{tern}}\)$ where \(\odot\) denotes the element-wise Hadamard product, and \(\Delta_{\text{tern}} \in \{-1, 0, 1\}^{d_{\text{out}} \times d_{\text{in}}}\). For any coordinate \((i, j)\), the element \(\Delta_{ij}\) carries self-consistent discrete semantics: \(+1\) retains the original base weight, \(-1\) flips its sign, and \(0\) prunes the connection to zero. Because multiplication over \(\{-1, 0, 1\}\) exhibits complete algebraic closure, the adapted weights remain strictly ternary throughout training and inference, obviating dequantization and requantization entirely.
2. Low-rank Kronecker factorization: constructing expressive high-rank discrete masks from miniature factors
Allocating a full-sized ternary mask of dimension \(d_{\text{out}} \times d_{\text{in}}\) would match full fine-tuning parameter overhead, whereas classical SVD cannot maintain discrete ternary values while preserving sufficient rank. The proposed method decomposes the ternary adaptation matrix via a Kronecker product: $\(\Delta_{\text{tern}} = A \otimes B\)$ where the factor matrices \(A \in \{-1, 0, 1\}^{p \times q}\) and \(B \in \{-1, 0, 1\}^{r \times s}\) satisfy the dimensional compatibility criteria \(p \cdot r = d_{\text{out}}\) and \(q \cdot s = d_{\text{in}}\). In the Kronecker expansion, each entry of \(A\) scales the entire sub-block \(B\). By tensor product properties, the rank of the resulting update scales multiplicatively: $\(\mathrm{rank}(A \otimes B) = \mathrm{rank}(A) \cdot \mathrm{rank}(B) \le \min(p, q) \cdot \min(r, s)\)$ For standard square projection layers (\(d_{\text{out}} = d_{\text{in}} = d\)), selecting balanced factors \(p = q = r = s = \sqrt{d}\) collapses trainable parameter volume to \(p \cdot q + r \cdot s = 2d\). On Llama-3.2-1B, this accounts for merely ~0.06% of total backbone parameters, yet delivers effective rank capacity on the order of \(O(d)\), balancing extreme parameter economy with strong representational expressivity.
3. Real-valued proxies and straight-through optimization: continuous gradients driving discrete state transitions
Because discrete step functions cannot be optimized directly with gradient-based optimizers, the method maintains continuous real-valued proxy parameters \(\bar{A} \in \mathbb{R}^{p \times q}\) and \(\bar{B} \in \mathbb{R}^{r \times s}\). During the forward pass, discrete ternary factors are obtained via an adaptive projection operator based on mean absolute values: $\(A = \operatorname{Tern}(\bar{A}), \quad B = \operatorname{Tern}(\bar{B})\)$ In the backward pass, straight-through estimators (STE) bypass non-differentiable quantization thresholds, delivering gradient signals directly to the continuous proxies updated via AdamW. Furthermore, to ensure the adapted model exactly matches the pre-trained backbone at initialization, a balanced initialization strategy populates \(\bar{A}\) and \(\bar{B}\) with equal distributions of \(+1\) and \(-1\), accompanied by sign compensation on base weights, thereby preventing gradient vanishing and initialization bias.
Loss & Training¶
The training process employs standard cross-entropy loss for language modeling or supervised instruction fine-tuning. For post-training quantized (PTQ) backbones, models are fine-tuned on the Stanford Alpaca dataset for one epoch with AdamW, an initial learning rate of \(1.5 \times 10^{-3}\), linear learning rate decay, a warmup ratio of 0.03, and an effective batch size of 16 on a single NVIDIA A40 GPU (48GB). Upon completion of training, the adaptation factors are folded into the base weights via: $\(W_{\text{merged}} = W_{\text{tern}} \odot (A \otimes B)\)$ The real-valued proxies are discarded post-merge. Deployed layers retain identical tensor dimensions, activations, and pure 1.58-bit ternary formats with zero added latency or storage overhead.
Key Experimental Results¶
Main Results¶
On Llama-3.2-1B and Llama-3.2-3B backbones quantized via SpinQuant, performance across nine zero-shot reasoning benchmarks and WikiText-2 perplexity (PPL) is summarized below:
| Model | Method | Precision | ARC-c | ARC-e | BoolQ | HellaSwag | MMLU | PiQA | 9-Task Avg (%) ↑ | WikiText-2 PPL ↓ |
|---|---|---|---|---|---|---|---|---|---|---|
| Llama-3.2-1B | Full Precision (FP16) | 16-bit | 31.4 | 65.2 | 63.6 | 47.7 | 36.7 | 74.6 | 50.2 | 9.7 |
| RTN Baseline | 2-bit | 22.8 | 23.8 | 62.1 | 25.5 | 22.9 | 53.1 | 32.8 | \(1.5 \times 10^6\) | |
| GPTQ Baseline | 2-bit | 20.2 | 31.6 | 43.1 | 26.3 | 24.1 | 55.2 | 31.4 | 170.0 | |
| SpinQuant Baseline | 2-bit | 18.3 | 37.6 | 62.0 | 28.7 | 22.9 | 57.4 | 34.7 | 43.3 | |
| SpinQuant Backbone | 1.58-bit | 20.3 | 29.7 | 61.3 | 27.0 | 22.9 | 54.5 | 33.2 | 86.5 | |
| Ours (Adapter) | 1.58-bit | 19.5 | 39.7 | 61.2 | 29.4 | 23.1 | 58.2 | 35.3 | 44.6 | |
| Llama-3.2-3B | Full Precision (FP16) | 16-bit | 42.4 | 74.6 | 72.9 | 55.2 | 54.0 | 76.6 | 60.0 | 7.8 |
| RTN Baseline | 2-bit | 22.5 | 24.4 | 37.8 | 25.3 | 26.8 | 52.6 | 30.4 | \(7.6 \times 10^5\) | |
| GPTQ Baseline | 2-bit | 21.5 | 25.2 | 40.8 | 26.0 | 22.9 | 53.1 | 30.0 | 190.0 | |
| SpinQuant Baseline | 2-bit | 19.3 | 29.3 | 53.6 | 28.2 | 22.9 | 57.0 | 32.6 | 41.7 | |
| SpinQuant Backbone | 1.58-bit | 18.4 | 33.3 | 45.6 | 27.9 | 22.9 | 57.2 | 31.9 | 45.6 | |
| Ours (Adapter) | 1.58-bit | 23.8 | 46.0 | 62.2 | 34.0 | 24.8 | 64.6 | 38.3 | 22.3 |
Ablation Study & Baseline Comparisons¶
Downstream adaptation on GSM8K (exact-match accuracy), vision transformer evaluation on ImageNet-100 (Top-1 accuracy), and comparisons with requantized QLoRA are detailed below:
| Benchmark / Model | Configuration / Method | Bits (Base/Adapter) | Metric | Note |
|---|---|---|---|---|
| Llama-3.2-3B (Alpaca) | QLoRA (Requantized post-merge) | 1.58 / 1.58 | 37.5% / PPL 22.9 | Requantization causes severe truncation error |
| Ours (Merged ternary) | 1.58 / 1.58 | 38.3% / PPL 22.3 | Native domain adaptation avoids requantization noise | |
| GSM8K (Math Reasoning EM %) | BitNet b1.58 2B4T (Frozen) | 1.58-bit | 60.1% | Pre-trained native ternary LLM backbone |
| BitNet b1.58 2B4T + Ours | 1.58-bit | 63.0% | +2.9% improvement with 0 inference overhead | |
| Falcon Edge 1B (Frozen) | 1.58-bit | 52.0% | Pre-trained ternary edge model | |
| Falcon Edge 1B + Ours | 1.58-bit | 55.1% | +3.1% gains on multi-turn math reasoning | |
| Falcon Edge 3B (Frozen) | 1.58-bit | 65.4% | Pre-trained ternary edge model | |
| Falcon Edge 3B + Ours | 1.58-bit | 66.4% | +1.0% gain while maintaining pure ternary deployment | |
| ViT-B/16 (ImageNet-100 Top-1) | Full-FT (Full ternary tuning via STE) | 1.58-bit | 85.7% | Upper bound with heavy backpropagation cost |
| QLoRA (Retaining 16-bit adapters) | 1.58 / 16 | 85.3% | Requires dual-precision compute paths at runtime | |
| QLoRA (Requantized to ternary) | 1.58 / 1.58 | 78.9% | Requantization drops accuracy significantly (-6.4%) | |
| Ours (Merged ternary) | 1.58 / 1.58 | 83.0% | Outperforms requantized QLoRA by +4.1%, closing gap to Full-FT |
Key Findings¶
- Requantization degradation is systematically eliminated: At the sub-2-bit boundary, conventional PEFT approaches like QLoRA inevitably incur substantial accuracy collapses when merged and requantized to ternary (dropping 6.4% on ViT-B/16). By formulating updates natively in the ternary space, the proposed method bypasses this bottleneck and achieves 83.0% Top-1 accuracy.
- Weight transitions reveal a strict zero-locking dynamic: Transition matrix tracking shows that 100% of weights quantized to 0 in the base model remain exactly 0 post-adaptation, because \(0 \times \Delta_{ij} \equiv 0\). All empirical gains stem from redistributing and calibrating the signs among existing non-zero entries (\(\pm 1\)).
- Sub-0.1% parameter budgets surpass 2-bit PTQ: Updating only ~0.06% of model parameters on Llama-3.2-3B cuts perplexity nearly in half (from 45.6 down to 22.3), comfortably exceeding both the 1.58-bit SpinQuant baseline and the more parameter-heavy 2-bit SpinQuant counterpart across all benchmark tasks.
Highlights & Insights¶
- Algebraic paradigm shift from additive shift to multiplicative closure: By re-conceptualizing weight adaptation as an element-wise Hadamard product over \(\{-1, 0, 1\}\), this work circumvents the domain escape dilemma inherent to traditional additive LoRA methods.
- Kronecker factorization decouples parameter count from rank capacity: Exploiting the tensor product identity \(\mathrm{rank}(A \otimes B) = \mathrm{rank}(A)\mathrm{rank}(B)\), the adapter generates high-rank discrete modulation patterns using only \(O(d)\) trainable proxy parameters.
- Broad generality across architectures and pre-training schemes: The approach operates seamlessly across post-training quantized transformers, natively trained 1.58-bit models (BitNet, Falcon Edge), and visual architectures (Ternary ViT) without requiring structural alterations.
Limitations & Future Work¶
- Inability to reactivate pruned weights (Zero-Locking): Because multiplicative updates cannot transition elements out of zero, layers subjected to aggressive initial pruning cannot restore potentially useful connections, placing a structural ceiling on recovery.
- Dimensional divisibility constraints: Kronecker factorization requires layer dimensions \(d_{\text{out}}\) and \(d_{\text{in}}\) to factor into balanced integer pairs \((p, r)\) and \((q, s)\). While transformer projection dimensions are typically powers of two, asymmetric or non-standard architectures may induce factor imbalances.
- Unaddressed activation quantization: The current framework targets ternary weights while maintaining activations in 8-bit or 16-bit precision; developing fully unified discrete weight-activation multiplicative tuning remains an open direction.
Related Work & Insights¶
- vs QLoRA [Dettmers et al., 2023]: QLoRA dequantizes weights to FP16 to add continuous low-rank updates, incurring either dual-precision inference overhead or severe degradation upon post-hoc requantization. The proposed adapter operates natively in the ternary domain for zero-overhead, lossless merging.
- vs QA-LoRA [Xu et al., 2023]: QA-LoRA integrates additive updates into group-wise quantization scales and offsets, a technique fundamentally incompatible with ternary weights lacking continuous zero-points. This paper provides the first principled PEFT solution tailored to sub-2-bit ternary models.
- vs LoRMA [Bihany et al., 2025]: LoRMA utilizes continuous multiplicative transformations to rotate weight spaces in floating-point models, whereas this method tackles sign flips and zeroing in discrete \(\{-1, 0, 1\}\) algebraic spaces.
Rating¶
- Novelty: ⭐⭐⭐⭐⭐ [Pioneering multiplicative Kronecker adaptation within the discrete ternary domain, solving the sub-2-bit PEFT merge dilemma]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Validated across 6 language and vision backbones, encompassing PTQ, native ternary LLMs, ViT, and detailed weight transition audits]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, mathematically rigorous formulation, and crisp architectural illustrations]
- Value: ⭐⭐⭐⭐⭐ [Critical milestone for the practical fine-tuning, domain specialization, and zero-overhead deployment of 1.58-bit green edge models]