SEED: Self-Speculative Decoding via Implicit Encoder–Decoder¶
Conference: NeurIPS 2026
arXiv: 2609.36590
Code: https://github.com/lhk2004/SEED
Area: LLM Efficiency / Self-Speculative Decoding
Keywords: self-speculative decoding, implicit encoder–decoder, deep KV cache, block-causal training, tree verification
TL;DR¶
SEED trains the last two layers of an existing decoder-only model as a drafter that reuses deep KV caches from the previous full-model verification, achieving average decoding speedups of 2.6×/2.7× over AR on Qwen3-1.7B/4B while maintaining or improving task quality under task-specific fine-tuning and adaptive tree verification.
Background & Motivation¶
An autoregressive large language model processes every generated token through the entire network; in the memory-bound setting discussed in the paper, GPU parallel compute remains underused. Speculative decoding first proposes multiple candidates with a cheap drafter, then verifies them in one parallel target-model pass, amortizing a full forward pass over more outputs. A separate drafter requires additional weights, training, and deployment, motivating self-speculative approaches that generate candidates within the target model itself.
The challenge is not simply reducing the layer count, but retaining the contextual depth needed for prediction while doing less computation. Early-exit methods such as LayerSkip use the initial layers, reducing cost but accessing shallower contextual representations. MTP-style methods exploit deep features but still need full-model representations for a drafting round, potentially with additional sampling modules. SEED observes that the preceding verification has already processed the accepted prefix through the full model: KV caches in its final layers can serve subsequent drafting without recomputing the heavy initial stack for every draft token.
This does not mean the latest token already has a full-depth representation. The extra token emitted by verification has not yet been processed as a full-model input, and later drafts also begin as raw embeddings; the lightweight final layers must learn to combine deep caches of an old prefix with shallow inputs from a new segment. Core idea: make full verification provide the contextual encoding for the next drafting round, and train the existing final layers to generate consecutive candidates against a fixed deep cache, skipping expensive initial-layer computation without discarding deep context.
Method¶
Overall Architecture¶
SEED adds no explicit encoder, cross-attention module, or external drafter. It conceptually partitions the existing Transformer into an initial encoder and a final decoder. The Qwen3-1.7B example uses 26 encoder layers and 2 decoder layers out of 28; the 4B model likewise uses only its last 2 layers. The full model remains the verifier, and its final-layer drafter shares parameters with it.
Training first processes a ground-truth sequence through the full model, producing the standard next-token loss and decoder KV entries at each position. A block-causal mask then trains the final layers to predict a future segment from raw embeddings. At inference, the final layers draft a candidate tree using the verified prefix cache; the full model verifies and accepts a path, refreshes the cache, and starts another round. Training supervision consists of the auxiliary loss, whereas inference follows a repeated draft–verify loop.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Ground-truth sequence /<br/>inference prefix"] --> B["Implicit Final-Layer Drafter"]
B --> C["Block-Causal Joint Training"]
C -.->|Training supervision: next-token and draft losses| B
B -->|Inference: raw embeddings and deep cache| D["Adaptive Tree Verification"]
D --> E["Accepted path +<br/>bonus token"]
E -->|Refresh cache and resume drafting| B
Key Designs¶
1. Implicit Final-Layer Drafter: skip initial layers while reusing context they already encoded
In a normal full forward pass, a token's raw embedding passes through the initial encoder before entering the final decoder. During drafting, its raw embedding enters the final decoder directly, without computing its encoder representation; final-layer attention reads the deep prefix cache created by the preceding full verification. The new input is shallow, but the historical context it can access is deep, unlike early exiting through the initial layers.
These caches are the final decoder layers' own KV entries, not KV entries from the first 26 layers connected directly to the last two. They carry deep context because, when they were created, historical tokens had already passed through the initial encoder and entered the corresponding final layers. The paper's term cross-attention emphasizes shallow new queries attending to deep historical keys and values; implementation still uses the existing self-attention and KV cache, with no added conventional encoder–decoder cross-attention module.
In the paper's notation, \(H^0\) denotes raw token embeddings and \(n'\) marks the current draft segment's start. The drafting distribution can be summarized as:
The first condition contains existing tokens within the draft segment; the second is the deep cache preceding its boundary. Generation inside the segment remains token-by-token autoregressive, not independent parallel prediction of all future tokens. Temporary final-layer KV entries generated during drafting are also not equivalent to full-model verified KV entries. Until the next verification, the deep prefix stays fixed and the new segment runs only through the lightweight final layers.
2. Block-Causal Joint Training: adapt the final layers to mixed deep and shallow inputs
Unadapted final layers normally receive encoder outputs, so directly substituting raw embeddings changes their input distribution. SEED adds a speculative drafting loss to standard supervised fine-tuning: a full forward pass first obtains final-layer KV entries at every position, then inputs are divided into non-overlapping blocks. Each token can causally access lightweight drafting states in its own block and full-depth KV entries from preceding blocks, but not full-depth KV entries from the current block.
This restriction prevents training from using encoder features that have not been recomputed at deployment time. As all positions within a block share the same deep-cache boundary, all blocks can be processed in parallel in one final-layer pass; training need not sequentially simulate the entire verification loop. The main training block size is 10, a conditioning horizon rather than a requirement to draft exactly 10 tokens at inference.
Both paths share model parameters. The standard next-token loss retains task learning for the full verifier, while the speculative loss trains the final layers' new input mode and can affect contextual representations through the cache path. Future-token linear probes and cross-layer CKA support the authors' explanation of more future-predictive features, but this remains an experimentally supported mechanism hypothesis, not a mathematical demonstration of planning ability.
3. Adaptive Tree Verification: allocate candidates by confidence and maintain exact cache boundaries
The basic version generates a fixed-length draft and verifies it with one full-model call. Main experiments additionally use confidence-based stopping and a candidate tree: high confidence permits longer drafting, whereas confidence below a threshold stops it; candidate counts depend on the current top-1 probability. Probabilities above 0.95 retain 1 candidate, \((0.8,0.95]\) retains 3, \((0.5,0.8]\) retains 5, and the remaining case retains 10. Stopping thresholds are 0.7 for GSM8K, KodCode, and ScienceQA, and 0.5 for summarization.
Full verification packs the tree with a tree-attention mask. A node can attend to the entire verified prefix and its own ancestors, but not unrelated branches. Starting at the root, the algorithm accepts the path matching the verifier's greedy predictions, then appends one verifier-produced bonus token. At rejection, this is the correction at the first mismatched position; if all drafts are accepted, it is the following token. Other branches and cache entries beyond the accepted path must be discarded.
Although predicted by the verifier, the bonus token has not itself entered the full model as an input. The next round's deep cache therefore covers only positions before it. Drafting begins from its raw embedding, and full verification later recomputes it together with the candidate segment; temporary drafting caches must not be committed as full-depth caches. This boundary—one more output token than cached full-depth token—is central to connecting verification with the next drafting round.
Under the greedy setting actually evaluated, exact verification should reproduce AR outputs from the same jointly trained full verifier. Losslessness does not mean that SEED training leaves the original Base weights or a separately AR-fine-tuned model unchanged. Empirical preservation of task accuracy is not invariance of weights or output distributions; strict equivalence under random sampling would additionally require the corresponding sampling acceptance rule, which the experiments do not cover.
A Worked Example¶
The following illustration explains state boundaries and is not a measured case from the paper. Suppose the current full KV cache covers positions 1–100 and the output already contains the bonus token at position 101.
The final-layer drafter reads the raw embedding at 101 and the deep prefix, proposing four candidates at positions 102–105 without invoking the initial encoder.
Full verification first clears temporary lightweight drafting states, then uses deep cache 1–100 to process input positions 101–105 in parallel, obtaining target predictions for positions 102–106.
If only positions 102 and 103 match target predictions, with a mismatch at 104, the algorithm accepts 102 and 103 and appends the verifier's prediction at 104 as a correction. Three new tokens are committed, rather than retaining incorrect candidates at 104 and 105.
The full KV cache is retained only through position 103. The next round starts from the new token at 104 as a raw embedding; it cannot assume this correction already has a deep KV entry. Tree verification extends the same logic to multiple branches, while committing only one accepted path.
Loss & Training¶
Let \(\mathcal{L}_{CE}\) be the full model's standard next-token cross-entropy and \(b\) the block size. For target position \(j\), the boundary is determined by the block containing its input position \(j-1\). The main text and appendix algorithm give:
Following the appendix algorithm, the sum starts at the valid next-token position \(j=2\), avoiding a nonexistent predecessor for the first token in the boundary formula. The joint objective is:
The main configuration uses a 2-layer drafter, \(b=10\), and \(\lambda=1.0\). Task-specific fine-tuning uses 100 steps of linear warm-up followed by cosine decay, with validation-loss checkpoint selection and early stopping. The maximum budget is 30,000 steps, and learning rates for 1.7B/4B are \(10^{-5}\)/\(5\times10^{-6}\); this does not imply every run completes the full budget.
The source contains an unresolved attention-description ambiguity: the method text specifies block-causal masks preserving standard AR structure, and Appendix B.3 also says fully block-causal, but the Figure 2 caption explicitly gives context/prompt tokens bidirectional attention and response tokens causal attention. Causality for responses and draft blocks is clear; these passages alone do not establish whether prompts are bidirectional. Code inspection would be required to settle the precise implementation.
Key Experimental Results¶
Main Results¶
All methods use greedy decoding, with throughput measured on one 48 GB NVIDIA A6000. Generation stops at <|endoftext|> rather than continuing repetitive text to inflate acceptance. Training-based methods, including SEED and AR, are separately fine-tuned on task data. Main results use adaptive drafting and tree verification by default. The table selects 4B results from Table 1; throughput is in tokens/s, quality is accuracy (%) for the first three tasks and ROUGE-1/2/L for summarization.
| Dataset | AR quality | SEED quality | AR throughput | EAGLE-3 throughput | PARD throughput | SEED throughput |
|---|---|---|---|---|---|---|
| GSM8K | 75.1 | 76.0 | 28.3 | 64.4 | 57.7 | 74.9 |
| KodCode | 76.2 | 81.0 | 27.6 | 54.1 | 65.4 | 80.5 |
| ScienceQA | 95.4 | 96.8 | 26.0 | 56.0 | 62.7 | 89.3 |
| CNN/Daily Mail | 39.0 / 18.2 / 29.1 | 39.8 / 18.4 / 29.3 | 24.0 | 49.9 | 41.0 | 43.3 |
| Paper-reported average speedup | — | — | 1.0× | 2.1× | 2.1× | 2.7× |
SEED leads throughput on the first three 4B tasks but trails EAGLE-3 on summarization; average leadership does not imply being fastest everywhere. On 1.7B, average speedup is 2.6×, while GSM8K accuracy rises from 56.3% to 57.1% and throughput from 37.9 to 91.8 tokens/s.
The paper's approximately 28% advantage over EAGLE-3 summarizes average acceleration. The appendix explains that corresponding Base drafter weights were unavailable, so EAGLE-3 uses public Instruct models and paired drafters, whereas SEED uses task-specific Base fine-tuning. This is not a pure decoder comparison with identical target checkpoints. The 4B PARD comparison instead uses the corresponding task-specific AR-fine-tuned checkpoint.
Dataset scope matters: KodCode evaluation is restricted to easy online-judge problems; ScienceQA uses only text-only examples, merges the original train/validation splits for training, and samples 1,000 test examples. CNN/Daily Mail evaluates up to 1,000 examples with a maximum total length of 1,024 and a 90%/10% article/summary token allocation. Results do not automatically extend to difficult coding tasks, visual ScienceQA, or very long contexts.
Ablation Study¶
Table 6 analyzes tree verification on a GSM8K-fine-tuned Qwen3-1.7B model. Acceptance length counts accepted draft tokens per round, excluding the bonus token. Acceptance rate is a draft-acceptance statistic; simply dividing reported averages does not necessarily recover its aggregation convention.
| Drafting strategy | Tree verification | Throughput (tokens/s) | Average acceptance length | Acceptance rate |
|---|---|---|---|---|
| Fixed \(d=4\) | Yes | 82.3 | 3.16 | 79.80% |
| Fixed \(d=4\) | No | 76.9 | 2.67 | 69.45% |
| Fixed \(d=8\) | Yes | 82.5 | 4.65 | 59.10% |
| Fixed \(d=8\) | No | 75.2 | 3.67 | 46.64% |
| Dynamic \(\tau=0.7\) | Yes | 91.8 | 3.97 | 88.57% |
| Dynamic \(\tau=0.7\) | No | 85.1 | 3.28 | 77.59% |
Table 7 specifically removes parameter sharing. Both configurations use training block size \(b=4\), fixed draft length \(d=4\), and no tree verification. Its 74.3 tokens/s is therefore not a direct module-removal counterpart of the main experiment's 91.8.
| Config | Accuracy (%) | Throughput (tokens/s) | Acceptance rate | Average acceptance length |
|---|---|---|---|---|
| SEED, shared parameters | 57.5 | 74.3 | 73.0% | 2.9 |
| Without parameter sharing | 44.1 | 68.8 | 68.5% | 2.7 |
A source-number difference is retained: Table 6 reports 76.9 tokens/s for fixed \(d=4\) without trees, whereas Table 7 and Appendix F report 74.3 for their fixed-draft configuration. Table 7 explicitly uses training block size 4, versus 10 in the main configuration; Table 6 does not restate block size in its caption. The values should not be forced into agreement or labeled a typographical error without further evidence.
Key Findings¶
- Dynamic drafting with tree verification adds 6.7 tokens/s over dynamic single-candidate drafting, but the latter still reaches 85.1. Candidate expansion does not explain the entire core gain.
- Removing parameter sharing reduces task accuracy by 13.4 percentage points and acceptance rate by 4.5 points. This supports the importance of joint optimization, but it is not a control that changes only the inference path while freezing the verifier.
- Raising \(\lambda\) from 0.1 to 1.0 increases throughput from 81.5 to 91.8 but reduces accuracy from 58.5% to 57.1%; at 1.5 these become 92.1 and 55.5%. Stronger drafting supervision does not necessarily improve task quality.
- Appendix E.2 continues training from an AR-fine-tuned checkpoint, obtaining 56.5% accuracy and 90.6 throughput, close to direct SEED training's 57.1% and 91.8. This demonstrates adaptation through continued training, not training-free plug-and-play acceleration.
Highlights & Insights¶
- Verification also encodes: full-model verification provides more than acceptance decisions—it leaves context the next round can consume. Reusing computation already paid for is more direct than introducing another draft-feature network.
- A small model need not have shallow context: the 2-layer drafter accesses historical final-layer KV entries produced by a full forward pass. Draft-network depth and historical-feature depth can be designed separately, the central distinction from early exiting.
- Layer count alone does not determine cost: Appendix F estimates that two final layers plus the vocabulary head still touch about 411.9M weights on 1.7B, costing 0.239 full-forward units rather than simply \(2/28\). Vocabulary-head overhead explains why a layer-only calculation overstates savings.
Limitations & Future Work¶
- The authors evaluate only 1.7B, 4B, and task-specific fine-tuning, not 8B/14B/32B, large-scale pretraining, or general instruction post-training. Reported speedups are not universal model constants.
- Throughput comes from one A6000 and does not cover highly concurrent batching, different hardware, serving queues, or long contexts. Cache layouts and tree-attention kernels can change realized gains.
- The prompt bidirectional/causal attention-description conflict remains unresolved. Reproduction should inspect implementation masks, position numbering, and bonus-token cache-commit boundaries.
- The main text describes Apple MTP as needing a full pass at each step, but Appendix F explicitly amortizes one full-model call across a draft block with per-token sampling-head costs. It would be misleading to claim Apple MTP independently runs the entire network for every candidate.
- Linear probes support more distant future-token predictability, not a causal demonstration of real planning ability. Freezing different components and testing out of domain could separate representation changes from task-specific training bias.
Related Work & Insights¶
- vs LayerSkip / SWIFT / DEL: these retain earlier layers or dynamically skip layers; SEED concentrates drafting in the final layers to access deeper verified context. It requires training those layers for raw-embedding inputs and maintaining distinct full-verification and lightweight-drafting cache states.
- vs EAGLE-3: both exploit deep prefix information, but SEED shares final-layer parameters without a separate drafter. The appendix's rough EAGLE-3 per-step drafting estimate is about 0.080, below SEED's 0.239, so SEED's advantage is not lower absolute drafting cost alone; acceptance and verification efficiency matter.
- vs E2D2: SEED inherits amortization through expensive contextual encoding and lightweight generation, but uses causal drafts and full AR verification. In the appendix's controlled comparison with \(b=4\), fixed \(d=4\), and a 2-layer decoder, E2D2 achieves 23.7% / 75.9 tokens/s versus SEED's 57.5% / 74.3. The main distinction is quality, not raw speed.
- Research direction: final-layer partitions could adapt to hardware, context length, and acceptance statistics, provided unverified deep states never cross the causal boundary. This is a prospective direction, not an implemented capability in the paper.
Rating¶
- Novelty: 4/5. Combines verification caches, final-layer drafting, and block-causal training into self-speculation without adding architecture.
- Experimental Thoroughness: 4/5. Covers four task types, two scales, and several controlled ablations, but lacks larger models, serving batches, and an identical-target EAGLE comparison.
- Writing Quality: 3/5. The mechanism is intuitive, but prompt masks and MTP cost descriptions require appendix-level qualification.
- Value: 4/5. Useful for deployments willing to continue training and prioritizing single-request decoding efficiency, with benefits dependent on hardware and cache implementation.