Skip to content

Geometric Inductive Biases for Semi-Supervised Equalization: The Constellation-Aware Transformer

Conference: NeurIPS2026 (task archive; the cache is arXiv v1, acceptance not independently verified)
arXiv: 2609.33695
Area: Signal Processing & Communications
Keywords: semi-supervised equalization, constellation geometry, inter-symbol interference, bidirectional FIR, test-time training

TL;DR

CAT jointly processes the known constellation and received signals from the first layer, with a bidirectional FIR signal branch for inter-symbol interference, reducing pilot requirements through per-block semi-supervised adaptation; the benefit still requires sufficient pilots and receiver-side training, rather than enabling unsupervised or zero-shot decoding across channels.

Background & Motivation

A receiver knows the modulation constellation but may not know how the current channel changes its points. Fading and I/Q imbalance distort the constellation, while multipath superimposes neighboring symbols and produces inter-symbol interference (ISI). Pilots supply correspondences between known symbols and received observations, but every additional pilot displaces a symbol that could carry payload. Classical EM and decision-directed learning can exploit unlabeled payload, yet are sensitive to initialization and incorrect hard decisions. VAE-CNN jointly learns symbol inference and forward-channel reconstruction using a few pilots and a larger unlabeled payload.

Replacing the VAE encoder with a generic Transformer adds sequence interaction without fully exploiting the receiver's existing knowledge. Self-attention over signal tokens must implicitly develop a geometric constellation reference from limited data, and a tokenwise MLP lacks local temporal structure tailored to inverse filtering. The goal is not to add more offline labeled data, but to use each new block's pilots and payload more effectively: the known constellation supplies a template for symbol locations, and bidirectional filtering supplies temporal structure suitable for block equalization.

Core Idea: introduce early interaction between received signals and ideal constellation points in every layer, separating non-causal signal deconvolution from constellation refinement while learning the unknown channel under the same semi-supervised variational objective.

Method

Overall Architecture

Inputs are a received block and all ideal constellation points specified by the modulation scheme; outputs are per-position posterior probabilities over symbol classes. CAT replaces the inference encoder in a semi-supervised VAE, while retaining the forward-channel generative decoder used by the baselines. This VAE decoder reconstructs received observations; it is not the communication decoder that ultimately maps received signals to bits.

A block contains known pilots and unknown payload. Signals and constellation points are separately projected to the same hidden dimension. Signals receive fixed sinusoidal positional embeddings; constellation points, treated as a set, do not. The model then applies 3 TransFIRmer blocks, each consisting of joint bidirectional attention over all tokens followed by separate feed-forward branches for the two token types. Symbol posteriors are read out at the signal positions.

During adaptation, these posteriors and the forward-channel generative decoder use pilot supervision and payload reconstruction jointly. At inference time, the adapted CAT supplies symbol posteriors for hard decisions or soft-bit decoding. Solid edges below denote inference data flow; dashed edges denote supervision and generative branches used only during adaptation.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Received block + ideal constellation"] --> B["Constellation Joint Attention"]
    B --> C["Two-Stream TransFIRmer"]
    C -->|Repeat within the block for 3 layers| D["Per-position symbol posterior"]
    D --> E["Hard decisions / soft-bit output"]
    D -.-> F["Semi-Supervised Channel Reconstruction"]
    G["Pilot labels + received observations"] -.-> F
    F -.->|Training updates only| B

Key Designs

1. Constellation Joint Attention: bring the known geometric template into every layer of symbol inference

A standard Transformer attends only over received signals. CAT concatenates all \(N\) signal tokens and \(K\) ideal constellation tokens into a sequence of length \(N+K\). Queries, keys, and values are produced from this complete sequence, so the mechanism is more than signals querying a fixed constellation dictionary: signalโ€“signal, constellationโ€“constellation, and cross-group interactions are all possible, and constellation representations change with the current block. The full bidirectional mask reflects block equalization as smoothing with both earlier and later observations, rather than autoregressive prediction that must emit one symbol at a time.

Early interaction supplies candidate symbols' ideal locations from the first layer, while the current block teaches the model how the channel distorts them. Whole-block attention is useful even for a memoryless channel: fading and I/Q imbalance are shared within a block, so other positions help identify common distortions. Constellation points receive no positional embeddings because their coordinates have physical meaning, whereas their order in an input list carries no additional physical meaning.

The paper relates signalโ€“constellation correlations to matched filtering. The Gaussian-posterior expansion in Appendix A.3 can be summarized by the following normalization relation, where \(\mathbf z\) is an observation under Gaussian noise and \(\mathbf c_k\) is a candidate constellation point. It explains both why correlation is useful and why the dot product alone is not always sufficient.

\[ p(k\mid\mathbf z)\propto\exp\left(\frac{\mathbf z^T\mathbf c_k}{\sigma^2}-\frac{\|\mathbf c_k\|^2}{2\sigma^2}\right). \]

For constant-modulus constellations, candidate energy terms are equal and cancel during normalization. QAM inner and corner points have different energies, so this term does not cancel. Thus, attention equaling a Gaussian posterior is not an unconditional identity for all QAM constellations. The appendix reports that learned implicit energy normalization reduces the bias from 0.95 nats to 0.04 nats, but this is an empirical observation, not a general exact-posterior guarantee.

2. Two-Stream TransFIRmer: use different operations for temporal deconvolution and constellation refinement

ISI mixes several transmitted symbols into each received position. A tokenwise MLP does not directly supply the neighboring-position operations required for inverse filtering. TransFIRmer therefore splits the standard feed-forward network: signals use forward and backward 1D convolutions, while constellation tokens retain an MLP. The signal branch convolves the original sequence and separately convolves a reversed sequence before reversing that result back; the two outputs are summed. Each token stream has its own residual connection.

\[ \operatorname{FFN}_{\mathrm{sig}}(\mathbf Z)=\operatorname{Conv}_{\mathrm{fwd}}(\mathbf Z)+\operatorname{Flip}\left(\operatorname{Conv}_{\mathrm{bwd}}(\operatorname{Flip}(\mathbf Z))\right). \]

Bidirectionality is substantive: block MMSE equalization uses information on both sides of a position, and the backward branch provides a non-causal filtering bias. Ideal constellation points are not a temporal sequence, so convolving them would introduce an unnecessary ordering assumption; their parallel MLP instead refines geometric representations. Joint attention continues to exchange information between the streams, so these are not two independently trained equalizers.

Appendix F.1 specifies 3 layers, hidden dimension 10, single-head Multi-Query Attention, kernel size 12 for both convolutions, a constellation MLP with one hidden layer of width 10, and dropout 0.1. In the theoretical argument, attention can represent matched correlations and whitening, while bidirectional convolutions can approximate the non-causal filtering structure of a block Wiener receiver. These are arguments about representational capacity and inductive bias: they do not imply that training finds the exact Wiener solution, or that the architecture is an optimal receiver for nonlinear channels.

3. Semi-Supervised Channel Reconstruction: constrain the unknown channel with payload while retaining pilot-based symbol anchoring

The constellation template alone does not identify the current channel transformation. CAT supplies a symbol posterior, and the generative decoder predicts received observations from candidate transmitted symbols. If symbol inference and the channel model jointly explain a large unlabeled payload, a small number of pilot labels can become more effective. Pilots also directly supervise symbol classification, so this is neither purely unsupervised clustering nor repeated training that treats hard model-generated labels as ground truth.

For a memoryless channel, the generative decoder is an MLP with hidden widths 64, 32, and 16 that predicts Gaussian observation means and variances. For ISI, a two-layer MLP with hidden dimension 32 models the memoryless distortion, followed by a learnable complex FIR of length 12 for propagation, with learned noise variance. This training-time forward FIR is distinct from CAT's bidirectional inverse-filtering branch: one explains how the channel generates observations, while the other helps recover symbols from them.

A Worked Example

Consider the main experiment's \(h^{(1)}\) block with 16-QAM, 64 pilots, and 256 payload symbols. Section 2 defines \(N\) to include both pilots and payload, so this example has 320 signal positions plus 16 ideal constellation points; joint attention processes 336 tokens. This calculation illustrates the definition and is not an additional experimental result.

The pilot positions supply 64 known symbol classes; payload positions have no class labels. Each TransFIRmer block first jointly processes the 336 tokens, then applies bidirectional convolution to the 320 signal positions and an MLP to the 16 constellation positions. After 3 layers, signal positions output 16-class posteriors, and the generative decoder uses these latent symbols to explain the received block.

After optimization, SER is measured on the 256 payload symbols of this same block. It measures adaptation to the current block, not direct generalization to an unseen channel. A subsequent block with a changed channel requires adaptation again.

Loss & Training

Equation (2) combines pilot classification cross-entropy, pilot observation negative log-likelihood, and payload negative ELBO. To avoid interpreting the memoryless per-position notation as an independent-observation assumption for ISI, the expression below summarizes the original objective using individually normalized average-loss terms:

\[ \mathcal L_{\mathrm{SSL}}=\alpha\mathcal L_{\mathrm{pilot\text{-}CE}}+\gamma\mathcal L_{\mathrm{pilot\text{-}NLL}}+(1-\gamma)\mathcal L_{\mathrm{payload\text{-}NELBO}}. \]

The two pilot terms are normalized by the pilot count, and the payload term by the payload count. The negative ELBO comprises expected negative reconstruction log-likelihood under the posterior and a KL term relative to the uniform symbol prior. The ISI encoder conditions on the entire received sequence, and its generative decoder explicitly models sequence convolution; each observation cannot be treated as generated only by the current symbol. Appendix E uses a single-sample approximation and Gumbel-Softmax for differentiability.

The default cold-start protocol trains from random initialization for 5,000 steps on every block, using only that block's pilots and unlabeled payload, and measures SER on that same payload. This is transductive test-time adaptation, with neither true payload labels nor data from other channel realizations. The CAVIA baseline and meta-initialized experiments are explicit exceptions that use offline information from previous channels.

AdamW starts at learning rate 0.001 with weight decay 0.01, and the learning rate decays linearly to zero. Mini-batches contain 16 pilot symbols and 32 payload symbols. The classification weight is 0.2; the generative supervision weight is annealed to shift emphasis toward payload, while the Gumbel-Softmax temperature decreases to a minimum of 0.5. Both annealing schedules update every 100 steps. The appendix does not fully explain how mini-batch sampling interacts with whole-block attention; reproduction requires checking this rather than inventing implementation details.

Key Experimental Results

Main Results

Selected rows from Table 2 follow: 16-QAM, \(E_x/N_0=17\) dB, and payload 256. SER is the fraction of evaluated symbols decoded incorrectly, so lower is better. Methods adapt on the current block, and results average 1,000 independent Monte Carlo trials. The optimal column is a known-channel reference decoder, not a deployable baseline under the same information budget.

Channel Pilots VAE-CNN SER Vanilla Transformer SER CAT SER Optimal SER
\(h^{(1)}\), 5 taps 16 0.2900 0.3392 0.3580 0.0121
\(h^{(1)}\), 5 taps 64 0.0523 0.0610 0.0198 0.0121
\(h^{(1)}\), 5 taps 128 0.0494 0.0290 0.0156 0.0121
\(h^{(2)}\), 4 taps 64 0.1002 0.0750 0.0340 0.0101
\(h^{(2)}\), 4 taps 128 0.0869 0.0888 0.0257 0.0101
\(h^{(3)}\), 10 taps 64 0.1709 0.1692 0.1181 Intractable
\(h^{(3)}\), 10 taps 128 0.1087 0.1022 0.0846 Intractable

On the first two channels, CAT with 64 pilots achieves a lower SER than either baseline with 128 pilots. It still improves on the 10-tap channel, but this does not establish the same pilot-halving claim there. CAT also does not always win at the smallest pilot budget: with 16 pilots on \(h^{(1)}\), it is worse than both learned baselines. The claim of superiority across all three ISI channels must be qualified as requiring at least 32 pilots.

Ablation Study

Table 4 uses different conditions for its two columns: memoryless has 16 pilots, 64 payload symbols, and 22 dB; ISI uses \(h^{(1)}\), 64 pilots, 256 payload symbols, and 17 dB. The memoryless column retains the reported 95% confidence intervals without inventing unreported ones. Absolute values across the two columns do not represent equal task difficulty.

Config Memoryless SER ISI SER Note
Full CAT 0.0599 ยฑ 0.0005 0.0198 Joint attention and bidirectional FIR
Remove backward FIR 0.0608 ยฑ 0.0005 0.0287 Forward signal convolution only
Replace FIR with MLP 0.0615 ยฑ 0.0006 0.0351 Retains constellation and joint attention
Remove constellation, retain bidirectional FIR Not reported 0.0438 Isolates the geometric prior
Vanilla Transformer 0.0628 ยฑ 0.0007 0.0610 Signal-only self-attention
Rotate constellation prior by 45 degrees 0.1550 ยฑ 0.0015 Not reported Incorrect geometric reference
Add constellation positional embeddings 0.0631 0.0224 Artificial ordering of the set

Key Findings

  • The gain reflects task-specific structure, not merely extra inputs. Removing backward FIR increases ISI SER from 0.0198 to 0.0287; replacing all FIR with an MLP increases it to 0.0351. Corresponding memoryless changes are smaller, consistent with the absence of ISI to deconvolve.
  • The geometric prior must be correct. Retaining bidirectional FIR while removing constellation input gives 0.0438. An incorrectly rotated prior produces memoryless SER 0.1550, worse than the no-prior vanilla Transformer's 0.0628.
  • SNR and channel length constrain the benefit. In Table 3, CAT leads on \(h^{(1)}\) throughout 14โ€“26 dB. Appendix C.1 shows the smallest FIR gain at 14 dB, when noise limits the relative value of deconvolution structure.
  • Soft outputs also improve, still within link-level simulation. Section 5.2 uses a rate-1/2, length-1024 5G NR LDPC code. At target BLER 0.01, CAT improves over the vanilla Transformer by 2.5 dB and reduces ECE from 0.089 to 0.012, without temperature scaling.

The following computational analysis comes from Table 1: \(N=256\), target SER 0.01, channel \(h^{(1)}\), and SNR 16โ€“24 dB. The dB gain is an improvement in the SNR required for a fixed SER, not a reduction of SER itself in dB.

Modulation Constellation points Extra cost over vanilla attention SNR gain at target SER
16-QAM 16 About 13% 3.8 dB
64-QAM 64 About 56% 3.0 dB
256-QAM 256 About 300% 2.2 dB

Attention expands from the square of the signal length to the square of the combined signal and constellation length. Savings apply to transmitted pilots, not receiver computation. Higher-order modulation reduces the gain and increases computation; I/Q factorization and cheaper attention are proposed future directions, not validated extensions.

Highlights & Insights

  • Bring the known output space into inference. The constellation is not latent structure that must be rediscovered from labels; interactive constellation tokens can focus learning on channel changes. The transferable principle is to supply a correct candidate geometry, not simply add another token group.
  • Use object-specific feed-forward biases. Temporal sequences benefit from filtering, whereas constellation sets need pointwise geometric refinement; global attention complements local bidirectional convolution. The differing ablation effects under ISI and memoryless conditions support this explanation better than an aggregate win rate alone.
  • Separate theoretical capacity from trained performance. Hypothesis-space inclusion shows that access to constellation information cannot worsen the best population MMSE. Appendix A.2 further notes equality for a fixed constellation. Actual few-pilot gains arise through structure and optimization, not a finite-sample advantage automatically established by that inclusion.

Limitations & Future Work

  • Adaptation, not direct unseen-channel generalization. Training and testing both use observations from the same payload block, and the channel is fixed within that block. A changed channel requires retraining. There is no over-the-air, within-block high-mobility, HARQ, or scheduling-level evaluation.
  • Insufficient pilots still cause failure. A constellation prior cannot replace all information needed to learn the unknown channel. The 16-pilot ISI results do not support universal superiority at extremely small pilot budgets.
  • Cold-start cost remains substantial. Appendix C.2 reports SER 0.0156 for 5,000 cold-start steps on \(h^{(1)}\) with 128 pilots, 0.0230 for 50 cold-start steps, and 0.0159 for 50 CAVIA-initialized steps. The last result relies on offline meta-training excluded from adaptation latency, and neither vanilla Transformer nor VAE-CNN receives corresponding meta-training. It is not a matched-total-budget acceleration result.
  • Latency is source-reported only. Appendix C.3 reports 24.5 microseconds for a complete training step on a single RTX 4090, batch size 1, \(N=256\), and 16-QAM. This note does not reproduce the measurement locally or extrapolate it into a real-time deployment claim.
  • The theory does not guarantee finite-sample training performance. The PAC-Bayes pilot-efficiency argument is informal and assumption-dependent; the appendix leaves formal finite-sample separation open. Structural attention/FIR alignment likewise does not guarantee convergence to the optimal receiver.
  • The cache has presentation and reproduction gaps. Section 5.3 says TDL-A, TDL-C, and TDL-D were evaluated, but cached Table 6 contains only TDL-A and TDL-D rows. TDL-C absolute SER values cannot be reconstructed. Table 1's \(N\) is also not synonymous with โ€œpayload 256โ€; computational overhead must not mix these lengths. This note preserves each table's stated conditions.
  • Exact QAM geometry and higher-order extensions need further validation. The energy term cannot be dropped by applying the constant-modulus derivation to QAM. Dynamic phase-tracking gains also require additional PLL tracking and learned rotation. Explicit energy biases, matched meta-training budgets, and high-order constellation factorization merit controlled follow-up experiments.
  • vs VAE-CNN: retains pilot supervision plus payload ELBO and joint forward-channel learning, changing primarily the geometry-aware and temporal encoder. It does not originate the entire semi-supervised variational equalization framework.
  • vs Vanilla Transformer: adds constellation early interaction and two-stream FIR feed-forward structure at comparable scale. Retaining constellation input while reverting to an MLP, and retaining FIR while removing constellation input, isolate the two inductive biases.
  • vs CAVIA / Few-Pilot Meta-Learning: CAVIA obtains cross-task priors from historical channels, whereas cold-start CAT uses only the current block. Their information budgets differ. They can be combined when historical data are available, but baselines should receive the same meta-training opportunity.
  • vs DeepRx, ViterbiNet, DeepSIC: these approaches use offline or supervised protocols, whereas this paper adapts from current-block pilots and unlabeled payload. A stronger practical comparison would account jointly for online compute, offline data, and actual throughput in receiver cost.

Rating

  • Novelty: 4/5. A clear combination of constellation early interaction and object-specific two-stream feed-forward processing, within an inherited variational framework.
  • Experimental Thoroughness: 3/5. Includes ablations, coded links, and standardized-channel simulations, but lacks over-the-air validation, matched meta-training budgets, and complete reproduction details.
  • Writing Quality: 4/5. Clearly qualifies population capacity versus training guarantees; the cached TDL-C table omission still requires checking the original layout.
  • Value: 4/5. Provides reusable structural ideas for few-pilot neural equalization, with practical value contingent on online latency and total compute budget.