Skip to content

SemMSA: Latent Semantic-Aided Robust Multimodal Sentiment Analysis with Incomplete Data

Conference: NeurIPS 2026 (task-list placement; this note uses arXiv v1)
arXiv: 2609.30238
Area: Multimodal VLM (multimodal sentiment analysis)
Keywords: missing modalities, latent semantics, continuous hidden states, kernel spectral alignment, instance separation

TL;DR

SemMSA feeds incomplete language, visual, and acoustic evidence into a frozen LLM, recursively appends continuous hidden states as auxiliary semantics, and learns robust representations through within-instance kernel spectral alignment and cross-instance separation, achieving average Acc-2 scores of 74.36/73.91 on MOSI, 79.61/79.38 on MOSEI, and 75.46 on SIMS across ten intra-modal missing rates.

Background & Motivation

Multimodal sentiment analysis uses language content, visual expressions, and acoustic cues in video clips to predict annotated sentiment polarity and intensity. With complete inputs, the modalities can complement each other; incomplete data may instead lose individual words, frames, or acoustic segments, or an entire modality. This paper primarily addresses the former, intra-modal missingness: all three sequences may retain fragmented evidence, so the model cannot assume that one stream remains reliable. The task discussed here is sentiment-label prediction on standard datasets, not diagnosing personal mental health or inferring sensitive attributes from individual data.

Existing methods broadly follow two directions. Reconstruction methods such as TFR-Net attempt to recover missing features, but the same text can accompany different prosody and visual expressions: statistically plausible features need not preserve a clip's sentiment meaning. Methods such as LNLN and P-RMF emphasize fusion or dominant-modality guidance. They exploit observed evidence but may propagate noise when the guiding modality is severely corrupted. Both approaches face a shared difficulty: when low-level evidence is insufficient, how can task-relevant semantics be supplemented without relying on a fixed alignment anchor that may itself be incomplete?

LLM contextual representations offer another source of compensation, but decoding natural-language descriptions incurs additional cost, and those descriptions need not be accurate. SemMSA retains hidden-space computation and treats semantics as continuous representations for a trainable downstream model, rather than readable explanations. Core Idea: supplement incomplete multimodal evidence with continuous hidden-state task semantics, jointly constrain the four representations through anchor-free kernel spectral alignment within each instance, and prevent collapse through cross-instance separation.

Method

Overall Architecture

The input consists of incomplete language, visual, and acoustic sequences; the output is a continuous sentiment score, from which classification metrics are computed according to dataset protocols. Two complementary representation paths remain available. Original modalities pass through feature processors to produce three task representations, while Cross-modal Semantic Refinement (CSR) constructs a fourth, semantic representation from the same evidence. Cross-modal Spectral Alignment (CSA) then constrains the joint structure of these four representations during training.

Within CSR, visual and acoustic adapters compress variable-length non-language sequences into prefixes accepted by the frozen LLM; language uses that LLM's token embeddings directly. The frozen LLM repeatedly processes an updated prefix, appending its final-position hidden state at each iteration, and the collected continuous states are pooled. The four task representations are added element-wise and passed to a linear prediction head. The LLM does not directly produce the final sentiment answer.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["Incomplete language, vision, audio"] --> A["Evidence Adaptation<br/>Non-language prefixes and language embeddings"]
    A --> R["Continuous Semantic Refinement<br/>Recursive state appending with a frozen LLM"]
    I --> M["Original-modality task representations"]
    R --> F["Sum four representations<br/>Linear sentiment prediction"]
    M --> F
    R -.->|Training only: semantic representation| C["Kernel Spectral Alignment<br/>and Instance Separation"]
    M -.->|Training only: three modality representations| C
    Y["Training labels"] -.-> L["Regression loss and joint optimization"]
    F -.-> L
    C -.-> L
    F --> O["Inference: sentiment score"]

Dashed edges represent training constraints, whereas solid edges represent prediction data flow. CSA is not a required inference-time feature transformation: in the paper's equations, final fusion directly sums the four representations, and eigendecomposition computes training losses.

Key Designs

1. Evidence Adaptation: compress incomplete non-language sequences before entering LLM space

Visual and acoustic inputs cannot simply serve as language token embeddings. Each non-language adapter uses 8 learnable prompt vectors and adds learnable positional embeddings to the input sequence to preserve frame or segment order. The prompts query observed sequences through cross-attention, then exchange information through self-attention. After 2 lightweight Transformer blocks, the adapter outputs a fixed number of compact vectors and linearly projects them to the frozen LLM's embedding dimension.

These prompts are learnable continuous vectors, not manually written instructions. They convert variable-length, partially missing sequences into fixed-budget non-language prefixes and allow task supervision to learn which cues to aggregate, rather than restoring complete data frame by frame. Language is embedded directly by the LLM, and the initial prefix concatenates language, vision, and audio in that fixed order. Thus, โ€œanchor-freeโ€ describes the subsequent alignment objective, not a fully symmetric sequence ordering in CSR.

Two types of encoders should also be distinguished. The paper first uses features supplied by frozen BERT, OpenFace, and Librosa, then modality processors with linear and Transformer layers form downstream task representations. Language enters CSR through the frozen LLM's embeddings instead. BERT feature extraction and LLM token embedding are not the same operation.

2. Continuous Semantic Refinement: append the final-position hidden state as the next input

The initial prefix contains observed evidence from all three modalities. Each LLM pass produces a new continuous semantic state from the final position of the last layer, which is appended to the existing prefix. The next pass sees both the original evidence and the previous state; earlier appended states also remain. This process neither generates an explanation and re-encodes it nor produces a textual chain-of-thought with discrete word tokens.

Using the paper's notation, \(\mathcal{F}_{\theta}\) denotes the frozen LLM and \(\mathbf{U}^{(0)}\) the initial prefix. The essential recurrence is:

\[ \mathbf{z}_{k}=\mathcal{F}_{\theta}(\mathbf{U}^{(k-1)})_{|\mathbf{U}^{(k-1)}|},\qquad \mathbf{U}^{(k)}=[\mathbf{U}^{(k-1)};\mathbf{z}_{k}],\quad k=1,\ldots,O. \]

The default uses 4 refinement steps, each adding one continuous state with the same dimensionality as the LLM embeddings. All collected states are pooled and projected into a downstream semantic representation whose dimension matches the three modality representations. The source specifies pooling without identifying its type, so mean or maximum pooling should not be asserted as the authors' implementation.

Freezing LLM parameters does not eliminate LLM computation. The adapters and downstream modules still learn how to exploit the fixed transformation for the task, and each refinement step adds an LLM forward pass. The appendix reports improvements from 1 to 4 steps, followed by slight decreases at 5 and 6 steps. โ€œToken-efficientโ€ means avoiding long decoded natural-language descriptions, not cost-free semantic generation, nor a guarantee that each hidden state is a faithful and interpretable reasoning step.

3. Kernel Spectral Alignment and Instance Separation: build within-sample consensus without losing between-sample distinctions

Adding a semantic branch is insufficient unless it learns structure that can be used jointly with the original modalities. CSA normalizes the semantic, visual, acoustic, and language representations of the same sample, takes them as the four columns of \(\mathbf{V}=[\mathbf{h}_{s},\mathbf{h}_{v},\mathbf{h}_{a},\mathbf{h}_{l}]\), and constructs a \(4\times4\) Gram matrix with an RBF kernel. โ€œGlobalโ€ here refers to the joint relationships among these four representations, not a large kernel matrix mixing every modality in an entire batch.

\[ K_{pq}=\exp\left(-\frac{\|\mathbf{h}_{p}-\mathbf{h}_{q}\|_{2}^{2}}{2\sigma^{2}}\right),\qquad \mathcal{L}_{\mathrm{csa}}=-\frac{1}{N}\sum_{i=1}^{N}\log\frac{\exp(\lambda_{1}^{i}/\tau)}{\sum_{j=1}^{4}\exp(\lambda_{j}^{i}/\tau)}. \]

Each sample's kernel matrix is separately decomposed into four descending nonnegative eigenvalues. Training encourages the largest eigenvalue to dominate a softmax, rather than designating language or semantics as an anchor and pulling the other branches toward it individually. The RBF kernel has unit diagonal, so the eigenvalues sum to 4. The appendix proves that complete alignment corresponds to a rank-one matrix with spectrum \((4,0,0,0)\). This motivates enhancing the dominant spectral component, but the implemented loss is a temperature-scaled eigenvalue softmax. It is not identical to directly minimizing residual spectral energy or maximizing \(\lambda_1/4\), and it does not guarantee that rank one has been reached.

If only within-sample consistency is encouraged, different samples may collapse to similar representations. The authors take each kernel matrix's dominant eigenvector as combination coefficients over the four representations, construct a direction back in the original representation space, and average squared inner products between directions from different samples in the batch:

\[ \mathbf{u}_{1}^{i}=\frac{\mathbf{V}^{i}\mathbf{q}_{1}^{i}}{\|\mathbf{V}^{i}\mathbf{q}_{1}^{i}\|_{2}},\qquad \mathcal{L}_{\mathrm{sep}}=\frac{1}{N(N-1)}\sum_{i=1}^{N}\sum_{j\ne i}\left[(\mathbf{u}_{1}^{i})^{\top}\mathbf{u}_{1}^{j}\right]^2. \]

The dominant eigenvector contains only four combination coefficients and cannot itself be compared as a high-dimensional semantic vector. More precisely, those coefficients reflect structure in the reproducing kernel Hilbert space (RKHS), but the paper uses them to weight the four original-space vectors. This is not an explicit RKHS principal direction, and no exact inverse mapping from kernel space to original space is demonstrated. Squared inner products penalize strongly correlated directions with either sign, encouraging approximate orthogonality across instances. This is not a supervised contrastive loss organizing positive and negative samples by sentiment class.

A Worked Example

Consider a clip with only part of its review text, several video frames, and fragmented acoustic segments remaining. This is a schematic walkthrough, not a measured example published in the paper. The visual and acoustic adapters each produce 8 prefix vectors, appended after language embeddings. The first LLM pass takes the final-position hidden state and appends one continuous state; the second repeats the operation on the updated prefix, until 4 states have been collected. These are not 4 generated words or 4 explanatory sentences.

The states are pooled into one semantic representation, while the original-modality processors separately produce visual, acoustic, and language representations. At inference time, the four representations are simply summed before sentiment regression. Training additionally constructs this clip's own \(4\times4\) kernel matrix and encourages a relatively dominant largest eigenvalue. Its projected direction is also compared with those of other training samples to avoid a shared direction for all clips.

The example clarifies that compensation targets discriminative semantic representations, not the faithful restoration of missing frames, sounds, or words. If the remaining evidence is unreliable, the model can still predict an incorrect score; a hidden-state loop does not recover facts from nothing.

Loss & Training

The joint objective sums sentiment-score MSE, kernel spectral alignment loss, and instance separation loss, with no additional weights in the source equation. Final fusion adds the unnormalized task representations element-wise. Normalized representations are used for CSA regularization, so final fusion should not be described as spectrally weighted voting.

Training uses AdamW, a learning rate of \(10^{-4}\), batch size 64, at most 200 epochs, warm-up, cosine annealing, and early stopping. The frozen LLM is Qwen3-1.7B. Defaults are \(M=8\), \(B=2\), \(O=4\), \(\sigma=1.0\), and \(\tau=0.1\), with an NVIDIA RTX 6000 Ada GPU.

Instance-wise Bernoulli masking is applied during training, with 50% of samples kept complete. Missing visual and acoustic positions are replaced by zero vectors, and missing language positions by [UNK]. Testing independently samples missing positions per modality at rates from 0.0 to 0.9 in steps of 0.1, averaging results over seeds 1111, 1112, and 1113.

Section 3 uses \(d\) for the LLM embedding dimension, whereas the experimental configuration lists \(d=128\) as the hidden dimension. This notation is ambiguous, so 128 should not be asserted as Qwen3-1.7B's token embedding dimension. Reproduction also requires verification of internal feature dimensions, pooling, and any caching used in the refinement loop.

Key Experimental Results

Main Results

MOSI and MOSEI labels lie in \([-3,+3]\), and SIMS labels in \([-1,+1]\). Their train/validation/test splits contain 1,284/229/686, 16,326/1,871/4,659, and 1,368/456/457 samples, respectively. The table below selects results from source Tables 1โ€“2. Every score is averaged across ten intra-modal missing rates, not measured on complete data alone.

MOSI and MOSEI Acc-2 and F1 follow two protocols: before the slash is negative/non-negative, and after it is negative/positive. Classification metrics are percentages; lower MAE and higher Corr are better.

Dataset Method Acc-2 F1 MAE Corr
MOSI LNLN 70.94/72.55 71.25/72.73 1.046 0.527
MOSI P-RMF 71.53/72.81 71.69/72.93 1.038 0.525
MOSI TF-Mamba 73.46/72.54 73.59/72.57 1.035 0.548
MOSI SemMSA 74.36/73.91 74.18/73.82 1.011 0.550
MOSEI P-RMF 78.83/78.14 80.39/79.33 0.658 0.589
MOSEI TF-Mamba 77.34/77.61 77.18/77.43 0.673 0.578
MOSEI SemMSA 79.61/79.38 80.62/79.87 0.648 0.601
SIMS P-RMF 73.64 74.65 0.500 0.414
SIMS TF-Mamba 74.68 72.20 0.512 0.386
SIMS SemMSA 75.46 77.50 0.474 0.501

Under the first Acc-2 protocol, MOSI improves by 0.90 percentage points over TF-Mamba, MOSEI by 0.78 points over P-RMF, and SIMS by 0.78 points over TF-Mamba. SIMS Corr improves by 0.087 over P-RMF. The strongest baseline can differ by metric, so these differences should not be summarized as one uniform relative percentage improvement.

Inter-modal missingness is evaluated separately on MOSEI. The sets in source Table 3 denote retained, not removed, modalities: SemMSA F1 is 83.61 with language only, 66.59 with audio only, 65.91 with vision only, and 72.68 with audio and vision but no language. Its average across seven retained combinations is 77.89 versus CorrKD's 75.18, a 2.71-point difference. The three-modality score in that table is 86.32 and should not be mixed with the preceding ten-rate average.

Ablation Study

The following table selects results from source Table 4. MOSI Acc-2 uses only the pre-slash protocol; SIMS Acc-3/Acc-5, whose column labels conflict elsewhere, are omitted in favor of unambiguously corresponding metrics.

Config MOSI Acc-2 MOSI F1 MOSI MAE SIMS Acc-2 SIMS F1 SIMS Corr
No CSR or spectral losses 67.49 63.28 1.085 66.83 57.92 0.387
CSR only 73.36 73.04 1.040 74.18 76.21 0.450
CSR + kernel spectral alignment 73.94 73.83 1.024 74.92 76.88 0.493
CSR + kernel spectral alignment + instance separation 74.36 74.18 1.011 75.46 77.50 0.501

CSR provides the largest initial gain: MOSI Acc-2 increases by 5.87 points and SIMS by 7.35 points. Adding kernel spectral alignment raises SIMS Corr from 0.450 to 0.493, and instance separation raises it to 0.501. This supports an interpretation in which semantic compensation contributes most of the information gain and alignment/separation further organize representations. Incremental ablation is not a complete causal analysis of module interactions.

The efficiency table selects source Table 6. Trainable parameters and Task GFLOPs count only modules optimized for sentiment analysis; memory and latency count the entire pipeline, including the frozen LLM.

Method Trainable parameters Task GFLOPs Memory GB Latency ms
LNLN 116M 9.0 12.8 25.1
P-RMF 117M 9.7 13.5 67.0
SemMSA 113M 4.8 13.0 30.6

SemMSA has fewer task FLOPs than both comparators and lower latency than P-RMF, but higher latency than LNLN. It is therefore not the fastest end-to-end method among the three, and 4.8 GFLOPs does not describe total computation including the frozen LLM.

Key Findings

  • More refinement is not always better. In appendix Table 11, SIMS F1 at 1/4/6 steps is 75.94/77.50/76.84, and Corr is 0.468/0.501/0.489. The authors suggest excess states introduce redundancy or noise, but do not directly validate that explanation.
  • Anchor-free alignment does not lead every metric. In Table 5(b), SemMSA's SIMS F1 is 77.50 versus Volume's 76.35, but its Corr of 0.501 is below Volume's 0.506.
  • An average lead does not imply a lead at every missing rate. At MOSI missing rate 0.9, SemMSA Acc-2 is 58.97/55.95 versus TF-Mamba's 60.20/60.37. Semantic compensation cannot eliminate extreme evidence scarcity.
  • Source numbers and headers conflict. Tables 2 and 11 report SIMS Acc-3/Acc-5 as 58.84/35.68, whereas some headers and corresponding row orders in Tables 4 and 5 are reversed; this note does not silently repair those tables. The main text also attributes MOSI Acc-5/Acc-7 improvements of 2.43/2.23 to comparison with TF-Mamba. Table 1 gives TF-Mamba 37.74/33.95 and SemMSA 40.93/36.49, yielding 3.19/2.54 percentage points instead. The stated 2.43/2.23 more closely matches differences from the strongest baseline for each metric, so that textual attribution should not be reused.

Highlights & Insights

  • Semantic compensation need not mean natural-language generation. The frozen LLM participates in continuous representation computation while a task model retains responsibility for prediction, suggesting a transferable way to use language priors under limited decoding budgets.
  • Within-sample consistency and between-sample discrimination are separated into two geometric constraints. A small kernel matrix captures consensus across four branches, while cross-instance direction regularization suppresses collapse; modality agreement is not treated as sufficient for discrimination.
  • An adapter does more than match dimensions. SIMS F1 in Table 5(a) is 73.62 for linear projection versus 77.50 for the proposed adapter, suggesting that evidence aggregation deserves consideration alongside LLM size.

Limitations & Future Work

  • The authors acknowledge that evaluation primarily covers video opinion clips. It does not establish generalization to conversational emotion recognition, sarcasm detection, or long-form interactions. Separate training and evaluation on three datasets is not evidence of zero-shot cross-language transfer.
  • Continuous semantic states lack natural-language explanations, making evidence attribution difficult. When most remaining cues are unreliable, neither CSR nor CSA guarantees successful compensation.
  • Fixed prefix order, independently sampled random missingness, and 50% complete training samples are explicit protocol assumptions. Real systems can experience contiguous loss, correlated sensor failures, or nonrandom missingness, motivating structured-missingness evaluation.
  • CSA encourages a single shared direction and may compress complementary modality-specific information. Instance separation also ignores sentiment class and may push same-label samples apart. Class structure, spectral distributions, and their relationship to task performance merit examination beyond classification gains.
  • Some baseline scores are adopted from previous papers without unified reproduction or statistical-significance evidence. Ambiguous dimension notation, SIMS column labels, and improvement attribution increase reproduction costs. Task GFLOPs and end-to-end overhead require separate auditing.
  • Dataset scores establish label-prediction capability only. They should not be extended to personal mental-health assessment, sensitive-attribute inference, or high-stakes individual decisions; real data use also requires informed consent, privacy safeguards, and bias auditing.
  • vs TFR-Net: TFR-Net reconstructs missing modality features, whereas SemMSA supplements hidden-space task semantics. The latter avoids equating reconstruction quality with sentiment-information quality, but does not establish that its latent semantics faithfully reflect the original evidence.
  • vs LNLN / P-RMF: These methods emphasize dominant-modality or proxy guidance; SemMSA's CSA does not preselect a single branch as an anchor. This distinction concerns the alignment objective, since CSR still uses a fixed languageโ€“visionโ€“audio prefix order.
  • vs InfoNCE / Volume: The paper compares particular implementations substituted as alignment objectives. Its kernel spectral loss jointly observes four within-instance representations and adds instance separation. This does not establish that all contrastive learning requires a single anchor or that every volume objective cannot model joint structure.
  • vs direct multimodal LLM prediction: In Table 7, Qwen2.5-Omni-7B directly predicts MOSI F1 of 69.92/71.58, whereas Qwen3-1.7B as SemMSA's semantic source reaches 74.18/73.82. Model choice and task adaptation both change, so the comparison supports the auxiliary-representation route in this setup, not a universal disadvantage of direct LLM prediction.

Rating

  • Novelty: 4/5 โ€” Continuous semantic refinement combined with a within-instance kernel spectral objective gives the method a distinct design.
  • Experimental Thoroughness: 4/5 โ€” Three datasets, two missingness granularities, ablations, and efficiency analysis are included, but unified reproduction and realistic missingness validation remain absent.
  • Writing Quality: 3/5 โ€” The main pipeline is clear, but spectral interpretation versus the actual objective, table headers, and numerical attribution need clarification.
  • Value: 4/5 โ€” A practical route for leveraging a frozen LLM under incomplete multimodal evidence, with deployment value constrained by reliability and total computation.