title: >- [Paper Note] SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark description: >- [ECCV 2026][Human Understanding][Multimodal Emotion Recognition] A balanced multimodal benchmark comprising 30K speaker-segments across seven emotions with strict movie-level splits. tags: - ECCV 2026 - Human Understanding - Multimodal Emotion Recognition - Benchmark Dataset - Class Imbalance date: 2026-09-19 content_hash: a65a49ffb90583fd
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark¶
Conference: ECCV 2026
Paper: ECCV Official Link
Code: https://github.com/xxx
Area: Multimodal VLM / Human Understanding
Keywords: Multimodal Emotion Recognition, Benchmark Dataset, Speaker Segments, Class Imbalance, Cross-Dataset Generalization
TL;DR¶
Addressing the severe class imbalance and actor/scene leakage in existing conversational multimodal emotion benchmarks, SpEmoC introduces a large-scale benchmark of 30,000 speaker-segment clips across seven balanced emotion categories from 3,100 movies with strict movie-level partitioning, delivering robust cross-dataset and minority-emotion generalization.
Background & Motivation¶
Understanding human emotions in spoken conversations is fundamental to affective computing, empathetic AI agents, and conversational human-computer interaction. Emotion is intrinsically multimodal, expressed synergistically across facial expressions, vocal prosody, and linguistic context. However, advancing multimodal emotion recognition (MER) has been severely constrained by foundational benchmark design flaws. Early datasets like IEMOCAP and RAVDESS were captured under controlled studio or acted conditions, falling short in natural expressive richness. Subsequent benchmarks such as MELD (from the TV sitcom Friends) and CAER (from TV shows) expanded towards dialogue settings, but remain constrained in scale (typically ~13k utterances) and suffer from acute emotion distribution skew. In both MELD and CAER, the "Neutral" category dominates 35% to 48% of samples, while critical minority classes such as "Fear" and "Disgust" constitute less than 3% to 5%.
This extreme class disparity induces severe majority-class bias in deep learning architectures, causing trained models to collapse on minority negative emotions with F1-scores frequently dropping near zero. Moreover, current benchmarks rely on random or session-level data partitioning where the same characters, costumes, visual backgrounds, and actor mannerisms appear across training and test splits. Consequently, the test set is not truly unseen, leading to over-optimistic evaluation metrics that fail to reflect in-the-wild generalization. Additionally, untrimmed conversational video segments often include background interruptions, overlapping speech, and absent active speakers, introducing significant noise into multimodal representation alignment.
To overcome these structural limitations, this work rethinks benchmark curation from the ground up rather than proposing superficial model tweaks. Core idea: by anchoring conversational units to natural speaking segments and enforcing strict movie/franchise-level splits, SpEmoC employs an automated text-audio logit fusion pipeline with KL-divergence consistency regularization paired with expert multi-rater verification to deliver the first large-scale, tightly synchronized, class-balanced seven-emotion benchmark.
Method¶
Overall Architecture¶
The curation and benchmarking pipeline of SpEmoC processes raw long-form videos through an automated multimodal extraction, alignment, and filtering workflow, followed by triple-blind expert verification. Starting from 306,544 raw dialogue segments harvested from 3,100 full-length feature films and series (\(\ge 40\) min), the pipeline enforces speech-unit integrity, logit-based cross-modal pseudo-labeling with divergence penalties, dual-threshold neutral filtering, and face visibility checks, culminating in 30,000 verified, balanced speaking clips partitioned under a strict movie-independent protocol.
%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
A["Long-Form Video Library<br/>3,100 movies & TV series"] --> B["Utterance-Aligned Speaking Segment Extraction"]
B --> C["Logit-Based Multimodal Pseudo-Labeling with KL Regularization"]
C --> D["Dual-Threshold Neutral Filtering & Face Verification"]
D --> E["Tri-Modal Expert Blind Validation & Movie-Level Partitioning"]
E --> F["SpEmoC Balanced Benchmark<br/>30,000 clips · 7 balanced emotions"]
Key Designs¶
1. Utterance-Aligned Speaking Segment Extraction: Extracting Complete Semantic Units
To avoid arbitrary temporal windowing that fragments emotional semantics and facial trajectories, the framework transcribes 3,100 full-length videos using the Whisper ASR model equipped with word-level timestamps. Each candidate segment \(v_k^i\) must satisfy two rigid criteria: (1) containing at least 12 words to guarantee sufficient linguistic and emotional context, and (2) terminating with definitive punctuation (., !, ?) to capture a grammatically complete utterance. Clips typically span 3 to 6 seconds, naturally focusing on a single active speaker. Using FFmpeg, the aligned audio waveform \(\mathcal{A}_k^i\), visual sequence \(v_k^i\), and transcript \(\mathcal{T}_k^i\) are isolated across the interval \([t_{\text{start}}, t_{\text{end}}]\). Temporal consistency is verified by enforcing: $\(|\text{Duration}(\mathcal{A}_k^i) - (t_{\text{end}}(v_k^i) - t_{\text{start}}(v_k^i))| \le \epsilon\)$ with maximum permissible drift \(\epsilon = 0.1\) s, systematically eliminating desynchronized footage.
2. Logit-Based Multimodal Pseudo-Labeling with KL Regularization: Robust Multimodal Agreement
To generate scalable and reliable pseudo-annotations without relying on noisy visual backgrounds, sentiment logits are extracted via DistilRoBERTa (fine-tuned on text) and Wav2Vec 2.0 (pretrained on speech), yielding unnormalized confidence vectors \(L_t, L_a \in \mathbb{R}^7\) over Ekman's seven emotion classes. Under a uniform prior \(P(e_i) = 1/|E|\), the posterior fuses Softmax-normalized probability distributions \(\tilde{P}_t\) and \(\tilde{P}_a\). A symmetric Kullback-Leibler (KL) divergence penalty is integrated to penalize inter-modal contradiction: $\(S(e_i) = \log \tilde{P}_t(e_i) + \log \tilde{P}_a(e_i) - \lambda D_{\text{KL}}(\tilde{P}_t \parallel \tilde{P}_a)\)$ where \(\lambda = 0.5\) balances the regularization strength. The predicted pseudo-label is given by \(e^* = \arg\max_{e_i \in E} S(e_i)\), accompanied by confidence score \(F(e^*) = \frac{1}{1 + \exp(-S(e^*))}\). Samples are retained only when predictions strictly align (\(\arg\max L_t = \arg\max L_a\)) and confidence exceeds calibrated criteria, eliminating ambiguity through distributional consensus.
3. Dual-Threshold Neutral Filtering & Face Verification: Restoring Category Balance
The initial 306,544 segments exhibited intense concentration in the neutral class. To eliminate weak-affect noise and restore balance, three threshold filters are deployed: (1) YOLOv8 runs frame-by-frame to compute human face presence \(f_k^i \ge \theta_f = 0.9\), requiring the active speaker's face to appear in at least 90% of frames; (2) Sigmoid-transformed neutral logits \(w_{t,k}^i = \sigma(L_{t,\text{neutral}})\) and \(w_{a,k}^i = \sigma(L_{a,\text{neutral}})\) must satisfy \(w_{t,k}^i < \theta_t = 0.05\) and \(w_{a,k}^i < \theta_a = 0.05\). By discarding unexpressive neutral clips while retaining a calibrated 15% representative neutral baseline, the candidate pool is refined to 50,000 highly expressive clips with an evenly distributed class profile.
4. Tri-Modal Expert Blind Validation & Movie-Level Partitioning: Preventing Subjective Bias and Data Leakage
The 50,000 filtered clips undergo rigorous human verification by 20 English-proficient expert annotators trained on Ekman guidelines. Each sample is evaluated independently across all three modalities (video, audio, text) by at least three annotators without access to model-generated pseudo-labels. Consensus is reached via majority voting, achieving substantial inter-annotator agreement (Fleiss' \(\kappa = 0.62\)). Inconclusive clips are discarded, securing 30,000 gold-standard instances. Data splitting strictly adheres to movie- and series-level isolation: all clips belonging to the same film, sequel, or TV franchise are assigned exclusively to either train (70%), validation (10%), or test (20%). This completely eliminates actor, lighting, scenery, and acoustic leakage across splits.
Key Experimental Results¶
Main Results¶
All benchmarks are evaluated under consistent optimization protocols and 70/10/20 source-level splits across five representative architectures: MISA, EmotionCLIP, TCL-MAP, EMOE, and MulT. The metrics comprise Weighted F1 (W-F1), Macro-F1, and the class imbalance gap \(\Delta = \text{W-F1} - \text{Macro-F1}\) (where lower \(\Delta\) signifies fairer, unskewed predictions across all classes).
| Methods | Dataset | Surprise | Joy | Fear | Disgust | Anger | Neutral | Sadness | W-F1 (%) ↑ | Macro-F1 (%) ↑ | \(\Delta\) (Gap ↓) |
|---|---|---|---|---|---|---|---|---|---|---|---|
| MISA [ACM MM'20] | MELD | 13.30 | 18.93 | 0.74 | 0.00 | 23.52 | 37.95 | 14.33 | 25.09 | 18.61 | 6.48 |
| MISA | CAER | 18.72 | 58.08 | 0.00 | 15.71 | 17.80 | 49.33 | 21.92 | 32.82 | 25.94 | 6.88 |
| MISA | SpEmoC (Ours) | 41.30 | 66.40 | 36.00 | 49.80 | 50.90 | 32.40 | 51.70 | 50.78 | 47.50 | 3.28 |
| EmotionCLIP [CVPR'23] | MELD | 26.28 | 39.72 | 0.00 | 0.00 | 28.44 | 58.63 | 17.21 | 34.59 | 24.33 | 10.26 |
| EmotionCLIP | CAER | 11.94 | 34.88 | 0.00 | 0.00 | 21.37 | 46.52 | 13.76 | 27.45 | 18.35 | 9.10 |
| EmotionCLIP | SpEmoC (Ours) | 52.63 | 66.93 | 50.49 | 55.82 | 51.60 | 31.76 | 48.47 | 51.30 | 49.50 | 1.80 |
| TCL-MAP [AAAI'24] | MELD | 56.04 | 56.44 | 17.07 | 22.22 | 46.05 | 77.75 | 33.93 | 62.68 | 45.46 | 17.22 |
| TCL-MAP | CAER | 13.24 | 28.73 | 10.85 | 7.51 | 23.37 | 44.79 | 12.98 | 28.26 | 20.99 | 7.27 |
| TCL-MAP | SpEmoC (Ours) | 74.36 | 75.55 | 71.13 | 70.84 | 73.76 | 75.51 | 72.25 | 73.67 | 73.34 | 0.33 |
| EMOE [CVPR'25] | MELD | 54.59 | 54.55 | 6.06 | 6.38 | 43.95 | 73.86 | 27.92 | 59.04 | 38.19 | 20.85 |
| EMOE | CAER | 0.00 | 28.66 | 0.00 | 0.00 | 24.34 | 51.95 | 3.82 | 36.52 | 15.54 | 20.98 |
| EMOE | SpEmoC (Ours) | 80.94 | 84.73 | 66.42 | 66.78 | 65.85 | 61.01 | 66.05 | 70.30 | 70.25 | 0.05 |
| MulT [ACL'19] | MELD | 49.74 | 55.42 | 7.27 | 9.09 | 42.46 | 74.08 | 32.42 | 58.85 | 38.64 | 20.21 |
| MulT | CAER | 9.84 | 22.99 | 0.00 | 0.00 | 4.02 | 52.16 | 0.00 | 36.03 | 12.72 | 23.31 |
| MulT | SpEmoC (Ours) | 47.10 | 75.26 | 40.44 | 60.06 | 60.05 | 35.13 | 60.47 | 53.37 | 52.72 | 0.65 |
Ablation & Transfer Studies¶
The tables below show direct cross-dataset generalization (evaluating TCL-MAP without adaptation) and transfer learning gains using SpEmoC pretraining under full and low-resource downstream fine-tuning (using EMOE).
Table 1: Cross-dataset generalization across benchmarks (TCL-MAP)
| Train \(\rightarrow\) Test | Surprise | Joy | Fear | Disgust | Anger | Neutral | Sadness | W-F1 (%) ↑ | Macro-F1 (%) ↑ | \(\Delta\) (Gap ↓) |
|---|---|---|---|---|---|---|---|---|---|---|
| CAER \(\rightarrow\) CAER | 13.24 | 28.73 | 10.85 | 7.51 | 23.37 | 44.79 | 12.98 | 28.26 | 20.21 | 8.05 |
| CAER \(\rightarrow\) MELD | 19.46 | 17.44 | 5.13 | 1.14 | 16.09 | 50.71 | 13.56 | 34.62 | 17.65 | 16.97 |
| CAER \(\rightarrow\) SpEmoC | 21.14 | 34.79 | 3.49 | 5.47 | 40.94 | 25.89 | 17.07 | 24.64 | 19.83 | 4.81 |
| MELD \(\rightarrow\) CAER | 53.27 | 57.32 | 15.58 | 17.82 | 44.59 | 79.91 | 34.06 | 64.21 | 43.22 | 20.99 |
| MELD \(\rightarrow\) MELD | 56.04 | 56.44 | 17.07 | 22.22 | 46.05 | 77.75 | 33.93 | 62.68 | 45.46 | 17.22 |
| MELD \(\rightarrow\) SpEmoC | 44.83 | 53.47 | 3.92 | 18.21 | 51.28 | 34.91 | 30.77 | 36.00 | 32.17 | 3.83 |
| SpEmoC \(\rightarrow\) CAER | 21.31 | 17.92 | 10.78 | 13.30 | 20.30 | 46.69 | 11.36 | 30.75 | 20.82 | 9.93 |
| SpEmoC \(\rightarrow\) MELD | 45.19 | 29.37 | 3.96 | 16.72 | 39.71 | 69.59 | 25.24 | 50.73 | 32.83 | 17.90 |
| SpEmoC \(\rightarrow\) SpEmoC | 75.51 | 74.36 | 71.13 | 72.25 | 75.55 | 70.84 | 73.76 | 73.67 | 73.34 | 0.33 |
Table 2: Transfer learning improvements and low-data regime adaptation (EMOE)
| Setup & Transfer Configuration | Surprise | Joy | Fear | Disgust | Anger | Neutral | Sadness | Macro-F1 ↑ | W-F1 (%) ↑ | Gain \(\Delta\) |
|---|---|---|---|---|---|---|---|---|---|---|
| MELD Baseline Training | 54.59 | 54.55 | 6.06 | 6.38 | 43.95 | 73.86 | 27.92 | 38.19 | 59.04 | Baseline |
| MELD Fine-tuning (+SpEmoC Pretraining) | 58.22 | 55.20 | 16.33 | 17.65 | 41.68 | 78.39 | 21.90 | 41.34 | 63.26 | +4.22 |
| CAER Baseline Training | 0.00 | 28.66 | 0.00 | 0.00 | 24.34 | 51.95 | 3.82 | 15.64 | 36.52 | Baseline |
| CAER Fine-tuning (+SpEmoC Pretraining) | 13.92 | 26.98 | 0.00 | 2.30 | 23.35 | 52.47 | 10.70 | 19.81 | 38.16 | +1.64 |
| 10% MELD Data (Baseline / +SpEmoC) | - | - | - | - | - | - | - | - | 57.13 / 59.16 | +2.03 |
| 10% CAER Data (Baseline / +SpEmoC) | - | - | - | - | - | - | - | - | 31.40 / 34.70 | +3.30 |
Key Findings¶
- Elimination of majority bias: On MELD and CAER, model predictions exhibit extreme imbalance gaps (\(\Delta = 15\) to 23 points) due to Neutral dominance. On SpEmoC, \(\Delta\) consistently shrinks to 0.05 - 3.28 points across all architectures (e.g., EMOE attains \(\Delta = 0.05\)).
- Dramatic recovery of minority emotions: Fear and Disgust, which suffer near-zero F1 on legacy datasets, achieve robust F1 scores of 36.0% to 71.1% on SpEmoC (e.g., TCL-MAP scores 71.13% on Fear and 70.84% on Disgust).
- Exceptional low-resource transfer: Initializing models with SpEmoC pretraining produces notable improvements when downstream training data is constrained to 10%, boosting W-F1 on MELD by +2.03% and CAER by +3.30%.
Highlights & Insights¶
- KL divergence regularized logit fusion: Penalizing inter-modal divergence between text and audio logit distributions prevents pseudo-label noise far more effectively than hard argmax voting.
- Strict movie-level dataset isolation: Prevents actor, scene, lighting, and acoustic background leakage across splits, establishing a rigorous standard for true in-the-wild generalization benchmarking.
- Reverse pruning of unexpressive neutral clips: Actively discarding weakly emotional segments forces models to prioritize discriminative affective representations rather than collapsing into trivial majority-class predictions.
Limitations & Future Work¶
- Single discrete categorical labeling: Restricting annotations to Ekman's seven discrete classes fails to capture mixed emotions, continuous valence-arousal shifts, and subtle conversational subtexts like irony or sarcasm.
- Threshold-based exclusion of non-frontal poses: The 90% facial visibility requirement discards natural segments where emotion is communicated through body posture, head turns, or acoustic tone alone.
- Linguistic and cultural scope: The dataset is derived exclusively from English-language movies, limiting its applicability to cross-lingual and culturally diverse conversational domains.
Related Work & Insights¶
- vs MELD [ACL'19]: MELD comprises 13k utterances exclusively from Friends, with Neutral dominating 47% of samples and recurrent characters leaking across splits; SpEmoC expands to 30k clips from 3,100 movies, balances all seven classes, and enforces strict movie-level partitioning.
- vs CAER [ICCV'19]: CAER lacks aligned text transcripts and exhibits extreme class skew with near-zero recognition on minority classes; SpEmoC incorporates synchronized tri-modal alignment and provides balanced supervision.
- vs EmotionCLIP [CVPR'23]: While EmotionCLIP explores contrastive pretraining, SpEmoC demonstrates that a clean, balanced, and leak-free dataset design enables standard architectures to outperform complex contrastive formulations.
Rating¶
- Novelty: ⭐⭐⭐⭐☆ [Introduces the first balanced large-scale speaking-segment benchmark with divergence-regularized multi-stage filtering]
- Experimental Thoroughness: ⭐⭐⭐⭐⭐ [Extensive evaluations across five architectures, three benchmarks, transfer setups, and low-data regimes]
- Writing Quality: ⭐⭐⭐⭐⭐ [Clear motivation, rigorous mathematical formulation, and thorough ablation reporting]
- Value: ⭐⭐⭐⭐⭐ [Resolves the persistent class-imbalance bottleneck in MER, providing a pivotal foundation for empathetic conversational AI]