Skip to content

iSyncTab: Learning Cross-Modal Feature Sequencing for Image-Tabular Data via Neural Synchrony

Conference: ECCV 2026
Paper: ECCV 2026
Code: https://github.com/zadid6pretam/iSyncTab
Area: Multimodal VLM
Keywords: Neural Synchrony, Image-Tabular Multimodal Learning, Cross-Modal Feature Sequencing, NeuroAI, Memory-augmented Transformer

TL;DR

Addressing the optimization instability of self-attention caused by arbitrary feature permutations in image-tabular fusion, iSyncTab models feature ordering as a Column Permutation Problem via neural synchrony matching and couples it with a linear-complexity memory-augmented Transformer, achieving top classification accuracy across six benchmarks with favorable computational efficiency.

Background & Motivation

Joint multimodal learning combining visual and structured tabular data plays a vital role across critical domains such as clinical healthcare, autonomous perception, and scientific modeling. However, unlike natural images or text corpora that inherently possess canonical 2D spatial lattices or sequential temporal topologies, tabular columns exhibit no natural or intrinsic order. Prevailing multimodal fusion architectures (e.g., STiL, TIP, MMCL) flatten and concatenate extracted visual and tabular embeddings into a combined vector, subsequently feeding them directly into self-attention layers or cross-modal attention blocks. This standard pipeline forces models to discover pairwise dependencies across arbitrary feature permutations, obscuring cross-modal correspondences and exposing sequence-sensitive operators to high inter-feature dispersion.

The core tension lies in the sensitivity of sequence-aware attention operators to permutation order versus the total absence of structural sequence priors in concatenated multimodal features. Permutation-invariant formulations (such as Deep Sets) discard the inductive benefits that structured attribute ordering provides in reducing redundancy and exposing hierarchical dependencies. Meanwhile, traditional attribute ordering heuristics in tabular learning often rely on greedy single-variable metrics or computationally intractable combinatorial searches over full feature sets, failing to address the structural alignment and heterogeneous semantics across distinct modalities.

This paper's angle of attack is inspired by the Communication Through Coherence (CTC) hypothesis in neuroscience, which posits that distributed neuronal assemblies transfer information with high bandwidth and minimal cost when their rhythmic oscillations are aligned in both phase and energy. Core idea: formulate cross-modal feature sequencing as a Column Permutation Problem (CPP), align visual and tabular feature clusters via an energy-phase neural synchrony matrix solved with Hungarian bipartite matching, and feed the resulting ordered sequence into an Order-aware Memory-augmented Transformer (OMT) regularized by an auxiliary order-consistency objective.

Method

Overall Architecture

The end-to-end iSyncTab pipeline encompasses modality-specific feature extraction, Neural Synchrony-guided Paired Feature Sequencing (NS-PFS), and an Order-aware Memory-augmented Transformer (OMT) backbone. First, a ResNet-50 backbone extracts visual features while a specialized tabular Transformer encodes structured attributes. Next, NS-PFS clusters transposed feature matrices from both modalities, constructs an energy-phase synchrony matrix, and computes an optimal bipartite cluster assignment using the Hungarian algorithm. Inside each matched joint cluster, local dispersion is minimized to assemble a global permutation \(\pi_{\mathrm{NS-PFS}}\). Finally, the synchronized token sequence is prepended with learnable global memory readout tokens and passed into a linear-complexity Linformer encoder, trained jointly with cross-entropy classification and an auxiliary feature sequencing consistency loss.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Multimodal Inputs<br/>X_vis ∈ R^{N×d_vis}, X_tab ∈ R^{N×d_tab}"] --> B["Cross-Modal Feature Clustering & Energy<br/>Transposed clusters {I_a} and {T_b}"]
    B --> C["Neural Synchrony-guided Paired Feature Sequencing (NS-PFS)<br/>Energy-Phase synchrony matrix + Hungarian matching"]
    C --> D["Local Dispersion Minimization & Global Concatenation<br/>Solve local CPP for each J_{(r)} to yield π"]
    D --> E["Order-aware Memory-augmented Transformer (OMT)<br/>Prepend memory tokens T_mem + Linformer encoding"]
    E --> F["Joint Multitask Supervision<br/>Classification loss L_CE + Order-consistency loss L_FS"]

Key Designs

1. Neural Synchrony-guided Paired Feature Sequencing (NS-PFS): cross-modal bipartite cluster matching to curb feature dispersion To bypass the combinatorial \(\mathcal{O}(m!)\) complexity of globally ordering \(m = d_{\mathrm{vis}} + d_{\mathrm{tab}}\) features, NS-PFS clusters transposed sample matrices \(G_{\mathrm{vis}} = X_{\mathrm{vis}}^\top \in \mathbb{R}^{d_{\mathrm{vis}} \times N}\) and \(G_{\mathrm{tab}} = X_{\mathrm{tab}}^\top \in \mathbb{R}^{d_{\mathrm{tab}} \times N}\) into \(k_I\) visual clusters \(\{I_1, \ldots, I_{k_I}\}\) and \(k_T\) tabular clusters \(\{T_1, \ldots, T_{k_T}\}\). For each cluster \(C\), cumulative metric energy \(\xi_C\) under metric \(M(\cdot)\) (e.g., KL divergence, variance, or cosine distance) and centroid embedding \(\mu_C\) are defined as: $\(\xi_C = \sum_{f \in C} M(f), \quad \mu_C = \frac{1}{|C|} \sum_{f \in C} \phi(f)\)$ Drawing upon the neurobiological CTC theory, cross-modal synchrony \(\Psi_{ab} = E_{ab} S_{ab}\) is computed from energy balance \(E_{ab} = 1 - \frac{|\xi_{I_a} - \xi_{T_b}|}{\xi_{I_a} + \xi_{T_b} + \epsilon}\) and phase-analogue centroid similarity \(S_{ab} = \frac{\langle \mu_{I_a}, \mu_{T_b} \rangle}{\|\mu_{I_a}\| \|\mu_{T_b}\|}\). The Hungarian algorithm then identifies the optimal one-to-one assignment \(P^\star\) in \(\mathcal{O}(K^3)\) time, where \(K = \max(k_I, k_T) \le 15\), forming matched joint clusters \(J_{(a,b)} = I_a \cup T_b\). This design groups functionally correlated cross-modal features without combinatorial explosion.

2. Regime-aware modulation and local-to-global sequence assembly: hierarchical routing and long-range coordination Matched visual-tabular cluster pairs exhibit heterogeneous coupling strengths. NS-PFS introduces a four-state modulation scheme \(\rho_{ab} \in \{\mathrm{HH}, \mathrm{HL}, \mathrm{LH}, \mathrm{LL}\}\) parameterized by energy and synchrony cutoffs \((\tau_L, \tau_H)\), mirroring gamma-band binding (HH), feedforward/feedback predictive routing (HL/LH), and large-scale contextual coherence (LL) in cortical circuits. Under modulation weight \(\gamma(\rho_{ab}, \Psi_{ab})\), each joint cluster \(J_{(r)}\) undergoes local dispersion minimization over sparsified feature pairs \(\mathcal{E}_{J_{(r)}}\): $\(\pi^\star_{J_{(r)}} = \arg\min_{\pi} \sum_{(u,v) \in \mathcal{E}_{J_{(r)}}} \gamma(\rho_{b_r}, \Psi_{a_r b_r}) |M(f_u) - M(f_v)| |\pi(u) - \pi(v)|\)$ Ranked by normalized joint energy score \(\alpha_{J_{(r)}}\), local sequences are concatenated into the global permutation \(\pi_{\mathrm{NS-PFS}} = [\pi^\star_{J_{(1)}} \parallel \ldots \parallel \pi^\star_{J_{(R)}}]\), ensuring tightly coupled, highly informative cross-modal features are sequenced contiguously.

3. Order-aware Memory-augmented Transformer (OMT): global readout tokens with auxiliary order-consistency regularization The synchronized feature sequence is embedded into \(L = n_{\mathrm{tab}} + n_{\mathrm{vis}}\) tokens. To circumvent standard attention's quadratic \(\mathcal{O}(L^2)\) cost, OMT deploys Linformer linear-complexity attention with rank \(r_{\mathrm{LF}}\). A set of \(n_{\mathrm{mem}}\) learnable memory tokens \(C_{\mathrm{mem}} \in \mathbb{R}^{n_{\mathrm{mem}} \times d_h}\) is prepended as global information accumulators. After encoding, the pooled memory state \(h = \frac{1}{n_{\mathrm{mem}}} \sum_{\ell=1}^{n_{\mathrm{mem}}} H_{\mathrm{mem},\ell}\) feeds the classification head, while ordered data tokens \(H_\pi\) feed an auxiliary order prediction head matching target normalized ranks \(\beta_{\pi,\ell} = \frac{\ell - 1}{L - 1}\): $\(\mathcal{L}_{\mathrm{total}} = \mathcal{L}_{\mathrm{CE}}(y, \hat{\mathbf{y}}) + \lambda_{\mathrm{FS}} \frac{1}{BL} \sum_{i=1}^B \sum_{\ell=1}^L (q_{i\ell} - \beta_{\pi,\ell})^2\)$ The auxiliary loss \(\mathcal{L}_{\mathrm{FS}}\) anchors feature representations to their synchronized sequential structure throughout training, preventing latent representations from drifting into permutation disorder.

Key Experimental Results

Main Results

On six diverse image-tabular benchmarks (DVM, HAM10000, DeepLesion, Pokémon, CheXpert, PetFinder), iSyncTab is compared against tabular GBDTs, deep tabular models, image-only models, and multimodal fusion baselines:

Model Modality (I/T) DVM (%) HAM (%) DLes (%) Pok (%) CheX (%) Pet (%) Avg. Rank ↓ Avg. Regret ↓
LGBM T 26.89 71.80 76.71 36.10 52.25 39.70 8.67 ± 3.14 36.66 ± 21.12
CatBoost T 27.30 72.05 77.18 36.59 53.19 39.40 7.67 ± 2.69 35.35 ± 21.20
TabSeq T 26.66 70.14 83.46 15.61 51.10 37.33 11.08 ± 4.25 41.46 ± 25.81
ResNet50 I 32.07 82.13 69.17 31.22 48.35 35.52 10.25 ± 3.16 39.10 ± 21.92
ViT I 88.00 82.33 69.17 35.61 48.70 34.94 8.58 ± 3.83 29.05 ± 18.93
DAFT I+T 74.22 74.88 77.07 18.05 67.75 45.45 7.42 ± 2.71 29.27 ± 19.86
Interact Fuse I+T 78.58 84.87 70.68 11.71 50.90 52.44 8.33 ± 4.71 30.65 ± 22.37
MMCL I+T 85.79 70.09 67.34 38.05 49.00 29.72 10.67 ± 5.19 27.75 ± 16.03
TIP I+T 98.27 70.39 69.17 37.56 87.50 83.86 6.00 ± 4.07 10.08 ± 8.09
STiL I+T 99.27 78.48 81.35 27.32 88.60 87.68 4.42 ± 2.83 11.73 ± 20.49
iSyncTab (Ours) I+T 99.60 86.82 85.63 42.44 88.60 88.02 1.08 ± 0.19 0.00 ± 0.00

Ablation Study

Component-wise ablations on DVM and HAM10000 isolate the impact of sequencing, memory tokens, auxiliary loss, and visual backbone:

Config / Variant DVM (w/o tuning) DVM (w/ tuning) HAM (w/o tuning) HAM (w/ tuning) Note & Mechanism Observation
Full model (ResNet-50 + NS-PFS + OMT) 99.80 99.60 78.72 86.82 Combines neural synchrony sequencing and memory tokens
Variant w/ ResNeXt-50 backbone 99.10 97.90 78.14 84.28 Heavy visual backbone fails to improve fusion performance
w/o sequencing-loss term 97.80 98.70 78.06 84.36 Missing auxiliary order loss degrades HAM accuracy by 2.46%
w/o memory tokens 92.68 87.60 77.90 80.12 Omitting global memory readout drops tuned DVM by 12.00%
w/o feature sequencing 96.80 96.50 78.16 78.60 Standard unordered concatenation drops tuned HAM by 8.22%
w/ random column shuffle 86.72 82.31 62.06 62.16 Destroys local neighborhood structure; HAM collapses by 24.66%
Concat Fuse baseline 96.50 96.20 76.58 78.42 Naive fusion lacking both sequencing and memory mechanisms

Key Findings

  • Structural sequencing is an essential inductive bias: randomly shuffling feature columns causes accuracy on HAM to plummet from 86.82% to 62.16% (-24.66%), disproving the assumption that Transformers are inherently permutation-invariant across feature sets.
  • Memory tokens provide critical global aggregation: removing memory tokens drops DVM tuned performance from 99.60% to 87.60%, demonstrating that mean-pooling long multimodal sequences without dedicated readouts introduces severe representational dispersion.
  • Superior compute-memory Pareto efficiency: leveraging Linformer attention and sub-quadratic Hungarian clustering, iSyncTab reduces inference GPU memory by 30-40% compared to full self-attention baselines while dominating the accuracy-complexity Pareto frontier.

Highlights & Insights

  • Biologically inspired cross-modal binding: maps neural phase-locking values and oscillatory energy balance to feature cluster similarity and energy metrics, providing concrete theoretical grounding for multimodal feature alignment.
  • Bipartite cluster-level formulation: collapses an intractable feature permutation problem into a small Hungarian matching problem over representative clusters (\(k_I, k_T \le 15\)), achieving mathematical rigor without computational bottlenecks.
  • Broad cross-modal transferability: successfully generalizes beyond image-tabular data to audio-visual emotion recognition on RAVDESS (97.56% best accuracy), showing that neural synchrony sequencing is a universal multimodal inductive bias.

Limitations & Future Work

  • Fixed cluster granularity: hyperparameters such as cluster counts \(k_I, k_T\) and modulation thresholds \((\tau_L, \tau_H)\) require empirical tuning or Optuna search rather than dynamic data-driven discovery.
  • Offline discrete sequencing: the current Hungarian matching and CPP steps operate on pre-extracted feature clusters as a discrete preprocessing stage rather than a fully differentiable continuous relaxation.
  • Future explorations could integrate continuous relaxation operators (such as Gumbel-Sinkhorn) to realize end-to-end differentiable neural synchrony sequencing.
  • vs STiL & TIP: STiL uses semi-supervised hierarchical cross-attention and TIP uses tabular masking pretraining, both treating multimodal tokens as unordered collections. iSyncTab introduces structured neural synchrony sequencing, securing higher accuracy (Avg. Rank 1.08 vs 4.42 / 6.00).
  • vs COPER: COPER applies permutation-based correlation to unsupervised multi-view clustering, whereas iSyncTab designs an end-to-end framework specifically optimized for supervised multimodal classification.
  • vs TabSeq & Mambular: While TabSeq and Mambular confirmed the significance of column ordering in single-modality tabular learning, iSyncTab pioneers this concept across heterogeneous visual-tabular feature spaces.

Rating

  • Novelty: ⭐⭐⭐⭐⭐ Groundbreaking intersection of neuro-inspired synchrony (CTC) and combinatorial feature sequencing (CPP).
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Evaluated across six benchmarks, label noise corruption, efficiency Pareto frontiers, and audio-visual transfer.
  • Writing Quality: ⭐⭐⭐⭐⭐ Rigorous mathematical formulations, clear pipeline visual mappings, and exhaustive ablations.
  • Value: ⭐⭐⭐⭐⭐ Establishes structured feature sequencing as a foundational design choice for multimodal tabular-vision systems.