Skip to content

MulTaBench: Benchmarking Multimodal Tabular Learning with Text and Image

Conference: NeurIPS2026 Spotlight (from the supplied list metadata)
arXiv: 2605.10616
Paper: https://arxiv.org/abs/2605.10616
Code: https://github.com/alanarazi7/MulTaBench
Area: Multimodal VLM
Keywords: multimodal tabular learning, target-aware representations, benchmark curation, frozen embeddings, tabular foundation models

TL;DR

MulTaBench constructs a benchmark of 20 image-tabular and 20 text-tabular tasks by requiring both complementary multimodal signal and gains from target-aware representations over frozen embeddings, finding that adaptation gains generalize to new learners but are not consistently significant on every selected dataset.

Background & Motivation

Tabular foundation models can handle numerical and categorical columns, but generally cannot directly consume images or long text. A common workaround extracts fixed vectors with a pretrained encoder and supplies them as additional feature columns to TabPFN or tree models. This preserves a strong tabular learner but compresses the unstructured input into a summary independent of the prediction target: the same image may need to preserve different information for detecting pathology, estimating age, or identifying the scan type.

Existing multimodal tabular benchmarks often start by asking whether text or images are present. They do not systematically distinguish modality availability from additional predictive signal, or sufficient generic embeddings from representations that must be reorganized for the target. If structured columns already explain almost all labels, or a frozen encoder already retains the necessary information, jointly adapting a model may offer little advantage. Stronger encoders and lightweight adaptation now make these distinctions empirically measurable rather than purely conceptual.

The paper therefore does not introduce a new fusion network. It changes how evaluation tasks are assembled: different learners compete under controlled input conditions, and only tasks requiring both modality complementarity and representation adaptation are retained. Core idea: use a reproducible two-part filter to isolate multimodal tabular problems where frozen embeddings lose target-relevant information, rather than combine every dataset with multiple input types into a general leaderboard.

Method

Overall Architecture

MulTaBench is a benchmark and curation protocol, not an end-to-end model. Each candidate task pairs numerical/categorical columns with text or images, and the output remains a row-level classification label or regression target. The authors compare four conditions on the same task: structured only, unstructured only, joint frozen representations, and joint target-aware representations. Agreement across learners determines inclusion.

The protocol distinguishes whether frozen representations supplement structured columns, whether structured columns supplement unstructured inputs, and whether label-adapted representations further improve joint prediction. These checks respectively exclude tasks that collapse into purely tabular prediction, purely text/visual prediction, or simple fusion requiring only generic embeddings. Since the contribution is task curation and analysis, the experimental conditions are not drawn as a network architecture.

Target-Aware Representations (TAR) provide a measurable adaptation probe. The encoder is first adapted using only unstructured inputs and training labels. Its outputs are then concatenated with structured columns for a downstream tabular learner. Structured columns do not participate in encoder adaptation, and there is no joint back-propagation from the downstream learner.

Key Designs

1. Complementary signal test: joint prediction must beat the strongest unimodal baseline

A table containing photographs, descriptions, and prices does not establish a need for multimodal prediction. A description might merely repeat an existing category column, or structured columns might already reveal the target. Holding the unstructured representation fixed, the authors compare the first three conditions: structured only retains numerical and categorical columns; unstructured only retains vectors from the frozen encoder; joint frozen uses both. The same downstream learner is used across conditions so that changing learners does not get mistaken for a fusion gain.

Joint Signal does not mean outperforming an arbitrarily weak baseline. Joint prediction must outperform the stronger of the two unimodal conditions. This establishes that removing either modality hurts predictive performance, without requiring evidence for a particular cross-modal interaction mechanism; complementary additive signal can also qualify. With \(S_m\) denoting learner \(m\)'s mean AUC or \(R^2\) across runs, the paper defines joint gain as:

\[ \Delta_{\text{Joint}}(m)=S_m(\text{Joint Frozen})-\max\big(S_m(\text{UnimodalStructured}),S_m(\text{UnimodalUnstructured})\big). \]

This distinction matters when images already predict the label independently and the structured columns are irrelevant metadata. A high joint image-tabular score alone would not make that dataset a representative multimodal tabular task.

2. Target-aware representation probe: label adaptation exposes blind spots in frozen compression

Text uses e5-small-v2 and images use DINO-v3-small; the small encoders produce 384-dimensional representations. Frozen conditions directly extract generic embeddings. TAR adds LoRA to the encoder's last 3 layers and trains a linear classification head on the prediction target. Adaptation takes place only within the training split, with a 90/10 training/validation split for checkpoint selection; the test set does not participate in representation learning. This preprocessing is shared across downstream learners to isolate the effect of replacing representations.

Regression targets are not used for direct continuous regression during encoder adaptation. Instead, the authors divide training targets into 20 equal-frequency bins and adapt the encoder with cross-entropy classification over the bins, reducing instability from outliers. The final tabular learner still predicts the original regression target. Label discretization is therefore a TAR preprocessing strategy, not a conversion of the benchmark's regression evaluation into classification.

Text may occupy multiple columns, so a single e5 encoder is shared for efficiency. Each row-column pair becomes a training example consisting of the column name and value, supervised by that row's target. The encoder adapts jointly across all text columns. This can capture column context but may introduce interference between fields; it is not equivalent to assigning each column its own specialized encoder.

Both frozen and TAR representations are reduced to 30 dimensions using PCA by default, then supplied as features to the tabular learner. Joint TAR retains the same structured columns and replaces only the unstructured representation. Target-awareness gain is defined as:

\[ \Delta_{\text{Awareness}}(m)=S_m(\text{Joint TAR})-S_m(\text{Joint Frozen}). \]

This is an algorithm-dependent operational definition. It measures the limitations of the chosen generic encoder and adaptation recipe on a task, rather than proving that every possible frozen encoder must fail. The authors' longer-term goal includes conditioning representations on both the target and other modalities, but the probe in this paper does not implement the latter.

3. Joint acceptance across learners: no single algorithm defines task value

The curation committee consists of LightGBM, CatBoost, TabM, TabPFNv2, and TabPFN-2.5, spanning trees, an MLP, and tabular foundation models. Each candidate is evaluated under 4 conditions with 5 learners and 5 random seeds. Runs are capped at 10,000 examples; the appendix more specifically describes a cap on training examples per fold. Classification uses AUC and regression uses \(R^2\).

Both requirements must hold for the same learner before it contributes a vote. Three models passing Joint Signal and a different three passing Task-awareness cannot be combined into acceptance. The default minimum gain is \(\delta=0.001\), with support required from at least 3/5 learners. Mean scores are rounded to three decimal places, making this threshold the smallest positive gain expressible by the rule.

\[ \text{Accept}(\mathcal{D})\iff\left|\left\{m\in\mathcal{M}:\Delta_{\text{Joint}}(m)\geq\delta\;\land\;\Delta_{\text{Awareness}}(m)\geq\delta\right\}\right|\geq\rho\cdot|\mathcal{M}|, \qquad \delta=0.001,\quad\rho=3/5. \]

Passing this rule means exceeding an empirical gain threshold with majority agreement. It is not equivalent to a separate statistical significance test for every candidate. The appendix later performs paired tests with a broader learner pool and finds three selected tasks without significant TAR gains, illustrating why the two decisions should not be conflated.

4. Layered task collection: separate target-aware challenges from ordinary fusion tasks

Text candidates come from 4 existing benchmarks, yielding 56 datasets after deduplication. Of these, 23 pass both criteria, an acceptance rate of 41%, and 20 are retained. Existing image-tabular studies yield 16 available, unique tasks, only 5 of which pass. Additional public datasets are curated to reach 20 image-tabular tasks. The 40 core tasks range from 400โ€“114,000 rows and 1โ€“245 structured features; experiments do not all train on these full dataset sizes.

Image datasets use one image per row. Multi-image records are reduced to a single image, and rows with missing or corrupt images are removed. Curation can also change targets or features, such as binning PetFinder age into 8 classes or removing structured columns that leak labels or overwhelmingly dominate the image signal. Released tasks must therefore be understood together with their preprocessing definitions, rather than treated as the original datasets' default tasks. Tables and images are uploaded to Kaggle with relative image paths, reducing external-link failures and ambiguity about the original processing pipeline.

The authors release 40 additional candidates satisfying Joint Signal, producing an extended collection of 80 tasks. These are not all tasks lacking target-awareness: some weakly pass or were omitted to balance the core collection. The extended collection lets future models test both gains on difficult adaptation tasks and whether performance regresses where frozen embeddings already suffice.

Trimodality uses stricter requirements: text and images must each add complementary signal; adapting either must improve the joint frozen condition; and adapting both must outperform adapting either alone. Among the core image subset's 9 datasets containing text, only PetFinder and Amazon Packages meet all requirements. The presence of three modality fields does not make all 9 a strict trimodal benchmark. Although the entire 80-task release includes 22 trimodal candidates, they have not all undergone the same curation.

Loss & Training

LoRA uses rank 16, scaling factor 32, and dropout 0.1. Training uses AdamW with weight decay 0.01 and batch size 256. e5 uses a learning rate of \(10^{-4}\) for up to 50 epochs; DINO uses 0.001 for up to 100 epochs. Training stops after 3 epochs without validation-loss improvement. Both classification labels and discretized regression labels supervise the encoder through cross-entropy.

The paper does not perform per-dataset hyperparameter optimization for encoders or downstream learners. Frozen/TAR comparisons test representation value under a fixed configuration, not optimally tuned rankings for every method. Cross-task summary figures use minโ€“max score normalization; per-dataset result tables report raw AUC or \(R^2\), with negative \(R^2\) clipped before averaging. These presentations should not be mixed.

Key Experimental Results

Main Results

The following selection from appendix Tables 11โ€“12 compares joint frozen and joint TAR per-dataset results. Image results average 12 learners and text results average 10 learners, each with 5 runs. These are not limited to the curation committee and are not comparisons between a new model and previous SOTA. Gain values reproduce the paper's reported column.

Dataset Metric Joint Frozen Joint TAR Reported gain Adjusted p-value
Mango Mass \(R^2\) 0.533 0.653 +0.120 <0.001
CheXpert AUC 0.762 0.803 +0.041 <0.001
CS:GO Skins AUC 0.871 0.870 โˆ’0.001 0.763
Jigsaw Toxicity AUC 0.806 0.926 +0.119 <0.001
Video Games Sales \(R^2\) 0.348 0.385 +0.036 0.004
Fake Job Postings AUC 0.916 0.918 +0.002 0.459

The displayed values contain rounding inconsistencies: Jigsaw Toxicity's 0.926โˆ’0.806 equals 0.120, but the gain column reports 0.119; Video Games Sales' displayed difference is 0.037, but the reported gain is 0.036. The columns are preserved rather than adjusted for consistency, and reported gains should not be described as exact differences of displayed values.

For each task, TAR and Frozen runs are paired by learner and fold for a one-sided paired t-test. All 40 tests receive Benjaminiโ€“Hochberg correction at significance level 0.05. Significant gains occur on 37/40 tasks: 18/20 image tasks and 19/20 text tasks. Painting Price, CS:GO Skins, and Fake Job Postings are the exceptions. This supports the overall trend, not a claim that every task improves.

Ablation Study

PetFinder's eight-class age prediction analysis compares frozen trimodal inputs with different adaptation combinations. The table retains all five learner rows from the relevant columns of main-paper Table 2. Values are AUC percentages, not cross-dataset normalized scores.

Learner S+I+T all frozen Image TAR only Text TAR only Image and text both TAR
LightGBM 81.1 82.8 84.2 85.7
CatBoost 83.2 83.9 85.2 86.4
TabM 84.2 84.8 86.3 87.0
TabPFNv2 83.9 84.5 86.3 87.1
TabPFN-2.5 84.9 85.3 87.3 88.0

S, I, and T denote structured columns, images, and text. Each TAR encoder is adapted using row labels; this does not mean the encoders are trained end-to-end together with structured columns. Adapting both performs best, indicating that their adaptation gains are not fully redundant on this task.

Curation-threshold sensitivity is reported in Table 15. The entries below count retained tasks among the core 40, analyzing benchmark composition rather than ablating a network component.

Minimum gain \(\delta\) Consensus 3/5 Consensus 4/5 Consensus 5/5
0.001 40 33 23
0.002 36 30 20
0.005 30 21 13
0.01 21 12 2
0.02 7 3 2

Key Findings

  • TAR gains extend to new learners but vary in strength. Run-level win rates across all tasks are 91.5% for CatBoost, 87.5% for RandomForest, and 65.0% for TabICLv2. These are neither AUC scores nor rankings of the strongest models.
  • Larger encoders do not replace target adaptation. Large variants have approximately 10 times more parameters and 1024-dimensional outputs, yet TAR still outperforms Frozen; the paper also reports Small TAR outperforming Large Frozen.
  • PCA is not the sole explanation. Gains occur with 15, 30, and 60 components. The no-PCA check covers only CatBoost, LightGBM, and 33 tasks, excluding datasets with more than 5 text features; it is not a full evaluation of every learner.
  • Committee replacement is analyzed only on the 56 text candidates. Across 252 five-model committees drawn from 10 learners, 17 tasks always pass and 16 always fail, but 4 have acceptance rates of 50.0%; borderline cases remain.

Highlights & Insights

  • Make multimodality empirically testable. Multiple input types alone no longer justify inclusion; removing either modality must actually reduce predictive performance. This avoids judging fusion models on tasks that do not need fusion.
  • Distinguish representation limitations from tabular learning limitations. Tree models also benefit when shared adapted embeddings replace frozen ones, so the bottleneck is not exclusively downstream neural fusion. Improving a tabular foundation model also requires checking whether information was lost before reaching it.
  • Use the core and extended collections together. Winning only on a subset selected for adaptation gains may reflect targeted optimization. Maintaining strong performance on the extended collection provides a stronger test of general multimodal tabular capability.

Limitations & Future Work

  • Curation depends on particular encoders, a LoRA probe, and learners, deliberately entangling task properties with algorithmic capabilities. Rankings of curation learners suffer from selection bias; MulTaBench is not an unbiased general SOTA leaderboard.
  • Representation adaptation is expensive. On one A100 40GB GPU with 8 CPU cores, LightGBM's median total runtime with the small image encoder rises from 141 to 287 seconds. With the small text encoder it rises from 223 to 2,417 seconds, and with the large text encoder from 1,284 to 10,867 seconds. Cross-validation and tuning multiply repeated adaptation costs.
  • Single-image reduction, target binning, and removal of dominant structured features shape task difficulty rather than reproducing an untouched deployment distribution. Medical tasks are benchmark predictions, not evidence of clinical generalization.
  • Autoregressive LLMs/VLMs are not evaluated. TIME and MultiModalPFN are also omitted because of code availability or interface issues. Attention visualizations support changes in spatial focus, but do not causally establish information retention.
  • Future work could condition image/text representations through in-context learning without frequent parameter updates, evaluating both core and extended collections under multiple splits. Stricter curation and a dedicated trimodal collection could further reduce borderline cases.
  • vs CARTE / TextTabBench: CARTE emphasizes short strings and high-cardinality categories, while TextTabBench requires signal from both text and structured columns. MulTaBench additionally requires gains from target adaptation, so leadership on older benchmarks does not guarantee leadership here.
  • vs TabSTAR / AutoGluon-Multimodal: These methods provide joint modeling or adaptation capabilities. MulTaBench supplies tasks for testing such capabilities rather than proposing another model. Its TAR preprocessing should not be conflated with these systems' end-to-end training.
  • vs ConTextTab / MultiModalPFN: Static representations can retain generic semantics without preserving details needed by the current target. A transferable lesson is to identify what the encoder summary loses before deciding whether more complex fusion architectures are worthwhile.

Rating

  • Novelty: 4/5. Adds target-awareness to dataset selection rather than emphasizing modality coexistence alone.
  • Experimental Thoroughness: 4/5. Covers learners, scale, dimensions, and thresholds, but omits some relevant systems and extensive tuning.
  • Writing Quality: 4/5. Concepts and operational criteria are clear, with rounding differences between some displayed scores and reported gains.
  • Value: 4/5. A targeted test for multimodal tabular representation research, not a replacement for general tabular leaderboards.