Skip to content

Mapping the Concept Landscape: Structural Perception of Global Distributions for Transparent Data Pruning

Conference: ECCV 2026
Paper: ECCV 2026
Area: Interpretability
Keywords: data pruning, concept graph, multimodal instruction tuning, interpretability, greedy marginal coverage

TL;DR

Addressing the core limitation of opaque embedding-based data pruning that obscures fine-grained semantic coverage and lacks human interpretability, this paper proposes Mapping the Concept Landscape (MCL), which decouples image-text samples into explicit concept graphs of entities, events, and attributes, aggregates them into a dataset-level semantic landscape, and applies a greedy marginal coverage maximization strategy for efficient and transparent pruning.

Background & Motivation

The proliferation of web-scale vision–language datasets, such as LAION and CC12M, has fueled the remarkable capabilities of modern vision-language models (VLMs). However, scaling up training corpora introduces prohibitive computational and storage overhead, while automated data scraping pipelines inevitably inject massive amounts of redundant, low-quality, and noisy samples. To enhance data efficiency, prior data pruning frameworks estimate sample importance through training dynamics, loss profiles, and uncertainty, or select representative subsets using geometric clustering and distance metrics within compressed feature embedding spaces. These approaches implicitly treat each sample as an indivisible unit compressed into a single dense vector.

This conventional paradigm leads to a fundamental Embedding–Concept Misalignment. Geometric proximity in continuous representation spaces does not reliably correspond to redundancy in underlying semantic concepts. Because multi-concept semantics are collapsed into coupled numerical vectors, samples that appear close or redundant in geometric space may still preserve distinct fine-grained, long-tail concepts; conversely, geometrically distant samples may heavily duplicate identical conceptual contents. As a result, embedding-based selection optimizes geometric dispersion rather than true semantic coverage, systematically pruning away low-frequency yet essential concepts. Furthermore, filtering data inside opaque embedding spaces denies practitioners any ability to inspect, audit, or reason about which concepts are retained or discarded, leaving the pruning pipeline as an untrustworthy black box.

This paper tackles the challenge by moving beyond monolithic vector representations toward explicit, fine-grained concept coverage. The core idea is to decouple each image-text pair into an interpretable concept graph spanning entities, events, and attributes, aggregate them to perceive the global dataset semantic landscape and quantify concept rarity alongside structural relational degrees, and iteratively select samples via greedy marginal coverage maximization to produce an efficient, anti-redundant, and transparently auditable subset.

Method

Overall Architecture

MCL shifts the pruning unit from monolithic numerical embeddings to explicit semantic concepts and tracks their global distribution across the entire dataset. The overall pipeline proceeds in three distinct phases: first, semantic dependency parsing extracts atomic concepts and co-occurrence relations from each image-text pair to form sample-level concept graphs; second, these sample graphs are merged into a unified dataset-level concept graph to compute category-normalized importance scores based on empirical frequency and topological degree; finally, data selection is formulated as a set-coverage maximization problem and solved via parallelized greedy selection guided by dynamic marginal coverage gain.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Image-Text Pairs (I_i, T_i)"] --> B["Sample-level Concept Graph Construction<br/>Extract entity/event/attribute concept nodes & dependency edges"]
    B --> C["Dataset-level Concept Graph & Structural Importance<br/>Aggregate global topology, quantify rarity & degree with normalization"]
    C --> D["Marginal Gain-driven Greedy Concept Coverage Pruning<br/>Dynamically evaluate uncovered concepts & greedily select samples in parallel"]
    D --> E["Pruned Core Subset with Transparent Semantic Landscape"]

Key Designs

1. Sample-level Concept Graph Construction: Decoupling Monolithic Samples into Fine-grained Tripartite Semantics

To eliminate semantic obscurity caused by coupled feature embeddings, MCL maps each image-text pair \(x_i = (I_i, T_i)\) into a transparent, structured concept graph \(G_i = (V_i, E_i)\). The textual description \(T_i\) can be either the original caption or synthetic captions generated by off-the-shelf vision models. Using spaCy, the framework applies dependency parsing and part-of-speech (POS) tagging to filter stop words and categorize semantic tokens into three atomic concept types: Entities (nouns) denoting visual objects and subjects, Events (verbs) capturing actions and interactions, and Attributes (adjectives and adverbs) characterizing properties and modifier states. Edges \(E_i\) preserve contextual dependencies and co-occurrences. This symbolic decomposition requires minimal computational overhead while exposing fine-grained semantic components for downstream aggregation and human inspection.

2. Dataset-level Concept Graph and Structural Importance: Integrating Global Rarity with Relational Connectivity

While individual concept extractions may carry minor linguistic noise, their relative frequencies and coverage patterns exhibit high stability when aggregated across the corpus. MCL merges all \(N\) sample graphs into a dataset-level concept graph \(G_D = (V_D, E_D)\) where \(V_D = \bigcup_{i=1}^N V_i\), assigning each node a sample frequency weight \(w(v) = \sum_{i=1}^N \mathbf{1}[v \in V_i]\). However, importance should not depend solely on rarity; an isolated low-frequency concept that never interacts with other semantics offers limited generalizability compared to a rare concept bridging diverse semantic contexts. To balance semantic rarity and structural participation, MCL defines the unnormalized concept importance score as:

\[\tilde{\phi}(v) = \log\left(\frac{N_{\text{type}(v)}}{w(v)} \cdot d(v)\right)\]

where \(N_{\text{type}(v)}\) denotes the total number of concepts in category \(\text{type}(v)\) and \(d(v)\) is the node degree in \(G_D\). The inverse-frequency factor prioritizes underrepresented concepts while the degree term introduces a lightweight relational boost. Because the total vocabulary counts across entities, events, and attributes differ substantially, Min-Max normalization is performed independently within each category \(V_{\text{type}(v)}\) to derive the final importance \(\phi(v)\), ensuring rare events and attributes compete fairly with abundant entities.

3. Marginal Gain-driven Greedy Concept Coverage Pruning: Dynamic Redundancy Elimination via Chunked Parallelism

Ranking samples based on static, independent importance sums inevitably selects clusters of redundant instances that duplicate identical high-value concepts. MCL formalizes dataset pruning as a constrained set-coverage maximization problem: under a sample budget \(b = (1-p) \cdot |D|\), the objective maximizes the total importance of distinct covered concepts \(F(\mathcal{S}) = \sum_{v \in V_{\mathcal{S}}} \phi(v)\), where \(V_{\mathcal{S}} = \bigcup_{x_i \in \mathcal{S}} V_i\). Under this monotone submodular formulation, once a concept is incorporated into \(\mathcal{S}\), subsequent samples sharing that concept provide zero additional gain, naturally preventing semantic duplication.

To solve this efficiently, MCL employs an iterative greedy selection process. At step \(t\), the marginal gain for any unselected candidate \(x_i \in D \setminus \mathcal{S}_t\) measures the value of previously uncovered concepts it brings:

\[\Delta(x_i \mid \mathcal{S}_t) = F(\mathcal{S}_t \cup \{x_i\}) - F(\mathcal{S}_t) = \sum_{v \in V_i \setminus V_{\mathcal{S}_t}} \phi(v)\]

The algorithm greedily adds \(x^* = \arg\max \Delta(x \mid \mathcal{S}_t)\) to the retained subset and updates the coverage set. To scale across web-scale multimodal corpora, the dataset is partitioned into multiple chunks (e.g., 8 chunks for LLaVA-1.5-mix-665k), and greedy selection runs concurrently across chunks using fast set-difference and scalar summation, completing selection on over 660K samples in approximately 1.7 hours.

Key Experimental Results

Main Results

The framework is evaluated on the multimodal vision-language instruction tuning benchmark LLaVA-1.5-mix-665k (Backbone: LLaVA-v1.5-7B with LoRA fine-tuning) across 8 evaluation suites, as well as the classical pure vision object detection benchmark COCO 2017 (Backbone: DETR-ResNet50).

On LLaVA-1.5-mix-665k under 50k (7.5%) and 133k (20%) data budgets:

Method Kept Data MME-P MME-C SEED-I POPE MMMU SQA GQA TextQA Rel.
Full Dataset 665k 1476.9 267.9 67.4 86.4 32.8 70.0 63.0 58.2 100.0%
Random 50k (7.5%) 1387.5 287.5 59.7 85.7 32.2 68.4 55.0 53.1 95.7%
EL2N 50k (7.5%) 1077.3 252.5 59.3 80.8 33.6 71.0 61.0 41.7 90.1%
InsTag 50k (7.5%) 1317.1 345.0 57.4 82.1 34.0 69.3 52.5 53.3 97.0%
LESS 50k (7.5%) 1344.8 281.8 61.2 79.4 33.0 71.0 53.4 52.0 94.4%
TIVE 50k (7.5%) 1434.8 291.5 61.6 84.9 33.3 71.2 56.3 52.0 97.2%
DataTailor 50k (7.5%) 1447.4 322.3 60.6 82.4 33.9 70.4 57.1 52.9 98.6%
MCL (Ours) 50k (7.5%) 1455.8 318.2 60.3 84.8 33.8 70.0 57.6 53.9 99.0%
COINCIDE 133k (20%) 1496.0 298.1 62.9 86.1 32.7 69.2 59.8 55.6 99.3%
ICONS 133k (20%) 1487.1 295.3 63.0 87.5 33.0 70.8 60.7 55.6 99.9%
MCL (Ours) 133k (20%) 1501.5 308.4 63.4 85.7 33.4 70.7 60.3 56.0 100.5%

On COCO 2017 object detection with 70% data selection ratio (DETR-ResNet50):

Method Ratio \(\text{AP}_{50:95}\) \(\text{AP}_{50}\) \(\text{AP}_{75}\) \(\text{AP}_{\text{small}}\) \(\text{AP}_{\text{mid}}\) \(\text{AP}_{\text{large}}\)
Full Dataset 100% 40.06 61.12 42.03 19.27 43.39 59.24
Random 70% 38.26 (-1.80) 58.91 39.77 17.83 41.19 56.69
EL2N 70% 34.35 (-5.71) 53.29 36.14 16.58 37.64 48.54
InfoBatch 70% 34.95 (-5.11) 53.86 36.55 17.00 38.74 49.05
DivBS 70% 39.74 (-0.32) 60.84 41.98 18.07 43.14 58.78
PFB 70% 39.69 (-0.37) 60.85 41.74 18.45 43.25 59.07
MCL (Ours) 70% 40.12 (+0.06) 61.06 42.27 18.95 43.78 59.13

Ablation Study

  1. Greedy Marginal Coverage vs. Static Ranking (COCO 2017 Detection AP):
Selection Ratio Static Ranking (\(\sum \phi(v)\)) MCL Greedy Marginal Coverage AP Margin
40% ~36.8% 38.1% +1.3%
50% ~37.9% 39.0% +1.1%
60% ~38.8% 39.6% +0.8%
70% ~39.4% 40.12% +0.72%
  1. Concept Coverage Ratio Comparison (LLaVA-1.5-mix-665k \(|V_{\mathcal{S}}| / |V_D|\)):
Selection Ratio Random Sampling Coverage MCL Greedy Coverage Absolute Coverage Gain
5% 32.83% 92.76% +59.93%
10% 46.31% 99.12% +52.81%
15% 56.42% 99.99% +43.57%
20% 64.05% 100.00% +35.95%
  1. Wall-clock Pruning and Training Time (4 \(\times\) NVIDIA RTX 3090 GPUs):
  2. Full Model Training: 100.0 hours training, 0 hours selection.
  3. TIVE: 100.0 hours pre-warmup/gradient computation + 8.0 hours selection + 7.5 hours training = 115.5 hours.
  4. DataTailor: 15.0 hours multimodal feature clustering + 7.5 hours training = 22.5 hours.
  5. MCL (Ours): 1.7 hours total selection (concept extraction + greedy selection) + 7.5 hours training = 9.2 hours, slashing total pipeline time by over 90%.

Key Findings

  • Concept Coverage Governs Downstream Generalization: At a strict 7.5% pruning budget, MCL already attains 92.76% semantic concept coverage, delivering 99.0% of the full-data baseline score across comprehensive benchmarks, and outperforming full-data training at 20% budget (100.5%).
  • Static Ranking Biases Toward Long-Tail Noise: Relying solely on static rarity scores oversamples obscure long-tail concepts while starving dominant structural categories needed for robust representation learning; dynamic marginal gain naturally balances dominant semantic foundations and tail coverage.
  • Task-Agnostic Versatility: Even on pure computer vision tasks like DETR object detection, concept graphs derived from paired descriptions effectively guide representative sample selection, surpassing full data performance at 70% retention.

Highlights & Insights

  • From Continuous Vector Diversity to Symbolic Concept Coverage: The paper uncovers the critical flaw of vector-based data pruning—embedding-concept misalignment—and demonstrates that discrete symbolic concept graphs provide a much more faithful and interpretable abstraction of dataset diversity.
  • Relational Degree as an Anti-Noise Anchor: Rather than naively up-weighting all rare terms, MCL factors in topological node degree in the global concept graph, effectively filtering out isolated noise while prioritizing structurally influential rare concepts.
  • Extreme Computational Efficiency: Bypassing expensive gradient calculations, model forward passes, and high-dimensional clustering, the algorithm relies on fast linguistic dependency parsing and parallel set updates, reducing pruning wall-clock time from days to 1.7 hours.

Limitations & Future Work

  • Dependency on Text Captions: The approach fundamentally requires textual descriptions. For uncaptioned pure vision datasets, an upfront vision-language model must synthesize captions, which introduces potential hallucination or error propagation.
  • Absence of Spatial Geometry Modeling: The graph structure reflects grammatical dependency rather than fine-grained spatial bounding boxes, object occlusion, or 3D scene geometry.
  • Future Directions: Integrating open-vocabulary object detectors or scene graph generation models to bridge 2D visual coordinate topology directly with linguistic concept graphs.
  • vs DataTailor / COINCIDE: DataTailor and COINCIDE operate in continuous feature embedding spaces prone to vector collapse; MCL models discrete semantic graphs to preserve fine-grained concept frontiers without representation compression loss.
  • vs TIVE / LESS: TIVE and LESS incur massive computation costs by tracking training dynamics and gradient influence; MCL operates completely offline without training models, cutting selection time by an order of magnitude.
  • vs EL2N / InfoBatch: Loss-based pruning is susceptible to noisy labels and optimization instability; MCL relies on global corpus-level concept distributions, yielding superior stability and cross-task generalization.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Conceptually shifts multimodal dataset pruning from opaque embedding distances to interpretable concept graph coverage.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous evaluations on multimodal instruction tuning and pure vision object detection, supported by detailed runtime benchmarks, coverage curves, and graph redistributions.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear logical progression, persuasive motivation regarding embedding-concept misalignment, and well-designed visualizations.
  • Value: ⭐⭐⭐⭐⭐ High practical utility for accelerating vision-language pretraining and instruction tuning while providing human-interpretable data curation audits.