Skip to content

Causal Intervention in Concept Bottleneck Models

Conference: ECCV 2026
Paper: ECCV 2026 Official Page
Code: https://github.com/LMBTough/CI
Area: Interpretability
Keywords: Concept Bottleneck Models (CBMs), Causal Intervention, Explainable AI, Human-in-the-Loop Interaction, Concept Realignment

TL;DR

Proposes Causal Intervention (CI), a training-free intervention approach that leverages concept-correlation structures already encoded inside trained Concept Bottleneck Models by treating the input image as a shared upstream carrier to propagate human feedback through the \(x \to c\) pathway, achieving state-of-the-art performance with minimal intervention effort.

Background & Motivation

Concept Bottleneck Models (CBMs) construct an explicit intermediate bottleneck layer composed of human-understandable concepts between input features and final classification decisions. This architecture offers transparent reasoning paths and provides an intuitive interface for human-in-the-loop intervention: whenever the model displays high uncertainty on specific concepts, human domain experts can supply ground-truth values to rectify downstream predictions. However, in high-stakes domains such as clinical diagnosis or automated driving, each human intervention step incurs real-world cognitive effort and manual labeling costs. Maximizing predictive performance gains under a strict budget of minimal intervention rounds is therefore a central challenge in interactive interpretability.

Early concept intervention schemes (e.g., UCP, CCTP, ECTP) evaluate and update individual concepts independently, failing to capture the strong statistical dependencies that naturally bind real-world concepts together (for instance, "gray abdomen" and "gray plumage" strongly co-occur on bird species). Recent work on Concept Realignment (CR) attempted to address this gap, but relied on training an auxiliary neural network to predict adjustments for remaining concepts based on edited concept vectors. This separate model introduces additional architectural complexity, optimization instability, and hyperparameter tuning overhead. More fundamentally, because it operates solely on the low-dimensional concept bottleneck, it completely overlooks the rich concept-correlation representations already learned across the upstream encoder network.

Empirical analysis reveals that the parameter weights of a trained CBM intrinsically encode usable concept-correlation structures. In the CBM computational graph, concepts do not directly cause one another; rather, their correlations are mediated by the shared upstream input representation acting as a common cause. Core idea: eliminate the need for extra realignment models by treating the original input as a shared upstream carrier, using the trained \(x \to c\) pathway to back-propagate human feedback gradients into input space and adaptively realign non-intervened concepts.

Method

Overall Architecture

Causal Intervention (CI) operates during inference as a plug-and-play, training-free interaction mechanism. Given an input sample \(x\), an encoder \(f\) produces predicted concept representations \(\hat{c} = f(x)\), followed by a concept-based classifier \(g\) that outputs the label prediction \(\hat{y}\). At each interaction step, an intervention policy \(\pi\) selects the concept index requiring correction, and human feedback provides the true value. Instead of merely overwriting the selected concept dimension in isolation, CI propagates the concept loss backwards to compute sign gradients with respect to the input \(x\). A small, bounded perturbation updates \(x\), driving linked non-intervened concepts toward consistent states through forward pass propagation, followed by ground-truth clamping on the intervened set.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Input Sample x"] --> B["Concept Encoder f(x)"]
    B --> C["Concept Prediction & Policy Selection<br/>Identify concept to intervene & query ground-truth"]
    C --> D["Input Carrier Gradient Propagation<br/>Backpropagate sign gradient to update x"]
    D --> E["Cross-Concept Cooperative Realignment<br/>Forward propagate updated x to realign remaining concepts"]
    E --> F["Ground-Truth Clamping & Final Prediction<br/>Clamp intervened concepts and classify via g(c)"]

Key Designs

1. Verification of Inherent Concept Correlations: Validating encoded structure in trained weights Auxiliary realignment modules operate under the assumption that the main network cannot effectively propagate inter-concept dependencies. This paper disproves that assumption empirically. On the CUB dataset, the correlation matrix \(M_1 \in \mathbb{R}^{n \times k}\) derived from ground-truth concept annotations was compared with matrix \(M_2 \in \mathbb{R}^{m \times k}\) formed by the learned linear weights connecting backbone features to the concept space. The signed Spearman rank correlation between \(M_2\) and \(M_1\) reaches 0.849 (absolute Spearman 0.635) with a top-50 correlation pair overlap of 0.86, whereas random control weights show correlations near zero. This confirms that the trained backbone parameters already faithfully mirror the empirical concept co-occurrence matrix, providing a sound basis for intrinsic realignment.

2. Input Carrier Gradient Propagation: Leveraging input representations as the common cause Within the CBM computational graph, individual concepts lack lateral edges connecting one another; their mutual dependencies stem entirely from sharing the upstream input carrier \(x\) (analogous to how student effort serves as the common cause for both exam performance and homework quality). When policy \(\pi\) selects concept \(i\) at round \(t\) and human feedback provides ground truth \(c_i\), CI formulates a first-order Taylor expansion on the encoder and constructs a bounded update in input space via the sign gradient: $\(\Delta x = -\eta \cdot \operatorname{sign}\left(\sum_{j \in S_t} \nabla_x \mathcal{L}(\hat{c}_j(x), c_j)\right)\)$ where \(S_t\) denotes the cumulative set of intervened concepts up to step \(t\), and \(\mathcal{L}\) represents binary cross-entropy loss. Through the chain rule, updating \(x\) produces an effect on any non-intervened concept \(j\) proportional to \(\operatorname{sign}(\nabla_x \hat{c}_i) \cdot \nabla_x \hat{c}_j\), thereby automatically aligning positively correlated concepts without separate parameter training.

3. Perturbation Constraints and Ground-Truth Clamping: Preserving semantics and guaranteeing convergence To prevent input updates from degenerating into unconstrained adversarial artifacts, CI bounds the step size \(\eta\) (typically 0.03125 to 0.25) under an \(L_\infty\) norm restriction. After \(T\) intervention rounds, the total deviation satisfies \(\|x_T - x\|_\infty \le T \cdot \eta\). Empirically on standard CBMs, the perturbation remains minimal, yielding an \(L_\infty\) distance of roughly 0.002 and an image Structural Similarity Index (SSIM) of 0.999. Prior to computing the final prediction through \(g\), ground-truth values are strictly clamped onto all queried concept indices (\(\hat{c}_j \leftarrow c_j, \forall j \in S\)). This guarantees that expert feedback is preserved with 100% fidelity while unvisited concepts benefit from data-driven cooperative realignment.

Loss & Training

The framework is completely training-free during intervention. The underlying neural network parameters remain frozen. At test time, CI executes the following iterative update: 1. Initialize the intervened set \(S \leftarrow \emptyset\); 2. At each round, policy \(\pi\) queries concept \(i\), appends it to \(S\), and obtains human label \(c_i\); 3. Compute the accumulated concept loss gradient with respect to \(x\): \(\nabla_x \sum_{j \in S} \mathcal{L}(\hat{c}_j(x), c_j)\), and update \(x \leftarrow x - \eta \cdot \operatorname{sign}(\nabla_x \mathcal{L})\); 4. After completing \(T\) interactions, run a forward pass \(\hat{c} = f(x)\), enforce \(\hat{c}_j \leftarrow c_j\) for all \(j \in S\), and output classification logits \(\hat{y} = g(\hat{c})\).

Key Experimental Results

Main Results

Experiments were conducted on Animals with Attributes 2 (AwA2, 85 attributes) and Caltech-UCSD Birds-200-2011 (CUB, 112 concepts across 28 groups) using three representative architectures: CBM, Concept Embedding Models (CEM), and Intervention-aware CEM (Int-CEM). Cumulative performance across interaction rounds is quantified via Intervention-AUC using trapezoidal integration.

The table below summarizes the early-stage intervention performance at step 10 (AUC@10):

Method AwA2 (CEM) AwA2 (Int-CEM) AwA2 (CBM) CUB-Ind (CEM) CUB-Ind (Int-CEM) CUB-Ind (CBM) CUB-Grp (CEM) CUB-Grp (Int-CEM) CUB-Grp (CBM)
CCTP 791.35 782.35 820.06 719.12 729.91 618.45 732.02 750.88 636.50
COOP 789.00 777.62 815.62 714.37 707.21 610.29 738.61 775.13 655.75
ECTP 826.87 842.91 836.08 757.68 788.23 679.32 763.54 801.87 706.31
EUDTP 814.81 787.70 819.58 731.64 751.59 616.15 742.64 772.47 626.49
Random 799.87 794.80 820.06 720.49 726.21 622.18 733.21 762.12 643.16
UCP 828.70 821.79 832.60 750.16 781.11 665.16 770.16 811.45 697.75
CR (ECCV24) 838.76 842.63 840.37 757.10 781.15 682.99 784.10 815.07 729.09
CI (Ours) 863.93 844.96 859.28 801.91 814.01 743.31 811.87 836.43 761.63

Across overall intervention rounds, CI maintains consistent superiority: on CUB-Individual, CI scores 10661.75 (CEM), 11087.43 (Int-CEM), and 10500.56 (CBM), outperforming competing baselines by an average margin of 5.07%.

Ablation Study & Computational Efficiency

To examine computational efficiency and resource footprints, models were benchmarked on Energy-based CBM (E-CBM, evaluated on MNIST-Add) across 33,000 samples and 32 interaction rounds on an NVIDIA L40S GPU:

Method / Metric Overall AUC Early AUC(10) Total Time (s) Memory (MB) Characteristic Note
CCTP 1916.3 583.8 15.6 864 Gradient-based concept attribution
COOP 1915.6 582.2 43.1 864 Joint cooperative scoring; highest runtime
ECTP 1933.6 600.4 13.2 864 Expected KL divergence change
EUDTP 1885.0 586.3 14.4 1296 Entropy reduction policy
Random 1916.6 587.8 25.6 1422 Random selection; higher memory usage
UCP 1942.6 601.8 14.3 1296 Uncertainty heuristic baseline
CI (Ours) 1957.8 605.1 28.7 1022 Training-free, best AUCs with moderate memory

Key Findings

  • Rapid Early Convergence: CI reaches target accuracy plateaus within roughly 20 intervention steps, whereas baseline methods typically require exhausting nearly all concept annotations (85 or 112 steps). CI delivers an 8.59% average improvement over all baselines and a 3.90% gain over CR within the first 10 steps.
  • Robustness to Step Size: Testing learning rates from 0.03125 to 0.25 demonstrates modest sensitivity, with total AUC varying within less than 1% across most settings, eliminating tedious hyperparameter tuning.
  • Synergy with Uncertainty Policies: Pairing CI with Uncertainty of Concept Predictions (UCP) yields peak performance on individual attribute tasks, while EUDTP serves as a viable lightweight candidate under grouped concept settings.

Highlights & Insights

  • From External Patching to Intrinsic Re-use: Demonstrates that standard CBM parameters inherently capture semantic co-occurrence matrices, dispensing with the need for separate, error-prone realignment networks.
  • Upstream Common-Cause Formulation: Conceptualizes the input \(x\) as the shared causal mediator across concepts; updating \(x\) naturally propagates inter-concept corrections with minimal mathematical overhead.
  • Practical High-Stakes Utility: Delivers substantial accuracy improvements in the critical early intervention regime (first 10–20 steps) while maintaining lightweight GPU memory usage and zero architectural modifications.

Limitations & Future Work

  • Differentiability Requirement: CI requires computing backpropagation gradients from the concept layer back to the input image. It cannot be applied directly to non-differentiable hardware pipelines or black-box API setups.
  • Cumulative Visual Degradation: Under prolonged individual interventions (>100 steps), accumulated perturbations can slightly degrade input fidelity (e.g., individual CEM SSIM declining to 0.934), suggesting future research into adaptive momentum clipping or projected gradient bounds.
  • Absence of SCM Identification: The method exploits empirical statistical dependencies learned by neural weights rather than structural causal identifiability in the strict Judea Pearl SCM sense.
  • vs Concept Realignment (CR, ECCV 2024): CR trains an additional neural network solely over the concept bottleneck layer, adding training instability; CI is entirely training-free, propagates feedback through the full \(x \to c\) pathway, and improves AUC@10 by 3.90%.
  • vs Concept Selection Policies (UCP, ECTP, EUDTP, CCTP): Existing policies solely answer "which concept to query" and leave remaining concepts unchanged; CI focuses on "how to propagate feedback across concepts after query," forming an orthogonal, complementary mechanism.

Rating

  • Novelty: ⭐⭐⭐⭐☆ Repositions concept realignment as an input-mediated causal propagation problem, bypassing external models.
  • Experimental Thoroughness: ⭐⭐⭐⭐⭐ Rigorous coverage across 2 benchmarks, 3 core CBM architectures, E-CBM, and 7 baseline policies with latency/memory profiling.
  • Writing Quality: ⭐⭐⭐⭐⭐ Clear motivational narrative, tight mapping between intuition and gradient equations, and precise empirical evaluation.
  • Value: ⭐⭐⭐⭐⭐ Delivers an efficient, plug-and-play solution for human-in-the-loop decision support in high-stakes interpretable AI.