Interpretable but Fragile? Robustness of Concept Bottlenecks under Geometric-Semantic Perturbations¶
Conference: NeurIPS 2026 (acceptance information supplied in task metadata)
arXiv: 2609.38625v1
Area: Interpretability
Keywords: concept bottleneck models, generator-based evaluation, semantic perturbations, randomized smoothing, certified robustness
TL;DR¶
Using a shared generator to separate latent geometric changes from generator-concept changes, the paper compares prediction stability and smoothed-classifier certificates for standard and concept bottleneck models, finding that interpretability offers no uniform robustness advantage but changes where sensitivity appears; the conclusions remain conditional on generator-native inputs and specific task settings.
Background & Motivation¶
A Concept Bottleneck Model (CBM) separates image classification into interpretable concept prediction and concept-based class prediction. This interface exposes which attributes inform a decision, but does not itself guarantee stable attributes or stable final predictions. Previous studies have reported both that bottlenecks filter irrelevant variation and improve prediction robustness, and that concept representations can be unstable. These observations need not conflict: explanation stability and class-prediction stability are different questions.
Conventional pixel perturbations rarely identify which semantic factors have changed. The paper therefore uses an existing concept bottleneck generator to modify continuous latent variables or discrete concept states before image generation, then passes the same images to different classifiers. The generator is a measurement instrument, not a new classification network introduced by this paper. This arrangement controls external changes, but does not automatically equate a generator concept with a classifier's internal concept, or ensure that every semantic change should preserve the original class.
The question is therefore not whether CBMs are generally safer, but where a bottleneck filters or amplifies changes once input provenance, perturbation space, evaluation level, and task structure are specified. Core Idea: use a shared perturbation–generated image–prediction interface to measure empirical sensitivity separately from geometry-matched smoothing certificates, then examine class similarity and concept vocabulary changes to explain reversals in robustness comparisons.
Method¶
Overall Architecture¶
The experimental starting point is a pretrained generator containing an existing concept bottleneck autoencoder (CB-AE). A latent input passes through the first generator stage, is encoded into an explicit concept representation, and is decoded through the remaining generator stages into an image. CUB uses a StyleGAN-based generator and 15 frequently occurring concepts by default; RIVAL-10 uses BigGAN and all 18 attributes by default. Once trained, the generator is frozen and shared by the downstream classifiers.
The primary comparison uses a standard classifier and a Post-hoc Concept Bottleneck Model (PCBM) with a shared pretrained ResNet-50 backbone. The standard classifier predicts classes directly from image features; PCBM predicts through a concept-activation-vector-style representation and a decision head. The study focuses on post-hoc models with frozen backbones and does not replace evaluation of jointly trained or end-to-end CBMs.
The generator supports two types of change: modifying the continuous latent input and regenerating an image, or swapping positive and negative coordinates in paired generator concepts before decoding. Classifiers always receive images, not the modified generator concepts directly. Consequently, even a standard classifier without an explicit concept interface can be evaluated under the same external semantic changes.
Empirical evaluation records whether predictions change, whether generator concept states change, and how far PCBM's internal representation moves. Certified evaluation separately constructs randomized-smoothed classifiers and derives stability neighborhoods matched to latent-space or concept-space noise. These are distinct evaluations, and the smoothed predictor is not the unsmoothed base classifier.
This is generator-based evaluation and mechanism analysis, not a new multi-module network; datasets, metrics, and statistical tests are therefore not drawn as an invented network architecture. The discussion below follows the shared reference, sensitivity decomposition, geometry-matched certificates, and task-condition controls.
Key Designs¶
1. Shared generator and agreement-based reference: control external changes without treating them as natural-image ground truth
The paper first selects unperturbed generated images on which the standard classifier and PCBM agree, forming an agreement set, and uses their shared prediction as the reference label. This reduces confounding from different initial predictions: a subsequent change is not merely a disagreement that already existed. However, both classifiers may initially be wrong. The reference is not human ground truth, and initial agreement on this set is not natural-image accuracy.
This also introduces conditional selection: the results describe neighborhoods of generated images where both models already agree, rather than the entire generated or real test distribution. The official RIVAL-10 split informs training and dataset setup; stability numbers on the agreement set should not be rewritten as classification accuracy on that natural test set. Semantic changes can also change the appropriate class, so disagreement with the initial prediction measures instability and cannot universally be called a recognition error.
Latent changes are controlled by a Euclidean-norm budget, describing continuous neighborhoods in generator latent space rather than a fixed pixel distance or physical transformation. Empirical semantic evaluation uniformly selects a fixed number of distinct concepts and swaps each concept's paired positive and negative coordinates. Its semantic interpretation depends on generator and annotation quality: semantic control does not imply perfect disentanglement, strict geometric preservation, or class preservation.
Internal representations need not align across models. The appendix adds PixPNet to the same external interface and measures its prediction stability after generator-concept changes; it does not directly modify PixPNet's localized prototypes. Generated images and prediction-level evaluation can be shared, but this does not make generator concepts, PCBM concepts, and PixPNet prototypes equivalent representations.
2. Layered sensitivity measures: separate concept changes from their decision consequences
Concept Flip Rate (CFR) uses binary concept classifiers to detect changes in generator concept states. Prediction Flip Rate (PFR) measures changes in a classifier's final class. Decision Function Sensitivity (DFS) measures the Euclidean distance between PCBM's predicted concept representations for the two images. They describe external concept states, output classes, and an internal continuous representation, respectively. They are not interchangeable, and DFS is not a representation-distance metric with a shared scale across the standard classifier and PCBM.
For an original image and its changed counterpart, the normalized CFR, PFR, and DFS definitions in the paper can be written as:
Here the concept classifier describes generator concept states, whereas the concept extractor describes PCBM's internal representation. CFR is a proxy for concept changes, not newly annotated semantic ground truth for natural images. PFR tables use percentages, while the equation defines a probability.
A unit discrepancy must be retained: the methodology defines CFR as an average fraction of changed concepts, but the appendix's semantic CFR entries directly report the fixed flip count, and geometric entries can exceed 1. These entries cannot be substituted directly into the fractional definition, nor should counts and percentages be conflated. The numerical tables in this note extract only PFR, whose units are clear, rather than silently converting the paper's CFR values.
Why include DFS? Two inputs can change the same number of concepts while affecting attributes that the decision head uses to very different degrees. Continuous concept coordinates can also move without changing their binary states. DFS captures this internal drift and can therefore relate more closely to some prediction changes than a count alone. The paper nevertheless shows associations across budgets, not that DFS is a sufficient statistic or a causal law connecting concept drift to class flips.
3. Geometry-matched randomized smoothing: certificates constrain a new predictor in a specified neighborhood
Latent smoothing adds isotropic Gaussian noise to the latent input, repeatedly generates images, and aggregates the base classifier's class probabilities. The most probable class defines the smoothed classifier. The certificate concerns the latent-space smoothed version of the generator–classifier composite. With exact probabilities, its Gaussian radius depends on the top-class probability and the largest competing probability; practical evaluation uses lower and upper probability bounds.
The standard interpretation requires the top-class lower bound to exceed a valid upper bound on every competing probability; otherwise the predictor should abstain. The radius is measured in latent Euclidean space, not pixels or physical geometric tolerances. It does not establish that the base classifier remains constant everywhere in that neighborhood. Changing smoothing strength also changes the smoothed predictor, so a larger radius must be interpreted alongside its agreement-reference performance.
Concept smoothing uses a different probability model: every concept independently swaps its positive and negative coordinates with a specified probability, making the total flip count random. This differs from empirical evaluation's distribution with exactly a fixed number of selected concepts. For Bernoulli swap noise, the paper uses Neyman–Pearson probability transfer bounds on a Hamming cube: transfer the current top-class lower probability bound to its worst-case neighborhood lower bound, and each competitor's upper bound to its worst-case upper bound, requiring separation throughout.
The supported domain is the discrete orbit reachable by paired-coordinate swaps from the same continuous concept representation. Hamming distance counts which concepts have been swapped; it does not certify arbitrary continuous concept vectors. Two vectors with identical binary concept labels can still have different magnitudes and decode into different images. The theorem's wording in terms of binary-label distance over all concept representations is broader than the paired-swap mechanism directly supports; this note adopts the conservative swap-orbit interpretation.
The paper estimates probabilities with one-sided Clopper–Pearson intervals and claims an overall certification confidence. Individual class interval coverage, however, is not automatically simultaneous multiclass coverage. Selecting the top class from the same samples, selecting competitors, and comparing multiple upper bounds require joint-error accounting or appropriate post-selection reasoning. Neither the text nor the algorithm explains how these errors are combined, so an individual 95% interval is not treated here as a verified simultaneous multiclass 95% certificate.
Another implementation discrepancy must not be merged away: Appendix B.4 and the main certification table use independent Bernoulli flips, whereas Appendix C.4.2 explicitly uses fixed-count flips and cites Def. 5.7 / Cor. 5.6. These references cannot be resolved to the displayed certification results within the supplied text, and no equivalence between fixed-count noise and Bernoulli transfer bounds is established. That appendix's statistical table is not an exact replication under the same noise law.
Average certified radii retain only samples whose predictions match the reference and do not abstain. Such conditional averages can rise while coverage falls, so coverage must be reported alongside them. SmAcc counts abstentions as disagreement and measures the fraction that is reference-consistent and non-abstaining; overall non-abstention should still be distinguished from reference-consistent non-abstention. Appendix D's CA@1 additionally uses all evaluated samples as its denominator and measures the fraction that is reference-consistent and certified against at least one additional concept swap, providing population-level certificate coverage rather than an isolated conditional radius.
4. Task-condition controls: expose reversals through class structure and vocabulary design
The original RIVAL-10 task has almost no prediction flips at several budgets. A comparison restricted to this setting could mistake an easy task for an architectural advantage. Holding 18 concepts fixed, the authors add semantically similar or dissimilar classes to vary class count and interclass distances. Similarity is estimated by Euclidean distances between ResNet-50 class-centroid features, not an independent human semantic scale.
Another controlled experiment holds the original classification task fixed while expanding the vocabulary from 18 to 27 and 36 concepts, with additional attributes proposed using a large language model. This evaluates sensitivity under particular vocabularies and generators, not solely an abstract concept-dimensionality effect. Redundancy, correlations, supervision quality, and generator expressiveness may change together, so more concepts cannot automatically mean more information of unchanged quality.
PCBM outperforms the standard classifier in some expanded-class settings, but not every advantage survives a backbone change. Vocabulary experiments also reverse direction: an intermediate vocabulary is slightly favorable under large semantic changes, while a larger vocabulary is slightly unfavorable. Filtering weakly used concepts and sensitivity from correlated concepts are plausible explanations, but require per-concept usage and independent interventions; the observed reversal is not an established causal mechanism.
Loss & Training¶
The paper does not introduce a new robustness-training loss. Its core comparison uses standard classification or a PCBM concept pathway on frozen pretrained backbones, then measures their responses with a fixed generator. Generator training objectives or generic smoothing principles should not be presented as a new regularizer proposed by this work.
The main setup states 100 images per run, while the subsequent empirical statistics and appendix tables describe five paired subsets of 200 images each. This is a protocol conflict, not 100 repeated runs; the current version does not support assigning all main-table values to one reconciled sampling protocol. The certification section explicitly uses 1000 Monte Carlo samples per input and five 200-image subsets, sampled without replacement within each subset but potentially overlapping across subsets.
Appendix C.4.1 first describes a pool of 1000 generator-native inputs, then describes drawing subsets from a 400-image pool without explaining the relationship. Paired subsets reduce shared sample differences between methods, but overlapping subsets also mean that subset variability is not uncertainty from independently retrained models. The paper does not provide sufficient information to resolve these protocol ambiguities.
Key Experimental Results¶
Main Results¶
The following entries extract PFR from Appendix Table 1. Values are percentages, with lower values indicating greater stability relative to the initial prediction, not natural-image ground-truth accuracy. Error terms are retained as reported rather than forcing all tables into one aggregation protocol.
| Dataset | Perturbation budget | Standard classifier PFR (%) | PCBM PFR (%) |
|---|---|---|---|
| CUB-200-2011 | \(\epsilon=0.3\) | 74.6 ± 1.6 | 76.6 ± 0.8 |
| CUB-200-2011 | \(\tau=5\) | 77.7 ± 0.9 | 80.1 ± 0.6 |
| RIVAL-10 | \(\epsilon=0.3\) | 0.00 | 0.00 |
| RIVAL-10 | \(\tau=5\) | 0.72 ± 0.07 | 0.73 ± 0.09 |
PCBM changes predictions slightly more often in these CUB entries. RIVAL-10 is nearly saturated, so its similarly low flip rates cannot establish a bottleneck advantage. Cross-dataset values also depend on the generator, class structure, and vocabulary; they are not rankings under matched task difficulty.
The following Appendix Table 3 entries show randomized smoothing. Each “SmAcc / radius” cell first gives SmAcc as a percentage under the agreement reference, followed by the conditional average radius over reference-consistent, non-abstaining samples. Latent and concept radii have different units and cannot be compared by magnitude.
| Dataset | Smoothing setting | Standard classifier: SmAcc (%) / radius | PCBM: SmAcc (%) / radius |
|---|---|---|---|
| CUB-200-2011 | \(\sigma=0.10\) | 81.70 ± 2.39 / 0.1034 ± 0.0030 | 77.20 ± 2.14 / 0.0918 ± 0.0038 |
| CUB-200-2011 | \(\rho=0.10\) | 59.13 ± 1.46 / 0.909 ± 0.151 | 57.60 ± 0.59 / 0.751 ± 0.061 |
| RIVAL-10 | \(\sigma=0.10\) | 100.00 ± 0.00 / 0.2463 ± 0.0000 | 100.00 ± 0.00 / 0.2463 ± 0.0000 |
| RIVAL-10 | \(\rho=0.10\) | 99.90 ± 0.01 / 6.560 ± 0.002 | 99.89 ± 0.01 / 6.558 ± 0.002 |
These representative entries do not support a claim that PCBM generally has larger certified radii. At the smallest latent noise on CUB, PCBM can have slightly higher SmAcc, reinforcing that reference-consistent prediction coverage and conditional radius are separate dimensions, not strict dominance at every noise level.
Ablation Study¶
The table combines representative paired tests from Appendix Table 2, class-condition analysis from Table 6, and vocabulary analysis from Table 8. The paired difference is standard-classifier PFR minus PCBM PFR, so negative values favor the standard classifier. Displayed method means are not used to recompute or replace the source's paired statistics.
| Analysis setting | Standard classifier PFR (%) | PCBM PFR (%) | Paired difference (pp) | Raw p | Holm p |
|---|---|---|---|---|---|
| CUB, \(\epsilon=0.30\) | 74.6 ± 1.6 | 76.6 ± 0.8 | -2.00 ± 1.21 | 0.021 | 0.098 |
| CUB, \(\tau=5\) | 77.7 ± 0.9 | 80.1 ± 0.6 | -2.33 ± 0.84 | 0.004 | 0.032 |
| CUB, \(\tau=10\) | 90.0 ± 0.6 | 90.4 ± 0.4 | -0.84 ± 0.49 | 0.018 | 0.098 |
| RIVAL, added similar classes, \(\epsilon=0.8\) | 12.80 ± 1.41 | 9.55 ± 1.22 | Not reported | Not reported | Not reported |
| RIVAL, intermediate vocabulary, \(\tau=10\) | 8.74 ± 0.99 | 8.69 ± 0.88 | Not reported | Not reported | Not reported |
| RIVAL, larger vocabulary, \(\tau=10\) | 2.62 ± 0.18 | 3.06 ± 0.19 | Not reported | Not reported | Not reported |
The class-condition entry adds 5 similar classes, yielding 15 classes with 18 concepts. Intermediate and larger vocabularies use 27 and 36 concepts, respectively. “Not reported” explicitly means the source table does not provide these tests, not that a test passed or a difference is absent.
Appendix Table 2 applies Holm correction across all 9 budgets. Only the semantic budget of 5 is significant at the 0.05 threshold, with a corrected value of 0.032. Multiple raw p-values below 0.05 do not establish significant differences across all budgets.
Key Findings¶
- Original and expanded tasks can yield different directions. Similar-class expansion exposes gaps between standard classification and PCBM, whereas the saturated original RIVAL-10 task hardly distinguishes them. Some similar-class settings with DenseNet-161 instead give PCBM higher PFR, so the ResNet-50 trend is not a backbone-independent law.
- Vocabulary benefits are non-monotonic. For 27 concepts at the large semantic budget, the main text reports PCBM PFR 9.29 and standard PFR 9.42; for 36 concepts it reports 3.11 and 2.65. The table above retains a different set of values from Appendix Table 8. Directions agree, but numbers differ and should not be silently reconciled.
- Concept counts alone do not explain prediction stability. DFS adds continuous representation information, but its association with PFR does not prove a mechanism. Descriptions of results at different budgets as “similar” should be read against the actual values, not as a strict monotonic relationship.
- Appendix D supports applying the external interface to PixPNet, not universal superiority of prototype models. Gaussian results in Table 14 match the main table, but semantic Table 15 gives ResNet SmAcc 75.25 at noise probability 0.10, versus 59.13 in main Table 3. Concept-radius magnitudes in Tables 16 and 3 also differ substantially without a clear explanation in the current version.
Highlights & Insights¶
- Separating interpretability from robustness is more informative than labeling an entire model stable or fragile. Generator concept states, PCBM representations, class outputs, and smoothing certificates answer different questions; comparisons require explicit measurement conventions.
- Sharing external changes rather than forcing internal explanations to align allows PCBM and localized-prototype models to be measured under the same input variations. The reusable component is the evaluation interface, not a mathematical equivalence between concept and prototype spaces.
- Saturated tasks can hide differences, making class structure and vocabulary part of evaluation design. Population-level certificate coverage should accompany conditional mean radii to avoid highlighting only the easier remaining samples.
Limitations & Future Work¶
- Generator-native inputs, agreement-based selection, and reference predictions constrain external validity. Shared predictions are not natural ground truth, and semantic changes need not preserve classes. Real inputs and human concept/class checks are needed, but existing certificates cannot simply be extrapolated to them.
- The paired-swap orbit is the supported concept-certification domain; broader claims about arbitrary continuous representations with identical binary labels require additional assumptions. Full probability-transfer constructions and joint confidence-error control should also be specified rather than treating named tools as sufficient evidence of overall coverage.
- Sampling conflicts include 100 images versus five 200-image subsets, and 1000-image versus 400-image pools. Fixed-count smoothing in C.4.2 must not be merged with Bernoulli smoothing in B.4; its Def. 5.7 / Cor. 5.6 references remain unresolved. These issues affect reproducibility and require explicit version and data provenance.
- Displayed means and paired differences disagree: at CUB semantic budget 10, the displayed PFR difference is -0.4 pp while the paired table gives -0.84 pp; at latent noise 0.05, the main-table SmAcc difference is 3.50 pp while the statistical table gives 0.80 pp. The text does not establish whether different aggregations or runs explain this. Both reports are retained rather than corrected speculatively.
- Model-assisted vocabulary construction does not fully disentangle attribute correlations, supervision quality, and generator expressiveness. Future evaluations should hold task and generator quality fixed, report per-concept contributions, and distinguish subset variability from training variability and semantic annotation errors.
- Stress-test results offer supplementary evidence that prediction sensitivity differs across spaces. No operational procedure is provided here, and neither a single stress test nor improvements after smoothing establish general robustness guarantees.
Related Work & Insights¶
- vs Koh et al.'s CBM / Yuksekgonul et al.'s PCBM: CBM establishes concept-mediated prediction, while PCBM builds a post-hoc concept pathway on frozen backbones. This paper primarily studies PCBM's response to shared generator changes, not identical trade-offs across every CBM training regime.
- vs Sinha et al. / Rasheed et al. on robustness: separating concept stability from final-prediction stability explains why studies can observe different phenomena. This clarifies applicability conditions rather than experimentally reproducing every historical disagreement.
- vs Cohen et al. / Lee et al. on randomized smoothing: existing certificate ideas are applied to generator-composite latent space and discrete concept-swap space. The transferable lesson is to define the noise law and allowed change set before discussing a certificate, not use one certificate for every semantic change.
- vs PixPNet: prototypes provide another interpretable internal representation. The appendix keeps that representation intact while applying shared external evaluation, supporting cross-family measurement rather than one-to-one alignment between concepts and local prototypes.
Rating¶
- Novelty: 4/5. A clear joint perspective on perturbation spaces, evaluation levels, and task conditions; certification largely reuses existing theory.
- Experimental Thoroughness: 3/5. Multiple tasks, vocabularies, and model families are covered, but sampling protocols, numerical provenance, and certification assumptions need clarification.
- Writing Quality: 3/5. The main argument is accessible; inconsistent appendix units, statistical differences, and noise laws weaken reproducibility.
- Value: 4/5. Useful for designing robustness evaluations of interpretable models, not a general safety or natural-image robustness promise.