Universal Cross-Prompt Adversarial Attacks on Promptable Concept Segmentation¶
Conference: NeurIPS 2026 (acceptance information supplied in the task package)
arXiv: 2609.39265
Area: Segmentation
Keywords: promptable concept segmentation, cross-prompt robustness, universal perturbation, temporal memory, robustness evaluation
TL;DR¶
AdvPCS examines vulnerabilities in promptable concept segmentation through prompt variation, concept perception, and temporal memory, reducing three-model average video mIoU on SA-CO with text prompts from 70.24% to 4.61% in controlled digital-input experiments, without establishing failure guarantees for arbitrary prompts, models, or real-world inputs.
Background & Motivation¶
SAM shifts segmentation from fixed-category prediction to an interactive task in which points or boxes specify the target; SAM2 adds video memory to maintain object identity using earlier observations. SAM3 extends this to Promptable Concept Segmentation (PCS), accepting concept descriptions and attempting to identify multiple matching instances. Segmentation quality consequently depends not only on pixel boundaries, but also on whether the model recognizes that the target exists, associates the prompt with the correct objects, and continues tracking the same objects over time.
This transition changes the robustness question. Degradation under one point prompt does not establish that a box prompt or synonymous text will also fail; conversely, temporary recovery after changing the prompt does not establish stable resistance to shared visual abnormalities. Earlier evaluations of SAM and SAM2 primarily address spatial prompts or visual representations, with less coverage of concept perception and multiple prompt modalities together. The paper therefore studies a reusable universal perturbation and asks whether its effects extend across prompt types, prompt instances, and video frames, rather than documenting only an isolated input failure.
Video introduces a two-sided factor: historical memory can buffer unreliable current observations, but it can also retain erroneous associations and influence later decisions. The paper examines prompt-conditioned perception together with temporal state, avoiding an explanation that attributes every low-overlap result to boundary prediction. Core Idea: treat cross-prompt segmentation robustness as a system property jointly determined by concept recognition, local correspondence, and temporal state, and examine its degradation boundaries through controlled perturbation reuse and stratified evaluation.
Method¶
Overall Architecture¶
The paper introduces AdvPCS to study SAM3 and its PCS variants, rather than to train a more accurate segmentation model. During normal inference, an image encoder supplies visual representations, prompts participate in target recognition and localization, a detection branch judges concept presence, and a tracking branch combines historical information to produce segmentation outputs. Current-frame video predictions therefore depend on both present observations and previous state.
From a non-operational research perspective, the method has three connected concerns: cross-prompt coverage, globalโlocal perception, and temporal state deviation. The first asks whether different prompts expose a shared visual vulnerability; the second distinguishes frame-level concept recognition from local candidate-object correspondence; the third asks whether errors remain confined to the current frame or accompany state changes that affect later predictions.
The experiments contrast segmentation quality on normal and controlled perturbed inputs, separating image and video settings, point and box and text prompts, datasets, and models. Universality here primarily means reuse across inputs within the corresponding experimental configuration, rather than individual adaptation to each test sample. The paper does not sufficiently establish that one completely unchanged experimental artifact covers every task, model, and mode, so โuniversalโ should not be expanded into unconditional universality.
The threat model assumes access to open-source SAM3 and public research data. This supports observation of internal model behavior under white-box conditions, which is stronger access than a closed service exposing only final masks. The paper also reports cross-model transfer, but model transfer and strictly black-box evaluation are not equivalent: the latter requires explicit target-access permissions, query constraints, and independent testing conditions.
This note covers mechanism interpretation, evaluation boundaries, and defensive implications without reproducing perturbation construction, objectives, prompt-selection recipes, update procedures, or memory-interference implementation. It also omits a generation-flow diagram to avoid presenting a conceptual analysis as an executable method.
Key Designs¶
1. Cross-prompt coverage: treat prompt variation as an independent evaluation dimension
A concept description, an interior point, and a bounding box can indicate the same object, but they provide different information. Text emphasizes semantic categories and descriptions, points provide local positions, and boxes supply a fuller spatial extent. Testing only one prompt may chiefly reveal that promptโs information bottleneck rather than stable model understanding of the visual content.
At a conceptual level, AdvPCS considers multiple prompts and their variations together, seeking conclusions less dependent on one prompt instance. A crucial distinction is between transfer across prompt types and transfer across prompt instances: the former includes point-to-box variation, whereas the latter includes different textual descriptions of the same target. Both require explicit test-coverage boundaries; failure across a few instances does not establish coverage of the entire prompt space.
The source reports prompt-instance transfer plots and evaluates multi-point inputs in the appendix. Additional spatial information partially improves results on perturbed inputs, indicating that prompts retain a compensatory role. โChanging the prompt is ineffectiveโ is therefore too strong; the supported conclusion is that the tested prompt variations do not fully eliminate the performance gap.
2. Globalโlocal perception: distinguish concept presence from object correspondence
PCS differs from a task that generates masks solely from positions: the model must first understand what the prompt denotes and judge whether matching objects exist in the scene. Frame-level concept-presence judgments and local candidate-object correspondence can fail differently. The former may omit objects that should be detected, whereas the latter may produce erroneous associations or incomplete coverage.
The paper considers these perception scales together, emphasizing that segmentation failure can occur before mask output rather than merely through rougher boundaries. For defensive research, checking mask shape or edge smoothness alone may consequently miss concept-level omissions; concept presence, object localization, and final mask quality should be observed separately.
The exploratory experiments report an association between perception-confidence changes and mIoU degradation. This is empirical support for examining the perception branch, but an association does not establish perception as the sole causal source. Spatial constraints also differ across prompts, and final overlap can jointly depend on visual representations, localization, and tracking.
3. Temporal state deviation: distinguish current-observation errors from persistent errors
A video model does not simply run image segmentation independently on every frame. It retains information about object identity and previous appearances to handle motion, occlusion, and temporary disappearance. The same current-frame abnormality may therefore yield different predictions under different historical states, and single-frame mIoU does not fully describe video failure dynamics.
At an abstract level, the temporal component of AdvPCS examines whether state evolution under controlled inputs departs from the normal sequence. Its research intent is to evaluate immediate perception degradation together with cross-frame identity continuity, rather than assume that memory necessarily protects against abnormalities. Historical information can support recovery or perpetuate incorrect correspondence, depending on how the model trusts and updates its state.
This explanation does not require operational instructions for interfering with internal state. The direct defensive implication is to measure identity consistency, failure duration, and recovery after abnormalities alongside average mask overlap. The paper mainly supplies mIoU and qualitative descriptions, without sufficient numbers to quantify these additional metrics independently.
A Worked Example¶
Consider a short video containing the same target, indicated separately by a concept description, a point, and a box. Under normal conditions, the model should recognize the target and maintain correspondence over successive frames; under controlled abnormal conditions, an evaluator observes whether the prompts still produce similarly correct masks and whether errors persist into later frames.
Partial recovery with a box prompt indicates that additional spatial information helps in that condition, not that all other abnormalities have been resisted. If image-mode performance degrades strongly while video retains some overlap, historical information may have a compensatory role, but independent temporal analysis is needed to establish the reason.
This is a conceptual example for interpreting evaluation results, not a measured case from the paper or a perturbation-generation procedure. Its purpose is to connect prompt variation, perception scales, and temporal state in one scenario.
Key Experimental Results¶
Main Results¶
The paper uses SA-CO, YouTube-VOS (called YouTube in the tables), MOSE, and DAVIS, evaluating SAM3, SAM3.1, and EfficientSAM3 (E-SAM3). Image tests use frames extracted from videos, rather than a completely independent static-image source. The paper states that perturbation development uses SA-CO before transfer to other datasets; the cache does not sufficiently specify separation between fitting and testing samples within SA-CO, so all results cannot automatically be described as strictly unseen-sample generalization.
BoU denotes mIoU on normal inputs, and AoU denotes mIoU on perturbed inputs. mIoU measures average overlap between predicted and ground-truth masks; a lower AoU indicates more severe degradation in this controlled evaluation, not the percentage of failed samples. The following selection from source Table 1 uses three-model averages, all in %, rather than SAM3-only results.
| Dataset | Prompt | Video BoU | Video AoU | Image BoU | Image AoU |
|---|---|---|---|---|---|
| SA-CO | Text | 70.24 | 4.61 | 78.42 | 3.96 |
| YouTube | Point | 75.78 | 17.48 | 60.38 | 12.81 |
| YouTube | Box | 77.88 | 37.27 | 82.24 | 25.41 |
| MOSE | Point | 62.91 | 15.75 | 56.99 | 12.09 |
| MOSE | Box | 65.87 | 27.26 | 76.72 | 27.28 |
| DAVIS | Point | 67.62 | 25.90 | 51.71 | 11.58 |
| DAVIS | Box | 76.97 | 42.01 | 80.81 | 23.92 |
Average AoU with SA-CO text prompts is indeed below 5%, but not every model is below 5%. E-SAM3 has text AoU of 12.50% for video and 11.19% for images; the corresponding SAM3 values are 0.00% and 0.07%. Averages should not obscure model differences, and a rounded 0.00% should not be interpreted as exactly zero overlap on every instance.
Ablation Study¶
Source Table 2 compares SAM3: its normal-input values match the SAM3 columns in Table 1, not the three-model averages. The following table preserves the aggregate mean over seven datasetโprompt conditions: SA-CO text, and point and box prompts on the other three datasets.
| Method | Video average mIoU (%) | Image average mIoU (%) | Interpretation boundary |
|---|---|---|---|
| Normal inputs | 77.74 | 81.65 | SAM3 normal-input reference |
| AttackSAM | 72.90 | 70.24 | Baseline adapted to the paperโs shared settings |
| DarkSAM | 73.67 | 74.30 | Baseline adapted to the paperโs shared settings |
| S-RA | 40.02 | 23.01 | Baseline adapted to the paperโs shared settings |
| UAD | 32.31 | 27.15 | Baseline adapted to the paperโs shared settings |
| UAP-SAM2 | 29.83 | 11.58 | Baseline with the second-lowest aggregate values here |
| AdvPCS | 18.33 | 2.49 | Lowest aggregate values in this table |
Checking individual columns in Table 2 shows that AdvPCS is below the listed baselines in all seven video conditions and all seven image conditions. This conclusion applies only to the tableโs SAM3 evaluation and baseline adaptations; it does not establish comprehensive superiority across three models, arbitrary prompts, or every existing method.
Module ablations appear in Figure 5(a): the prose states that removing each component weakens the degradation effect, but the cache contains no readable exact plot values. No component-removal magnitudes are invented, and the largest contributor cannot be determined. The next table instead uses the numerically supported multi-point evaluation in appendix Table A1, averaging SAM3 over YouTube, MOSE, and DAVIS.
| Input points | Normal average mIoU (%) | Perturbed average mIoU (%) | Supported observation |
|---|---|---|---|
| 1 | 74.71 | 7.45 | Large gap under single-point input |
| 2 | 78.88 | 15.71 | More spatial information accompanies partial recovery |
| 3 | 82.52 | 18.53 | Perturbed performance remains far below normal performance |
| 4 | 83.09 | 20.58 | Highest perturbed average in the table |
| 5 | 83.35 | 19.71 | Recovery is not strictly monotonic |
Key Findings¶
- Box prompts often retain more overlap, but still exhibit substantial degradation; this supports separate prompt-level reporting rather than a single overall mean.
- Multi-point evaluation shows that additional prompt information can partially compensate for abnormalities without restoring normal performance in the tested conditions.
- The source prose places YouTubeโs 17.48% and DAVISโs 25.90% near a discussion of text prompts, but Table 1 explicitly assigns them to point prompts; this note follows the table labels.
- The appendix claims evaluation with 1โ6 points, but readable Table A1 contains only 1โ5 points; no sixth row is reconstructed.
- Seed stability and model transfer are mainly supported by plots without readable exact values in the cache; no confidence intervals or transfer-matrix values are reported from them.
Highlights & Insights¶
- Prompt robustness is not one score. Different prompts yield different degradation levels for the same visual content, making prompt information structure part of evaluation. Defensive benchmarks should retain prompt-group results and normal-performance references.
- Concept recognition and segmentation boundaries require separate diagnosis. Low mIoU can arise from omissions, incorrect correspondence, or poor contours. Separating these failure types helps identify the actual system bottleneck.
- Memory provides both compensation and persistent risk. Video state may support target recovery or perpetuate errors. Better evaluations should report recovery capability together with instantaneous performance.
Limitations & Future Work¶
- The authors acknowledge additional computational overhead from cross-prompt processing and dependence on concept-presence mechanisms in PCS; the method may not apply to conventional segmentation or SAM and SAM2.
- The main study relies on open-source model access; transfer results do not replace a strictly defined black-box test or establish reliable failure in physical environments.
- The observation that text is more vulnerable is confounded by dataset differences: Table 1 reports text only on SA-CO, and points and boxes on other datasets. These rows alone cannot identify a causal effect of prompt type.
- The defense prose summarizes preprocessing results as normal mIoU falling from approximately 70% to 40%, with perturbed mIoU remaining below 15%; pruning also severely reduces normal performance. Defensive evaluation must measure retained normal utility rather than count destruction of model capability as successful protection.
- The defense section initially describes experiments on SA-CO, but the pruning paragraph explicitly specifies DAVIS. The two defense studies are not an unambiguous single-dataset experiment, and neither represents all defenses.
- Future research could examine prompt-consistency checks, concept-omission warnings, and temporal-state trust monitoring. These are defensive directions, not solutions already validated by this paper.
Related Work & Insights¶
- vs SAM / SAM2 / SAM3: The first two establish spatial prompting and video memory, while SAM3 adds concept recognition. This paper examines robustness after these capabilities are combined, rather than improvements in normal segmentation accuracy.
- vs AttackSAM / DarkSAM: Earlier methods provide references for prompt-related and cross-prompt visual robustness; this paper expands attention to PCS perception and video state. Adapted baseline results do not represent the complete capabilities reported in their original papers.
- vs UAP-SAM2: Both examine cross-frame reuse and memory-related failure, but this paper additionally discusses concept perception and multiple prompt types. Comparison should focus on controlled experimental coverage rather than claim that a mechanism is universally more vulnerable.
- Defensive research implications: Semantically consistent prompt groups and temporal-recovery evaluations can help distinguish input abnormalities, prompt ambiguity, and persistent state errors, reducing dependence on a single average metric.
Rating¶
- Novelty: 4/5 โ Brings PCS perception, cross-prompt variation, and video state into one robustness question.
- Experimental Thoroughness: 3/5 โ Covers four datasets and three models, but prompts and datasets are intertwined and some plots lack readable numbers.
- Writing Quality: 3/5 โ Clear research direction, with inconsistencies in prompt labeling, point-evaluation coverage, and defense datasets.
- Value: 4/5 โ Useful evidence for stratified defensive evaluation of segmentation foundation models, not a guarantee of universal failure or defensive effectiveness.