Watermarking Should Be Treated as a Monitoring Primitive¶
Conference: NeurIPS2026 Position Track (position paper)
arXiv: 2605.13095
Area: AI Safety / Watermark Security & Governance
Keywords: watermarking, monitoring primitive, entity linkability, attribution access, privacy governance
TL;DR¶
Through an observer-based threat model and controlled text and image experiments, this paper argues that persistent entity bindings and reliable inference can make watermarking support repeated attribution, motivating governance of both internal attribution access and design-dependent external source identification beyond per-sample robustness.
Background & Motivation¶
Generative watermarking is commonly presented as provenance infrastructure: downstream systems determine whether content was generated by a model or recover embedded information. Research consequently emphasizes detectability after rewriting or editing and resistance to forgery. These evaluations establish signal reliability, but do not fully characterize who can infer the generating entity or what privacy implications arise when the same capability is used repeatedly. More reliable detection does not necessarily make every deployment purpose safer.
The connection between attribution and monitoring comes from deployment, not just the watermark payload. A zero-bit watermark indicates whether a particular signal is present, yet distinct keys persistently bound to accounts can still support account attribution when an actor with attribution access reliably distinguishes those signals. Another exposure requires no observer access to keys: some configurations leave persistent, learnable source differences. Whether they support external identification depends on design and data conditions; watermark presence alone does not imply that every user can be identified.
This is a NeurIPS2026 Position Track paper, not a new universal identification algorithm or the first observation of watermark privacy risks. Its contribution is to bring access, binding granularity, and persistence into one analytical framework, supported by feasibility evidence in limited configurations. Core Idea: extend the unit of watermark evaluation from isolated samples to entity-level inference capabilities, while distinguishing internal observers with attribution access from external observers without keys or detectors.
Method¶
Overall Architecture¶
Here, โmethodโ refers to the threat model and organization of evidence, rather than a new network. The paper separates watermark-presence detection, entity attribution, and cross-output linkability; analyzes the conditions for these goals under different observer access levels; and supports its position with internal attribution experiments, external source identification experiments, and controls.
The internal pathway depends on detector or decoder access, entity mappings, and reliable decisions. The external pathway depends on suitable source labels, persistent bindings, and learnable watermark-related differences. Both concern observed outputs only and do not automatically reveal an entity's total model usage. The experiments do not reconstruct complete activity histories.
Four connected questions structure the analysis: access categories determine who can attribute, binding granularity and persistence determine the attribution unit and temporal scope, observable structure determines external exposure, and reliability and governance boundaries determine how results should be interpreted. The key designs below summarize this analytical framework; they are not four new algorithmic modules introduced by the authors.
Because the contribution is primarily a position and threat-model analysis, the note does not turn its argument outline into an algorithm diagram or provide monitoring procedures, identity-mapping implementations, or source-label acquisition methods.
Key Designs¶
1. Classify observers by access: organizational affiliation does not determine capability
โInternalโ does not mean employee, and โexternalโ does not designate a fixed type of institution. An internal observer can use the relevant detectors and associate their results with entities. This capability may come from possessing keys or receiving access to a detection service. In multi-bit watermarking, attribution may also depend on a decoder recovering entity-related information. A third-party institution granted these capabilities remains an internal observer in this model.
Consequently, a presence-detection interface alone, or an entity mapping alone, does not guarantee reliable attribution. Detection must effectively distinguish the relevant candidates, and the information must genuinely correspond to the entity under discussion. Signal failure after transformations, withdrawn access, or changed bindings can interrupt this inference chain. The analysis specifies capability conditions rather than assuming identical capabilities across platforms.
An external observer has neither keys nor detector access. The experiments assume suitable source-labeled reference outputs and evaluate source classification on unseen outputs. This setting requires trustworthy labels, sufficient data, consistent entity bindings between reference and test sets, and learnable output differences. It does not establish that ordinary users necessarily have enough reference material available in practice.
Publication accounts, generating accounts, and natural persons must be distinguished. A source label may designate only an account or organization; it need not identify the generating account, much less a verified real-world identity. โIdentificationโ here means prediction of experimental entity labels, not confirmation of a private person's identity.
2. Separate payload capacity from deployment granularity: zero-bit does not mean zero attribution capability
Zero-bit watermarking indicates whether a particular signal is present, whereas multi-bit watermarking can recover additional messages. This distinction describes the payload, but cannot independently determine privacy properties. Under persistent per-entity multi-key deployment, the correspondence between detection signals and entities can support attribution without explicitly encoding names or account identifiers in content.
Determining that an output contains a watermark, attributing it to an entity, and establishing that two outputs share an entity are three different conclusions. The first does not automatically imply the second; reliable attribution with persistent bindings may support the third. Even successful attribution associates an output with an entity in the deployment, rather than establishing that the entity is a particular real-world person.
Persistence further determines whether the capability extends across time. Long-lived per-account bindings can increase linkability, while group-level or shared bindings reduce the distinction provided by the watermark at the cost of entity-specific attribution. Key rotation may weaken persistent linkage only when the observer cannot retain cross-rotation associations. Merely changing a key's name does not constitute a privacy guarantee.
This explains the paper's emphasis on bindings and access in governance. Supporting legitimate auditing is one question; allowing access granted for an investigation to persist for other purposes is another. The paper does not label every attribution query abusive monitoring, nor does it establish that providers or authorities have committed such misuse.
3. Examine design-dependent observable structure: external exposure is not universal
The central external-identification question is not whether a key can be recovered, but whether outputs contain persistent entity-related differences that distinguish sources without a detector. The authors use KGW text and Tree-Ring images as feasibility cases. The 12 methods covered internally must not be described as externally identified, and two cases do not establish equivalent exposure across all watermarking designs.
Tree-Ring requires a particularly important qualification: an entity's experimental โkeyโ includes combinations of watermark pattern, radius, mask shape, and channel, rather than only a random seed under a fixed configuration. Its external results therefore concern learnable configuration-level differences, not a demonstration that changing only the seed necessarily distinguishes users. The text results likewise apply only to the evaluated KGW configuration; the paper does not systematically sweep other strengths or context settings.
To avoid mistaking content preferences for watermark differences, entities use shared prompt pools, while prompts and samples do not overlap across data splits. No-watermark and shared-key controls further examine whether source classification retains a substantial advantage after entity-specific watermark distinctions are removed. These controls support the watermark's contribution under the experimental conditions, rather than ruling out every real-world source cue.
The figures' โsamples per entityโ axis measures the amount of labeled reference data, not monitoring duration or joint inference over multiple anonymous outputs. The authors predict the source of individual held-out outputs; they do not evaluate joint aggregation of anonymous probes or demonstrate reconstruction of long-term activity trajectories.
4. Interpret reliability and authorization together: accuracy is not identity proof
The main internal metric is the correct-and-detected attribution rate: success requires the correct entity and passage of the corresponding calibrated threshold, while incorrect attribution and abstention both count as failures. TPR@1%FPR refers to per-key false-positive calibration on non-matching outputs, not a global 1% false-attribution guarantee for multi-key selection. Candidate-set size and correlations between detector scores affect such extrapolation.
External top-1 and top-3 measure source classification within a finite candidate set, not open-world identity authentication. The binary experiment also shows that high overall accuracy can coexist with consequential false positives. When genuine target outputs are rare, false positives may form a larger share of positive predictions. Experimental accuracy cannot replace analysis of deployment base rates, rejection of unknown sources, and the consequences of errors.
Governance therefore needs separate statements of attribution purposes, authorized actors, mapping retention scope, binding persistence, result retention, and secondary-use restrictions, together with review and contestability. Restricting detector access primarily addresses the internal pathway and does not automatically eliminate external exposure from existing learnable structure. Changing observable structure likewise does not automatically remove internal capabilities held by actors retaining attribution access.
These are proposed deployment trade-offs and governance directions, not experimentally established mitigations. Distribution-preserving or undetectable constructions may reduce externally learnable differences, but the paper neither evaluates these mitigations nor establishes failures of their theoretical guarantees. Properties of single outputs must also be considered separately from repeated-observation conditions.
Key Experimental Results¶
Main Results¶
Text generation uses Qwen2.5-14B with C4 prompts; image generation uses Stable Diffusion v2.1 with the Stable Diffusion Prompt dataset. Prompts are matched across entities, and training and test prompts and samples are disjoint. External testing uses 100 held-out outputs per entity.
The internal experiments cover 12 zero-bit methods, with entity counts increasing from 1 to 16. Figure 3 and the accompanying prose report generally high attribution with modest declines for some methods as entity counts grow. The cache does not provide exact values at every plotted point, so percentages are not inferred from extracted figure text.
Table 1 summarizes explicitly reported external-identification values. It describes only the two evaluated configurations, not an algorithm ranking.
| External experimental condition | Entities | Labeled reference samples per entity | Metric | Result | Corresponding random baseline |
|---|---|---|---|---|---|
| KGW text, smaller reference set | 16 | 100 | top-1 | 11.3% | 6.25% |
| KGW text, larger reference set | 16 | 4000 | top-1 | 73.0% | 6.25% |
| KGW text, larger reference set | 16 | 4000 | top-3 | 90.0% | 18.75% |
| Tree-Ring images, configuration-level entity distinctions | 16 | 4000 | top-1 | 91.0% | 6.25% |
Source: ยง5.2 and Figure 4, full-text cache lines 439โ459 and 521โ525. Random baselines are calculated from \(1/n\) and \(\min(3/n,1)\); 100 and 4000 are the reference-data endpoints in Figure 4, and the prose reports KGW's change from 11.3% to 73.0%.
Ablation Study¶
These are source-distinction controls rather than ablations of new model modules. Table 2 retains only qualitative results supported by the paper and does not invent precise values from extracted Figure 5 text.
| Config | Text / image scope | Source-inference performance | Supported interpretation | Unsupported interpretation |
|---|---|---|---|---|
| Per-entity multi-key, internal observer | KGW / Tree-Ring control figure | Near-perfect | Access enables attribution using the tested signals | Reliability at every scale or a global 1% false-attribution rate |
| Per-entity multi-key, external observer | KGW / Tree-Ring | Strong but below internal performance | Entity-related structure is learnable in selected designs | External validation of all 12 methods |
| No watermark | Corresponding KGW / Tree-Ring generation settings | Near random | No substantial watermark distinction benefit in the matched setting | Real-world unwatermarked content cannot be identified |
| Shared key across entities | KGW / Tree-Ring | Near random | Sharing removes the tested entity-specific watermark distinction | Shared keys guarantee universal anonymity |
Source: ยง5.2 Controls and Figure 5, cache lines 532โ542 and 590โ596. Model, prompt distribution, and generation settings are controlled here; other source cues may still exist.
Table 3 presents counts from the paper's Table 1 binary experiment and their direct implications, illustrating the distinction between accuracy and false attribution.
| Binary experiment item | Value | Interpretation |
|---|---|---|
| Target outputs correctly classified as target | 100 / 100 | Target recall is 100% |
| Background outputs incorrectly classified as target | 48 / 1500 | Target false-positive rate is 3.2% |
| Background outputs correctly classified as background | 1452 / 1500 | Background recall is 96.8% |
| Overall accuracy | 97.0% | Mixed test set of 100 target and 1500 background outputs |
| Always-background baseline | 93.75% | Target recall is 0% |
| Precision of target-positive predictions | Approximately 67.6% | Note calculation: 100 / (100 + 48), not an additional experiment reported by the paper |
Source: the paper's Table 1 and ยง5.2 Targeted Identification, cache lines 414โ427, 460โ463, and 526โ531. The experiment includes only one randomly selected target and a background comprising 15 other entities; it cannot be generalized to larger, more diverse populations.
Key Findings¶
- External source classification improves with more labeled reference data in the tested configurations, but the paper does not establish that ordinary users have that amount of suitable material available.
- No-watermark and shared-key controls support the contribution of entity-specific watermark structure. This is a controlled-experiment conclusion, not a real-world anonymity guarantee.
- Internal evidence across 12 methods is broader than external evidence for two configurations; the two cannot be substituted. Internal reliability still depends on access, distinguishability, and bindings.
- The binary experiment's 48 background false positives call for caution under low base rates. Its 97.0% accuracy is not the confidence of an individual identity conclusion.
Highlights & Insights¶
- Separating payload from deployment is important. Zero-bit limits explicit message capacity, not the entity-level information obtainable from persistent key mappings.
- Defining observers by access places providers, platforms, and authorized third parties in one capability framework. This supports discussion of continuing permissions and purpose expansion without presuming malicious intent.
- Shared-key controls distinguish possible source differences in content from differences added by watermarking. Conditions and boundaries are more informative than an unqualified claim that watermarks identify sources.
- Privacy evaluation cannot stop at a single detection. Even when each attribution has a legitimate purpose, persistent access and cross-context retention of results require separate governance.
Limitations & Future Work¶
- External experiments cover only KGW text and configuration-level Tree-Ring image differences, not all watermarking designs, all generative models, or seed-only key distinctions.
- External identification after paraphrasing, summarization, image editing, and decoding changes is not evaluated, nor are context-width or watermark-strength variations systematically studied.
- Experiments assume trustworthy source labels, stable entity bindings, and sufficient reference data. Semi-supervised, clustering, and unlabeled settings remain future work.
- Population-scale identification, complete activity-history reconstruction, actual surveillance deployment, and real misuse are not established. Illustrated monitoring scenarios are hypothetical uses of capabilities, not measured facts.
- Per-key 1% false-positive calibration is not a global false-attribution guarantee. Open-world unknown sources, changing base rates, and error consequences still require evaluation.
- Group keys, key rotation, access restrictions, and distribution-preserving designs are governance or technical directions awaiting validation. Defensive evaluation should jointly report legitimate attribution benefits, external linkability, and residual internal access capabilities.
Related Work & Insights¶
- vs multi-bit identity-payload research: These schemes can recover entity-related messages; this paper emphasizes that zero-bit multi-key deployments may also support attribution. The distinction concerns deployment correspondences, not a new identity-encoding algorithm.
- vs watermark removal, forgery, and rule-inference research: Those objectives concern disrupting, forging, or recovering signals; this paper concerns observers using signals for entity inference. Successful detection or rule recovery does not itself establish successful entity identification.
- vs neural authorship attribution and stylometry: Those studies examine source cues in content. This paper uses prompt matching and two controls to assess watermark-related differences, without denying other real-world cues.
- vs distribution-preserving and undetectable watermarks: The paper treats these as possible directions for reducing external exposure and does not establish failure of their guarantees. Internal attribution capability and external learnability require separate evaluation.
- Governance insight: The paper notes that existing transparency frameworks already address purpose, retention, and access safeguards, rather than starting from an absence of governance. Its additional demand is explicit entity-linkability assessment and enforceable boundaries on continuing attribution and secondary use. This summarizes the authors' argument and is not independent legal advice.
Rating¶
- Novelty: 4/5. Unifies two observer models and deployment persistence as a governance issue without claiming the first discovery of watermark privacy risks.
- Experimental Thoroughness: 3/5. Controlled comparisons and internal coverage are valuable, while external coverage, scale, and transformation conditions remain limited.
- Writing Quality: 4/5. Clearly separates feasibility, hypothetical scenarios, and actual deployment evidence, while specifying the limits of per-key false-positive calibration.
- Value: 4/5. Adds entity-level privacy assessment to provenance systems while preserving legitimate auditing and accountability uses.