Skip to content

Witness Overlap: Directional Provenance Inside Open-Weight Model Families

Conference: NeurIPS 2026
arXiv: 2609.31784
Code: https://anonymous.4open.science/r/witness-llm-DCA3/
Area: Model Checkpoint Provenance / Interpretability / AI Safety
Keywords: directional provenance, witness checkpoint, update geometry, singular subspaces, lineage auditing

TL;DR

Within known same-family open-weight models with aligned parameters, Witness Overlap introduces a third checkpoint, compares each candidate endpoint's update overlap toward the target and witness, and predicts the lower-overlap endpoint as the parent; Frobenius cosine achieves 95.3% orientation accuracy over 1,542 single-witness triplets from 16 LLM families.

Background & Motivation

Open-weight models frequently undergo supervised fine-tuning, preference alignment, and re-release, so an auditor may need to determine not only whether two models are related but also which checkpoint preceded the other. Model cards and repository descriptions provide provenance clues, but their completeness depends on release practices. Existing representation fingerprints, weight distances, and behavioral tests can detect relatedness, yet often assume a known source or produce similarities unchanged by swapping the endpoints. Such scores alone cannot distinguish parent from child.

This paper separates family discovery from family-internal orientation and assumes that an upstream procedure has already supplied a same-family candidate set. Within this boundary, weight norm or kurtosis need not change monotonically during fine-tuning. The authors instead examine local angles between updates. Updates from a common parent to children fine-tuned on different tasks are often weakly correlated. Moving the anchor to one child makes the directions toward the parent and another child share a reverse component toward the parent. The geometry is consequently asymmetric across anchors.

Core Idea: use a third same-family checkpoint as a witness, replacing symmetric two-model similarity with anchor-conditioned update overlap and extracting directional lineage evidence from local branching geometry. The objects being analyzed are model checkpoints, not personal identities; the output supports research auditing rather than proving license violations or the actual training history.

Method

Overall Architecture

The inputs are an upstream same-family checkpoint set and corresponding two-dimensional dense weight matrices whose parameter coordinates can be matched across models. A query selects endpoints A and B plus another checkpoint W as a witness. It constructs two weight differences at each endpoint, computes matrix-level overlap, and aggregates scores by parameter block. The core consists of anchor overlap and task-specific decisions, rather than training a large lineage-recovery network.

The LLM experiments use dense projection blocks and exclude embeddings, the lm head, and normalization weights. VLM and diffusion experiments use the analogous shared projection blocks for each architecture. Equal parameter counts alone are insufficient: matrix names, shapes, and coordinates must align. If no additional same-family witness exists, the method should abstain rather than substitute an arbitrary model.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    I["Upstream same-family set<br/>Aligned 2D weights"] --> G["Anchor Overlap"]
    G --> A["Block Aggregation"]
    A --> D["Task Decisions"]
    D --> O["Root / edge direction<br/>Pair type / chain order"]
    T["Pair-type labels<br/>from training families"] -.-> C["Pair-Type Calibration"]
    A -.->|Training scores| C
    C -.->|Pair-type threshold only| D

Solid edges show inference data flow; dashed edges show pair-type classifier training and threshold calibration. Root identification, known parent–child orientation, and controlled single-chain ordering do not use this supervised branch. This does not make every experiment training-free.

Key Designs

1. Anchor Overlap: a third checkpoint makes endpoint swapping informative

Let \(W_A,W_B,W_W\) denote the corresponding matrix at the three checkpoints. With A as anchor, the updates are \(\Delta_{AB}=W_B-W_A\) and \(\Delta_{AW}=W_W-W_A\). Frobenius cosine uses the matrix inner product to measure alignment of the full update directions, with values in \([-1,1]\):

\[ \operatorname{overlap}_{\cos}^{(\ell)}(A;B,W)=\frac{\langle\Delta_{AB}^{(\ell)},\Delta_{AW}^{(\ell)}\rangle_F}{\|\Delta_{AB}^{(\ell)}\|_F\|\Delta_{AW}^{(\ell)}\|_F}. \]

Changing the anchor to B compares A−B with W−B, rather than merely exchanging the two arguments of a cosine. Thus, although cosine itself is symmetric, the two endpoint scores can differ. A parent whose outgoing updates are nearly independent has lower overlap. A child's two directions share a reverse-update component and have higher overlap. The evidence comes from update geometry, not release timestamps or model performance.

The second implementation applies truncated singular value decomposition (SVD) to each update matrix, retains orthonormal bases of the top right singular vectors with default \(k=16\), and compares the dominant subspaces:

\[ \operatorname{overlap}_{\mathrm{svd}}^{(\ell)}(A;B,W)=\frac{1}{k}\left\|V_{AW}^{(\ell)\top}V_{AB}^{(\ell)}\right\|_F^2. \]

This is the average squared cosine of the principal angles, with range \([0,1]\). It is insensitive to singular-vector signs and changes of basis within a subspace, and does not distinguish a matrix from its global negation at the subspace level. Directional evidence still arises from the difference structure induced by changing the anchor, not from singular-vector signs. Restricting attention to dominant update subspaces can suppress some fine-grained perturbations, but does not guarantee superiority over full cosine on every task.

2. Block Aggregation: average local scalars rather than computing cosine after concatenating all weights

Matrix scores are first averaged within each parameter-block role, then averaged across roles to obtain one triplet-level overlap. This block-aware aggregation allows different projection types to contribute evidence. It is neither a single cosine over concatenated layer updates nor a parameter-count-weighted whole-model cosine. Exchanging these operations changes the statistic used in the paper.

This distinction also limits the theoretical interpretation. The proposition uses flattened updates to explain anchor asymmetry, while practical decisions average across matrices and parameter blocks. A condition for one matrix cannot be treated as a proof of correct recovery for an entire family. Coordinate alignment is a prerequisite; architecture changes, reparameterization, or cross-family merging may make identically named matrices incomparable.

3. Task Decisions: compare endpoints, witness pairs, or chain scores according to the question

For A and B already known to form a direct parent–child pair, single-witness orientation compares \(s_A=\operatorname{overlap}(A;B,W)\) and \(s_B=\operatorname{overlap}(B;A,W)\) and predicts the lower-scoring endpoint as parent. The main experiment uses one witness per decision. Only the pooling ablation averages endpoint scores over multiple W checkpoints. This rule orients an edge; it does not first establish that an arbitrary pair is a direct edge.

Root identification uses a different aggregation. For every candidate M, it evaluates all unordered pairs of other checkpoints as target–witness pairs, averages the scores anchored at M, and selects the minimum. The two non-anchor roles are symmetric and all admissible combinations are used, rather than one randomly selected witness. Single-witness edge orientation and 100% root identification therefore refer to different task settings.

Pair-type discrimination uses the smaller endpoint score \(\min\{s_A,s_B\}\) to distinguish parent–child from sibling pairs. Parent–child pairs usually exhibit low–high scores, while siblings usually exhibit high–high scores, but a cross-family decision boundary must be learned. This covers the two relations defined in the paper, not arbitrary multigeneration ancestors, merge contributors, or general graph edges.

Single-chain ordering assumes a known base r. For each SFT checkpoint A, it averages \(\operatorname{overlap}_{\cos}(A;r,W)\) over other SFT checkpoints W and orders checkpoints by increasing average. Later checkpoints share more reverse-update components in their directions toward the base and earlier witnesses. Evidence comes from five controlled three-stage chains, not general directed acyclic graph recovery with an unknown base.

4. Pair-Type Calibration: learn the additional boundary only on training families

The actual pair-type experiment feeds the smaller endpoint score to logistic regression under leave-one-family-out (LOFO) evaluation. Training families provide parent–child/sibling labels and determine the threshold; the held-out family does not participate in calibration. The geometric score is training-free, but final pair-type classification is not, and its performance is substantially below orientation accuracy for known edges.

This separates two uses of the same feature: unlabeled endpoint ranking and supervised classification. Being prompt-free means the method does not need model-generated responses or query prompts. It does not eliminate the need to read weights, select aligned parameter blocks, construct same-family sets, or prepare calibration labels for pair typing.

A Worked Example

Consider an illustrative single-matrix family with parent p and children i and j whose nonzero updates have equal norms and are orthogonal. With p as anchor, the cosine toward i and j is 0. With i as anchor, the parent-directed difference is \(-u_i\) and the witness-directed difference is \(u_j-u_i\), giving cosine \(1/\sqrt{2}\approx0.7071\). For the known parent–child pair p and i with j as witness, the lower score of 0 predicts p as parent. Real models score and aggregate individual matrices; these illustrative values are not experimental measurements.

Proposition 4.1 and Appendix C.1 give the more general two-update explanation. Let \(\rho\) be the cosine between the parent-to-child updates and \(\lambda=\|u_i\|/\|u_j\|>0\). For nonzero, non-collinear updates:

\[ C(p;i,j)=\rho,\qquad C(i;p,j)=\frac{\lambda-\rho}{\sqrt{\lambda^2+1-2\lambda\rho}}. \]

Under these definitions, the paper's statement \(C(i;p,j)>C(p;i,j)\iff\rho<\lambda/2\) can be checked against the explicit expressions. The equal-norm boundary does not imply success for every positively correlated pair. If \(\rho<0\), the child score is positive and the parent score negative, so the inequality holds automatically. If \(\rho\ge0\), squaring requires preserving positivity of the numerator. The condition \(\lambda>2\rho\) itself ensures this, rather than permitting an unconditional squaring step.

For equal norms and \(\rho=0.75\), the child score is \(\sqrt{(1-0.75)/2}\approx0.3536\), lower than the parent score of 0.75; the equal-norm boundary is 0.5. Conversely, \(\rho=0.8,\lambda=1.9\) gives a child score of approximately 0.878, greater than 0.8 and consistent with the condition. This preserves the result within the specified update model without turning it into a necessary-and-sufficient guarantee for practical averages across parameter blocks.

Loss & Training

Witness Overlap does not update the audited models and has no training loss for root identification or known-edge orientation. Frobenius cosine and SVD scores come directly from checkpoint differences, with 16 dominant directions retained by default. The pair-type branch separately fits one-dimensional logistic regression and calibrates its threshold.

Fine-tuning in the controlled-chain experiment creates test objects with known order: one epoch each on MMLU, UltraChat, and PubMedQA produces three successive SFT checkpoints, using full-weight fine-tuning or LoRA. Training the test lineages must not be confused with training the provenance scorer.

Key Experimental Results

Main Results

The LLM collection contains 176 checkpoints across 16 families: 16 bases, 130 SFT descendants, and 30 RLHF/alignment descendants from more than 80 HuggingFace organizations. Labels come from model cards and repository descriptions. Direct-edge labels follow publicly declared metadata, rather than independently verified training logs. The full LLM evaluation takes approximately 8 GPU·h on one 48 GB L40S.

The following table selects root identification and single-witness orientation results from Tables 1–2. Root accuracy is measured by family. LLM orientation accuracy of 95.3% corresponds to 1,469/1,542 correct triplets, not 176 independent model decisions.

Method LLM root ID LLM SFT orientation LLM RLHF orientation LLM combined orientation VLM orientation Diffusion orientation
Witness Overlap, Frobenius cosine 93.8% 96.2% 91.4% 95.3% 100.0% 85.0%
Witness Overlap, SVD, \(k=16\) 100.0% 95.2% 86.4% 93.5% 87.5% 90.0%
L2 medoid 81.3% Not reported Not reported Not reported Not reported Not reported
Neural Phylogeny, norm rule 18.8% 79.2% 66.7% 76.9% 50.0% 50.0%
MoTHer, kurtosis for orientation 18.8% 79.2% 53.3% 74.4% 68.8% 70.0%

LLM combined orientation aggregates SFT and RLHF examples; it does not refer to weight-merged models. Both Witness metrics achieve 100.0% VLM root accuracy and 80.0% diffusion root accuracy. LLM root results are 16/16 for SVD and 15/16 for cosine. Scalar baselines do not use witnesses. Neural Phylogeny's supervised version is excluded from this training-free setting, so these results do not compare against all of its capabilities.

The source has a cross-domain counting inconsistency. The contribution paragraph reports 16 VLM checkpoints and 20 diffusion checkpoints, while Appendix N lists four VLM families and five diffusion families with one base and four descendants each, implying 20 and 25 total checkpoints. The earlier 16/20 counts match descendant counts. This note preserves the inconsistency rather than silently changing the experimental convention.

Ablation Study

Section 5.3 randomly samples witnesses before averaging endpoint scores. Its single-witness starting points differ from the main table; the original values are retained below and should not be forced into agreement or treated as the same sampled result.

Metric 1 witness 3 witnesses All witnesses Note
Frobenius cosine, LLM orientation 95.6% 97.9% 99.4% Different statistical setting from the main 95.3%
SVD, LLM orientation 94.1% 97.9% 100.0% Different statistical setting from the main 93.5%

Pair-type evaluation uses LOFO logistic regression, not the training-free orientation rule above. Table 3 reports:

Metric LLM AUROC LLM balanced accuracy VLM AUROC VLM balanced accuracy Diffusion AUROC Diffusion balanced accuracy
Frobenius cosine 0.898 82.4% 0.903 79.9% 0.855 80.3%
SVD, \(k=16\) 0.780 68.6% 0.715 54.9% 0.855 53.3%

Key Findings

  • SVD performs better for LLM root identification, while cosine performs better for LLM single-witness orientation and pair typing. Dominant subspaces are not universally superior. Single-witness failure counts are 73/1,542 for cosine and 100/1,542 for SVD; shared checkpoints create correlated triplets, not independent trials.
  • After averaging all witnesses per edge, Appendix J reports LLM parent/child cosine means of 0.0232/0.5527 and SVD means of 0.0725/0.3961. This separation supports the mechanism but cannot replace confidence calibration for an arbitrary new family.
  • All five controlled chains with known bases and three SFT stages are ordered correctly. This is chain-ordering evidence, not recovery of arbitrary trees, multiple bases, or merge nodes.
  • Perturbation results have clear boundaries: child-side noise is tested only on GPT-2 and TinyLlama. Under 10% pruning of all GPT-2 descendants, single-witness SVD/cosine accuracy is 77.8%/53.3%, and all-witness accuracy is 70.0%/50.0%. Pooling does not necessarily compensate for shared perturbations.
  • INT8 uses a specific per-tensor quantize–dequantize rule and a common recovered dense parameterization. Quantizing all checkpoints in three families preserves 3/3 roots and 37/37 all-witness edges for both metrics. This does not establish robustness to every quantization format; same-family merge tests likewise concern root-to-merge placement, not contributor recovery.

Highlights & Insights

  • A third witness changes the reference frame of a similarity problem. Common components in local updates provide asymmetric evidence without assigning every model a globally monotonic scalar.
  • Block averaging transfers the mechanism to shared projection spaces in different architectures. What transfers is the comparison of aligned updates, not permission to subtract arbitrary cross-architecture weights.
  • One geometric statistic supports several auditing questions, but their aggregation and supervision requirements differ. Separating feature construction from decision calibration prevents one accuracy result from being interpreted as complete lineage recovery.

Limitations & Future Work

  • The same-family set is an external input; family discovery is not solved here. A limited test with one contaminating member does not establish arbitrary contamination robustness. Missing witnesses, unaligned coordinates, or mixed origins call for abstention or explicit uncertainty.
  • Highly correlated updates or severe norm imbalance can reverse the geometry. The single-update proposition for nonzero, non-collinear vectors is not a global guarantee for block averages. Future work could study block reliability, abstention for close endpoint scores, and family-level calibration.
  • Most public-checkpoint experiments use metadata-declared base–descendant edges. Missing intermediate checkpoints can affect direct-edge labels. Audits should incorporate verifiable training records rather than deriving provenance facts or license liability directly from weight scores.
  • Complex branching, cross-family merging, contributor recovery, and adversarial changes are outside the validated scope. This note presents research-audit evidence, not procedures for altering checkpoints to conceal their origins.
  • Cross-domain samples are small and the main-text counts differ from the appendix convention. Main-table, witness-pooling, and pruning-subset results should not be merged into one universal accuracy. Broader multigeneration families and clearly defined statistical units would strengthen evaluation.
  • vs REEF, LLM DNA, and black-box provenance tests: these primarily identify related models and can support upstream family screening. Witness Overlap assumes the family is known and infers internal direction. The two stages should complement one another rather than letting orientation replace relatedness screening.
  • vs Neural Phylogeny: the comparison covers its training-free parameter-norm rule, not its supervised detector. Anchor-conditioned geometry reduces dependence on monotonic norm changes but requires an additional same-family witness and white-box weights.
  • vs MoTHer: MoTHer uses distances for edge placement and kurtosis for direction while recovering trees. This paper compares local update overlap and does not experimentally establish general tree recovery. Directional evidence and graph-structure search remain separable research problems.
  • Research implication: for model supply-chain audits with trustworthy family boundaries, geometric scores can accompany model-card evidence while separating orientation confidence, edge existence, and pair type. Reliability and abstention are useful next steps, not interpreting high scores as legal conclusions.

Rating

  • Novelty: 4/5 — Extracts direction from a third checkpoint's anchor-conditioned geometry rather than another model-level scalar.
  • Experimental Thoroughness: 4/5 — Sixteen LLM families, multiple tasks, and stress tests provide substantial coverage, but labels depend on metadata and cross-domain/complex-lineage coverage remains limited.
  • Writing Quality: 4/5 — Score definitions and task boundaries are clear, but cross-domain counts and single-witness sampling results need clearer explanation.
  • Value: 4/5 — Provides concise directional evidence for white-box model supply-chain research, not standalone proof of actual origin or license violations.