Skip to content

Position: Let’s Strengthen Verifiability If We Can’t Enforce Reproducibility

Conference: NeurIPS 2026 — Position Paper Track (a position paper, not a regular main-track method paper)
arXiv: 2609.35854
Code: https://github.com/giddyyupp/position-enforce-verifiability
Area: Others (research reproducibility and peer-review policy)
Keywords: verifiability, reproducibility, experiment logs, metric verification, peer review

TL;DR

Drawing on a code-availability survey of five leading ML/CV conferences from 2021–2025, this paper proposes integrating experiment-log checks and metric recomputation from prediction files into submission, review, and publication when full reproduction cannot be enforced; it is a proposal for partial consistency verification, not a demonstrated fraud-detection system.

Background & Motivation

Public code is often considered the entry point to reproduction in computer science, but a paper containing a link, a repository being cloneable, training or test code being present, and the reported results being reproducible are four different conditions. Model weights let outsiders run a trained model without revealing which data actually entered training. Even with source code, undisclosed test-time augmentation, ensembling, checkpoint selection, and expensive computation can create gaps between the paper and its execution. The authors are concerned not with issuing another code-release appeal, but with making such gaps visible before acceptance.

The survey finds declining code-sharing proportions in some conferences as early as 2023 and a broader decline in 2025; meanwhile, papers with official code receive more than twice the average citations of papers without code. This association has neither made code release universal nor established that releasing code causes more citations. Waiting for motivated researchers to reproduce a published paper requires additional time and compute, while unverifiable SOTA results may remain barriers against which subsequent work is judged.

Full code review offers stronger assurances but faces commercial confidentiality, licensing, privacy, and limited reviewer resources. The paper therefore narrows the object of review: instead of rerunning training and inference, reviewers inspect submitted experiment traces and the computation from fixed predictions to metrics. Core idea: make consistency between experimental material and the paper a formal status checked before acceptance and visible after publication, while explicitly distinguishing it from full reproducibility or result authenticity.

Method

Overall Architecture

The method here is a conference verification mechanism, not a neural architecture. Authors submit logs, prediction files, evaluation code, and necessary ground truth for their main experiments; reviewers inspect logs at Level 1 and inspect and execute evaluation code at Level 2, after which area chairs (ACs) and program chairs (PCs) consolidate consistency statuses. Following acceptance, an optional software link is checked, and statuses and material are displayed on the publication page.

Log verification, metric verification, status governance, and software verification form four key designs. The first two inspect different evidence: logs expose visible traces of training and execution, whereas metric verification checks the computational path from prediction outputs to reported numbers. A venue can introduce Level 1 first and add Level 2 later; the diagram's serial order represents this staged policy and report consolidation, not a requirement that logs must pass before metrics can be recomputed.

%%{init: {'flowchart': {'rankSpacing': 24, 'nodeSpacing': 28, 'padding': 6, 'wrappingWidth': 400}}}%%
flowchart TD
    A["Author submission<br/>main experiments and material"] --> B["Log verification<br/>reviewer checks Level 1"]
    B --> C["Metric verification<br/>assigned reviewer runs Level 2"]
    C --> D["Status governance<br/>AC synthesis and PC decision"]
    D -->|after acceptance| E["Software verification<br/>optional repository and publication"]
    E --> F["Readers see statuses and material<br/>not full reproduction certification"]

The mechanism has no model-training objective, inference-time data flow, or loss function. InternLM identifies code links and analyzes repositories in the survey; it is not a mandatory automated verification model for reviewers. The proposed secure LLM auditing platform is also only a future direction.

Key Designs

1. Log verification: trace each main empirical claim to a run

A final metric usually reveals neither its training run nor how much data was used or how the model was selected. The authors require existing logs for experiments supporting main empirical claims, explicitly connecting curves or records to a specific experiment in the paper, such as a row in a SOTA table; ablation logs are optional. Rather than making reviewers infer the correspondence from assorted screenshots, authors first establish a claim-to-run evidence link.

Appendix B.1 specifies shared fields: the experiment reference, dataset size, GPU/CPU utilization, parameter count, and FLOPs or a comparable compute estimate. Training logs additionally report the learning scheme, training parameters, and losses over steps or epochs; validation logs report final evaluation metrics. TensorBoard, MLflow, and Weights & Biases can export this material, so the proposal reuses development records rather than introducing a proprietary logging system.

Reviewers assess whether the training procedure, convergence, and performance variation match the paper, using epochs, batch size, and iteration counts to look for signs of additional data, and requesting clarification during rebuttal when necessary. Logs may also reveal checkpoint-selection issues, but these remain consistency checks on visible evidence: logs can be forged, and plausible curves do not exclude data leakage or recovering answers from sample identifiers.

2. Metric verification: recompute output-to-metric results, not training or output generation

Releasing a complete method implementation can be costly, whereas evaluation code is often much smaller. The proposal limits verification to outputs the authors have already generated: authors provide JSON predictions for the corresponding experiment, metric definitions or links to standard evaluation code, necessary ground truth, and download and execution instructions. An original metric requires anonymized evaluation code and, when needed, the model used for evaluation. An assigned reviewer inspects sample coverage, ground truth, and code validity, then executes evaluation and compares against the paper.

The authors recommend an anonymous Colab-like notebook packaging outputs, ground truth or its download steps, and evaluation code to reduce installation barriers. Independent verification allows acceptable stochastic variation in metric computation; it seeks sufficient agreement rather than mechanically identical floating-point values. A small main-text study with five experienced reviewers reports about 5 minutes on average for the short script and 45 minutes to 1 hour for the long script. This sample does not establish low review costs for every paper.

This separation can expose metric implementation or reporting errors and allow checks for omitted difficult samples, but cannot establish that the submitted predictions came from the claimed model. Sensitive outputs or private ground truth require separate treatment of access and anonymity: the paper specifies that proprietary data should not require a signed license agreement during verification, because this could reveal author or reviewer identities. Cases that cannot meet these requirements should receive a justified exemption, not have unsuccessful evaluation treated as a pass.

3. Status governance: incorporate verification into formal review, not an additional honorary badge

After authors submit material, reviewers' consistency comments and ratings become part of the formal review and affect recommendations. ACs summarize evidence and explain their final consistency recommendation in the meta-review; PCs assign a status accordingly. A venue may appoint a single experiment or metric reviewer per paper to avoid redundant execution. Unlike a reproducibility badge awarded only after acceptance, this places verification before acceptance, without stipulating automatic rejection of every paper lacking material.

Table 1 separately records logs, metrics, and software rather than compressing them into a universal trust score. Logs can be consistent, inconclusive, inconsistent, or unavailable; metrics can be similar enough, questionably different, definitely different, or unavailable; software can have a repository link, be proprietary, or be unmaintained. These are categorical statuses, without a continuous scoring formula or a universal numerical tolerance across tasks.

Theoretical papers, hidden-test servers, private data, and excessively large outputs may receive exemptions with reasons assessed by reviewers and ACs. The authors also encourage subsequent papers to distinguish comparisons with verified work or reproducible software, but do not forbid citing codeless papers. Older papers default to lacking verification status; unverified does not mean unreliable, a particularly important boundary when applying new labels to prior work.

4. Software verification: check actual repositories after acceptance rather than reward release promises

A promise that code is coming soon does not establish that software will exist when a conference opens. The proposal asks authors to declare software status and an optional repository link before the camera-ready deadline, then has an assigned reviewer check for the expected software before the conference. Confirmed links appear on the proceedings page. Automated repository checks can assist, with human follow-up for failures, but the process does not inspect implementation completeness or require full training execution.

Publication information includes experimental consistency ratings, software availability statuses, and experimental material, allowing readers to judge the evidence behind comparisons. Software access may involve a license agreement at this stage, unlike the earlier anonymous metric-verification stage. Repository existence still does not guarantee complete data, training procedures, augmentation, or ensembling configurations, so this check must not be presented as comprehensive code certification.

A Worked Example

The following is an illustrative scenario constructed from the proposal, not an additional experiment or an implemented conference case in the paper.

Suppose a paper supports a performance claim with a main result on a public dataset. Authors connect that result to a specific run log, submit prediction JSON for all target samples, and construct an anonymous notebook using standard evaluation tools. A reviewer first inspects curves, data quantities, and training configuration, then checks prediction sample coverage and ground-truth provenance, and finally executes the notebook.

If recomputation agrees sufficiently with the reported number, the metric status can be similar enough; if the logs cannot establish the training procedure, the log status remains inconclusive. The AC must explain these judgments separately, and the PC incorporates them into the final decision rather than letting metric agreement conceal insufficient log evidence.

After acceptance, if authors supply a repository link, a checker confirms that software is present before the corresponding status becomes public. Readers receive a limited conclusion such as these files reproduce the reported metric, not a guarantee that the model was trained on appropriate data or that its outputs were never fabricated.

Key Experimental Results

Main Results

The paper introduces no trained model and does not compare accuracy against SOTA methods. The table summarizes the code-availability survey in Appendix Tables 2–6: availability means detected training or test code, not successful reproduction.

Conference and comparison years Main-track accepted papers: earlier year → 2025 Training or test code available: earlier year → 2025 Interpretation
CVPR, 2024 → 2025 2719 → 2872 1255 → 1267 Code counts rise slightly, but acceptances grow faster
ICCV, 2023 → 2025 2160 → 2701 1035 → 1108 Biennial conference; not a year-on-year 2024–2025 comparison
ICLR, 2024 → 2025 2296 → 3827 1199 → 1964 Higher absolute counts do not imply a higher proportion
ICML, 2024 → 2025 2634 → 3339 1150 → 1485 The proportion does not decline between these years; the general narrative cannot be applied to every entry
NeurIPS, 2024 → 2025 4037 → 5290 2056 → 2608 Acceptances grow faster than code counts

Numbers are retained from the original tables. In particular, ICML's 1150/2634 and 1485/3339 indicate a slightly higher proportion in 2025, whereas the main text claims a marked decline for every conference in 2025. This note preserves the conflict rather than correcting the authors' data. The paper also explicitly notes discrepancies between acceptance and successful-download counts, so different denominators affect proportions.

Ablation Study

There is no standard model ablation. The table collects survey analyses and verification-burden evidence, without treating small timing samples or experiential observations as controlled tests of policy effectiveness.

Evidence item Original number or conclusion What it supports; what it does not
Automated survey scale 55,377 papers; 100 papers used for prompt comparison Broad coverage; no full manual labels or meaningful error bars
Code and citations, Figure 3 Papers with official code receive more than twice the average citations of papers without code Association, not a causal benefit of code release
Acknowledged repositories, Figure 1(c) The most acknowledged 20% of repositories account for 81.5% of acknowledgements Concentrated reuse, not 81.5% of performance improvements
Metric-verification user study 5 experienced reviewers; short script about 5 minutes; long script 45 minutes–1 hour Feasibility for two scripts, not total review costs for all tasks
AC experience, Appendix C.1 Of 74 submissions, 16 included code; among the associated 48 reviews, only 2 mentioned code Code submission is not actual review; incomplete and topically biased sample
Reproducibility-related issues in public repositories, Appendix C.2 2023: 1.9%; 2024: 2.3%; 2025: 4.0% The paper's issue-based signal, not an audited reproduction-failure rate

Key Findings

  • The survey extracts links from official paper PDFs and supplementary material, resolves project pages, clones repositories, and assesses training code, test code, and weights. Abstract-only checks or counts of GitHub links can miss artifacts or count empty repositories as code.
  • Prompts for InternLM (internlm3-8b-instruct) were adjusted to match human judgments on 100 randomly selected papers. The authors report a slight residual tendency to overestimate code and weight availability; this is not an independent held-out test or a quantified full-survey classification error.
  • The appendix says repository acknowledgements increase every year for every venue, but NeurIPS Table 6 reports 765 → 758 from 2024 → 2025. The table should be retained rather than repeating an exception-free growth claim.
  • Verification saves resources by not rerunning training and inference. Whether it reduces erroneous publications, improves reviewer agreement, or changes author behavior still requires a conference pilot.

Highlights & Insights

  • The proposal divides reproduction-related evidence into levels rather than choosing between full open sourcing and no checking. Logs and metrics each have explicit verification targets and retain separate gaps in authenticity.
  • Reviewing a minimal computational chain is transferable: when generating outputs is costly but metric computation is relatively light, a venue can verify results on fixed outputs without claiming to reproduce their generation.
  • Publishing statuses alongside papers helps readers distinguish evidence strength across comparisons. Useful labels require explicit scope, exemption reasons, and inconclusive statuses rather than a generalized reputation score.

Limitations & Future Work

  • The authors acknowledge that logs and predictions can be forged; metric agreement cannot exclude training leakage, lookup outputs, or inconsistencies between the method and its implementation.
  • The scope mainly covers repeatable virtual experiments with mathematically defined metrics, not directly physical robot operations or subjective human evaluations.
  • Timing two script types with five reviewers does not cover complicated dependencies, enormous outputs, multiple datasets, or costly evaluation models. The claim that verification generally takes less than one hour is primarily the authors' burden estimate.
  • The survey depends on link extraction, repository retrieval, and LLM judgments, while papers from different years have had different amounts of time to release artifacts. Without independent manual auditing and error bars, changes in code availability cannot be equated with changes in actual reproduction success.
  • Public statuses may unfairly penalize unverified older papers or justified exemptions. A pilot should track exemption distributions, actual verification time, reviewer agreement, and error types, rather than only the fraction receiving labels.
  • A secure third-party execution platform and LLM code auditing are future proposals involving hosting costs, confidentiality, and operations; the paper does not demonstrate an available platform.
  • vs NeurIPS reproducibility checklists and challenges: checklists organize author disclosure, and challenges rely on independent reproduction; this paper requires review-time inspection of logs and metrics but provides weaker guarantees.
  • vs JAIR checklists, structured abstracts, badges, and reproducibility reports: the emphasis here is consistency verification before acceptance and its role in formal decisions, without equating artifact release with verification.
  • vs ICPR's RRPR Badge: the paper describes evaluation by dedicated reviewers after acceptance, without affecting acceptance decisions; this proposal moves lightweight verification into pre-acceptance review.
  • vs Paper2Code: Paper2Code investigates generating implementations from papers and assessing paper–code consistency. This paper proposes conference governance, with a survey that additionally examines full-text links and actual repositories, not an automated reproduction system.

Rating

  • Novelty: 4/5 — Connects lightweight evidence verification, formal decisions, and publication statuses with explicit policy granularity.
  • Experimental Thoroughness: 2/5 — Valuable large descriptive survey, but narrative conflicts, small burden samples, and no policy pilot remain.
  • Writing Quality: 4/5 — Roles, materials, and boundaries are clear; some general trend statements conflict with appendix tables.
  • Value: 4/5 — A pilotable option when conferences cannot require full open sourcing; its value depends on implementation costs and label governance.