NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
A position paper submitted to arXiv argues that the Conference on Neural Information Processing Systems (NeurIPS) must enforce strict reproducibility standards for papers making frontier AI safety claims. The authors highlight an 'evidential inversion' where the most consequential safety assertions regarding model deployment and governance are often the least reproducible due to withheld evaluation artifacts. Citing the 2026 International AI Safety Report and the 2025 Foundation Model Transparency Index, the paper notes that reliable pre-deployment testing is increasingly difficult as models distinguish between test and deployment contexts. To address this, the authors propose a three-tier disclosure framework distinguishing public, controlled, and claim-restricted access. This system includes mandatory claim inventories and utilizes a federated colloquium of secure-review hosts for confidential assessments. The proposal treats non-reproducibility as a methodological failure rather than a transparency preference, aiming to ensure that scrutiny for high-stakes AI claims matches the rigor applied to less significant research, thereby enhancing public trust and scientific integrity in AI safety evaluations.
Wire timeline
NeurIPS Should Require Reproducibility Standards for Frontier AI Safety Claims
A position paper submitted to arXiv argues that the Conference on Neural Information Processing Systems (NeurIPS) must enforce strict reproducibility standards for papers making frontier AI safety claims. The authors highlight an 'evidential inversion' where the most consequential safety assertions regarding model deployment and governance are often the least reproducible due to withheld evaluation artifacts. Citing the 2026 International AI Safety Report and the 2025 Foundation Model Transparency Index, the paper notes that reliable pre-deployment testing is increasingly difficult as models distinguish between test and deployment contexts. To address this, the authors propose a three-tier disclosure framework distinguishing public, controlled, and claim-restricted access. This system includes mandatory claim inventories and utilizes a federated colloquium of secure-review hosts for confidential assessments. The proposal treats non-reproducibility as a methodological failure rather than a transparency preference, aiming to ensure that scrutiny for high-stakes AI claims matches the rigor applied to less significant research, thereby enhancing public trust and scientific integrity in AI safety evaluations.
cs.AI updates on arXiv.org