Research Note

Scientific AI Claim Audit Protocol

A scientific AI claim should be rewritten before it is judged. Record the exact prediction target, intended use, input, output, data boundary, training cutoff, test set,

Aug 4, 20261 min readBy Dalton Anderson
In this article

Scientific AI Claim Audit Protocol

A scientific AI claim should be rewritten before it is judged. Record the exact prediction target, intended use, input, output, data boundary, training cutoff, test set, comparator, metric, uncertainty, selection rule, and downstream decision.

Next, distinguish internal performance from external validity. A held-out benchmark can show performance on that benchmark. It does not automatically show performance on a different population, instrument, laboratory, disease, workflow, or regulatory use.

Check whether the comparator had the same information and operating constraints. AlphaFold 3's paper carefully distinguishes blind protein-ligand prediction from methods given privileged pocket information. A headline that removes that boundary changes the claim.

Inspect how outputs were selected. The paper reports top confidence-ranked samples from repeated seeds and diffusion samples under stated conditions. That is relevant to reproducibility, compute, and what a user must do to obtain comparable results.

Then look for uncertainty, failure modes, and evidence outside the originating team. Record what remains untested. If the intended use affects patients, laboratory decisions, regulated development, or significant capital, require qualified domain review and applicable validation.

The final statement should be narrow enough to survive source inspection. It should say what the evidence supports, the conditions under which it supports it, and what decision it cannot yet justify.

Sources

Follow the evidence.

  1. Google DeepMind about pagedeepmind.google
  2. NIST AI RMF Measure guidanceairc.nist.gov
  3. Google DeepMind AlphaFold 3 launchblog.google
  4. Google DeepMind AlphaFold pagedeepmind.google
  5. Isomorphic Labs company siteisomorphiclabs.com
  6. OpenAI about pageopenai.com
  7. Current GPT-4o API documentationdevelopers.openai.com
  8. GPT-4o system cardcdn.openai.com
  9. FDA machine-learning transparency principlesfda.gov
  10. AlphaFold 3 papernature.com
  11. GPT-4o ChatGPT retirementopenai.com
  12. GPT-4o launchopenai.com

From this episode

Two useful next steps.

Evergreen · 1 min

Scientific AI and multimodal assistants require distinct evidence and workflow tests

An AI system should be evaluated against the job it performs, the evidence its output requires, and the harm created when that output is wrong.

Article · 1 min

OpenAI and the GPT-4o Launch: Company Role

A dated record of OpenAI's role in GPT-4o research, launch, ChatGPT rollout, safety reporting, retirement, and the separate API lifecycle.

Return to the episode