Research Note
Different AI Systems Evidence Matrix
| Question | Scientific prediction system | Multimodal assistant | |---|---|---| | Claimed job | Predict a defined scientific object or relationship | Help a person compl
Different AI Systems Evidence Matrix
| Question | Scientific prediction system | Multimodal assistant |
|---|---|---|
| Claimed job | Predict a defined scientific object or relationship | Help a person complete an interaction or knowledge task |
| Core evidence | Data boundary, comparator, metric, uncertainty, reproducibility, external validation | Representative tasks, error classes, latency, privacy, consent, recovery, accepted work |
| Typical overclaim | Benchmark becomes clinical or commercial outcome | Demonstration becomes universal availability or workflow reliability |
| Domain reviewer | Scientist with relevant subject and method expertise | Qualified task owner plus privacy, security, accessibility, and domain reviewers |
| Deployment evidence | Performance under intended scientific or regulated conditions | Complete human-system workflow under realistic conditions |
| Stop condition | Uncertain, non-generalizable, or decision-inadequate prediction | Unacceptable failure, review burden, data risk, or poor recovery |
The shared word "performance" hides incompatible claims. AlphaFold 3's paper evaluates structure-prediction tasks against defined datasets, metrics, and comparators. GPT-4o's launch materials combine model evaluations, product demonstrations, staged rollout, and vendor latency measurements.
The comparison can still be useful if it starts with the decision. A scientific claim needs evidence that the target, test set, metric, comparison, uncertainty, and generalization match the intended scientific use. An assistant claim needs evidence that the actual user can complete the actual task with acceptable error, review, data handling, latency, accessibility, and recovery.
NIST's AI RMF supports context-specific testing, evaluation, verification, and validation. The AlphaFold 3 paper and GPT-4o system card supply system-specific evidence. None of those sources certifies a deployment.
Sources
Follow the evidence.
- Google DeepMind about pagedeepmind.google
- NIST AI RMF Measure guidanceairc.nist.gov
- Google DeepMind AlphaFold 3 launchblog.google
- Google DeepMind AlphaFold pagedeepmind.google
- Isomorphic Labs company siteisomorphiclabs.com
- OpenAI about pageopenai.com
- Current GPT-4o API documentationdevelopers.openai.com
- GPT-4o system cardcdn.openai.com
- FDA machine-learning transparency principlesfda.gov
- AlphaFold 3 papernature.com
- GPT-4o ChatGPT retirementopenai.com
- GPT-4o launchopenai.com