Research Note
Training Data Disclosure Evidence Framework
Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time
Training Data Disclosure Evidence Framework
Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time cutoff, synthetic-data description, post-training description, or undisclosed.
Record the publisher, artifact, version, date, quotation or faithful paraphrase, scope, exclusions, uncertainty, and what cannot be inferred.
A phrase such as publicly available online data does not identify domains, documents, rights, geography, language distribution, persistence, or removals. A filtering description does not prove every prohibited or low-quality item was removed.
Use missing rather than estimated when the source does not disclose a fact.
Sources
Follow the evidence.
- Introducing Llama 3.1ai.meta.com
- ai.meta.com: the llama 3 herd of modelsai.meta.com
- arxiv.org: 1810arxiv.org
- crfm.stanford.edu: indexcrfm.stanford.edu
- arxiv.org: 2203arxiv.org
- open.spotify.com: 0iRBPcPw9iYjpUVAVWSkRCopen.spotify.com
- NIST AI Risk Management Frameworknist.gov
- github.com: MODEL CARDgithub.com
- daltonanderson.ghost.io: metas llama 3 1 inside the ai research paperdaltonanderson.ghost.io
- Meta Llama models repositorygithub.com
- arxiv.org: 2001arxiv.org
- youtu.be: UMhmWCor1kYyoutu.be
- github.com: LICENSEgithub.com