Back to the episode map

Research Note

Training Data Disclosure Evidence Framework

Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time

Aug 4, 20261 min readBy Dalton Anderson

Training Data Disclosure Evidence Framework

Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time cutoff, synthetic-data description, post-training description, or undisclosed.

Record the publisher, artifact, version, date, quotation or faithful paraphrase, scope, exclusions, uncertainty, and what cannot be inferred.

A phrase such as publicly available online data does not identify domains, documents, rights, geography, language distribution, persistence, or removals. A filtering description does not prove every prohibited or low-quality item was removed.

Use missing rather than estimated when the source does not disclose a fact.

Sources

Follow the evidence.

  1. Introducing Llama 3.1ai.meta.com
  2. ai.meta.com: the llama 3 herd of modelsai.meta.com
  3. arxiv.org: 1810arxiv.org
  4. crfm.stanford.edu: indexcrfm.stanford.edu
  5. arxiv.org: 2203arxiv.org
  6. open.spotify.com: 0iRBPcPw9iYjpUVAVWSkRCopen.spotify.com
  7. NIST AI Risk Management Frameworknist.gov
  8. github.com: MODEL CARDgithub.com
  9. daltonanderson.ghost.io: metas llama 3 1 inside the ai research paperdaltonanderson.ghost.io
  10. Meta Llama models repositorygithub.com
  11. arxiv.org: 2001arxiv.org
  12. youtu.be: UMhmWCor1kYyoutu.be
  13. github.com: LICENSEgithub.com
Training Data Disclosure Evidence Framework