Research Note

Training Data Disclosure Evidence Framework

Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time

Aug 4, 20261 min readBy Dalton Anderson
In this article

Training Data Disclosure Evidence Framework

Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time cutoff, synthetic-data description, post-training description, or undisclosed.

Record the publisher, artifact, version, date, quotation or faithful paraphrase, scope, exclusions, uncertainty, and what cannot be inferred.

A phrase such as publicly available online data does not identify domains, documents, rights, geography, language distribution, persistence, or removals. A filtering description does not prove every prohibited or low-quality item was removed.

Use missing rather than estimated when the source does not disclose a fact.

Sources

Follow the evidence.

  1. Introducing Llama 3.1ai.meta.com
  2. ai.meta.com: the llama 3 herd of modelsai.meta.com
  3. arxiv.org: 1810arxiv.org
  4. crfm.stanford.edu: indexcrfm.stanford.edu
  5. arxiv.org: 2203arxiv.org
  6. open.spotify.com: 0iRBPcPw9iYjpUVAVWSkRCopen.spotify.com
  7. NIST AI Risk Management Frameworknist.gov
  8. github.com: MODEL CARDgithub.com
  9. daltonanderson.ghost.io: metas llama 3 1 inside the ai research paperdaltonanderson.ghost.io
  10. Meta Llama models repositorygithub.com
  11. arxiv.org: 2001arxiv.org
  12. youtu.be: UMhmWCor1kYyoutu.be
  13. github.com: LICENSEgithub.com

From this episode

Two useful next steps.

Research Note · 1 min

Scaling Law Decision Framework

Scaling laws are empirical relationships estimated from a defined model family, dataset regime, metric, compute range, and training procedure. They can guide allocation a

Research Note · 1 min

Open Weight Deployment Evaluation Framework

Begin with one use case and record the model artifact and hash, tokenizer, prompt format, license, acceptable-use policy, source, hardware, runtime, precision, quantizati

Return to the episode