Research Note

E029 Historical Paper Reading Record

E029 preserves Dalton's August 2024 reading of the first half of Meta's `The Llama 3 Herd of Models` paper. Its value is the learning process: model families, dense trans

Aug 4, 20261 min readBy Dalton Anderson
In this article

E029 Historical Paper Reading Record

E029 preserves Dalton's August 2024 reading of the first half of Meta's The Llama 3 Herd of Models paper. Its value is the learning process: model families, dense transformers, data preparation, scaling, pretraining, post-training, infrastructure, long context, and evaluation became a connected system rather than a model-size headline.

The transcript also contains corrections that public copy must surface. Llama 3.1 uses a custom Community License, not an unrestricted free-use or standard open-source software license. Publisher benchmarks are reported evidence, not independent superiority. Public descriptions of a data pipeline do not disclose the underlying corpus.

Current successors belong in a separate living release record. They do not rewrite the July 2024 artifact.

Sources

Follow the evidence.

  1. Introducing Llama 3.1ai.meta.com
  2. ai.meta.com: the llama 3 herd of modelsai.meta.com
  3. arxiv.org: 1810arxiv.org
  4. crfm.stanford.edu: indexcrfm.stanford.edu
  5. arxiv.org: 2203arxiv.org
  6. open.spotify.com: 0iRBPcPw9iYjpUVAVWSkRCopen.spotify.com
  7. NIST AI Risk Management Frameworknist.gov
  8. github.com: MODEL CARDgithub.com
  9. daltonanderson.ghost.io: metas llama 3 1 inside the ai research paperdaltonanderson.ghost.io
  10. Meta Llama models repositorygithub.com
  11. arxiv.org: 2001arxiv.org
  12. youtu.be: UMhmWCor1kYyoutu.be
  13. github.com: LICENSEgithub.com

From this episode

Two useful next steps.

Research Note · 1 min

Training Data Disclosure Evidence Framework

Classify each statement as exact item disclosure, named dataset, source category, proportion, collection method, filtering method, deduplication method, token count, time

Research Note · 1 min

Scaling Law Decision Framework

Scaling laws are empirical relationships estimated from a defined model family, dataset regime, metric, compute range, and training procedure. They can guide allocation a

Return to the episode