Back to the episode map

Research Note

Llama 3.1 Safety Tooling Source Record

| Artifact or practice | Source-era job | Evidence boundary | |---|---|---| | Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish appl

Aug 4, 20262 min readBy Dalton Anderson

Llama 3.1 Safety Tooling Source Record

Source-era stack

Artifact or practiceSource-era jobEvidence boundary
Llama 3.1 Instruct alignmentGeneral model behavior and refusalsDoes not establish application safety
Llama Guard 3Input and output content classificationClassifier scope, taxonomy, language, version, and threshold matter
Prompt GuardPrompt-injection and jailbreak classificationDoes not cover every instruction or system path
Code ShieldInference-time insecure-code filtering and execution protectionsApplies to code-oriented risks and integration choices
CyberSecEval 3Cybersecurity evaluation resourcesBenchmark results remain versioned and bounded
Internal and external red teamsAdversarial discoveryCoverage is not completeness
Uplift studiesRelative capability comparisonResults remain tied to population, baseline, task, access, and metric

Current lineage

The official Purple Llama repository currently includes Llama Guard 4, Llama Guard 3 variants, Llama Prompt Guard 2, Code Shield, and cybersecurity benchmarks. The current Llama 4 model card continues to recommend system-level protections and use-case-specific evaluation.

This lineage is successor context, not evidence that a Llama 3.1 deployment automatically receives later safeguards.

Required living fields

Every public tooling entry should record artifact name, exact version, owner, repository or model card, license, input, output, taxonomy, supported modalities, supported languages, intended placement, configuration, measured results, known limits, successor, last verification date, and system owner.

Boundary

Publisher material establishes intended use and reported results. It does not prove control effectiveness inside another architecture, against a later threat, or under another language, modality, policy, threshold, runtime, or operating environment.

Sources

Follow the evidence.

  1. ai.meta.com: the llama 3 herd of modelsai.meta.com
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. owasp.org: www project top 10 for large language model applicationsowasp.org
  4. youtu.be: 1KNOcY e9Tsyoutu.be
  5. github.com: PurpleLlamagithub.com
  6. crfm.stanford.edu: indexcrfm.stanford.edu
  7. NIST AI Risk Management Frameworknist.gov
  8. mlcommons.org: jailbreak 0 7mlcommons.org
  9. mlcommons.org: safety faqmlcommons.org
  10. github.com: MODEL CARDgithub.com
  11. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  12. github.com: MODEL CARDgithub.com
  13. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  14. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  15. NIST Generative AI Profilenvlpubs.nist.gov
  16. mlcommons.org: safety methodologymlcommons.org
  17. huggingface.co: concept guidehuggingface.co
  18. github.com: MODEL CARDgithub.com
  19. csrc.nist.gov: red teamingcsrc.nist.gov
  20. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
Llama 3.1 Safety Tooling Source Record