Research Note

Llama 3.1 Safety Tooling Source Record

| Artifact or practice | Source-era job | Evidence boundary | |---|---|---| | Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish appl

Aug 4, 20262 min readBy Dalton Anderson
In this article

Llama 3.1 Safety Tooling Source Record

Source-era stack

Artifact or practiceSource-era jobEvidence boundary
Llama 3.1 Instruct alignmentGeneral model behavior and refusalsDoes not establish application safety
Llama Guard 3Input and output content classificationClassifier scope, taxonomy, language, version, and threshold matter
Prompt GuardPrompt-injection and jailbreak classificationDoes not cover every instruction or system path
Code ShieldInference-time insecure-code filtering and execution protectionsApplies to code-oriented risks and integration choices
CyberSecEval 3Cybersecurity evaluation resourcesBenchmark results remain versioned and bounded
Internal and external red teamsAdversarial discoveryCoverage is not completeness
Uplift studiesRelative capability comparisonResults remain tied to population, baseline, task, access, and metric

Current lineage

The official Purple Llama repository currently includes Llama Guard 4, Llama Guard 3 variants, Llama Prompt Guard 2, Code Shield, and cybersecurity benchmarks. The current Llama 4 model card continues to recommend system-level protections and use-case-specific evaluation.

This lineage is successor context, not evidence that a Llama 3.1 deployment automatically receives later safeguards.

Required living fields

Every public tooling entry should record artifact name, exact version, owner, repository or model card, license, input, output, taxonomy, supported modalities, supported languages, intended placement, configuration, measured results, known limits, successor, last verification date, and system owner.

Boundary

Publisher material establishes intended use and reported results. It does not prove control effectiveness inside another architecture, against a later threat, or under another language, modality, policy, threshold, runtime, or operating environment.

Sources

Follow the evidence.

  1. ai.meta.com: the llama 3 herd of modelsai.meta.com
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. owasp.org: www project top 10 for large language model applicationsowasp.org
  4. youtu.be: 1KNOcY e9Tsyoutu.be
  5. github.com: PurpleLlamagithub.com
  6. crfm.stanford.edu: indexcrfm.stanford.edu
  7. NIST AI Risk Management Frameworknist.gov
  8. mlcommons.org: jailbreak 0 7mlcommons.org
  9. mlcommons.org: safety faqmlcommons.org
  10. github.com: MODEL CARDgithub.com
  11. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  12. github.com: MODEL CARDgithub.com
  13. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  14. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  15. NIST Generative AI Profilenvlpubs.nist.gov
  16. mlcommons.org: safety methodologymlcommons.org
  17. huggingface.co: concept guidehuggingface.co
  18. github.com: MODEL CARDgithub.com
  19. csrc.nist.gov: red teamingcsrc.nist.gov
  20. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com

From this episode

Two useful next steps.

Research Note · 1 min

Quantized Model Artifact Evaluation Framework

A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity,

Research Note · 1 min

Open Model Threat-Control Framework

Begin with one use case and record users, identities, data, retrieval, model artifact, system prompts, tools, external services, outputs, human decisions, logging, operat

Return to the episode
Llama 3.1 Safety Tooling Source Record