Research Note

Llama 3.1 Safety Tooling Source Record

| Artifact or practice | Source-era job | Evidence boundary | |---|---|---| | Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish appl

Aug 4, 20262 min readBy Dalton Anderson
In this article

Llama 3.1 Safety Tooling Source Record

Source-era stack

Artifact or practiceSource-era jobEvidence boundary
Llama 3.1 Instruct alignmentGeneral model behavior and refusalsDoes not establish application safety
Llama Guard 3Input and output content classificationClassifier scope, taxonomy, language, version, and threshold matter
Prompt GuardPrompt-injection and jailbreak classificationDoes not cover every instruction or system path
Code ShieldInference-time insecure-code filtering and execution protectionsApplies to code-oriented risks and integration choices
CyberSecEval 3Cybersecurity evaluation resourcesBenchmark results remain versioned and bounded
Internal and external red teamsAdversarial discoveryCoverage is not completeness
Uplift studiesRelative capability comparisonResults remain tied to population, baseline, task, access, and metric

Current lineage

The official Purple Llama repository currently includes Llama Guard 4, Llama Guard 3 variants, Llama Prompt Guard 2, Code Shield, and cybersecurity benchmarks. The current Llama 4 model card continues to recommend system-level protections and use-case-specific evaluation.

This lineage is successor context, not evidence that a Llama 3.1 deployment automatically receives later safeguards.

Required living fields

Every public tooling entry should record artifact name, exact version, owner, repository or model card, license, input, output, taxonomy, supported modalities, supported languages, intended placement, configuration, measured results, known limits, successor, last verification date, and system owner.

Boundary

Publisher material establishes intended use and reported results. It does not prove control effectiveness inside another architecture, against a later threat, or under another language, modality, policy, threshold, runtime, or operating environment.

Sources

Follow the evidence.

  1. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  4. ai.meta.com: the llama 3 herd of modelsai.meta.com
  5. crfm.stanford.edu: indexcrfm.stanford.edu
  6. csrc.nist.gov: red teamingcsrc.nist.gov
  7. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  8. github.com: MODEL CARDgithub.com
  9. github.com: MODEL CARDgithub.com
  10. github.com: PurpleLlamagithub.com
  11. github.com: MODEL CARDgithub.com
  12. huggingface.co: concept guidehuggingface.co
  13. mlcommons.org: safety faqmlcommons.org
  14. mlcommons.org: jailbreak 0 7mlcommons.org
  15. mlcommons.org: safety methodologymlcommons.org
  16. NIST Generative AI Profilenvlpubs.nist.gov
  17. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
  18. owasp.org: www project top 10 for large language model applicationsowasp.org
  19. NIST AI Risk Management Frameworknist.gov
  20. youtu.be: 1KNOcY e9Tsyoutu.be

From this episode

Two useful next steps.

Research Note · 1 min

Quantized Model Artifact Evaluation Framework

A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity,

Research Note · 1 min

Open Model Threat-Control Framework

Begin with one use case and record users, identities, data, retrieval, model artifact, system prompts, tools, external services, outputs, human decisions, logging, operat

Return to the episode