Research Note
Llama 3.1 Safety Tooling Source Record
| Artifact or practice | Source-era job | Evidence boundary | |---|---|---| | Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish appl
In this article
Llama 3.1 Safety Tooling Source Record
Source-era stack
| Artifact or practice | Source-era job | Evidence boundary |
|---|---|---|
| Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish application safety |
| Llama Guard 3 | Input and output content classification | Classifier scope, taxonomy, language, version, and threshold matter |
| Prompt Guard | Prompt-injection and jailbreak classification | Does not cover every instruction or system path |
| Code Shield | Inference-time insecure-code filtering and execution protections | Applies to code-oriented risks and integration choices |
| CyberSecEval 3 | Cybersecurity evaluation resources | Benchmark results remain versioned and bounded |
| Internal and external red teams | Adversarial discovery | Coverage is not completeness |
| Uplift studies | Relative capability comparison | Results remain tied to population, baseline, task, access, and metric |
Current lineage
The official Purple Llama repository currently includes Llama Guard 4, Llama Guard 3 variants, Llama Prompt Guard 2, Code Shield, and cybersecurity benchmarks. The current Llama 4 model card continues to recommend system-level protections and use-case-specific evaluation.
This lineage is successor context, not evidence that a Llama 3.1 deployment automatically receives later safeguards.
Required living fields
Every public tooling entry should record artifact name, exact version, owner, repository or model card, license, input, output, taxonomy, supported modalities, supported languages, intended placement, configuration, measured results, known limits, successor, last verification date, and system owner.
Boundary
Publisher material establishes intended use and reported results. It does not prove control effectiveness inside another architecture, against a later threat, or under another language, modality, policy, threshold, runtime, or operating environment.
Sources
Follow the evidence.
- ai-challenges.nist.gov: ariaai-challenges.nist.gov
- ai-challenges.nist.gov: genaiai-challenges.nist.gov
- ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
- ai.meta.com: the llama 3 herd of modelsai.meta.com
- crfm.stanford.edu: indexcrfm.stanford.edu
- csrc.nist.gov: red teamingcsrc.nist.gov
- daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
- github.com: MODEL CARDgithub.com
- github.com: MODEL CARDgithub.com
- github.com: PurpleLlamagithub.com
- github.com: MODEL CARDgithub.com
- huggingface.co: concept guidehuggingface.co
- mlcommons.org: safety faqmlcommons.org
- mlcommons.org: jailbreak 0 7mlcommons.org
- mlcommons.org: safety methodologymlcommons.org
- NIST Generative AI Profilenvlpubs.nist.gov
- open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
- owasp.org: www project top 10 for large language model applicationsowasp.org
- NIST AI Risk Management Frameworknist.gov
- youtu.be: 1KNOcY e9Tsyoutu.be