Research Note
Llama 3.1 Safety Tooling Source Record
| Artifact or practice | Source-era job | Evidence boundary | |---|---|---| | Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish appl
Llama 3.1 Safety Tooling Source Record
Source-era stack
| Artifact or practice | Source-era job | Evidence boundary |
|---|---|---|
| Llama 3.1 Instruct alignment | General model behavior and refusals | Does not establish application safety |
| Llama Guard 3 | Input and output content classification | Classifier scope, taxonomy, language, version, and threshold matter |
| Prompt Guard | Prompt-injection and jailbreak classification | Does not cover every instruction or system path |
| Code Shield | Inference-time insecure-code filtering and execution protections | Applies to code-oriented risks and integration choices |
| CyberSecEval 3 | Cybersecurity evaluation resources | Benchmark results remain versioned and bounded |
| Internal and external red teams | Adversarial discovery | Coverage is not completeness |
| Uplift studies | Relative capability comparison | Results remain tied to population, baseline, task, access, and metric |
Current lineage
The official Purple Llama repository currently includes Llama Guard 4, Llama Guard 3 variants, Llama Prompt Guard 2, Code Shield, and cybersecurity benchmarks. The current Llama 4 model card continues to recommend system-level protections and use-case-specific evaluation.
This lineage is successor context, not evidence that a Llama 3.1 deployment automatically receives later safeguards.
Required living fields
Every public tooling entry should record artifact name, exact version, owner, repository or model card, license, input, output, taxonomy, supported modalities, supported languages, intended placement, configuration, measured results, known limits, successor, last verification date, and system owner.
Boundary
Publisher material establishes intended use and reported results. It does not prove control effectiveness inside another architecture, against a later threat, or under another language, modality, policy, threshold, runtime, or operating environment.
Sources
Follow the evidence.
- ai.meta.com: the llama 3 herd of modelsai.meta.com
- ai-challenges.nist.gov: genaiai-challenges.nist.gov
- owasp.org: www project top 10 for large language model applicationsowasp.org
- youtu.be: 1KNOcY e9Tsyoutu.be
- github.com: PurpleLlamagithub.com
- crfm.stanford.edu: indexcrfm.stanford.edu
- NIST AI Risk Management Frameworknist.gov
- mlcommons.org: jailbreak 0 7mlcommons.org
- mlcommons.org: safety faqmlcommons.org
- github.com: MODEL CARDgithub.com
- ai-challenges.nist.gov: ariaai-challenges.nist.gov
- github.com: MODEL CARDgithub.com
- daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
- ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
- NIST Generative AI Profilenvlpubs.nist.gov
- mlcommons.org: safety methodologymlcommons.org
- huggingface.co: concept guidehuggingface.co
- github.com: MODEL CARDgithub.com
- csrc.nist.gov: red teamingcsrc.nist.gov
- open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com