Research Note
AI Red Team Exercise Framework
An AI red-team exercise needs a named system version, owner, authorization, objective, threat model, assets, actors, access paths, exclusions, safe handling rules, stop c
AI Red Team Exercise Framework
Exercise contract
An AI red-team exercise needs a named system version, owner, authorization, objective, threat model, assets, actors, access paths, exclusions, safe handling rules, stop conditions, evidence format, triage path, disclosure path, retest plan, and residual-risk owner.
Evidence record
| Field | Required record |
|---|---|
| System | Model, prompts, retrieval, tools, identities, permissions, filters, runtime, and date |
| Test | Threat, actor, precondition, input, context, settings, action, output, and observed impact |
| Finding | Reproduction, severity, affected boundary, uncertainty, and evidence location |
| Treatment | Mitigation, control owner, expected effect, side effects, and deployment state |
| Closure | Retest result, regression case, remaining exposure, acceptance owner, and review trigger |
Safety boundary
Tests that could create harmful instructions, sensitive data exposure, unauthorized access, unsafe code, or real-world impact need controlled environments and qualified review. Public editorial material should explain method and accountability without publishing operational bypass recipes.
Completion rule
Discovery is not completion. A finding becomes operationally useful when it is reproducible, owned, treated or explicitly accepted, covered by regression evidence, and connected to monitoring and incident response.
Sources
Follow the evidence.
- ai.meta.com: the llama 3 herd of modelsai.meta.com
- ai-challenges.nist.gov: genaiai-challenges.nist.gov
- owasp.org: www project top 10 for large language model applicationsowasp.org
- youtu.be: 1KNOcY e9Tsyoutu.be
- github.com: PurpleLlamagithub.com
- crfm.stanford.edu: indexcrfm.stanford.edu
- NIST AI Risk Management Frameworknist.gov
- mlcommons.org: jailbreak 0 7mlcommons.org
- mlcommons.org: safety faqmlcommons.org
- github.com: MODEL CARDgithub.com
- ai-challenges.nist.gov: ariaai-challenges.nist.gov
- github.com: MODEL CARDgithub.com
- daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
- ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
- NIST Generative AI Profilenvlpubs.nist.gov
- mlcommons.org: safety methodologymlcommons.org
- huggingface.co: concept guidehuggingface.co
- github.com: MODEL CARDgithub.com
- csrc.nist.gov: red teamingcsrc.nist.gov
- open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com