Back to the episode map

Research Note

AI Red Team Exercise Framework

An AI red-team exercise needs a named system version, owner, authorization, objective, threat model, assets, actors, access paths, exclusions, safe handling rules, stop c

Aug 4, 20261 min readBy Dalton Anderson

AI Red Team Exercise Framework

Exercise contract

An AI red-team exercise needs a named system version, owner, authorization, objective, threat model, assets, actors, access paths, exclusions, safe handling rules, stop conditions, evidence format, triage path, disclosure path, retest plan, and residual-risk owner.

Evidence record

FieldRequired record
SystemModel, prompts, retrieval, tools, identities, permissions, filters, runtime, and date
TestThreat, actor, precondition, input, context, settings, action, output, and observed impact
FindingReproduction, severity, affected boundary, uncertainty, and evidence location
TreatmentMitigation, control owner, expected effect, side effects, and deployment state
ClosureRetest result, regression case, remaining exposure, acceptance owner, and review trigger

Safety boundary

Tests that could create harmful instructions, sensitive data exposure, unauthorized access, unsafe code, or real-world impact need controlled environments and qualified review. Public editorial material should explain method and accountability without publishing operational bypass recipes.

Completion rule

Discovery is not completion. A finding becomes operationally useful when it is reproducible, owned, treated or explicitly accepted, covered by regression evidence, and connected to monitoring and incident response.

Sources

Follow the evidence.

  1. ai.meta.com: the llama 3 herd of modelsai.meta.com
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. owasp.org: www project top 10 for large language model applicationsowasp.org
  4. youtu.be: 1KNOcY e9Tsyoutu.be
  5. github.com: PurpleLlamagithub.com
  6. crfm.stanford.edu: indexcrfm.stanford.edu
  7. NIST AI Risk Management Frameworknist.gov
  8. mlcommons.org: jailbreak 0 7mlcommons.org
  9. mlcommons.org: safety faqmlcommons.org
  10. github.com: MODEL CARDgithub.com
  11. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  12. github.com: MODEL CARDgithub.com
  13. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  14. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  15. NIST Generative AI Profilenvlpubs.nist.gov
  16. mlcommons.org: safety methodologymlcommons.org
  17. huggingface.co: concept guidehuggingface.co
  18. github.com: MODEL CARDgithub.com
  19. csrc.nist.gov: red teamingcsrc.nist.gov
  20. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
AI Red Team Exercise Framework