Research Note

AI Red Team Exercise Framework

An AI red-team exercise needs a named system version, owner, authorization, objective, threat model, assets, actors, access paths, exclusions, safe handling rules, stop c

Aug 4, 20261 min readBy Dalton Anderson
In this article

AI Red Team Exercise Framework

Exercise contract

An AI red-team exercise needs a named system version, owner, authorization, objective, threat model, assets, actors, access paths, exclusions, safe handling rules, stop conditions, evidence format, triage path, disclosure path, retest plan, and residual-risk owner.

Evidence record

FieldRequired record
SystemModel, prompts, retrieval, tools, identities, permissions, filters, runtime, and date
TestThreat, actor, precondition, input, context, settings, action, output, and observed impact
FindingReproduction, severity, affected boundary, uncertainty, and evidence location
TreatmentMitigation, control owner, expected effect, side effects, and deployment state
ClosureRetest result, regression case, remaining exposure, acceptance owner, and review trigger

Safety boundary

Tests that could create harmful instructions, sensitive data exposure, unauthorized access, unsafe code, or real-world impact need controlled environments and qualified review. Public editorial material should explain method and accountability without publishing operational bypass recipes.

Completion rule

Discovery is not completion. A finding becomes operationally useful when it is reproducible, owned, treated or explicitly accepted, covered by regression evidence, and connected to monitoring and incident response.

Sources

Follow the evidence.

  1. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  4. ai.meta.com: the llama 3 herd of modelsai.meta.com
  5. crfm.stanford.edu: indexcrfm.stanford.edu
  6. csrc.nist.gov: red teamingcsrc.nist.gov
  7. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  8. github.com: MODEL CARDgithub.com
  9. github.com: MODEL CARDgithub.com
  10. github.com: PurpleLlamagithub.com
  11. github.com: MODEL CARDgithub.com
  12. huggingface.co: concept guidehuggingface.co
  13. mlcommons.org: safety faqmlcommons.org
  14. mlcommons.org: jailbreak 0 7mlcommons.org
  15. mlcommons.org: safety methodologymlcommons.org
  16. NIST Generative AI Profilenvlpubs.nist.gov
  17. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
  18. owasp.org: www project top 10 for large language model applicationsowasp.org
  19. NIST AI Risk Management Frameworknist.gov
  20. youtu.be: 1KNOcY e9Tsyoutu.be

From this episode

Two useful next steps.

Research Note · 1 min

Quantized Model Artifact Evaluation Framework

A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity,

Research Note · 1 min

Open Model Threat-Control Framework

Begin with one use case and record users, identities, data, retrieval, model artifact, system prompts, tools, external services, outputs, human decisions, logging, operat

Return to the episode