Research Note

Open Model Threat-Control Framework

Begin with one use case and record users, identities, data, retrieval, model artifact, system prompts, tools, external services, outputs, human decisions, logging, operat

Aug 4, 20261 min readBy Dalton Anderson
In this article

Open Model Threat-Control Framework

System boundary

Begin with one use case and record users, identities, data, retrieval, model artifact, system prompts, tools, external services, outputs, human decisions, logging, operators, and downstream consumers.

Threat-control record

FieldRequired decision
ThreatActor, precondition, path, affected asset, and plausible impact
PreventionLeast privilege, isolation, validation, policy, limits, and approval
DetectionLogs, signals, thresholds, reviewer, and alert path
RecoveryDisablement, rollback, revocation, containment, notification, and restoration
EvidenceTest, version, result, false-positive and false-negative behavior, and limitation
OwnershipControl owner, risk owner, responder, approver, and review date

Layer rule

Model alignment and guard classifiers are controls inside a larger system. They cannot replace identity, authorization, data minimization, retrieval boundaries, tool permissions, output validation, human review, monitoring, incident response, or governance.

Completion rule

A control map is ready for accountable review when every material threat has prevention, detection, recovery, evidence, owners, residual risk, and a retest trigger. It is not a certification or a universal checklist.

Sources

Follow the evidence.

  1. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  4. ai.meta.com: the llama 3 herd of modelsai.meta.com
  5. crfm.stanford.edu: indexcrfm.stanford.edu
  6. csrc.nist.gov: red teamingcsrc.nist.gov
  7. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  8. github.com: MODEL CARDgithub.com
  9. github.com: MODEL CARDgithub.com
  10. github.com: PurpleLlamagithub.com
  11. github.com: MODEL CARDgithub.com
  12. huggingface.co: concept guidehuggingface.co
  13. mlcommons.org: safety faqmlcommons.org
  14. mlcommons.org: jailbreak 0 7mlcommons.org
  15. mlcommons.org: safety methodologymlcommons.org
  16. NIST Generative AI Profilenvlpubs.nist.gov
  17. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
  18. owasp.org: www project top 10 for large language model applicationsowasp.org
  19. NIST AI Risk Management Frameworknist.gov
  20. youtu.be: 1KNOcY e9Tsyoutu.be

From this episode

Two useful next steps.

Research Note · 1 min

Quantized Model Artifact Evaluation Framework

A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity,

Guide · 1 min

How to Design Layered Controls for an Open LLM

Secure an open-model system by mapping threats to controls across identity, data, retrieval, prompts, tools, outputs, monitoring, incident response, and ownership.

Return to the episode