Research Note

Multimodal Claim Evidence Ladder

| State | What it establishes | |---|---| | Concept | A proposed architecture, task, or research direction | | Internal experiment | A tested setup under reported conditi

Aug 4, 20261 min readBy Dalton Anderson
In this article

Multimodal Claim Evidence Ladder

Evidence states

StateWhat it establishes
ConceptA proposed architecture, task, or research direction
Internal experimentA tested setup under reported conditions
Benchmark resultA measured result for a named artifact, dataset, metric, and baseline
DemonstrationA selected interaction or prototype path
Released artifactWeights, code, or another reproducible object with terms
Accessible interfaceAn API or product surface available to a defined audience
Supported capabilityA documented, maintained feature with operational boundaries

Claim record

Record exact input and output modalities, task, model path, component versions, training and evaluation data, metric, comparator, limitation, artifact identity, access, terms, product surface, support state, geography, account, and verification date.

Historical application

The Llama 3 paper reported compositional image, video, and speech experiments. Its research page explicitly said the resulting models were still under development and not broadly released. The July 2024 Llama 3.1 model card identified the released family as text in and text out.

Later multimodal Llama releases are descendant evidence. They do not make the 2024 text artifacts multimodal retroactively.

Sources

Follow the evidence.

  1. ai-challenges.nist.gov: ariaai-challenges.nist.gov
  2. ai-challenges.nist.gov: genaiai-challenges.nist.gov
  3. ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
  4. ai.meta.com: the llama 3 herd of modelsai.meta.com
  5. crfm.stanford.edu: indexcrfm.stanford.edu
  6. csrc.nist.gov: red teamingcsrc.nist.gov
  7. daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
  8. github.com: MODEL CARDgithub.com
  9. github.com: MODEL CARDgithub.com
  10. github.com: PurpleLlamagithub.com
  11. github.com: MODEL CARDgithub.com
  12. huggingface.co: concept guidehuggingface.co
  13. mlcommons.org: safety faqmlcommons.org
  14. mlcommons.org: jailbreak 0 7mlcommons.org
  15. mlcommons.org: safety methodologymlcommons.org
  16. NIST Generative AI Profilenvlpubs.nist.gov
  17. open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com
  18. owasp.org: www project top 10 for large language model applicationsowasp.org
  19. NIST AI Risk Management Frameworknist.gov
  20. youtu.be: 1KNOcY e9Tsyoutu.be

From this episode

Two useful next steps.

Research Note · 1 min

Quantized Model Artifact Evaluation Framework

A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity,

Research Note · 1 min

Open Model Threat-Control Framework

Begin with one use case and record users, identities, data, retrieval, model artifact, system prompts, tools, external services, outputs, human decisions, logging, operat

Return to the episode