Research Note
Quantized Model Artifact Evaluation Framework
A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity,
Quantized Model Artifact Evaluation Framework
Artifact contract
A quantized model is a new deployable artifact. Record the reference model, exact weights, tokenizer, prompt format, license, quantization method, precision, granularity, calibration data and settings, builder, hash, runtime, hardware, kernels, and generation configuration.
Matched comparison
| Dimension | Evidence |
|---|---|
| Task behavior | Representative and difficult cases at the intended workload |
| Safety behavior | Policy, refusal, injection, tool, and sensitive-output cases relevant to the system |
| Quality distribution | Per-task regressions, not only an average score |
| Operations | Load time, resident memory, time to first token, inter-token latency, throughput, power, and failures |
| Context | Short, typical, and long inputs under realistic batching |
| Reproducibility | Frozen prompts, seeds where applicable, settings, logs, versions, and hashes |
Decision rule
Accept a candidate only when its measured resource benefit matters for the deployment and regressions remain inside use-case-specific thresholds. Preserve rollback to the reference or prior accepted artifact.
Boundary
Results do not transfer automatically across quantization methods, hardware, runtimes, kernels, context lengths, batches, prompt formats, or tasks. File size alone is not an evaluation.
Sources
Follow the evidence.
- ai.meta.com: the llama 3 herd of modelsai.meta.com
- ai-challenges.nist.gov: genaiai-challenges.nist.gov
- owasp.org: www project top 10 for large language model applicationsowasp.org
- youtu.be: 1KNOcY e9Tsyoutu.be
- github.com: PurpleLlamagithub.com
- crfm.stanford.edu: indexcrfm.stanford.edu
- NIST AI Risk Management Frameworknist.gov
- mlcommons.org: jailbreak 0 7mlcommons.org
- mlcommons.org: safety faqmlcommons.org
- github.com: MODEL CARDgithub.com
- ai-challenges.nist.gov: ariaai-challenges.nist.gov
- github.com: MODEL CARDgithub.com
- daltonanderson.ghost.io: metas llama 3 safety scaling and simple solutionsdaltonanderson.ghost.io
- ai.meta.com: meta llama 3 1 ai responsibilityai.meta.com
- NIST Generative AI Profilenvlpubs.nist.gov
- mlcommons.org: safety methodologymlcommons.org
- huggingface.co: concept guidehuggingface.co
- github.com: MODEL CARDgithub.com
- csrc.nist.gov: red teamingcsrc.nist.gov
- open.spotify.com: 44o5OPSumaZJcvRkXutorBopen.spotify.com