Research Note
Open-Weight Model Evaluation Record
The object is an exact checkpoint or API model in a defined runtime and configuration, not a brand or family name.
Open-Weight Model Evaluation Record
Evaluation object
The object is an exact checkpoint or API model in a defined runtime and configuration, not a brand or family name.
Record repository and revision, artifact hash, base model, license chain, quantization, tokenizer, prompt template, runtime, hardware, context, sampling, reasoning controls, tools, adapters, date, and serving path.
Workload record
Use representative tasks and data with a documented sampling method. Preserve expected results, scoring rules, reviewer instructions, disagreement handling, and confidence.
Include easy, typical, difficult, adversarial, ambiguous, out-of-scope, refusal, and failure-recovery cases in proportion to use.
Author benchmarks and public leaderboards can inform hypotheses. They do not replace this workload.
Evaluation domains
| Domain | Evidence |
|---|---|
| Task quality | Accuracy, completeness, groundedness, usefulness, and calibration |
| Failure behavior | Hallucination, repetition, language mixing, unsafe completion, refusal, and recovery |
| Robustness | Prompt variation, long context, malformed input, adversarial input, and tool failure |
| Security | Artifact provenance, code execution, deserialization, tool permissions, supply chain, and isolation |
| Privacy | Input, output, logs, telemetry, retention, model access, and deletion |
| Fairness and impact | Relevant groups, errors, allocation, explanation, correction, and human review |
| Performance | Latency distribution, throughput, tokens, memory, utilization, and concurrency |
| Cost | Hardware or API, energy, network, storage, engineering, monitoring, and support |
| Operations | Deployment, scaling, patching, observability, incident, rollback, and retirement |
| License | Model, base, code, tokenizer, adapters, data, runtime, and intended-use obligations |
Risk framework
NIST's AI RMF provides voluntary govern, map, measure, and manage functions for AI risk:
https://www.nist.gov/itl/ai-risk-management-framework
The evaluation should be proportional to the system's impact and deployment context. High-stakes use requires qualified domain, legal, security, privacy, human-factors, and affected-party review.
Local versus API
Local control can improve artifact and data-path visibility but creates operational responsibility. Hosted service can reduce infrastructure burden but adds provider terms, retention, availability, version drift, rate limits, and dependency.
Test the delivery model the organization will actually operate.
Decision
Record whether the artifact is accepted, accepted with controls, limited to research, rejected, or requires more evidence.
The decision applies only to the tested revision, configuration, workload, time, and risk boundary.
Sources
Follow the evidence.
- github.com: LICENSEgithub.com
- bis.gov: commerce strengthens restrictions advanced computing semiconductors enhance foundry due diligence preventbis.gov
- arxiv.org: 2501arxiv.org
- NIST AI Risk Management Frameworknist.gov
- daltonanderson.ghost.io: deepseek vs nvidia the future of ai chip economicsdaltonanderson.ghost.io
- investor.nvidia.com: defaultinvestor.nvidia.com
- api-docs.deepseek.comapi-docs.deepseek.com
- bis.gov: 740bis.gov
- daltonanderson.net: deepseek vs nvidia the future of ai chip economicsdaltonanderson.net
- github.com: DeepSeek R1github.com
- open.spotify.com: 6jLI1bNwyoxI449vXJXzBVopen.spotify.com
- youtu.be: Qp24TkfT9XEyoutu.be
- bis.gov: 742bis.gov
- bis.gov: department commerce revises license review policy semiconductors exported chinabis.gov
- arxiv.org: 2412arxiv.org
- docs.nvidia.com: cudadocs.nvidia.com
- github.com: DeepSeek V3github.com