Back to the episode map

Research Note

Open-Weight Model Evaluation Record

The object is an exact checkpoint or API model in a defined runtime and configuration, not a brand or family name.

Aug 4, 20262 min readBy Dalton Anderson

Open-Weight Model Evaluation Record

Evaluation object

The object is an exact checkpoint or API model in a defined runtime and configuration, not a brand or family name.

Record repository and revision, artifact hash, base model, license chain, quantization, tokenizer, prompt template, runtime, hardware, context, sampling, reasoning controls, tools, adapters, date, and serving path.

Workload record

Use representative tasks and data with a documented sampling method. Preserve expected results, scoring rules, reviewer instructions, disagreement handling, and confidence.

Include easy, typical, difficult, adversarial, ambiguous, out-of-scope, refusal, and failure-recovery cases in proportion to use.

Author benchmarks and public leaderboards can inform hypotheses. They do not replace this workload.

Evaluation domains

DomainEvidence
Task qualityAccuracy, completeness, groundedness, usefulness, and calibration
Failure behaviorHallucination, repetition, language mixing, unsafe completion, refusal, and recovery
RobustnessPrompt variation, long context, malformed input, adversarial input, and tool failure
SecurityArtifact provenance, code execution, deserialization, tool permissions, supply chain, and isolation
PrivacyInput, output, logs, telemetry, retention, model access, and deletion
Fairness and impactRelevant groups, errors, allocation, explanation, correction, and human review
PerformanceLatency distribution, throughput, tokens, memory, utilization, and concurrency
CostHardware or API, energy, network, storage, engineering, monitoring, and support
OperationsDeployment, scaling, patching, observability, incident, rollback, and retirement
LicenseModel, base, code, tokenizer, adapters, data, runtime, and intended-use obligations

Risk framework

NIST's AI RMF provides voluntary govern, map, measure, and manage functions for AI risk:

https://www.nist.gov/itl/ai-risk-management-framework

The evaluation should be proportional to the system's impact and deployment context. High-stakes use requires qualified domain, legal, security, privacy, human-factors, and affected-party review.

Local versus API

Local control can improve artifact and data-path visibility but creates operational responsibility. Hosted service can reduce infrastructure burden but adds provider terms, retention, availability, version drift, rate limits, and dependency.

Test the delivery model the organization will actually operate.

Decision

Record whether the artifact is accepted, accepted with controls, limited to research, rejected, or requires more evidence.

The decision applies only to the tested revision, configuration, workload, time, and risk boundary.

Sources

Follow the evidence.

  1. github.com: LICENSEgithub.com
  2. bis.gov: commerce strengthens restrictions advanced computing semiconductors enhance foundry due diligence preventbis.gov
  3. arxiv.org: 2501arxiv.org
  4. NIST AI Risk Management Frameworknist.gov
  5. daltonanderson.ghost.io: deepseek vs nvidia the future of ai chip economicsdaltonanderson.ghost.io
  6. investor.nvidia.com: defaultinvestor.nvidia.com
  7. api-docs.deepseek.comapi-docs.deepseek.com
  8. bis.gov: 740bis.gov
  9. daltonanderson.net: deepseek vs nvidia the future of ai chip economicsdaltonanderson.net
  10. github.com: DeepSeek R1github.com
  11. open.spotify.com: 6jLI1bNwyoxI449vXJXzBVopen.spotify.com
  12. youtu.be: Qp24TkfT9XEyoutu.be
  13. bis.gov: 742bis.gov
  14. bis.gov: department commerce revises license review policy semiconductors exported chinabis.gov
  15. arxiv.org: 2412arxiv.org
  16. docs.nvidia.com: cudadocs.nvidia.com
  17. github.com: DeepSeek V3github.com
Open-Weight Model Evaluation Record