Research Note
AI Training Cost Claim Record
A training-cost claim is interpretable only when the artifact, stage, unit, scope, source, exclusions, and conversion assumptions are explicit.
AI Training Cost Claim Record
Core rule
A training-cost claim is interpretable only when the artifact, stage, unit, scope, source, exclusions, and conversion assumptions are explicit.
DeepSeek's V3 report is the primary source for the model authors' compute accounting:
https://arxiv.org/abs/2412.19437
The report describes a 671-billion-total-parameter mixture-of-experts model with 37 billion parameters activated per token. It reports 2.664 million H800 GPU-hours for pretraining and about 0.1 million GPU-hours for subsequent stages, summarized as 2.788 million H800 GPU-hours for full training.
Claim ledger
| Field | Required entry |
|---|---|
| Claimant | Person or organization making the statement |
| Artifact | Exact model, checkpoint, or training run |
| Stage | Data work, pretraining, supervised tuning, reinforcement learning, distillation, evaluation, or serving |
| Metric | GPU-hours, accelerator-hours, FLOPs, tokens, energy, dollars, or another unit |
| Hardware | Exact accelerator and relevant configuration |
| Scope | Runs and operations included |
| Exclusions | Research, failed runs, earlier models, data, labor, facilities, serving, and other omitted categories |
| Source | Paper, repository, filing, invoice, benchmark, or estimate |
| Reproduction | Independent reproduction status and conditions |
| Conversion | Ownership, rental, utilization, power, facility, network, storage, labor, depreciation, period, and region |
| Uncertainty | Range, missing inputs, sensitivity, and confidence |
| Allowed wording | Exact public statement the evidence supports |
Compute is not dollars
GPU-hours describe hardware time under a defined accounting method. A dollar amount requires a rate.
A cloud rental rate may include hardware, facility, power, network, maintenance, and margin. An owned-fleet estimate needs purchase cost, useful life, utilization, financing, power, cooling, facility, operations, spares, network, and storage.
Multiplying a public hourly price by reported GPU-hours creates a scenario, not an audited company cost.
Training run is not company development
Research and development can include architecture work, data pipelines, failed experiments, ablations, earlier checkpoints, staff, software, evaluation, safety, deployment, and overhead beyond a reported final run.
R1 also includes post-training and distillation relationships that should not be collapsed into the V3 base training number.
The public claim should say exactly what the report counted and exactly what the conversion adds.
Sensitivity
Show how the result changes when utilization, power price, rental or ownership, useful life, network, storage, labor, and excluded experiments change.
Avoid false precision. If inputs are uncertain, publish a range and name the largest drivers.
Decision boundary
A lower reported training-compute figure can be meaningful technical evidence. It does not by itself establish the lowest total cost, model superiority, sustainable price, company profitability, hardware-market decline, or investment outcome.
Sources
Follow the evidence.
- github.com: LICENSEgithub.com
- bis.gov: commerce strengthens restrictions advanced computing semiconductors enhance foundry due diligence preventbis.gov
- arxiv.org: 2501arxiv.org
- NIST AI Risk Management Frameworknist.gov
- daltonanderson.ghost.io: deepseek vs nvidia the future of ai chip economicsdaltonanderson.ghost.io
- investor.nvidia.com: defaultinvestor.nvidia.com
- api-docs.deepseek.comapi-docs.deepseek.com
- bis.gov: 740bis.gov
- daltonanderson.net: deepseek vs nvidia the future of ai chip economicsdaltonanderson.net
- github.com: DeepSeek R1github.com
- open.spotify.com: 6jLI1bNwyoxI449vXJXzBVopen.spotify.com
- youtu.be: Qp24TkfT9XEyoutu.be
- bis.gov: 742bis.gov
- bis.gov: department commerce revises license review policy semiconductors exported chinabis.gov
- arxiv.org: 2412arxiv.org
- docs.nvidia.com: cudadocs.nvidia.com
- github.com: DeepSeek V3github.com