Evergreen
How to Read an AI Training Cost Claim
Interpret AI training cost claims by naming the model, stage, metric, hardware, scope, exclusions, source, reproduction, conversion assumptions, and uncertainty.
How to Read an AI Training Cost Claim
An AI training-cost claim is useful only when it identifies the exact artifact, training stage, metric, hardware, scope, source, exclusions, and dollar-conversion assumptions.
GPU-hours are not a company invoice. A reported final run is not total research and development. A cloud-rate multiplication is a scenario, not an audited cost.
Start with the artifact
"This model cost six million dollars" is incomplete before cost enters the discussion.
Which model? Which checkpoint? Which base? Which training stage? Which repository revision? Is the number about pretraining, supervised fine-tuning, reinforcement learning, distillation, evaluation, or serving?
DeepSeek R1 shows why this matters. Full R1 builds from DeepSeek V3 architecture and adds its own post-training path. R1-Zero and smaller distilled checkpoints are different artifacts.
A V3 base-training number cannot be relabeled as the complete R1 development cost without additional evidence.
Preserve the primary unit
DeepSeek's V3 technical report reports 2.664 million NVIDIA H800 GPU-hours for pretraining and about 0.1 million GPU-hours for subsequent stages, summarized as 2.788 million H800 GPU-hours for full training.
That is the first claim to preserve. It names a hardware-time unit under the authors' accounting.
The report also describes 14.8 trillion pretraining tokens and a 671-billion-total-parameter mixture-of-experts architecture with 37 billion parameters activated per token. These details help interpret the run. They do not turn it into dollars.
flowchart LR
A["Artifact and stage"] --> B["Primary compute metric"]
B --> C["Scope and exclusions"]
C --> D["Conversion model"]
D --> E["Dollar range"]
E --> F["Sensitivity and uncertainty"]
F --> G["Allowed editorial wording"]
Write a claim ledger
| Field | Required entry |
|---|---|
| Claimant | Who made the statement? |
| Artifact | Which exact model, checkpoint, or run? |
| Stage | Which data, training, tuning, distillation, evaluation, or serving stage? |
| Metric | GPU-hours, FLOPs, tokens, energy, dollars, or another unit? |
| Hardware | Which accelerator and configuration? |
| Scope | Which runs and operations are included? |
| Exclusions | Which experiments, staff, data, facilities, serving, and overhead are omitted? |
| Source | Paper, repository, filing, invoice, benchmark, or estimate? |
| Reproduction | Was it independently reproduced under comparable conditions? |
| Conversion | Which ownership, rental, utilization, energy, facility, labor, and time assumptions apply? |
| Uncertainty | Which inputs are ranges, missing, or sensitive? |
| Wording | What exact public sentence does the evidence support? |
Do not move to a headline until the ledger is complete enough to show what the number means.
Separate compute from dollars
GPU-hours describe time on accelerators. Dollars require a price model.
A cloud rental rate may bundle accelerator, host, network, power, facility, maintenance, support, and provider margin. An owned-fleet estimate needs purchase price, financing, useful life, utilization, energy, cooling, facility, networking, storage, staff, spares, and residual value.
The hourly rate also depends on date, region, contract, reservation, scale, utilization, and what else the service includes.
Multiplying 2.788 million hours by one public H800-equivalent rate creates one illustrative scenario. It does not show what DeepSeek paid, owned, negotiated, or allocated.
Separate final training from development
Model development can include data sourcing and cleaning, architecture research, failed experiments, ablations, earlier checkpoints, test runs, staff, software, safety work, evaluation, deployment, and overhead.
Some of these activities may share infrastructure with other projects. Some may never appear in a technical report. Some may be more expensive than the reported successful run.
DeepSeek's V3 repository is useful for the authors' architecture and compute statements. The R1 paper is useful for the later reasoning-model training path.
Neither source claims to be a complete audited company cost ledger.
Make exclusions visible
An exclusion is not proof that a claim is misleading. Technical papers reasonably report defined technical measurements.
The editorial failure occurs when a narrow number is widened without saying so.
For a training-cost page, state whether the number excludes research iterations, failed runs, data, labor, earlier models, hardware purchase, facility, power, network, storage, inference, company overhead, and opportunity cost.
If the source does not say, write "not established by this source" instead of guessing.
Show sensitivity instead of false precision
Create a range for uncertain inputs. Change utilization, power price, hourly rate, useful life, labor, network, storage, and omitted experiments. Show which assumptions drive the result.
If a small change in one assumption moves the estimate by hundreds of millions of dollars, the headline should not present a precise single number.
The best public result may remain the reported GPU-hours plus a bounded illustration of how different conversion models behave.
Keep technical, company, and market claims separate
A lower compute requirement for a defined run can be meaningful.
It does not by itself prove that the company has the lowest total cost, that the hosted API is profitable, that the model is superior, that a competitor's spending is wasteful, that fewer accelerators will be sold, or that a security should rise or fall.
Those require separate evidence about quality, workload, serving, price, adoption, company finances, supply, competition, customer behavior, and market expectations.
Write the sentence the evidence supports
A careful sentence for DeepSeek V3 is:
DeepSeek's V3 technical report states 2.788 million NVIDIA H800 GPU-hours for the defined full-training scope; that figure is not the total cost of DeepSeek, the complete development cost of R1, or an audited dollar bill.
The sentence preserves the technical achievement without borrowing claims the source does not make.
This page was developed with AI assistance from the E055 transcript and linked primary sources, then structured for human machine-learning, infrastructure, finance, quantitative, and editorial review. It does not audit DeepSeek, estimate a security's value, or provide investment advice.
Sources
Follow the evidence.
- github.com: LICENSEgithub.com
- bis.gov: commerce strengthens restrictions advanced computing semiconductors enhance foundry due diligence preventbis.gov
- arxiv.org: 2501arxiv.org
- NIST AI Risk Management Frameworknist.gov
- daltonanderson.ghost.io: deepseek vs nvidia the future of ai chip economicsdaltonanderson.ghost.io
- investor.nvidia.com: defaultinvestor.nvidia.com
- api-docs.deepseek.comapi-docs.deepseek.com
- bis.gov: 740bis.gov
- daltonanderson.net: deepseek vs nvidia the future of ai chip economicsdaltonanderson.net
- github.com: DeepSeek R1github.com
- open.spotify.com: 6jLI1bNwyoxI449vXJXzBVopen.spotify.com
- youtu.be: Qp24TkfT9XEyoutu.be
- bis.gov: 742bis.gov
- bis.gov: department commerce revises license review policy semiconductors exported chinabis.gov
- arxiv.org: 2412arxiv.org
- docs.nvidia.com: cudadocs.nvidia.com
- github.com: DeepSeek V3github.com