Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
DeepSeek released its R1 reasoning model on January 20, 2025, and reported benchmark results comparable to OpenAI’s o1-1217 on selected reasoning tasks. The often-cited $5.6 million figure, however, describes DeepSeek-V3’s reported final training run—not the full cost of developing R1. The distinction matters if you are weighing AI providers, estimating operating costs, or deciding whether an open-weight model could reduce your business expenses.
What DeepSeek released in January 2025
DeepSeek-R1 was designed to spend more computation working through difficult questions, particularly in mathematics, coding, and logic. DeepSeek released R1 alongside R1-Zero, an experimental model trained with large-scale reinforcement learning without conventional supervised fine-tuning as its initial step, as well as smaller distilled models based on Qwen and Llama families. The release included model weights, code, and technical material. DeepSeek’s release announcement and its technical paper describe the models and evaluations.
“Open-source” became common shorthand for the release, but “open-weight” is more precise. Public weights make it possible to download and run a model, subject to its applicable terms; they do not by themselves make all training data, data preparation, infrastructure, or development work public. DeepSeek says the R1 repository materials use an MIT license, but developers should check the terms for the specific checkpoint and any underlying or derived components before commercial redistribution. See the R1 repository and its license.
How comparable was R1 to OpenAI o1?
DeepSeek reported that R1 matched or approached OpenAI-o1-1217 on several mathematics, coding, and reasoning benchmarks, including AIME, MATH-500, GPQA Diamond, and Codeforces evaluations. That is meaningful evidence of competitiveness on selected tests, not proof that R1 was better at every task or interchangeable with OpenAI’s product.
#1 Best Overall
- Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
- Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
- Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
- Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
- Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.
The scores in DeepSeek’s paper are the company’s reported evaluations. Benchmark outcomes can depend on model version, prompts, sampling settings, answer scoring, and test-set exposure; a score on one test does not establish comparative quality in production. The paper’s results also should not be transferred automatically to smaller distilled models: a compact checkpoint may be cheaper to serve but is not the full R1 model.
For a buyer, “comparable” is only one input. The practical comparison also includes factual accuracy on your data, latency, reliability, tool use, context limits, privacy terms, support, and the cost of handling mistakes. R1’s benchmark performance did not establish parity across all those dimensions.
What the $5.6 million figure actually measures
DeepSeek’s approximately $5.576 million figure is its reported cost for DeepSeek-V3’s final training run, based on 2.788 million H800 GPU-hours and the paper’s assumed rental rate. It is not a verified all-in price for creating R1, building the company’s AI capability, or bringing a production service to market. The figure and its accounting basis appear in the DeepSeek-V3 paper; the Associated Press also explains why the headline number should be treated narrowly.
Recommended Free Tools
Rank #2
- Intel Core Ultra 9 285 Processor: Newly developed cores deliver ultra-smooth and responsive gameplay. AI accelerators prepare users for the next era of gaming on an AI PC
- Operating System: Enjoy the latest generation of Windows 11 Home for your everyday needs. MSI recommends Windows 11 Pro for business use
- NVIDIA GeForce RTX 5070 GPU: Experience cutting-edge graphics performance with the powerful NVIDIA GeForce RTX 5070 graphics card for immersive gaming and content creation
- Advanced Cooling System: In conjunction with an RGB CPU Air Cooler, the Aegis RS features four system cooling fans; three in the front and one in the rear to pull in cool air and push heat out of the PC
- Customizable RGB Lighting: With the built-in RGB lighting, take your gaming experience to the next level by pressing the MSI LED button to cycle through lighting options. Customize lighting even further with MSI Center software
- Final training run: the reported V3 compute figure and its assumed cost basis.
- Full development: not established by that figure. It does not represent a complete accounting of research, data preparation, failed experiments, infrastructure, personnel, or hardware ownership.
- R1 development: R1 built on prior DeepSeek models and work, including V3, reinforcement learning, supervised fine-tuning, and curated data. Its total development cost has not been publicly established with the same precision.
- Running the model: API inference and self-hosting are separate costs from training.
So the evidence supports a narrower point: DeepSeek reported an unusually low compute cost for a particular V3 training run. It does not show that a frontier reasoning model can be developed end to end for $5.6 million, or provide a like-for-like cost comparison with OpenAI.
Why the approach could use less compute
DeepSeek-V3 combines several efficiency techniques; no single one explains the reported cost. Its mixture-of-experts (MoE) design has 671 billion total parameters but activates about 37 billion per token. Total parameters describe the model’s full capacity; active parameters describe the portion used for a particular token. An MoE model therefore does not perform every token calculation using all 671 billion parameters.
- Multi-head Latent Attention reduces key-value-cache memory requirements.
- Multi-token prediction lets the model predict multiple future tokens during training and can improve training efficiency.
- FP8 mixed-precision training reduces memory and compute requirements.
- Reinforcement learning helped improve R1’s reasoning behavior. DeepSeek used Group Relative Policy Optimization (GRPO) as one part of its approach.
- Distillation transfers some reasoning behavior into smaller models, which can be easier to deploy but may sacrifice performance on harder tasks.
These methods help explain how DeepSeek pursued efficiency. They do not, on their own, validate a complete development-cost estimate or predict the cost of running a model for a particular workload.
Rank #3
Training cost, API price, and self-hosting are different budgets
A low training-run estimate does not tell you what a business will pay to use a model. API bills depend on input and output volume, cached versus uncached input, retries, and the provider’s current rates. Reasoning work can also generate substantial output, so comparing only input-token rates can understate the bill.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsAt R1’s release, DeepSeek’s listed API rates were $0.14 per million cached input tokens, $0.55 per million uncached input tokens, and $2.19 per million output tokens. Those are historical R1 prices, not current rates. DeepSeek’s historical pricing details document them.
As of August 18, 2026, DeepSeek’s official pricing documentation lists V4-Flash and V4-Pro as current API models, each with a 1-million-token context window. The documented rates are below; prices and service terms can change, so check the current pricing page before budgeting.
Rank #4
- Featuring NVIDIA DLSS 4 technology, high-performance Blackwell architecture, and NVIDIA ray tracing
- With its balanced dimensions of 4.4 inches high by 10.5 inches long, this graphics card fits perfectly into mid- to full-tower configurations, while offering optimized space for efficient cooling.
- 48GB GDDR7 (384-bit), 14,080 CUDA processing cores, and up to 1,344 GB/s of memory bandwidth to provide the memory needed to create stunning visual realism.
- PCI Express 5.0 interface - Offers compatibility with a range of systems. Also includes DisplayPort and HDMI outputs for expanded connectivity.
- DisplayPort 2.1 support enables displays up to 8K at 240Hz or 16K at 60Hz, providing ample bandwidth for multi-display setups, content creation, and demanding work environments.
| Model listed by DeepSeek | Cached input, per million tokens | Uncached input, per million tokens | Output, per million tokens | Context window |
|---|---|---|---|---|
| DeepSeek-V4-Flash | $0.0028 | $0.14 | $0.28 | 1 million tokens |
| DeepSeek-V4-Pro | $0.003625 | $0.435 | $0.87 | 1 million tokens |
The old API names deepseek-chat and deepseek-reasoner were scheduled for retirement on July 24, 2026, at 15:59 UTC, according to DeepSeek’s API updates. The current model list is documented at DeepSeek’s list-models page. Do not assume code written for the original R1 endpoint will continue to work unchanged.
Self-hosting replaces a per-token provider bill with infrastructure and operating costs. The total depends on checkpoint size, quantization, GPU memory, throughput, storage, electricity, engineering, and maintenance. Open weights create a deployment option; they do not make inference free.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Ways to access DeepSeek—and what each entails
Web or mobile chat
The consumer chat service is a managed product, distinct from both the API and downloadable weights. Its availability and responsiveness may differ from a self-hosted model, and its terms govern what happens to submitted data.
API
DeepSeek’s API uses an OpenAI-compatible format, which may reduce migration work for developers. Compatibility does not guarantee identical behavior for tool calls, JSON output, streaming, system prompts, errors, rate limits, reasoning-token accounting, or data handling. Test the features your application actually uses, and review current policies before sending confidential information. The API documentation is at api-docs.deepseek.com; API access is available through DeepSeek’s platform.
Local or hosted inference with open weights
The R1 repository includes distilled checkpoints and serving examples. For instance, it gives this vLLM command for serving a 32B distilled Qwen checkpoint:
vllm serve deepseek-ai/DeepSeek-R1-Distill-Qwen-32B
--tensor-parallel-size 2
--max-model-len 32768
--enforce-eager
This is a repository example, not a universal installation recipe. Compatibility depends on the vLLM release, CUDA stack, GPU, model files, and hardware. A local deployment also needs monitoring, security, capacity planning, and a way to evaluate output quality.
How to decide whether DeepSeek fits a budget
Estimate the cost of completing your task, not just the price of sending a token. A useful comparison includes:
- Input and output tokens, including reasoning output and retries.
- Cached versus uncached inputs, plus any applicable batch rates.
- Tool calls, embeddings, retrieval, storage, monitoring, and moderation.
- Rate limits, availability, and the effect of delays or outages.
- For self-hosting: GPU rental or ownership, utilization, power, serving labor, and maintenance.
- The cost of review or failure when an incorrect answer reaches a customer or business process.
DeepSeek may suit a price-sensitive text-reasoning workload if its quality, policies, and reliability pass your own checks. A smaller distilled model may make sense for a narrow, repeatable task where local control or latency matters more than maximum reasoning performance. A closed commercial service may be a better fit where contractual support, regional controls, integrated tools, or a mature enterprise relationship outweigh token savings. These are trade-offs to test against the workload and the organization’s data-governance requirements, not universal rankings.
What the release did—and did not—show
- It showed that DeepSeek could release an open-weight reasoning model with reported results competitive with OpenAI-o1-1217 on selected benchmarks, alongside techniques intended to improve training and serving efficiency.
- It did not show that R1’s total development cost was $5.6 million; that number refers to V3’s reported final training run.
- It did not show that R1 was superior across every task, that a distilled version matched the full model, or that benchmark performance guaranteed production reliability.
- It did not make open weights synonymous with free operation, full reproducibility, or suitability for sensitive data.
By August 2026, R1 is best understood as a consequential release in the shift toward lower-cost, open-weight reasoning models—not as DeepSeek’s current API endpoint. The company’s documentation now lists V4-Flash and V4-Pro as its current API offerings.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

