DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

DeepSeek’s Low-Cost AI Model Shook Silicon Valley. Here’s What It Actually Proved

DeepSeek’s reported low-cost V3 training run and competitive R1 reasoning model challenged assumptions about AI spending—but the $6 million figure was not the total cost of building R1.
From TheFinanceBase Team9 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DeepSeek’s January 2025 breakthrough rattled investors because it suggested that capable AI might not require the enormous computing budgets many expected. On January 27, Nvidia shares fell about 17%, erasing roughly $593 billion in market value. But DeepSeek did not prove that it built a frontier model for $6 million—or that demand for AI chips is over. The figure was a reported compute cost for one DeepSeek-V3 training run, not the total cost of developing the company’s models.

What happened, and when?

DeepSeek is a Hangzhou-based AI company founded in 2023 and associated with Liang Wenfeng, co-founder of the quantitative hedge fund High-Flyer. Its work did not begin with the viral app release: the company had been publishing models and technical work before January 2025.

  • November 20, 2024: DeepSeek released R1-Lite-Preview.
  • December 26, 2024: It announced DeepSeek-V3.
  • January 20, 2025: It released DeepSeek-R1, a reasoning model it said was competitive with OpenAI’s o1 on selected tasks.
  • January 27, 2025: Surging attention to the models and consumer app coincided with a broad technology-stock selloff. Nvidia fell about 17%, losing approximately $593 billion in market value that day.

The app briefly rose above ChatGPT in the U.S. Apple App Store rankings, amplifying the story. App-store position reflects attention and adoption, however, not a controlled comparison of model quality. DeepSeek’s R1 announcement, its release index, and contemporary coverage of the company document the release sequence and app surge.

What were V3 and R1?

V3: a large model that activates only part of its capacity

DeepSeek-V3 is a mixture-of-experts language model. Its reported 671 billion total parameters do not all run for every token: about 37 billion are activated for a given inference pass. That distinction matters because total model size and the computation needed to generate each token are not the same thing. DeepSeek’s V3 repository describes its architecture and training report.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

R1: reasoning through reinforcement learning

DeepSeek-R1 used reinforcement learning alongside supervised fine-tuning stages to encourage multi-step problem solving. The company first described R1-Zero, which applied reinforcement learning directly to a base model. That approach produced reasoning behaviors but also repetition, awkward readability, and language mixing. R1 added “cold-start” data before reinforcement learning to improve usability.

The original R1 repository lists 671 billion total parameters, 37 billion activated per inference pass, a 128K context length, and an MIT license. DeepSeek also released distilled versions at 1.5B, 7B, 8B, 14B, 32B, and 70B parameters. Smaller distilled models are more practical to experiment with than the full model, but they are not identical to it. The R1 repository includes model details and deployment notes.

What did “low cost” mean?

Several different costs were blurred together in the January 2025 headlines. They should be kept separate:

Cost measure What the figure means
V3 training compute DeepSeek reported 2.788 million GPU-hours and less than $6 million in computing power for V3’s official training run on Nvidia H800 GPUs. This is a reported compute estimate for that run, not a full accounting of the project. Source: DeepSeek’s V3 repository.
Total development cost Not established. The compute figure does not include a complete accounting of personnel, data, earlier experiments, failed runs, infrastructure, hardware access or acquisition, software engineering, or R1’s development. Reuters and analysts cautioned that the total cost was unknown. Source: Reuters coverage.
R1 API price at launch DeepSeek listed $0.14 per million input tokens for cache hits, $0.55 per million input tokens for cache misses, and $2.19 per million output tokens in January 2025. These were service prices, not a measurement of the model’s underlying operating cost. Source: DeepSeek’s announcement.
Inference efficiency Mixture-of-experts computation and engineering choices can reduce resources needed per token. That is a separate question from what it cost to train the model or what a provider charges for API access.

So “DeepSeek built R1 for $6 million” is not a supported reading of the headline figure. The reported amount concerned V3’s compute run; it does not establish the total cost of V3, R1, or DeepSeek’s research program.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which engineering choices helped?

DeepSeek’s results came from applying several techniques together, rather than a single magic trick. These approaches are not all unique to DeepSeek, but their implementation and combination matter:

  • Mixture of experts: A routing system sends each token through a subset of specialist components, rather than activating the entire model.
  • Multi-head latent attention: Designed to reduce memory and bandwidth requirements.
  • Low-precision computation: FP8 and related techniques can reduce memory and compute demands when implemented carefully.
  • Distributed-training optimization: Better coordination among GPUs can improve utilization and reduce communication overhead.
  • Reinforcement learning: R1 used feedback-based training to encourage reasoning behaviors.
  • Distillation: Smaller models derived from R1’s reasoning data let developers test and adapt related capabilities on less hardware.

Stanford computer scientist Christopher Manning characterized DeepSeek as an important advance for open-weight AI, not a wholly new paradigm. The distinction is useful: strong engineering can change the economics of building and using models without making hardware irrelevant. Stanford’s assessment discusses the significance and limits of the advance.

Did DeepSeek beat OpenAI?

DeepSeek said R1 was comparable with OpenAI o1, and its published evaluations showed strong results on selected math, coding, and reasoning tests. That supports calling R1 competitive on some benchmarks; it does not establish that it was universally better than OpenAI, Anthropic, Google, or Meta models.

Even DeepSeek’s reported results varied by test. Its table showed R1 below OpenAI o1-1217 on MMLU, while reporting strong outcomes on other measures, including DROP, FRAMES, and AlpacaEval. These are vendor-reported comparisons, not independent proof of broad superiority. Benchmark results also depend on model version, prompts, sampling, test-time compute, language, scoring method, and possible data contamination. A benchmark win is evidence about a particular evaluation, not a general ranking of intelligence. DeepSeek’s model card and evaluation table provide its reported comparisons.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why did Silicon Valley and investors react so strongly?

The concern was economic as much as technical. Investors had expected the AI boom to drive sustained growth in demand for Nvidia GPUs, networking equipment, data centers, electricity, cooling, cloud capacity, and large capital investments by technology companies. DeepSeek raised the possibility that software, architecture, and engineering efficiency could deliver more capability from a given amount of hardware.

That possibility challenged a widely held investment assumption: that progress would depend chiefly on ever-larger training runs and ever-greater spending. The January 27 selloff reflected a repricing of those expectations. It was not proof that Nvidia’s business had been permanently broken or that data-center demand had peaked. If cheaper inference makes AI useful in more products, total usage—and the computing needed to support it—could also grow. Stanford’s analysis and coverage of the market reaction describe the shock and its context.

How did U.S. chip restrictions fit in?

DeepSeek used Nvidia GPUs; the story was not that it built its models without Nvidia hardware. Its reported V3 training run used H800s, a product designed for the Chinese market with lower interconnect performance than the unrestricted H100. The striking claim was that DeepSeek achieved competitive results despite tighter access to leading chips and hardware that was less capable in some respects.

That makes the episode relevant to export-control policy: efficiency can partly offset hardware constraints, although it does not make advanced chips unimportant. Claims that DeepSeek secretly possessed tens of thousands of H100s were reported as allegations without publicly supplied evidence, not established fact. The Washington Post’s account and Reuters coverage republished by Investing.com provide context on the hardware and allegations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What does DeepSeek’s open-weight release allow?

“Open-weight” is often more precise than “open source” for this release. DeepSeek published R1’s weights, code, and technical report under an MIT license, enabling commercial use, modification, and distillation subject to applicable terms. The license and downloadable weights make it possible to inspect and run the model without relying solely on DeepSeek’s hosted chatbot.

That does not mean every part of the system is transparent. The release does not fully disclose training data, reproduce the entire training pipeline, reveal all company costs, or guarantee that a hosted chatbot has the same behavior or privacy properties as a locally run model. Particular distilled variants may also have relevant base-model license terms. Stanford researchers raised questions about privacy protection, data sourcing, copyright, and national-security implications. DeepSeek’s release announcement, its repository, and Stanford’s analysis explain these distinctions.

Privacy, censorship, and security: what should users consider?

The risks depend on where the model runs. Prompts sent to a hosted chatbot or API are processed by that service under its policies. A local deployment can keep prompts within an organization’s environment, but the operator then owns the security, access controls, monitoring, updates, and compliance work. DeepSeek’s official privacy policy is relevant for users of its service; organizations should review it alongside their own requirements.

  • For sensitive data: Do not submit confidential, regulated, or personal information to a hosted service until legal, security, procurement, and data-residency teams approve the arrangement.
  • For political or sensitive questions: Hosted versions may refuse or alter responses on politically sensitive topics, including those involving the Chinese government. Behavior can differ across model versions, system prompts, local quantizations, fine-tunes, and third-party providers; it should not be generalized to every deployment.
  • For local operation: Self-hosting changes the data path but is not a complete security solution. Operators remain responsible for infrastructure and model-use controls.
  • For business deployment: Assess privacy, copyright, security, vendor terms, and applicable national-security or export-control obligations before relying on the system.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Should a developer or business use DeepSeek?

The right route depends on the sensitivity of the workload, required control, and ability to operate models. A low token price or low reported training compute does not settle the total cost of deployment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Route Best suited to Main trade-off
Official chat service Casual experimentation and non-sensitive tasks. Prompts leave your environment; service policies, availability, and features may change. Check the current service at DeepSeek Chat.
Official API Prototyping or applications that need a hosted, OpenAI-compatible interface. Review data handling, reliability, and current model identifiers before production. The platform and documentation are the places to check current terms and integration details.
Self-hosted R1 or a distilled model Research, greater control over prompts and retention, or data-sensitive workloads with suitable infrastructure. The full R1 is very large: Hugging Face lists it at roughly 689 GB. Hardware needs vary with quantization, context length, throughput, and latency targets. Distilled variants require less hardware but are not equivalent to full R1. See the model page, vLLM, and SGLang.
Third-party API routing Comparing inference providers or using a common interface across models. Adds another data processor and potential failure point; price, latency, quantization, and behavior vary. OpenRouter’s R1 listing displayed $0.70 per million input tokens and $2.50 per million output tokens at the time observed; that is a dated provider-specific signal, not a universal rate.
Enterprise inference tooling Organizations prioritizing support and integration with existing infrastructure. Not interchangeable with the official hosted service or self-hosting; compare support, licensing, data handling, hardware, and service commitments. Nvidia described R1 deployment through NVIDIA NIM.

For API budgeting, the historical launch prices are not a current quote. DeepSeek’s documentation observed on August 18, 2026 listed V4-Flash at $0.0028 per million cache-hit input tokens, $0.14 per million cache-miss input tokens, and $0.28 per million output tokens; V4-Pro was listed at $0.003625, $0.435, and $0.87 per million tokens for those respective categories. These are documentation-listed API prices, not a full estimate of application costs. Check the official pricing page for current rates and model identifiers.

What DeepSeek changed—and what it did not prove

DeepSeek made the economics of AI efficiency harder to ignore. It showed that a research-oriented Chinese company could release competitive open-weight models and a popular assistant while reporting an unusually low compute figure for a particular training run. Its distilled models also gave developers more ways to experiment with reasoning without hosting the full model.

It did not show that R1 was universally superior to U.S. models, that the entire project cost less than $6 million, that DeepSeek used no advanced Nvidia hardware, or that AI infrastructure spending was permanently destined to fall. The more durable lesson is narrower: hardware scale remains important, but architecture and engineering can change how much capability organizations get from each dollar of compute.

Where DeepSeek stands now

As of August 18, 2026, DeepSeek’s official API documentation lists V4-Flash and V4-Pro, with a documented context length of 1 million tokens and maximum output of up to 384K tokens for those models. The older compatibility names deepseek-chat and deepseek-reasoner were scheduled for deprecation on July 24, 2026. R1 is therefore best understood as the model behind the January 2025 Silicon Valley shock, not as DeepSeek’s current flagship. DeepSeek’s API pricing and model documentation lists the current lineup and migration details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase09 OCT 267 minMortgage Escrow FAQs: Taxes, Insurance, Shortages, and Refunds
  2. The Money DeskBlogTheFinanceBase09 OCT 265 minHow Mortgage Escrow Accounts Work and What Homeowners Pay For
  3. The Money DeskBlogTheFinanceBase09 OCT 265 minHow to Read a Stock Chart, Volume and Market-Cap Data
Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.