Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversFall ResetAmazon USFall reset deals: check better picks before checkoutAmazon US: today's deals, useful picks and quick comparisons.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
Blog

How DeepSeek Built a Powerful AI Model Despite U.S. Chip Restrictions

By TheFinanceBase Team11 min read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

DeepSeek did not build its breakthrough models on a secret supply of unrestricted frontier chips—or prove that advanced AI can be created for only $5.6 million. Its achievement was more specific: the Chinese company combined Nvidia H800 GPUs that were initially export-compliant, extensive hardware-software optimization, a sparse model architecture, low-precision training, and reinforcement learning to produce highly capable systems under tighter hardware constraints.

The result showed that U.S. export controls could make advanced AI development more difficult without making it impossible. It also exposed how misleading headline cost comparisons can be when a final training-run estimate is presented as the total cost of developing an AI company and product.

The short version

DeepSeek’s breakthrough involved two closely related releases:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • DeepSeek-V3, released in December 2024, was a general-purpose model whose technical report described a 671-billion-parameter Mixture-of-Experts system trained on 2,048 Nvidia H800 GPUs.
  • DeepSeek-R1, released on January 20, 2025, built on V3 and used large-scale reinforcement learning during post-training to improve mathematical, coding, and reasoning performance.

DeepSeek reported approximately 2.788 million H800 GPU-hours for the V3 training run. At an assumed rental rate of $2 per GPU-hour, that implies about $5.576 million. That is an estimate for the direct GPU cost of one reported training run—not DeepSeek’s total research budget, hardware investment, staffing, data, infrastructure, failed experiments, or deployment expenses. The Congressional Research Service has cautioned against treating the figure as the full cost of building DeepSeek’s technology (CRS analysis).

There was not one simple “chip ban”

The phrase “U.S. chip ban” compresses several rounds of export controls into a misleading shorthand.

The United States introduced major restrictions on advanced AI-chip exports in October 2022. The rules were updated and tightened in October 2023, expanding the products and performance categories subject to licensing requirements. Nvidia later disclosed that products including the A100, A800, H100, H800, L4, L40, L40S, and RTX 4090 were affected by U.S. export-control requirements or related restrictions (Nvidia’s SEC filing).

Those measures did not prohibit China from obtaining every Nvidia GPU. They targeted specific performance classes, products, destinations, and transactions. Chinese organizations could still use some lower-performance chips, older inventory, domestically produced hardware, and potentially cloud-based computing capacity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Nvidia H800 is central to the story. It was designed for the Chinese market with reduced interconnect bandwidth so it could fit within earlier export thresholds. A lower-bandwidth chip is less suitable than an unrestricted H100 for some large-scale workloads, but it remains a powerful accelerator. A sufficiently large cluster—and software designed around its limitations—can still train an advanced model.

DeepSeek’s V3 report says it used 2,048 H800 GPUs. The company’s disclosed hardware therefore does not support the simple claim that it trained the model entirely without Nvidia technology. A more accurate description is that DeepSeek developed a high-capability model using a large but comparatively constrained GPU fleet.

DeepSeek-V3: the documented breakthrough

DeepSeek-V3 was a general-purpose base and chat model released in December 2024. Its technical report described a 671-billion-parameter architecture, but that number needs context.

V3 used a sparse Mixture-of-Experts architecture. Rather than sending every token through every parameter, a routing system selects a limited group of expert networks for each token. The model retains a large total capacity while performing substantially less computation per token than a dense model with the same total parameter count.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Term Meaning Why it matters
Total parameters All parameters stored in the model Indicates the model’s overall capacity and memory requirements
Activated parameters The subset used for a particular token More closely reflects computation per token in a sparse model
GPU-hours The number of GPU devices multiplied by operating time Measures compute used by a particular run, not total company spending

MoE was not invented by DeepSeek. Major AI laboratories had already researched sparse expert models. DeepSeek’s contribution was the way it implemented and scaled the architecture while optimizing the entire training system for restricted interconnect performance.

The engineering techniques that reduced the hardware penalty

Mixture-of-Experts routing

MoE reduced the amount of active computation required for each token. That can lower training and inference costs, although it does not make the full model small. The inactive experts still occupy memory, and routing tokens between experts can create significant communication demands across GPUs.

This distinction is why “671 billion parameters” and “cheap to run” are not equivalent statements. Actual costs depend on active computation, memory movement, communication, batch size, context length, latency requirements, and the number of generated reasoning tokens.

Multi-head Latent Attention

DeepSeek also reported using Multi-head Latent Attention, or MLA. It reduces the memory required for the key-value cache used during inference.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The key-value cache stores information from earlier tokens so the model does not have to recompute the entire context for every new token. As conversations and documents become longer, that cache can become a major memory constraint. Reducing it can support longer contexts, larger batches, or less expensive serving on a fixed GPU fleet.

DeepSeek’s V3 report and independent hardware-aware analysis describe MLA as part of a broader effort to reduce memory and communication bottlenecks (V3 technical report; hardware-aware analysis).

FP8 mixed-precision training

DeepSeek reported using FP8 mixed-precision training. Lower numerical precision can reduce memory use and increase throughput, but it introduces risks such as overflow, underflow, instability, and accuracy loss.

FP8 was not a magic switch. Making it work at large scale required numerical safeguards, scaling methods, software changes, and careful monitoring. The lesson is that hardware restrictions can increase the value of engineering expertise: the constraint becomes a systems problem rather than simply a request for more chips.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Communication-aware cluster design

The H800’s reduced interconnect bandwidth made communication between GPUs more difficult than on unrestricted H100-class systems. In a large MoE model, experts and training states may need to move between devices. If communication is slow, GPUs can sit idle waiting for data.

DeepSeek’s approach included custom parallelism, scheduling, memory management, and network-topology optimizations intended to keep the cluster productive. Nvidia’s discussion of DeepSeek-R1 likewise emphasized the importance of high-bandwidth networking and deployment design (Nvidia’s NIM overview).

This is the most important technical point: the bottleneck was not only raw floating-point capacity. It was the combined challenge of computation, memory, networking, synchronization, and keeping thousands of processors busy.

Why DeepSeek-R1 mattered

DeepSeek-R1, released on January 20, 2025, extended the V3 foundation with a focus on reasoning. DeepSeek said it used large-scale reinforcement learning during post-training and achieved strong results in mathematics, coding, and reasoning with relatively little labeled data (DeepSeek’s release announcement).

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The division of labor is important:

  • Pretraining gives a model broad language, code, and factual representations.
  • Post-training shapes how the model behaves and solves tasks.
  • Reinforcement learning rewards useful behavior, particularly where answers can be checked, such as mathematics and programming.
  • Distillation transfers useful behavior into smaller models.

DeepSeek described an R1-Zero path that applied reinforcement learning more directly, followed by a fuller R1 process that addressed issues including readability and language mixing. This suggested that some reasoning capability could be improved through post-training rather than requiring an endlessly larger pretraining run.

DeepSeek also released six smaller distilled models, including 32B and 70B variants. The R1 weights and code were released under MIT terms that permit commercial use and distillation, according to the project repository (DeepSeek-R1 on GitHub).

What the $5.6 million figure really measures

The $5.6 million estimate is valuable, but only if compared with the right thing. It represents the estimated GPU rental-equivalent cost of the reported V3 final training run: approximately 2.788 million GPU-hours multiplied by an assumed $2 per H800 GPU-hour.

It does indicate

  • The direct computational cost of the disclosed final run was unusually low compared with the capital spending often associated with frontier AI.
  • Architecture and systems engineering can substantially reduce the compute required for a capable model.
  • GPU-hours provide a useful measure for comparing individual training runs.

It does not indicate

  • That DeepSeek built its entire company for $5.6 million.
  • That it purchased its complete GPU fleet for that amount.
  • That earlier experiments, failed runs, evaluations, or data preparation were free.
  • That research staff, networking, facilities, storage, software, electricity, deployment, and maintenance cost nothing.
  • That V3 and R1 together cost only $5.6 million.
  • That other AI laboratories necessarily spent the same amount inefficiently.

A realistic cost taxonomy includes hardware acquisition, rented compute, pretraining, post-training, failed experiments, salaries, data, facilities, networking, evaluation, inference, security, and ongoing maintenance. A final-run estimate is one line in that budget.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The separate figure of approximately $294,000 sometimes associated with R1 must also be handled carefully. It refers to a reported estimate for a particular R1 training run using 512 H800 GPUs, not the complete development cost of DeepSeek-R1 or the wider V3-to-R1 program (reported estimate).

What remains unverified about DeepSeek’s chips

DeepSeek’s technical report documents H800 use for V3, but the public record does not establish the exact composition or provenance of its entire hardware fleet.

Several possibilities have been discussed:

  1. Export-compliant H800s: H800s could initially be sold under the rules then in effect, and DeepSeek said it used chips that could legally have been purchased in 2023.
  2. Pre-ban inventory: DeepSeek’s wider parent ecosystem may have obtained Nvidia hardware before restrictions tightened.
  3. Cloud access: Computing capacity could potentially be rented indirectly, although the scale and legality of any such access would need evidence.
  4. Intermediaries or smuggling: Reports and analysts have raised allegations about restricted chips reaching China through indirect channels.

These categories should not be conflated. The documented fact is DeepSeek’s reported use of 2,048 H800 GPUs. Additional pre-ban or indirectly sourced hardware is plausible but unproven in the public record. Claims that DeepSeek used illegally exported restricted chips remain allegations, not established facts (Reuters report; Time background).

What about claims that DeepSeek copied OpenAI?

OpenAI and other observers raised concerns that DeepSeek may have trained models using outputs from larger proprietary systems. That possibility is relevant, but it should not be presented as proven without direct evidence.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Distillation itself is a normal machine-learning technique. An authorized developer can use a teacher model to train a smaller student model. The concern arises when a developer systematically collects outputs from a proprietary API in violation of the provider’s terms or uses model-generated data without authorization.

There is also a separate issue: public internet data may contain answers generated by other models, sometimes without the knowledge of the model developer. That is different from deliberate, unauthorized extraction. Reporting on the allegations is available from Axios, but the allegation does not erase the documented engineering work in DeepSeek’s architecture and training system.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Did export controls fail?

There is no reliable binary answer.

If the goal was to prevent any Chinese company from producing a powerful AI model, DeepSeek shows that restrictions did not achieve that goal. A well-funded research group could still combine export-compliant or previously acquired hardware with efficient software and post-training.

If the goal was to limit access to the newest and fastest accelerators, make large-scale training more expensive, slow the construction of giant clusters, and restrict the pace of capability development, DeepSeek’s success does not by itself prove failure. The controls may still impose meaningful costs even when they do not stop progress entirely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The policy outcome should therefore be measured using several questions:

  • Did the controls reduce access to the most capable chips?
  • Did they increase the cost or time required to train large models?
  • Did they encourage more efficient architectures?
  • Did they shift demand toward Chinese-designed accelerators?
  • Did they increase incentives for stockpiling, cloud workarounds, or smuggling?

DeepSeek is best understood as evidence that export controls create constraints—not as proof that hardware no longer matters or that controls have no effect.

Why the story matters for AI economics

DeepSeek’s example has consequences beyond geopolitics. It challenged the assumption that capability always requires proportionally more compute and capital. Sparse architectures, better memory management, low-precision training, and reinforcement learning can improve the amount of useful capability obtained from a fixed hardware budget.

But training economics and serving economics are different. A model can be inexpensive to train yet costly to operate if it requires substantial memory, complex expert routing, or long reasoning traces. Conversely, attention improvements, quantization, batching, and efficient deployment can reduce the cost of serving.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For businesses, the relevant question is not simply “How cheap was DeepSeek’s training run?” It is:

What is the total cost of producing a useful answer at the required quality, latency, reliability, privacy level, and scale?

Open weights also shift costs rather than eliminate them. Developers can run, modify, and distill DeepSeek models themselves, but they must pay for GPUs, storage, networking, monitoring, security, upgrades, and engineering. A low-cost hosted API may be cheaper for occasional use; self-hosting may be preferable for sensitive data or predictable high-volume workloads.

Current status: the V3/R1 story is now historical

The chip-ban controversy primarily concerns DeepSeek-V3 and DeepSeek-R1. It should not be read as a claim that every later DeepSeek model used exactly the same hardware or methods.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

As of August 18, 2026, DeepSeek’s official transparency center lists DeepSeek-V3.2, released December 1, 2025, and DeepSeek-V4, released April 24, 2026 (DeepSeek transparency center). DeepSeek’s change log also says the legacy API names deepseek-chat and deepseek-reasoner were scheduled for discontinuation on July 24, 2026, with compatibility mapping to V4-Flash during the transition (official change log).

That update matters for anyone evaluating the technology today: the V3/R1 breakthrough explains the origin of DeepSeek’s disruption, while current product, pricing, hardware, and API decisions require checking current official documentation.

Bottom line

DeepSeek developed a powerful AI model despite U.S. chip restrictions by making the most of constrained but still substantial computing resources. Its disclosed V3 run used 2,048 Nvidia H800 GPUs—chips designed for the Chinese market under earlier export thresholds—not a completely chip-free alternative.

The company’s real advantage was hardware-aware optimization: sparse Mixture-of-Experts routing, Multi-head Latent Attention, FP8 training, communication engineering, and reinforcement-learning-heavy post-training. The frequently cited $5.6 million was a narrow estimate for one final training run, not the total cost of DeepSeek’s research and business.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The strongest conclusion is therefore nuanced: U.S. controls raised barriers and limited access to the newest hardware, but they did not prevent a capable Chinese team from advancing through efficiency, existing resources, and open-weight distribution.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Written by TheFinanceBase Team

The Team behind TheFinanceBase.

Add your note

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.