DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Huawei’s Ascend 910C Moves Into China’s AI Infrastructure—But It Is Not Yet an Nvidia Replacement

Huawei’s Ascend 910C is now deployed in 384-chip Chinese systems, but conflicting output estimates and HBM, packaging and software constraints make it a domestic substitute—not yet a global Nvidia replacement.
From TheFinanceBase Team8 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Huawei’s Ascend 910C has moved beyond the April 2025 promise of imminent shipments into early large-scale deployment inside China. Huawei says its Atlas 900 A3 SuperPoD can connect up to 384 Ascend 910C chips, and the CloudMatrix384 paper describes a 384-chip CloudMatrix384 system. Those are meaningful production systems, not a retail launch or proof of global parity with Nvidia.

The investment and supply-chain question is therefore narrower and more important: can Huawei obtain enough advanced manufacturing, high-bandwidth memory (HBM), packaging, software support and electricity to make Ascend a repeatable domestic platform? Public evidence says Huawei has built a credible constrained-market alternative, while shipment totals and manufacturing independence remain uncertain.

What changed after the original 910C shipment report?

On April 21, 2025, Reuters reported that Huawei planned to begin mass shipments of the 910C to Chinese customers as early as May. The report, based on two unnamed sources, said some chips had already shipped and that the product was intended to fill part of the gap left by Nvidia’s restricted China-market H20. It also linked production of major components to Semiconductor Manufacturing International Corp. (SMIC) and described low yields as a continuing problem. The account was not a Huawei disclosure of shipment volume. Reuters report via Investing.com

By August 2026, the most defensible description is “early large-scale domestic deployment.” Huawei’s own roadmap claims a 384-chip Atlas 900 A3 SuperPoD, while the CloudMatrix384 paper describes 384 Ascend 910C NPUs paired with 192 Kunpeng CPUs. These systems show that Huawei can assemble and operate very large clusters. They do not establish a transparent annual output figure, worldwide distribution or one-for-one replacement of Nvidia’s latest products.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HPE NVIDIA Tesla V100 32GB HBM2 PCIe 3.0 x16 Passive GPU Computational Accelerator for AI Machine Learning HPC Deep Learning 699-2G500-0216-400 (Renewed)
  • NVIDIA Volta GV100 Architecture — 4,608 CUDA Cores, 640 1st-Gen Tensor Cores delivering 14 TFLOPS FP32 and 112 TFLOPS deep learning performance for AI training, inference, HPC, and scientific computing workloads
  • 32GB HBM2 ECC Memory — 900 GB/s Bandwidth — High-bandwidth memory on a 4096-bit bus with ECC error correction provides the memory capacity and throughput required for the largest AI models, simulations, and datasets
  • PCIe 3.0 x16 Interface — 250W TDP — Standard PCIe Gen3 connectivity with passive cooling designed for enterprise rack server deployment in HPE ProLiant, Dell PowerEdge, and Supermicro platforms with adequate chassis airflow
  • NVLink — Scale to 96GB Unified Memory — Connect two V100 GPUs via NVLink at 300 GB/s bi-directional bandwidth to scale GPU memory from 32GB to 96GB for larger AI training and HPC workloads
  • Multi-Precision Computing — Supports FP64 (7 TFLOPS), FP32 (14 TFLOPS), FP16 (112 TFLOPS) and INT8 precision modes for flexible deployment across training, inference, and scientific simulation workloads

Reuters also reported in March 2026 that ByteDance and Alibaba planned orders for a newer Huawei AI chip after testing. That is evidence of customer interest in Huawei’s broader roadmap, not a disclosed order quantity for 910C. The same report said Huawei had struggled to persuade private companies to adopt the 910C in large quantities. Reuters report via Investing.com

What the Ascend 910C is

The 910C is Huawei’s flagship AI accelerator in this period. Public reporting generally describes it as combining two 910B-derived dies or chiplets in one package for training and inference. It is usually compared with Nvidia H100- or H20-class hardware, not with Nvidia’s newest Blackwell generation. Public figures for memory, bandwidth, power and process technology vary by configuration and source, so they should not be treated as a single verified specification sheet.

Huawei’s product significance is less about a single card than about a complete Ascend stack: accelerator hardware, Atlas servers, interconnects, CANN software and Huawei Cloud services. The company’s roadmap is available in its September 2025 keynote. Huawei roadmap announcement

“At scale” has five different meanings

Reports often use “mass production” and “at scale” interchangeably. For investors, cloud buyers and policymakers, they describe different milestones:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Scale level What must be demonstrated What is established for 910C
Commercial shipment Completed accelerators delivered to customers Reuters reported planned 2025 shipments; Huawei has not published a definitive total.
Cluster scale Hundreds of chips operating as one production system Huawei claims a 384-chip Atlas system; CloudMatrix384 is described in the CloudMatrix384 paper.
Market scale Enough supply for major clouds and independent AI developers Customer interest exists, but broad private-sector adoption is not publicly quantified.
Supply-chain scale Output large enough to materially reduce dependence on imported accelerators Possible in selected Chinese segments; public production estimates conflict.
Global scale Meaningful deployment outside China and influence on worldwide market share Not established.

That distinction matters because one 384-chip supernode is strategically important but is not equivalent to Nvidia’s global supply, software ecosystem or installed base.

Rank #2
MX3 M.2 AI Accelerator
  • High-Performance AI Processing: The MX3 is designed to handle the most demanding AI computer vision workloads, delivering exceptional performance and efficiency.
  • Flexible Integration: The MX3 can be easily integrated into your existing systems via its M.2 M-key form factor and support for Linux operating systems.
  • Energy Efficient: The MX3 is designed to provide high performance while minimizing power consumption.
  • Comprehensive Software Development Kit (SDK): The MX3 is supported by a comprehensive SDK that simplifies development and deployment.
  • Hardware compatability: The MX3 is compatible with the PCI-SIG M.2 M-key 2280 Specification. It can be used with the Raspberry Pi 5 with a M-key 2280 HAT.

How many 910Cs can Huawei make?

No company-confirmed production total is available. A U.S. Commerce Department official told Reuters that Huawei’s advanced AI-chip capacity was likely at or below 200,000 chips in 2025, with most or all delivered inside China. Capacity is not the same as completed shipments. Reuters report via Investing.com

Other 2026 analyses estimate roughly 250,000–300,000 equivalent 910C chips, while separate reporting has suggested output could exceed 700,000 units by the end of 2026. These figures cannot be averaged into a single forecast: sources may count complete packages, dies, chiplets, inventory, Huawei Cloud allocations or equivalent compute rather than physical accelerators.

  • Congressional testimony discusses estimates around 250,000–300,000 equivalent chips and HBM constraints.
  • The Wire China reports a materially higher estimate using a different methodology.

Until Huawei publishes shipment data or customers disclose repeat purchases, the prudent conclusion is that production is substantial enough for strategic Chinese deployments but not measurable with Nvidia-like transparency.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The bottlenecks are wider than chip design

Manufacturing and yield

Reuters reporting associated SMIC with production of Huawei accelerator components using an enhanced 7-nanometer process. The exact configuration is not fully disclosed, and low yield means wafer starts do not translate directly into sellable accelerators. Reuters report via Investing.com

HBM and advanced packaging

HBM is a critical constraint because modern AI accelerators need high-capacity, high-bandwidth memory mounted close to the compute dies. Congressional testimony and industry analyses identify HBM availability as a limiting factor for Huawei output. Packaging capacity, materials and testing are additional constraints.

Rank #3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
  • ✅Powered by 26 Tera-Operations Per Second (TOPS) Hailo-8 AI Processor. 2.5W typical power consumption
  • ✅Scalable, enabling simultaneous processing of multi-streams & multi-models
  • ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
  • ✅Supports TensorFlow, TensorFlow Lite, ONNX, Keras, Pytorch frameworks
  • ✅Supports Linux and Windows. Supports the temperature range of -40°C to 85°C

Foreign inputs

A Chinese-designed accelerator is not automatically a fully indigenous product. Design sovereignty means Huawei controls the architecture. Manufacturing sovereignty means key dies can be fabricated domestically. Packaging sovereignty requires China to assemble the complete high-bandwidth package. Supply-chain sovereignty would require domestic access to equipment, electronic-design automation, memory, materials and manufacturing know-how. The 910C demonstrates progress in the first two categories without proving the fourth.

Power and cooling

Large clusters can compensate for weaker per-chip performance by using more accelerators, but that raises electricity, cooling, networking and capital requirements. A system-level comparison is therefore incomplete without performance per watt, rack density and total cost of ownership.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How it compares with Nvidia

Decision factor Ascend 910C in China Nvidia alternatives
Availability Strategically favored and less exposed to U.S. export restrictions inside China, but allocation may be limited. Broad global availability varies by product and export rules; China access is product-specific.
Per-chip performance Generally positioned against H100/H20-class hardware; not established as equal to current Blackwell products. Leading products offer stronger mature benchmark and efficiency positions, depending on model.
System scaling Huawei emphasizes 384-chip SuperPoD and CloudMatrix designs. Nvidia has mature large-scale networking and reference systems.
Software CANN and vLLM-Ascend support require migration and retuning from CUDA. CUDA has the broadest libraries, tools and developer base.
Power and cost More chips can raise power and cooling costs; workload-specific data is required. Often stronger performance per accelerator, but purchase access and pricing depend on region.
Support footprint Strongest in China through Huawei’s integrated enterprise and cloud channels. Broader international cloud and system support.

The CloudMatrix384 paper documents a 384-NPU, 192-CPU system and workload-specific evaluations. Tom’s Hardware reported comparisons suggesting aggregate throughput advantages in some configurations alongside a substantial power penalty. Neither source proves a universal benchmark win. CloudMatrix384 paper · Tom’s Hardware analysis

Why Chinese buyers may adopt it anyway

  • Supply certainty: A domestic allocation may be more dependable than a restricted imported product.
  • Policy alignment: State and infrastructure procurement can favor Chinese hardware.
  • Data sovereignty: Domestic systems simplify requirements for workloads that must remain in China.
  • System optimization: Huawei can tune hardware, networking, software and support as one stack.
  • Strategic insurance: A dual-stack or Ascend-first strategy reduces exposure to future export-control changes.

These advantages can outweigh lower per-chip efficiency for inference, public-sector workloads and Chinese-hosted services. Training teams deeply invested in CUDA face a higher migration cost.

The software gap is a business cost

Moving from CUDA to CANN can require rewriting kernels, replacing libraries, changing memory layouts, retuning distributed training and revalidating numerical results. Huawei said it planned to open-source or open parts of CANN and Ascend virtual-instruction-set interfaces based on 910B/910C designs by the end of 2025. Opening interfaces may attract developers, but it does not immediately reproduce CUDA’s accumulated tooling and community.

Rank #4

A 2026 field study of Ascend inference deployments documents the practical role of CANN and vLLM-Ascend in large-model workloads. Ascend deployment study

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What U.S. export controls changed

Export controls have two opposing effects. They restrict Huawei’s access to advanced manufacturing, HBM, equipment and some foreign accelerators, which can slow volume growth. At the same time, they increase the value of a domestic substitute, encourage Chinese developers to move away from CUDA and give government buyers a reason to standardize on Ascend.

It is therefore inaccurate to say simply that controls failed or succeeded. They reduce China’s access to leading-edge foreign hardware while accelerating a more separate Chinese compute ecosystem. Policy details have changed over time and apply differently to products, companies and inputs; “the ban” is not one single measure. Associated Press background

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What this means for the global AI supply chain

The structural change is fragmentation, not immediate worldwide Nvidia replacement. China is building an Ascend-centered stack that may include domestic accelerators, Huawei Cloud, Chinese model frameworks and locally sourced infrastructure. The U.S. and allied market remains centered on Nvidia/CUDA and associated suppliers. Both stacks may still depend on overlapping foreign memory, equipment, materials and software inputs.

  • Chinese cloud providers may standardize more workloads on Ascend.
  • Model developers may maintain separate Nvidia and Huawei implementations.
  • HBM, advanced packaging and networking become more strategically valuable.
  • Nvidia’s China business may become structurally smaller even if some products regain export permission.
  • International buyers face more compliance, interoperability and support decisions.

Checks for investors and infrastructure buyers

For Chinese AI companies

  • Confirm guaranteed allocation, HBM-backed capacity and Huawei support terms.
  • Price model-porting labor and validation, not just accelerator hardware.
  • Measure total cost of ownership, including power, cooling and networking.
  • Decide whether a dual-stack deployment is needed for portability.

For cloud providers

  • Look for stable commercial API capacity rather than demonstration access.
  • Ask which frameworks, model formats and inference engines are supported.
  • Demand customer-workload benchmarks with power and latency assumptions.
  • Check whether workloads can move back to Nvidia or other hardware.

For international buyers

Ascend is generally a poor fit where CUDA-only libraries, global support, transparent public pricing, latest performance-per-watt or non-Huawei procurement rules are essential. Huawei Cloud’s ModelArts page is the relevant managed-service entry point, but 910C pricing and availability are contract- and region-dependent. Huawei Cloud ModelArts

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
ASRock Radeon AI PRO R9700 Creator 32GB Professional Graphics Card, 2920 MHz Boost Clock, GDDR6, AMD RDNA 4, AI-Accelerators, DisplayPort 2.1a, PCIe 5.0, Blower Cooler
  • Professional AI & Creator Workstation: AMD Radeon AI PRO R9700 GPU with 32GB GDDR6 is engineered for AI development, professional content creation, and compute-intensive workloads.
  • Massive 32GB Memory Capacity: 32GB of GDDR6 memory on a 256-bit bus provides ample bandwidth for large AI models, 8K video editing, and complex 3D rendering.
  • Advanced RDNA 4 with AI Accelerators: 64 Compute Units with 3rd Gen Ray Tracing and dedicated 2nd Gen AI Accelerators for groundbreaking AI performance and visual computing.
  • Professional Blower Cooling: Efficient single blower design exhausts heat directly out of the chassis, ideal for multi-GPU workstation and server configurations.
  • Enterprise-Grade Thermal Solution: Vapor chamber heatsink with industrial Honeywell PTM7950 thermal interface material ensures reliable cooling under sustained professional loads.

What would prove that the shift is durable?

  1. Huawei-confirmed shipment totals that distinguish packages, dies and systems.
  2. Repeat orders disclosed by major private cloud and AI companies.
  3. Stable 910C capacity through commercial cloud APIs.
  4. Independent benchmarks on representative training and inference workloads.
  5. Evidence of greater domestic HBM and packaging output.
  6. Lower, documented software-porting costs and broader operator coverage.
  7. Deployments beyond government-linked or Huawei-internal projects.
  8. A measurable decline in Nvidia’s China share rather than isolated substitutions.

Frequently Asked Questions

Is the Ascend 910C a direct replacement for Nvidia’s newest chips?

No. It is a credible China-market alternative for selected workloads, but public evidence does not establish parity with Nvidia’s newest Blackwell-generation accelerators.

Does a 384-chip Huawei system prove mass production?

It proves cluster-level deployment. It does not disclose annual output, market share or global supply.

Is the 910C fully Chinese?

Huawei controls the design and Chinese fabrication is associated with SMIC, but HBM, packaging, equipment, software and materials mean complete supply-chain independence has not been demonstrated.

The Bottom Line

Huawei has turned the 910C from a promised 2025 shipment into a real component of China’s emerging AI infrastructure. Its strategic advantage is availability and system-level integration under export restrictions, not proven parity with Nvidia per accelerator. For investors, the key indicators are repeat commercial orders, verified output, HBM and packaging capacity, software adoption and energy cost—not a single headline production estimate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

Bestseller No. 2
MX3 M.2 AI Accelerator
MX3 M.2 AI Accelerator
Software and Documentation can be accessed at the MemryX developer website
$169.00
Bestseller No. 3
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
waveshare Hailo-8 M.2 AI Accelerator Module, Compatible with Raspberry Pi 5, Supports Linux/Windows Systems, Based On The 26TOPS Hailo-8 AI Processor, Module Only
✅Scalable, enabling simultaneous processing of multi-streams & multi-models; ✅Enabling real-time, low latency and high-efficiency AI inferencing on the edge devices
$219.99
Bestseller No. 4
Tesla L40S 48GB AI HPC Graphics Accelerator
Tesla L40S 48GB AI HPC Graphics Accelerator
48GB AI graphics accelerator
$6,199.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.