Ant Group reportedly cut AI-model training costs by about 20% using accelerators associated with Alibaba and Huawei, combined with a Mixture of Experts (MoE) design. People familiar with the work told Bloomberg that internal tests produced results comparable to Nvidia H800 systems. That is a significant demonstration of hardware-and-software co-design, not proof that Chinese chips now match Nvidia across AI workloads or that Ant has stopped using Nvidia.
What Ant Group actually claimed
The company behind the reported result is Ant Group, the Alibaba-affiliated Chinese fintech founded by Jack Ma. Ant operates substantial cloud-computing and artificial-intelligence infrastructure, but it is not a chip manufacturer.
According to Bloomberg’s March 24, 2025 report, Ant used Chinese-made hardware linked to Alibaba and Huawei. Alibaba develops AI hardware and cloud infrastructure, while Huawei produces Ascend AI accelerators. The report said Ant’s approach could reduce model-training costs by approximately 20% and that people familiar with the testing considered the results comparable with Nvidia H800-based systems.
Those are reported internal results. The public account does not identify the exact accelerator models, quantities, cluster design, software versions, training run, or independent benchmark.
#1 Best Overall
- PLEASE NOTE: Exporting an NVIDIA RTX Pro 6000 GPU outside the US requires strict adherence to the U.S. Export Administration Regulations (EAR) and issuance of an export license from the Bureau of Industry and Security (BIS). Compliance and Know Your Customer (KYC) screening may be required as a condition of order acceptance. [NVIDIA Blackwell Streaming Multiprocessor] The new SM features increased processing throughput, and new neural shaders that integrate neural networks inside of programmable shaders | DLSS 4: Multi Frame Generation ensures ultra-smooth frame pacing for lifelike simulations.
- [Double-Flow-Through Design] The RTX PRO 6000 Blackwell features a double-flow-through cooling design, optimizing efficiency and airflow to sustain peak performance under 600W power loads. | [5th Gen Tensor Cores] Deliver up to 3X the performance of the previous generation and support for FP4 precision for faster AI model processing times with reduced memory usage, enabling local fine-tuning of LLMs and generative AI | [4th Gen Ray Tracing Cores] Double the ray-triangle intersection rate of the previous generation to create photoreal, physically accurate scenes and immersive 3D designs with RTX Mega Geometry, which enables up to 100X more ray-traced triangles.
- [PCIe Gen 5] Support for PCIe Gen 5 provides double the bandwidth of PCIe Gen 4, improving data-transfer speeds from CPU memory and unlocking faster performance for data-intensive tasks like AI, data science, and 3D modeling. | [GDDR7 Memory] With 96 GB of GPU memory and 1.8 TB ps bandwidth, it can tackle massive 3D and AI projects, fine-tune AI models locally, explore large-scale VR environments, and drive larger multi-app workflows.
- [DisplayPort 2.1] Achieve unparalleled visual clarity and performance, driving high resolution displays at up to 8K at 240 Hz and 16K at 60 Hz. Increased bandwidth enables seamless multi-monitor setups while HDR and higher color depth support ensures superior color accuracy for precision work, such as video editing, 3D design, and live broadcasting.
- [Universal MIG] Divide a single RTX PRO 6000 Blackwell into multiple isolated instances, each with dedicated resources, allowing for concurrent execution of multiple workloads, optimized GPU utilization, and secure isolation of different applications or users. [WARRANTY] 3 YR Manufacturer's Warranty. Bulk OEM Packaging. Retail Packaging is NOT included.
Where the 20% saving comes from
“A 20% reduction” refers to the cost of training an AI model under the reported setup. It does not mean Chinese chips cost 20% less than Nvidia products. Nor does it establish a 20% reduction in Ant’s total AI budget, data-center construction, electricity bill, inference costs, or every model the company operates.
Training cost can include accelerator time, cloud or depreciation charges, electricity, networking, storage, and engineering effort. Savings may come from cheaper or more available hardware, higher utilization, lower energy use, or avoiding constrained imports. The available reporting does not break the figure into those components.
A Bloomberg-based account in The Standard gave an illustrative comparison of about 6.35 million yuan to train one trillion tokens on high-performance hardware versus roughly 5.1 million yuan with an optimized lower-specification approach. Those figures are attributed examples, not an independently verified price list or universal cost benchmark.
Why Mixture of Experts matters
MoE is a model-architecture and systems technique, not a chip brand. An MoE model contains multiple specialized “experts” and a router. For each token, the router activates only a subset of experts rather than running every parameter.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Sparse activation
Because fewer parameters are active for each token, the system can perform less computation while retaining a large total parameter count. That can lower accelerator time and make a cluster of less powerful devices more useful.
Optimization, not automatic parity
The benefit comes from coordinating model design, routing, scheduling, compilers, and hardware. It cannot be attributed entirely to the silicon. MoE also introduces routing and communication overhead, possible load imbalance, more complicated checkpointing, and sensitivity to batch size and sequence length. A well-tuned MoE workload may favor a domestic accelerator; a dense or communication-heavy workload may not.
Rank #2
- VD8465 Japanese Authorized Distributor Product
- The speed of FP32 calculation is twice as fast as previous generations, which greatly improves the complex 3D processing and graphics simulation workflow
- Up to 2X the throughput compared to previous generations and significantly faster workloads such as video content rendering, architectural design assessments, and virtual prototypes of product design
- Achieve more than twice the previous generation AI performance improvement, support faster FP8 precision data and accelerate the execution of mixed flotation decimal and whole numbers
- It has a large capacity of memory necessary for working with a vast array of data sets and workloads such as rendering, data science, and simulation
What “only Chinese chips” gets wrong
The headline wording overstates the scope. Bloomberg-syndicated reporting said Ant was still using Nvidia for AI development, while relying more on alternatives—including AMD and Chinese chips—for newer models. The specific demonstration may have used Chinese hardware, but that is different from an Nvidia-free company-wide infrastructure.
It helps to separate four claims:
- Specific experiment: Reportedly used Chinese accelerators with an MoE-based approach.
- Production environment: May contain a mixture of Chinese, Nvidia, AMD, and other hardware.
- Company-wide development: Still included Nvidia, according to the Bloomberg-based account.
- Strategic direction: Greater use of domestic alternatives is plausible, but complete replacement has not been established.
Ant’s continuing Nvidia use is reported by Bloomberg Law and summarized by Moneycontrol.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
How strong was the Nvidia comparison?
The stated comparison target was Nvidia’s H800, a China-market product affected by U.S. export controls—not Nvidia’s unrestricted H100 or newer H200 and Blackwell families. “Comparable” is therefore a narrow and ambiguous claim.
| Reported item | What is established | What remains unknown |
|---|---|---|
| Hardware | Chinese chips associated with Alibaba and Huawei | Exact models, quantities, fabrication, memory and interconnect configuration |
| Architecture | Mixture of Experts was part of the approach | Expert count, routing policy and software implementation |
| Cost result | About 20% lower model-training cost was reported | Hardware, energy, cloud, networking and engineering-cost breakdown |
| Performance target | Results reportedly similar to Nvidia H800 systems | Throughput, time to convergence, energy, cost per token and final model quality |
| Evidence | People familiar with the matter, as reported by Bloomberg | Public methodology, independent replication and production-scale validation |
Similar final benchmark scores are not the same as equal speed, reliability, software maturity, total cost of ownership, or performance across models. A test can also be tuned for one precision format, batch size, model, or communication pattern.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Why the report mattered in 2025
The claim arrived amid U.S. restrictions on advanced AI-chip exports to China, China’s push for domestic substitutes, and intense interest in efficient-model techniques following DeepSeek’s resource-efficiency claims. Its significance is therefore both technical and geopolitical.
Technical signal
The report suggested that architecture and software can partially compensate for weaker or less readily available accelerators. Hardware, model design, compilers, and cluster scheduling can be optimized together instead of judged only by peak chip specifications.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRank #3
- Small in Size, Serious in Performance — a space-saving design delivering professional-class performance, enterprise-grade security and reliability, flexible deployment options, and a MIL-STD-810H–certified build engineered for demanding work environments.
- Extreme AI and professional graphics performance — The ThinkStation P3 Ultra SFF Gen 2 combines an integrated Intel NPU with NVIDIA RTX 4000 SFF Ada Generation graphics (20GB GDDR6) to deliver up to 335 TOPS of AI performance across CPU and GPU. Ideal for AI inferencing, deep learning, 3D animation, content creation, advanced imaging, 3D modeling, and BIM software—all in a compact, energy-efficient workstation.
- Fast, secure storage with next gen memory & business-ready OS — 2TB PCIe Gen 5 TLC Opal SSD for ultra fast boot and load times, MAXED OUT 128GB DDR5-6400MHz memory, and Windows 11 Professional preinstalled.
- Easy-access front connectivity — USB-A (USB 10Gbps), 2 x USB-C (USB4 20Gbps) – data transfer only, Headphone/mic combo
- Warranty — Factory Sealed. 1 Year Lenovo Warranty
Supply-chain signal
Domestic hardware can be valuable even when it is not the fastest option. Local availability, policy support, predictable procurement, and freedom from export restrictions can improve resilience and capacity planning.
Neither point proves that China has solved the semiconductor challenge. Large AI systems also depend on advanced manufacturing, high-bandwidth memory, packaging, networking, power delivery, cooling, storage, compiler support, reliability, maintenance, and enough engineers to operate the cluster.
What would verify the claim?
A stronger technical record would include:
- A published Ant paper or engineering report.
- Named chip models and accelerator counts.
- Training time, throughput, energy use and cost per token.
- Model size, dataset, token count, precision and convergence criteria.
- Details of networking, memory, compilers and frameworks.
- Results on dense as well as MoE models, and on inference as well as training.
- Independent replication by researchers, cloud customers or third-party benchmarkers.
- Evidence from sustained production deployments rather than a controlled demonstration.
The initial reports do not supply those details, so the 20% figure should be treated as a reported result rather than a settled industry benchmark.
The practical interpretation
Ant’s report is best understood as evidence that Chinese organizations can co-design models and software around available domestic hardware. That can lower training costs for selected workloads and reduce dependence on imported accelerators. It does not show that a particular Chinese chip independently delivers Nvidia-equivalent performance, that Chinese chips are universally cheaper, or that export controls have failed.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchThe narrow claim—an optimized Ant training setup reportedly approached H800 results at about one-fifth lower cost—is meaningful. The broad claim—that China has broadly caught Nvidia or that Ant now uses only Chinese chips—is not supported by the public evidence.
The Bottom Line
Ant Group’s reported demonstration points to the power of MoE architecture and systems optimization under hardware constraints. It is a strategically important proof point, but not independent evidence of broad Chinese-chip parity with Nvidia or a complete Nvidia replacement.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




