Free tools Windows power users keep installed
One-click scans. No signup required.
Cerebras did announce a major AI-inference expansion, but not “just now.” Its March 11, 2025 announcement projected more than 40 million Llama 70B tokens per second across a planned network of six new datacenters. That is aggregate, model-specific capacity—not the speed of one facility or proof that Nvidia’s broader data-center business is collapsing.
The credible threat is narrower and more interesting: Cerebras may pressure Nvidia in latency-sensitive inference, particularly the output-generation stage used by real-time agents, coding assistants, search and conversational services. Its 2026 partnerships with AWS and AMD make that challenge more practical, while also showing that Cerebras may be a specialized component in heterogeneous systems rather than a wholesale Nvidia replacement.
What Cerebras actually announced
In its March 11, 2025 release, Cerebras said it was launching six new AI-inference datacenters across North America and Europe. The facilities would use Cerebras CS-3 systems built around the company’s Wafer-Scale Engine.
Cerebras said the expansion would provide expected aggregate capacity of more than 40 million Llama 70B tokens per second—approximately 20 times its previous capacity, according to the company. It also said 85% of the total capacity would be in the United States.
#1 Best Overall
- EVOLUTION AMD RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
Ownership and operation
- Oklahoma City and Montreal were described as using hardware exclusively owned and operated by Cerebras.
- The remaining sites were described as jointly operated with strategic partner G42.
- The Oklahoma City facility was described as containing more than 300 CS-3 systems and using Scale Datacenter infrastructure.
- Montreal was described as being operated by Enovum, a Bit Digital division.
The location list needs careful reading
The release does not present a clean list of six newly completed facilities. It mixes online locations, planned facilities and broader regional targets:
| Location or region | Status in Cerebras’s release |
|---|---|
| Santa Clara, California | Online |
| Stockton, California | Online |
| Dallas, Texas | Online |
| Minneapolis, Minnesota | Targeted for the second quarter of 2025 |
| Oklahoma City, Oklahoma | Listed for the third quarter of 2025; another passage said June 2025 |
| Montreal, Canada | Targeted for the third quarter of 2025 |
| Midwest/Eastern United States | Regional deployment target for the fourth quarter of 2025 |
| Europe | Regional deployment target for the fourth quarter of 2025 |
That combination explains why the headline should be treated as an announced deployment plan, not independent confirmation that every listed site was operating at the promised capacity.
What “40 million tokens per second” means
The 40-million figure is an aggregate capacity projection for Llama 70B across the planned network. It is not the performance of one CS-3, one datacenter or one customer request. “Expected to serve” also makes clear that the number was forward-looking when announced.
It is tied to one model and one type of measurement
The qualification matters: Cerebras specified Llama 70B. Results for another model, quantization, context length or software stack could differ substantially. The release also does not establish whether the figure means per-request speed, sustained service capacity across many customers, or a particular concurrency level.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchAny comparison with Nvidia must match at least the model, quantization, input and output lengths, batch size, concurrency, latency target, networking, power and cost basis. Cerebras itself warns that performance comparisons vary by workload, configuration, model, date and testing method.
Rank #2
- Built for Local AI Development: AMD Ryzen AI Halo is designed for local AI development and inference, featuring 128GB unified memory and support for up to 200B parameter models to build and run intensive AI workloads locally.
- 128GB Unified Memory: Features 128GB LPDDR5x unified memory at 8000 MT/s with 256 GB/s memory bandwidth, providing a shared memory pool across the CPU, GPU, and NPU to support larger AI models.
- AMD Ryzen AI Max+ 395 Processor: Features 16 cores, 32 threads, and Zen 5 architecture, paired with AMD Radeon 8060S integrated graphics featuring 40 RDNA 3.5 compute units and an AMD XDNA 2 NPU with up to 50 TOPS.
- Linux AI Developer Platform: Purpose-built for Linux-based AI development with full AMD ROCm software support and preloaded tools, models, and workflows optimized for local AI development.
- Compact, Connected Design: Includes a 2TB M.2 SSD, 10GbE LAN, Wi-Fi 7, Bluetooth 5.4, USB-C connectivity, and HDMI 2.1b.
Output generation is not the whole inference job
Inference has two important stages:
- Prefill processes the prompt and context. It is highly parallel and often compute-intensive.
- Decode generates the answer one token at a time. It is more serial and depends heavily on memory bandwidth and latency.
Cerebras’s strongest proposition is fast decode. A service can therefore feel dramatically faster while still needing a separate system to process long prompts efficiently. The relevant production measures are time to first token, sustained output tokens per second, latency at realistic concurrency, cost per input and output token, and availability—not a headline throughput number alone.
Why Cerebras can challenge Nvidia in inference
Latency-sensitive applications
Fast sequential generation matters to interactive coding tools, real-time agents, search and answer engines, conversational applications and reasoning systems that produce long responses. Faster output can improve user experience, let an AI service complete more work in a given period and potentially reduce the conventional GPU capacity needed for decode-heavy traffic.
Specialized inference economics
A purpose-built accelerator could be attractive when the business is selling large volumes of generated tokens. Buyers should compare:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Cost per million input tokens and output tokens.
- Tokens per second per watt.
- Tokens per second per dollar of capital expenditure.
- Latency and utilization under production concurrency.
- Network, reservation, minimum-commitment and engineering costs.
Cerebras markets fast inference, queue priority and enterprise throughput through its pricing page. Its page also advertises “20× faster” language; that is a vendor claim, not a universal benchmark.
More alternatives for developers and clouds
Cerebras said its inference platform was used by organizations including OpenAI, Cognition, Mistral, Perplexity, Hugging Face and AlphaSense. Customer references do not by themselves prove workload migration or market share, but broader access can make an alternative architecture more visible to developers.
Rank #3
- EVOLUTION RYZEN AI MAX+ 395 MINI PC - GMKtec EVO-X2 is the next evolution in AI mini PC Ryzen Strix Halo series. Thanks to AMD Simultaneous Multithreading (SMT) the core-count is effectively doubled, to 32 threads. Ryzen AI Max+ 395 has 64 MB of L3 cache and can boost up to 5.1 GHz, depending on the workload. The Ryzen AI Max+ 395 is currently rated as the "most powerful x86 APU" on the market for AI computing.
- AI NPU with XDNA 2 ARCHITECTURE - Powered by 16 “Zen 5” CPU cores, 50+ peak AI TOPS XDNA 2 NPU and a truly massive integrated GPU driven by 40 AMD RDNA 3.5 CUs, the Ryzen AI MAX+ 395 is a transformative upgrade and delivers a significant performance boost over the competition. The Ryzen AI Max+ 395 excels in consumer AI workloads like the llama.cpp-powered application: LM Studio. Shaping up to be the must-have app for client LLM workloads, LM Studio allows users to locally run the latest language model without any technical knowledge required and unleash their creativity and productivity.
- AMD RADEON 8090S iGPU GAMING PC - The AMD Radeon RX 8060S offers all 40 CUs with up to 2.9 GHz graphics clock and uses the new RDNA 3.5 architecture. The powerful iGPU is positioned between an RTX 4060 and 4070 laptop GPU and therefore enables gaming in FHD at maximum details in most demanding games. The 8060S can also utilize the full 128GB pool, which is perfect for running LLMs such as Deepseek 70B Q8, which runs comfortably on this machine.
- EIGHT CHANNEL LPDDR5X - LPDDR5X is a new ground breaking memory small form factor installed on-board. With blazing speeds up to to 8000MT/s, it runs 1.5x faster than the DDR5 SODIMMs; 90% better performance over DDR5 SODIMMs in video conferencing and photo editing; 30% better performance in productivity apps; 12% better performance in digital content workloads.
- QUAD SCREEN 8K DISPLAY SUPPORT - EVO-X2 AI Mini PC support 4-screen 4K/8K output via HDMI 2.1 (8K@60Hz), DisplayPort 1.4 (4K@60Hz), and dual USB 4 40Gbps Transfer speed (supporting PD3.0/DP1.4/DATA). Ideal for gaming, video editing, and multitasking, it provides expansive and crisp multi-display support.
The strategic pressure on Nvidia is larger than one benchmark. If inference can be split across different processors, cloud providers may optimize each stage instead of buying one dominant accelerator for every task.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What the AWS and AMD deals change in 2026
AWS: Trainium for prefill, Cerebras for decode
AWS says its planned collaboration will use Trainium for prompt processing and Cerebras CS-3 systems for output generation, connected with Elastic Fabric Adapter networking. The service is intended to be available through Amazon Bedrock, with rollout described as occurring in the coming months and additional models later in 2026. Details and timing remain forward-looking in the AWS announcement.
This is important because customers could obtain Cerebras-style decode without designing a separate datacenter. It is also evidence that Cerebras can complement another processor rather than replace it.
AMD: Helios for throughput, Cerebras for low-latency decode
AMD and Cerebras announced a planned system pairing AMD Helios with Cerebras hardware. AMD positions Helios for throughput, prompting and large contexts, while Cerebras handles low-latency token generation. Initial availability was expected through Cerebras Cloud in the second half of 2026, according to the announcement.
The companies claim up to five times higher tokens per second per watt. That figure requires workload, baseline and configuration details before it can be treated as an apples-to-apples comparison with a specific Nvidia system.
Rank #4
Why this does not mean Nvidia is in immediate trouble
Nvidia sells a broader platform
Nvidia competes across training, fine-tuning, inference, networking, rack-scale systems, software and enterprise support. CUDA and its related libraries, installed base and developer ecosystem remain central advantages. Nvidia is also building token-scale inference infrastructure and financing or partnering with AI clouds, as described in its AI-infrastructure strategy.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Its cloud-partner program lists providers including Crusoe, GMI Cloud, Lambda, Nebius, SoftBank, Lightning AI and Yotta. That gives customers multiple ways to obtain Nvidia-compatible infrastructure without buying directly from Nvidia.
Cerebras is specialized
Cerebras may be a poor fit when a team needs to train frontier models, support arbitrary custom architectures, use mature CUDA libraries, run small or irregular workloads, or standardize on one platform for training, fine-tuning, inference and networking. Long-context workloads may benefit from disaggregated prefill and decode, but short prompts and short answers may not justify the added complexity.
Capacity is not market share
The 40-million-token target does not establish sustained utilization, revenue, profitability, paying-customer count, cost per token, current installed capacity or customer migration away from Nvidia. It demonstrates infrastructure ambition, not competitive displacement.
How to evaluate Cerebras for a real workload
- Check model support. Confirm that the required model, custom weights, quantization and fine-tuning workflow are available.
- Define the latency objective. Separate time to first token from output-token speed and total response time.
- Measure the prompt-to-output ratio. Long contexts may favor a prefill/decode design; short interactions may not.
- Benchmark production concurrency. Require results with the expected number of simultaneous users, not only a single-request demonstration.
- Calculate total cost. Include input and output tokens, network transfer, minimum commitments, reserved capacity and engineering effort.
- Verify availability. Check the actual region, model, rate limits, queue priority, failover and multi-region redundancy.
- Test software integration. Confirm API compatibility, streaming, tool calling, structured output, observability and deployment behavior.
- Review enterprise terms. Verify service-level commitments, incident handling, support and custom-model policies.
The investor takeaway
Cerebras is a credible threat to Nvidia’s dominance in selected inference workloads, especially low-latency decode. The 2025 datacenter announcement did not show that Nvidia’s training business, software ecosystem or overall data-center franchise was collapsing.
The more consequential development is the move toward heterogeneous inference systems. AWS is pairing Trainium with Cerebras; AMD is pairing Helios with Cerebras; Nvidia is expanding its own full-stack and cloud-partner strategy. The competitive question is shifting from “which chip has the fastest headline?” to which complete system delivers the best latency, throughput, cost, software compatibility and availability for a specific production workload.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




