A reported $20 billion price tag makes the Groq–NVIDIA transaction sound like a conventional acquisition. It was not publicly announced that way. Groq and NVIDIA described a non-exclusive inference-technology licensing agreement, alongside the transfer of Groq founder Jonathan Ross, president Sunny Madra and other employees to NVIDIA. Groq said it would remain independent and that GroqCloud would continue operating.
The deal grew from a technical experiment: NVIDIA GPUs handled the compute- and capacity-heavy parts of large-language-model inference, while Groq’s processors handled latency-sensitive token generation. That division convinced NVIDIA that Groq’s technology could extend, rather than replace, its GPU platform.
What was announced on December 24, 2025?
Groq’s announcement called the arrangement a non-exclusive inference-technology licensing agreement. It did not disclose a $20 billion purchase price or describe NVIDIA as acquiring Groq outright.
- Jonathan Ross, Groq’s founder, moved to NVIDIA.
- Sunny Madra, Groq’s president, moved to NVIDIA.
- Other Groq employees also joined NVIDIA.
- Groq said it remained an independent company.
- GroqCloud continued operating.
EE Times reported that the licensing and hiring arrangement was worth $20 billion. That figure should therefore be treated as reported consideration, not an officially confirmed acquisition price. Groq’s subsequent announcement of $650 million in growth capital in June 2026 further demonstrated that the company had not simply been dissolved into NVIDIA. Groq said at that time that it operated 13 data centers, served more than five million developers and thousands of AI-native companies, and was targeting expansion toward 200 megawatts by the end of 2027.
#1 Best Overall
Calling the event “NVIDIA bought Groq” is shorthand that obscures the legal structure. The public facts support a license, personnel transfers and continuing corporate independence; the full financial and contractual terms are not public.
The “why not” moment
According to EE Times’ account, NVIDIA made NVLink access available to partners in early 2025. Groq asked whether it could use the interconnect protocol to connect its accelerator to NVIDIA GPUs. Ross said Jensen Huang answered, “why not.”
That response removed a practical barrier. Groq obtained NVIDIA GPUs and tested a heterogeneous system in which two very different processors shared one inference workflow. Ross told EE Times that the teams produced a working demonstration in roughly three weeks after presenting the idea to Huang, after which discussions accelerated. The timeline is Ross’s account of the technical effort, not an independently audited description of the legal closing.
Why LLM inference can be split between processors
Generating an answer from a language model is not one uniform operation. It has two broad phases, with different bottlenecks.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #2
| Inference stage | What it does | Typical bottleneck | Reported hardware fit |
|---|---|---|---|
| Prefill | Processes the prompt and its existing context | Compute and memory capacity for long inputs | NVIDIA GPU |
| Attention and KV-cache work | Retains and moves context information during generation | Capacity and data movement | NVIDIA GPU, depending on implementation |
| Decode | Generates output tokens sequentially | Memory bandwidth and per-user latency | Groq LPU for selected decode work |
Prefill
Prefill processes the user’s prompt and context. It is generally compute-intensive and benefits from GPUs’ broad software support, high throughput and comparatively large memory capacity.
Decode
Decode produces the response one token at a time. A service can have excellent total throughput while still feeling slow to each individual user if decode latency is high. Groq’s proposition targets that interactive experience: rapid, predictable token generation for coding, agents, voice applications and other real-time uses.
EE Times described a more specific split in which Vera Rubin systems perform prefill and attention-related work while Groq LPUs perform the feed-forward-network portion of decode. The exact partition depends on the model, context length, batching and software implementation.
What makes Groq’s LPU different?
Groq’s architecture is built around a large on-chip SRAM resource and a compiler that statically schedules computation. Instead of relying on a general-purpose accelerator to discover much of the execution plan dynamically, the compiler orchestrates where and when operations run. The intended result is predictable timing and low-latency inference.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchEE Times reported approximately 500 MB of SRAM for Groq 3; that figure is a reported product detail and should not be generalized to every Groq chip. SRAM is fast, but on-chip capacity is limited compared with the memory available in a GPU-based system. An LPU therefore may not hold all model weights, context and cache data needed for every stage of a large model by itself.
That limitation explains the complementarity. GPUs supply broad capacity and flexibility; LPUs supply a specialized, tightly scheduled output stage. Groq is not simply “better than GPUs,” and an LPU-only design is not a universal replacement for a GPU platform.
What NVIDIA licensed
EE Times reported comments from NVIDIA executive Ian Buck that the company licensed Groq’s full software stack. The reported components include:
- The Groq compiler.
- Software for sharding inference across multiple chips.
- The LPU software stack and interconnect technology.
- Software that coordinates GPU and LPU execution.
Groq engineers reportedly joined NVIDIA’s Dynamo inference-software team. This software dimension helps explain the strategic value: a fast chip is difficult to deploy at scale without compilation, model partitioning, communication and orchestration tools. The precise legal scope of the license and any field-of-use restrictions have not been disclosed.
Rank #4
- Author: Guillebeau, Chris.
- Publisher: Currency
- Pages: 304
- Publication Date: 2012-05-08
- Edition: NO-VALUE
Groq 3 LPX and the Vera Rubin platform
NVIDIA later presented the Groq 3 LPX rack as part of its Vera Rubin platform in its platform announcement. EE Times reported that a rack contains 256 Groq chips, with eight LPUs in each compute tray. NVIDIA describes an MGX-based, liquid-cooled rack designed to operate beside Vera Rubin GPU racks.
The intended architecture is a heterogeneous AI factory: GPU racks provide model capacity and broad throughput, while Groq racks accelerate latency-sensitive decode. EE Times reported a deployment ratio ranging from one LPX rack to one through four Vera Rubin racks; that ratio is a reported design description, not a universal requirement.
NVIDIA claimed that combining Vera Rubin and Groq hardware could deliver up to 35 times higher inference throughput for certain workloads and an opportunity approaching $300 billion per gigawatt for AI-factory customers. In its own announcement, NVIDIA described up to 35 times higher inference throughput per megawatt and up to 10 times more revenue opportunity for trillion-parameter models. These are NVIDIA-presented claims, not independent benchmarks; results depend on model, batch size, context, traffic and the comparison system.
Why the economics focus on tokens per user
AI infrastructure is often discussed in aggregate tokens per second, but interactive services also care about tokens per second per user. Faster individual responses can improve coding assistance, voice interaction, tool-using agents and other applications where users notice pauses between tokens.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Best Value
A provider may accept a more specialized architecture if it increases useful revenue or retention per watt, even when a GPU cluster remains superior for aggregate throughput, training or irregular workloads. The Groq thesis is therefore about token economics and service quality, not a claim that every workload should move from GPUs to LPUs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Trade-offs of the three deployment models
| Approach | Advantages | Constraints |
|---|---|---|
| GPU-only | Large memory capacity, mature CUDA ecosystem, training and prefill support, broad model compatibility | Lowest-latency decode may require overprovisioning |
| LPU-only | Predictable execution and fast token generation | Limited SRAM capacity, greater model-placement demands, narrower workload fit and compiler dependence |
| GPU–LPU system | Assigns each inference phase to a suitable architecture and can improve interactive efficiency | Requires synchronization, partitioning, networking, cooling and additional software complexity |
When the hybrid approach may not help
- Short prompts and short answers: communication and scheduling overhead can outweigh decode gains.
- Large context windows: GPU memory and KV-cache handling may dominate the workload.
- Heavy batching: aggregate GPU throughput can be more valuable than per-user latency.
- Small models: a specialized rack may not justify its cost or operational complexity.
- Irregular or rapidly changing models: static compilation and operator support can be harder.
- Training: the Groq proposition described here is primarily inference, not a replacement for NVIDIA’s training platform.
What happened to Rubin CPX?
EE Times reported that NVIDIA had considered Rubin CPX, a design aimed at improving prefill or time-to-first-token economics. After the Groq agreement, NVIDIA reportedly put that product on the back burner while focusing on decode performance and dollars per token. Ian Buck told the publication the concept could return in a later generation, so this should not be read as a permanent cancellation.
What remains unknown about the $20 billion figure
- The exact financial consideration.
- How much reflects technology licensing, personnel commitments or other arrangements.
- The detailed scope and restrictions of the non-exclusive license.
- Customer pricing and deployment schedules for LPX systems.
- Independent measurements of performance, total cost and power efficiency.
Non-exclusive licensing also leaves open how broadly Groq can serve other customers and how much of its independent cloud business remains commercially separate. Public announcements do not provide every downstream term.
How readers can access the technology
Most developers will encounter this technology through a service rather than by purchasing a rack. GroqCloud remains available through Groq’s console, and Groq’s announcement offered a free way to try the service. Current quotas and paid prices can change and were not specified in the public material summarized here.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Enterprise buyers evaluating inference providers should compare time to first token, sustained tokens per second per user, batch throughput, model and context support, structured output, geography, privacy terms, uptime commitments and price per input and output token. The available claims do not establish that Groq or the hybrid NVIDIA platform is universally fastest or cheapest.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




