October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
AI infrastructure

Together AI’s $305M Bet: Why Reasoning Models Can Increase GPU Demand

DeepSeek-R1 challenged the idea that more capable AI always needs more costly training. It did not prove that AI will use fewer GPUs overall: lower costs can expand adoption while reasoning and agentic tasks consume more inference compute.

By TheFinanceBase Team 8 min read

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

DeepSeek-R1 made powerful reasoning models more accessible, but that does not mean AI needs fewer GPUs overall. Lower costs can bring in more users and applications, while long reasoning traces and agentic workflows consume more inference compute per task. Together AI’s $305 million Series B in February 2025 captured that thesis; its later $800 million Series C and more than 500 MW in compute-capacity commitments show the company’s subsequent infrastructure ambitions—not independent proof that DeepSeek-R1 itself caused a global GPU-demand surge.

Why DeepSeek-R1 seemed to threaten GPU demand

When DeepSeek-R1 appeared, its combination of strong reasoning performance, open weights and a disruptive cost narrative raised a question for AI infrastructure investors: if a capable model could be developed and accessed more cheaply, would the industry need fewer high-end accelerators?

The question blends three different economics. Training cost concerns the compute used to develop a model. Serving cost concerns the compute needed to answer users repeatedly. Total infrastructure demand depends on both efficiency and how much people use AI. A change in one does not determine the others.

DeepSeek’s technical paper describes a reinforcement-learning approach to developing reasoning capability. Its reported figures should not be treated as a complete accounting of all research, experimentation, infrastructure or deployment costs. The paper is useful background on the model family, not a full measure of its lifetime economics: DeepSeek-R1 technical paper.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • 0dB technology lets you enjoy light gaming in relative silence
  • Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
  • Dual ball fan bearings last up to twice as long as sleeve bearing designs

How reasoning can use more compute per task

A conventional response may require a relatively short generation. A reasoning model can produce a longer sequence of tokens before it reaches an answer. In a tool-using or agentic application, one user request can also trigger repeated model calls, retries and tool interactions. The result can be more accelerator time per completed task, even if each token becomes cheaper to produce.

Dimension Short, one-shot inference Reasoning or agentic inference
Generation Often a shorter response May involve a longer reasoning trace before the final answer
Calls per task Often one model call May involve multiple model and tool calls
Resource occupancy Typically shorter per request Can occupy compute and memory resources for longer
Serving considerations Can suit shared serving when traffic allows May require more capacity planning to meet interactive latency targets

This is a useful comparison, not a universal rule: request length, model, serving engine, batching, quantization and reasoning budget all affect actual performance. NVIDIA said on a May 28, 2025 earnings call that reasoning models can use thousands more tokens per task than earlier one-shot inference and are increasing inference demand. That is the view of a major chip supplier, not an independent market-wide measurement: NVIDIA earnings-call transcript.

Longer sequences also matter for memory. As a request progresses, the system must manage its growing context and key-value cache, alongside the compute needed to generate tokens. If requests hold resources for longer, a fleet may handle fewer simultaneous requests unless it adds capacity or changes its serving configuration.

Why full DeepSeek-R1 can be demanding to serve

Together AI and VentureBeat describe DeepSeek-R1 as a 671-billion-parameter model. That figure does not mean every serving configuration uses all parameters in the same way, nor does parameter count alone specify the hardware footprint. It does, however, signal why full-scale serving is not equivalent to running a small model on one ordinary server: the model must be distributed across accelerators, with networking and coordination between them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5070 Ti
  • Integrated with 16GB GDDR7 256bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Together AI says R1’s longer reasoning chains increase memory and compute requirements per request, reduce how many concurrent requests a GPU fleet can serve, and raise per-query costs relative to DeepSeek-V3. These are provider statements about serving economics; actual cost depends on model configuration, traffic, hardware and software stack. Its DeepSeek FAQ distinguishes the full model from smaller distilled variants, which may serve more cheaply but can differ in reasoning quality or task reliability.

Open weights lower barriers to access, customization and deployment; they do not remove the need for accelerator memory, fast interconnects, power, cooling, monitoring or reliability engineering. Nor does a large parameter count by itself establish how much aggregate GPU capacity a model will consume: that depends on the number and type of requests it attracts.

The rebound effect: cheaper inference can mean more total demand

When a model becomes cheaper or easier to access, organizations may use it for work they previously avoided: coding assistance, research, document analysis, planning and automation. Better capability can also move AI from occasional experimentation into production systems that run continuously. If each task uses more tokens or calls, total demand can grow even while cost per token falls.

Together AI’s CEO told VentureBeat that a single user request in agentic workflows could lead to thousands of API calls. That is an executive observation, not a typical-workload estimate. The same interview described reasoning clusters sized from 128 to 2,000 chips and reported requests that could take two to three minutes; those details describe the company’s account, not a universal R1 serving benchmark. Read the VentureBeat interview.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
  • Powered by the NVIDIA Blackwell architecture and DLSS 4
  • Powered by GeForce RTX 5060
  • Integrated with 8GB GDDR7 128bit memory interface
  • PCIe 5.0
  • WINDFORCE cooling system

Efficiency and demand can therefore move in opposite directions. Providers can improve tokens per second per GPU, requests per GPU, energy per token or cost per token while total GPU-hours, reserved capacity or power demand rises because more workloads are running. The word “demand” itself can mean accelerator purchases, cloud rentals, GPU-hours consumed, peak capacity reserved or data-center power; those measures need not move together.

What Together AI announced in February 2025

On February 20, 2025, Together AI announced a $305 million Series B led by General Catalyst and co-led by Prosperity7, at a company-reported valuation of approximately $3.3 billion. The company said the financing would support inference, training, fine-tuning and other enterprise infrastructure, including large-scale Blackwell GPU deployment. Its announcement described the following plans and company-reported figures:

  • 200 MW of secured power capacity.
  • A planned Hypertec partnership involving 36,000 NVIDIA GB200 NVL72 GPUs, alongside immediate access to HGX B200 clusters.
  • Support for more than 200 open models and products for inference, training, fine-tuning, agentic workflows and synthetic data.
  • More than 450,000 registered AI developers, as reported by the company.

These are statements from Together AI’s financing announcement; secured power, planned GPU deployments and operating compute capacity are distinct measures. A financing round and infrastructure plan show what the company intended to build, not how much of its later GPU use came specifically from R1. The company’s original announcement is at Together AI’s Series B announcement.

What “reasoning clusters” mean

Reasoning clusters are a dedicated-compute offering, not a new model architecture. Together AI described them as capacity for large-scale, low-latency reasoning inference, with no shared rate limits or resource sharing, custom optimization for customer traffic, and enterprise service-level agreements. VentureBeat reported capacity options from 128 to 2,000 chips.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
  • Powered by Radeon RX 9070 XT
  • WINDFORCE Cooling System
  • Hawk Fan
  • Server-grade Thermal Conductive Gel
  • RGB Lighting

Together AI’s FAQ advertises speeds of up to 110 tokens per second and a 99.9% uptime SLA. These are vendor claims; without a specified model, workload, quantization, batch size and latency metric, they should not be read as guaranteed per-user generation speed or a general comparison with other providers.

The financing timeline—and what it does not prove

Date Company announcement What the figure describes
February 20, 2025 $305 million Series B; approximately $3.3 billion valuation Financing and valuation reported by Together AI
July 1, 2026 $800 million Series C; more than 500 MW in compute-capacity commitments Later financing and capacity commitments reported by Together AI

The Series B is an important marker of the 2025 inference thesis, not Together AI’s latest announced financing. In July 2026, the company announced its Series C and more than 500 MW in compute-capacity commitments. A commitment is not the same thing as installed GPUs or operating capacity, and company announcements do not establish a market-wide causal link between R1 and GPU demand. See Together AI’s Series C announcement.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How the infrastructure business serves different buyers

Inference providers can sell access at different levels, from a shared API to dedicated hardware. Together AI’s offering illustrates the trade-offs, though the categories apply more broadly than this one company.

Shared serverless inference

A serverless API is often the practical starting point for prototypes, variable traffic and teams testing several models. Per-token billing and minimal cluster operations reduce commitment, but shared capacity can bring load-dependent rate limits or variability. Long reasoning traces can also make the total bill larger than expected from a short sample. Together AI says DeepSeek-R1 limits vary by user tier and load, with higher limits available to larger build tiers and enterprise customers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
  • Axial-tech fans now feature a smaller fan hub that facilitates longer blades and a barrier ring that increases downward air pressure
  • Phase-change GPU thermal pad helps ensure optimal heat transfer, lowering GPU temperatures for enhanced performance and reliability
  • 2.5-slot design allows for greater build compatibility while maintaining cooling performance
  • Dual-ball fan bearings last up to twice as long as standard conventional sleeve bearings designs
  • 0dB technology lets you enjoy light gaming in relative silence

Dedicated inference endpoints

A dedicated endpoint suits steadier production traffic, custom models, single-tenant deployment or tighter latency requirements. Together AI describes its dedicated inference as single-tenant GPU deployment with custom configurations, autoscaling and guaranteed performance. The trade-off is paying for reserved capacity during quiet periods and planning for the hardware a large model requires. See its dedicated endpoints documentation.

GPU clusters

A GPU cluster can make sense for sustained inference, training, fine-tuning, large models requiring model parallelism or teams that need control over the serving stack. It also shifts more responsibility to the buyer: utilization, networking, storage, orchestration, health checks and deployment all affect the economics. Hourly compute savings can disappear if the cluster sits idle or needs substantial engineering to operate.

Together AI’s pricing page lists HGX H100 at $3.99 per hour on demand, HGX H200 at $5.99 per hour and HGX B200 at $8.19 per hour. These are time-sensitive listed rates, not all-in estimates for every cluster configuration or workload; check the current pricing page before making a cost comparison. The same page lists DeepSeek-R1 fine-tuning at $10 per million tokens for supervised fine-tuning and $25 per million tokens for DPO, with a $20 minimum charge. Those are fine-tuning rates, not ordinary inference prices.

When to use an API, reserve hardware or self-host

  • Choose serverless inference when traffic is modest or unpredictable, you are prototyping, or you want to compare models without operating a cluster.
  • Consider a dedicated endpoint when traffic is predictable and production latency, single-tenant hardware or custom deployment matters enough to justify reserved capacity.
  • Consider a GPU cluster when utilization is sustained, throughput is high, models need to span multiple GPUs, or you need control for training and custom serving.
  • Benchmark a smaller or distilled model when cost or latency is a constraint and the task does not require full R1’s quality. Measure successful task completion, not token price alone.
  • Self-host when infrastructure control or specialized governance requirements outweigh the engineering and capacity-management burden.

Before comparing options, estimate input tokens, generated and reasoning tokens, calls per user task, retries and tool calls, cache hits, peak concurrency, latency targets and any capacity left idle. A cost-per-token comparison can mislead if one model needs more calls or produces more tokens to finish the same job.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where the demand thesis is strong—and where it remains uncertain

The mechanisms are clear: long outputs can hold resources longer, agentic workflows can multiply calls, and cheaper access can expand usage. Together AI has a commercial interest in selling inference and GPU capacity, while NVIDIA benefits from a larger accelerator market. Their statements are relevant, but they should be attributed rather than treated as neutral market measurements.

The available figures do not establish what share of Together AI’s demand came from DeepSeek-R1, a verified company-wide utilization rate, or a market-wide causal estimate showing that R1 increased global GPU consumption. Nor do they show that every reasoning model has the same serving profile or that lower-cost distilled models cannot reduce total compute for some workloads.

Efficiency improvements—distillation, quantization, better kernels, batching and custom silicon—can reduce compute for a given task. They may also make new applications economical, or shift workloads to smaller models, edge devices or other accelerators. Whether aggregate GPU demand rises depends on whether growth in use outweighs those savings; it is not a law that every efficiency gain increases demand.

Quick Recap

Bestseller No. 1
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
ASUS Dual Radeon RX 9060 XT 16GB GDDR6 Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$529.00
Bestseller No. 2
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
GIGABYTE GeForce RTX 5070 Ti Gaming OC 16G Graphics Card, 16GB 256-bit GDDR7, PCIe 5.0, WINDFORCE Cooling System, GV-N507TGAMING OC-16GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5070 Ti; Integrated with 16GB GDDR7 256bit memory interface
$1,249.99
Bestseller No. 3
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
GIGABYTE GeForce RTX 5060 WINDFORCE OC 8G Graphics Card, Cooling System, 8GB 128-bit GDDR7, PCIe 5.0, Manufactured by NVIDIA, DisplayPort & HDMI - Video Output Interface, GV-N5060WF2OC-8GD Video Card
Powered by the NVIDIA Blackwell architecture and DLSS 4; Powered by GeForce RTX 5060; Integrated with 8GB GDDR7 128bit memory interface
SaleBestseller No. 4
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
GIGABYTE Radeon RX 9070 XT Gaming OC 16G Graphics Card, PCIe 5.0, 16GB GDDR6, GV-R9070XTGAMING OC-16GD Video Card
Powered by Radeon RX 9070 XT; WINDFORCE Cooling System; Hawk Fan; Server-grade Thermal Conductive Gel
$842.14
Bestseller No. 5
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
ASUS Prime Radeon RX 9070 XT 16GB GDDR6 OC Edition Gaming Graphics Card
0dB technology lets you enjoy light gaming in relative silence; Dual BIOS switch lets you toggle between Quiet and Performance BIOS profiles
$829.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Money Desk

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.