October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
The Finance Base
The Money Desk · Blog
Re:

Nvidia Touts New Storage Platform and Confidential Computing for Vera Rubin NVL72 Server Rack

Nvidia's Vera Rubin NVL72 pairs 72 Rubin GPU packages with 36 Vera CPUs, an inference KV-cache storage platform, rack-scale confidential computing and serviceability features. Here is what is known, what remains unverified and how buyers should evaluate it.
From TheFinanceBase Team7 min to read
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Nvidia used its January 2026 CES keynote to present the Vera Rubin NVL72 as more than a faster GPU system. The rack combines 72 Rubin GPU packages, 36 Vera CPUs, high-speed scale-up and scale-out networking, an inference-context storage tier for reusing large-language-model (LLM) key-value (KV) cache, and a rack-scale confidential-computing design. Nvidia said Rubin was in full production, but partner systems were expected to become available in the second half of 2026—not as a generally priced, off-the-shelf product.

What Nvidia announced at CES 2026

The announcements center on three infrastructure changes:

  • Inference Context Memory Storage Platform: an AI-oriented storage tier intended to retain and reuse LLM KV cache.
  • Rack-scale confidential computing: Nvidia says the Vera CPU, Rubin GPUs and their NVLink domain can operate inside one trusted execution environment.
  • Serviceability and resiliency: cable-free modular trays, NVLink Intelligent Resiliency and a second-generation RAS Engine are designed to keep a rack operating through selected maintenance and diagnostic operations.

The primary reported specifications and claims come from CRN’s CES 2026 coverage. Nvidia’s terminology, benchmark claims and “first” descriptions should be treated as vendor statements until detailed architecture documents and independent tests are available.

What the Vera Rubin NVL72 is

NVL72 is Nvidia’s flagship rack-scale Rubin configuration for tightly coupled, scale-up AI workloads. It contains 72 Rubin GPU packages and 36 custom Vera CPUs, rather than the four- or eight-GPU layout common in conventional servers. Nvidia positions it for large distributed models and DGX SuperPOD clusters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Tecmojo 12U Open Frame Network Rack for IT & AV Gear, AV Rack Floor Standing or Wall Mounted,with 2 PCS 1U Rack Shelves & Mounting Hardware,Network Rack for 19" Networking,Audio and Video Device
  • 【Powerful Load-bearing】12U Network Rack Open Frame is constructed from durable cold rolled steel; Rack shelf supports enhance stability, wall-mounted capacity of 130lbs, the ground-mounted up to 260lbs
  • 【Considerate Designs】Open-frame layout, including a top panel adding space, anti-slip shelf stops fixing devices and compatible racks for stack and expansion to meet requirements of home server rack
  • 【Complete Accessories】A 12U open frame server rack, two ventilated shelves, four shelf stops, four velcro straps and a set of equipment mounting screws
  • 【Versatile Application】Ideal for space-efficient multi-device setups in warehouses, retail, classrooms, offices and more; Excellent choices as AV Rack/IT Rack
  • 【Effortless Setup】 Network Rack includes hardware, a comprehensive manual, mounting hole drilling template and an online assembly video to simplify setup

Nvidia previously called the design Vera Rubin NVL144. That number counted 144 GPU dies; each package contains two dies. NVL72 counts the 72 physical GPU packages. Nvidia also lists an HGX Rubin NVL8 configuration for eight-GPU servers using x86 CPUs, which is a different deployment class.

Inside one rack

Component Reported specification
Rubin GPUs 72 packages; 50 petaflops NVFP4 inference and 35 petaflops NVFP4 training per GPU; 22 TB/s HBM4 bandwidth; 3.6 TB/s NVLink bandwidth per GPU
Vera CPUs 36 CPUs; each has 88 Olympus cores, 176 threads using Nvidia spatial multi-threading, 1.5 TB LPDDR5X, 1.2 TB/s memory bandwidth and 1.8 TB/s NVLink chip-to-chip bandwidth
Rack totals 3.6 exaflops NVFP4 inference, 2.5 exaflops NVFP4 training, 54 TB LPDDR5X, 20.7 TB HBM4, 1.6 PB/s HBM4 bandwidth and 260 TB/s scale-up bandwidth

These are aggregate, low-precision throughput figures. NVFP4 results are not directly comparable with FP16, BF16, FP8 or full-precision performance, and application results depend on model, software and workload.

How the context-memory storage platform works

When an LLM reads a prompt, attention layers generate intermediate key and value tensors. A serving system keeps those tensors in a KV cache so it does not recompute the same context for every subsequent token. On a long conversation, a retrieval-augmented-generation (RAG) request or an agent that calls several tools, the cache can become a substantial working set.

Nvidia’s Inference Context Memory Storage Platform is intended to make that working set shareable and reusable instead of treating it as ordinary cold or general-purpose network storage. The architecture uses BlueField-4 DPUs and Spectrum-X Ethernet to create what Nvidia calls an AI-native storage tier. The important claim is about the lifecycle of inference context—placing, sharing, restoring and evicting KV cache—not simply using faster SSDs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Where it can help

  • Multi-turn assistants that repeatedly revisit a conversation.
  • Long-context applications with large prompts.
  • RAG systems that repeatedly access similar retrieved context.
  • Agentic workflows that perform several reasoning or tool-use steps.
  • Distributed serving in which multiple GPUs or inference workers can reuse a cache.

The benefit depends on cache-hit rate, sequence length, concurrency, eviction policy, network topology and whether the serving framework can externalize and restore KV cache efficiently. A short, stateless request with little reuse may gain nothing; a cache that is constantly evicted can add complexity without reducing computation.

Rank #2
Sale
StarTech 42U 4-Post Open Frame Rack, 19in, 22-40in, 1323lb/600kg
  • ADJUSTABLE DEPTH: 4-Post 42U open frame server rack with 4 vertical rails and adjustable mounting depth 22" to 40" (56,0cm to 101,7cm); Compatible with various servers / switches / data / AV and other IT equipment; EIA/ECA-310-E Compliant
  • EASY ASSEMBLY: Mobile network rack with easy-to-follow assembly instructions and online video; Compact flat-pack shipping to avoid damage and facilitate installation; Total product height of 80.3in (204 cm) with casters, 78in (198cm) without casters
  • COLD ROLLED STEEL: Durable 4 Post 19in open frame rack designed for ventilation with 42U mounting height and 1320lb (600kg) weight capacity (stationary); 3 install options included: casters, levelling feet, or base-plate to secure rack to the floor
  • HARDWARE INCLUDED: Rolling computer/data rack includes cage nuts and screws to mount equipment, easy to read Units (U) and depth adjustment markings, cable management hooks for organization, and required assembly tools
  • THE IT PRO'S CHOICE: Designed and built for IT Professionals, this 42U rack is backed for 2-years, including free lifetime 24/5 multi-lingual technical assistance

Operational edge cases

  • Cache entries can become invalid after a model, tokenizer, quantization setting or serving configuration changes.
  • Shared caches need coherency rules; stale or incompatible entries can force recomputation or affect output quality.
  • Local GPU memory may still be the best option when strict locality and very low latency matter.
  • Inference runtimes, schedulers, orchestration and observability tools must support the external or distributed cache.

What performance Nvidia claims

Nvidia executive Dion Harris said the context-memory platform can deliver, versus traditional network-storage approaches for inference context:

Claim How to read it
Up to 5× higher tokens per second Nvidia’s stated result; the report does not provide workload, cache-hit rate, storage media or measurement method
5× better performance per dollar A vendor economic comparison, not an independently verified total-cost result
5× better power efficiency A vendor claim whose test conditions and system boundaries were not disclosed

“Tokens per second” can mean aggregate rack throughput, per-user throughput or a selected benchmark. Buyers should request the prompt length, output length, concurrency, latency target, cache-hit rate, software versions, baseline hardware and whether capital, networking and power costs are included.

Rubin compute, memory and networking

Vera CPU

Each Vera CPU has 88 custom Olympus cores and 176 threads, 1.5 TB of LPDDR5X memory, 1.2 TB/s of memory bandwidth and 1.8 TB/s of NVLink chip-to-chip bandwidth. Nvidia says Vera delivers twice Grace’s performance for data processing, compression and code compilation; that comparison is Nvidia’s claim, not an independent benchmark.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Rubin GPU

Nvidia reports 50 petaflops of NVFP4 inference and 35 petaflops of NVFP4 training per GPU, plus 22 TB/s of HBM4 bandwidth and 3.6 TB/s of NVLink bandwidth. Against Blackwell, Nvidia claims 5× inference, 3.5× training, 2.8× HBM bandwidth and 2× NVLink bandwidth. Those ratios are format- and workload-dependent.

Scale-up and scale-out fabrics

Scale-up is communication among CPUs and GPUs inside the rack. The liquid-cooled NVLink 6 Switch uses 400G SerDes, provides 3.6 TB/s per-GPU bandwidth and 28.8 TB/s aggregate bandwidth, and offers 14.4 teraflops of FP8 in-network computing.

Rank #3
VEVOR 12U Open Frame Server Rack, 23-40 in Adjustable Depth, Free Standing or Wall Mount Network Server Rack, 4 Post AV Rack with Casters, Holds All Your Networking IT Equipment AV Gear Router Modem
  • Adjustable Depth: 23-40'' adjustable depth is used for servers and network equipment, ensuring enough space for AV equipment, components, and cabling, while allowing you to access ports and equipment from multiple sides.
  • Strong Load Capacity: Ground-Mounted Load Capacity: 500 lbs, Wall-Mounted Load Capacity: 150 lbs. The av rack is made of carbon steel for better weldability performance and can help save space while meeting your need to place multiple devices.
  • User-friendly Design: Ergonomic design makes the open frame av rack easier to use. The additional top panel is able to place other items with more available space. Roller design moves anywhere and anytime, is convenient, and is more energy-saving.
  • Complete Accessories: We provide the accessories you need, including 2 x Pallets, 145 x M5*10 Cross Head Screws, 4 x Casters, 4 x M10*50 Expansion Screws,10 x M6*12 Cage Nuts, 1 x Grounding Wire, 1 x User Manual.
  • Wide Application: The server rack wall mount maximizes the use of available space, suitable for retail venues, classrooms, offices, and other places where space is limited.

Scale-out connects racks, servers, storage and clusters. ConnectX-9 SuperNICs handle network traffic, while BlueField-4 DPUs offload infrastructure functions and underpin the context-memory architecture. Spectrum-X Ethernet is part of Nvidia’s AI-oriented storage and networking design.

What rack-scale confidential computing means

Nvidia says Vera Rubin provides the first rack-scale Trusted Execution Environment spanning the Vera CPUs, Rubin GPUs and the NVLink domain that connects them. The intended protection covers proprietary models, training data and inference data while computation is occurring across the CPU-GPU system.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Three data states

  • At rest: storage encryption and key management protect data on drives or other persistent media.
  • In transit: network or interconnect encryption protects data moving between components.
  • In use: hardware-isolated execution and cryptographic controls protect data during computation.

A CPU-GPU-NVLink boundary is not automatically an end-to-end security guarantee. Logs, telemetry, checkpoints, external KV-cache storage, firmware, drivers, hypervisors, host software, key-release services and operator access must be included in a verifiable trust model.

The available announcement does not specify the attestation protocol, measured firmware components, key-provisioning workflow, complete software coverage, threat model or confidential-mode performance limits. Security teams should obtain those details before treating the rack as suitable for regulated workloads.

Confidential computing also does not guarantee model integrity, correct outputs, availability or immunity to application-layer attacks. Encryption and isolation can complicate debugging, observability and low-level administration, and may introduce performance or key-management overhead.

Rank #4
AxcessAbles 12U Network Rack with Wheels - 500lb Capacity, 18" Depth | 19-Inch Open Frame AV Rack Case with 3” Caster Wheels | Screws, Spacer, Tool Included
  • Universal 19” Rack Mount Compatibility – Perfect for pro audio, video, IT, and network gear. Compatible with mixers, routers, patch panels, servers, power amps, and more.
  • Heavy-Duty Load Capacity – Built to support up to 550 lbs. Ideal for studio gear, DJ setups, server equipment, and AV components that demand serious stability.
  • Robust Steel Frame & Design – Made with 1.5mm thick steel and weighs 36 lbs for maximum durability, reduced vibration, and long-term reliability in any setting.
  • Mobile & Secure – Preinstalled with 3” industrial-grade caster wheels (lockable), making it easy to move and position your rack exactly where you need it.
  • All-In-One Setup Kit Included – Comes with 34 rack screws (5mm & 6mm), a 1U blank spacer, and an assembly tool—ready for fast installation out of the box.

Resiliency and maintenance features

Nvidia describes the NVL72 as having third-generation rack-resiliency features:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Cable-free modular trays.
  • A claimed 18× faster assembly and service process.
  • NVLink Intelligent Resiliency.
  • Claimed zero-downtime maintenance for switch trays, including removing or partially populating trays while the rack remains operational.
  • A second-generation RAS Engine and GPU diagnostics without taking the rack offline.

“Zero downtime” is a capability claim, not a promise that every component can be removed during every workload. It can depend on redundancy, workload placement, software version, supported tray-removal procedure and operational discipline.

Who can realistically deploy an NVL72?

The rack is aimed at hyperscalers, cloud providers, large enterprises and research organizations running high-concurrency or very large distributed models. It can be excessive for small models, low inference volumes, CPU-heavy applications, short prompts or workloads that cannot reuse KV cache.

Acquisition is likely to occur through Nvidia enterprise sales, approved OEMs, cloud capacity reservations or managed infrastructure—not a standard online checkout. Nvidia said Rubin was in “full production,” while related systems were expected through partners in the second half of 2026. That distinction matters: platform production does not mean every NVL72 configuration, storage component, OEM system or cloud service is generally available.

Questions to put in a vendor request

  1. What exact GPU-package, CPU, HBM and system-memory configuration is offered?
  2. What storage media, usable capacity, endurance, redundancy and persistence model support the context tier?
  3. Which inference runtimes and schedulers support external KV-cache placement, migration, invalidation and encryption?
  4. What attestation documents, measured components and key-management integrations are provided?
  5. What are the rack’s power, liquid-cooling, dimensions and facility requirements?
  6. What service-level agreement applies to switch-tray maintenance and GPU diagnostics?
  7. What regional delivery date and support model apply?
  8. What benchmark conditions support each 5× claim?
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What remains unanswered

The announcement does not establish public pricing, rack power draw, cooling specifications, storage capacity or endurance, exact OEM delivery dates, a standardized product SKU, or whether the context platform can be purchased separately. It also does not explain how KV cache is indexed, compressed, encrypted, migrated or invalidated, nor does it provide independent testing against local NVMe, distributed filesystems, object storage, CXL memory or competing memory-tiering designs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
VEVOR 9U Open Frame Server Rack, 23''-40'' Adjustable Depth, Free Standing or Wall Mount Network Server Rack, 4 Post AV Rack with Casters, Holds All Your Networking IT Equipment AV Gear Router Modem
  • Adjustable Depth: Depth adjustable from 23" to 40", this open frame server rack accommodates servers and network equipment while providing ample space for A/V gears and cable management. Enjoy easy access to ports and devices from multiple angles.
  • High Weight Capacity: Supports up to 300 lbs on the floor (200 lbs when adjusted to maximum depth) and 200 lbs when wall-mounted (depth cannot be adjusted in wall-mounted mode). Made from carbon steel for superior welding performance and durability, this open frame rack is designed to save space while accommodating multiple devices.
  • User-Friendly Design: Designed with your convenience in mind, this open frame server rack features an top shelf for extra storage and improved space utilization. The rolling casters let you move it effortlessly wherever you need it, making setup and movement a breeze.
  • Widely Applicable: Maximize your space with this adaptable open frame server rack, designed to make the most of every inch. Ideal for retail spots, classrooms, offices, and any area where space is at a premium, it delivers practical solutions for your storage needs.
  • Everything You Need: Our open-frame rack comes with fully equipped accessory kit for easy setup and secure installation: 2 x Trays, 4 x Casters, 1 x set of Screws, 16 x M6*12 Cage Nuts, 1 x Grounding Wire, 1 x Internal & External Hex Wrenches, and 1 x User Manual.

Cloud providers and infrastructure companies named in industry coverage—including AWS, Microsoft, Google Cloud and CoreWeave—may offer Rubin-related capacity, but exact instance names, regions, launch dates and prices require direct confirmation. Relevant starting points are AWS Machine Learning, Microsoft Azure AI, Google Cloud AI Infrastructure and CoreWeave Cloud.

How to evaluate the business case

Compare the complete serving system, not headline FLOPS. Estimate the proportion of requests with reusable context, expected cache residency, latency targets, concurrency, model-update frequency and the cost of power, cooling, networking, software and operations. Then compare the NVL72 with a smaller GPU cluster, local NVMe cache, a distributed filesystem, CXL memory or cloud capacity.

The strongest case is a large deployment where long-lived context and high cache reuse dominate inference cost, and where the organization can operate liquid-cooled rack-scale infrastructure. The weakest case is a small or stateless deployment with low reuse, strict data locality, limited facilities or no software support for distributed KV-cache management.

Bottom line

Vera Rubin NVL72 is Nvidia’s attempt to make AI infrastructure a coordinated compute, memory, networking, storage and security system. Its context-memory platform could reduce repeated prompt processing when KV-cache reuse is high; rack-scale confidential computing could strengthen protection for data in use; and modular resiliency features could simplify maintenance. The practical value will depend on cache-aware software, independently reproducible benchmarks, a complete attestation and key-management story, facility readiness and partner availability in the second half of 2026.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More post from the Money Desk

  1. The Money DeskBlogTheFinanceBase07 MAR 2625 minWhat Is a 457 Plan?
  2. The Money DeskBlogTheFinanceBase07 MAR 2621 minTime Value of Money: What It Is and How It Works
  3. The Money DeskBlogTheFinanceBase07 MAR 2627 minAre You Living in One of These Top 10 Most Expensive Cities to Retire?
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.