Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Supermicro announced a 2U all-flash storage system on October 15, 2024, designed to serve AI and high-performance computing (HPC) clusters. It combines PCIe Gen5 SSDs with up to four NVIDIA BlueField-3 data-processing units (DPUs)—not NVIDIA GPUs—to accelerate storage and networking tasks. Supermicro claimed more than 250 GB/s of SSD bandwidth and up to 1.105 petabytes (PB) of raw capacity; neither figure is an independent benchmark or a measure of usable capacity.
What Supermicro announced
The October 2024 product is a JBOF, short for “Just a Bunch of Flash”: a hardware platform that houses flash drives and connects them to servers and a network. It is intended for software-defined storage deployments supporting AI training and inference, HPC, analytics, object storage, and parallel file systems. It is not a consumer NAS or a complete, ready-to-use storage service by itself.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
Gvdlink NMFP7E20 Optical Multimode Splitter Fiber Cable 5m (16.4ft) MPO12 to 2xMPO12 LSZH OM4 for... | $49.00 | Buy on Amazon |
The key distinction is that the NVIDIA components are BlueField-3 DPUs, not GPUs. A DPU is an infrastructure processor designed to handle tasks such as networking and storage processing. Supermicro’s design places storage software on the DPU’s 16 Arm cores, aiming to move some work away from a conventional server CPU and memory subsystem. The company describes the system as dual-port and capable of active-active clustering, though actual availability depends on the storage software and deployment design.
How the storage path works
A simplified data path is:
NVMe SSDs → PCIe Gen5 → BlueField-3 DPU → 400-Gb Ethernet or InfiniBand → compute servers and GPUs
#1 Best Overall
- The MFP7E20-Nxxx cable for NVIDIA, is a multimode, 4-channel-to-two 2-channel splitter fiber cable. The Multiple Push On, 12 fiber, Angled Polished Connectors (MPO-12/APC) uses 8 active fibers to transmit light and 4 inactive fibers as strength members. The Angled Polished Connector has a 8-degree polished angle to deflect internal optical back reflections from entering the transceivers and distorting the signal quality
- The 4-channel end is inserted into a Twin port OSFP, 800Gb/s transceiver. The 2-channel ends are inserted into two, single-port 400Gb/s OSFP and/or QSFP112 transceivers which with only 2 fibers can output 200G rates. Two splitter fiber cables are used in the twin-port OSFP transceiver enabling four, 2-channel ends to four transceivers.
- The fibers are “crossover”, Type-B cables enable directly attaching two transceivers together and allow the transmit laser fiber on pin 1 to “crosses over” and align with pin 12 of the opposite fiber end transceiver photodetector.
- The typical usecase is linking OSFP switches to in ConnectX-7 network adapters and/or BlueField-3 Data Processing Units (DPUs) in compute and storage servers.
- Rigorous cable production testing ensures best out-of-the-box installation experience, performance, and durability. For NVIDIA’s optical solutions provide short, medium, and long reach scalability for all topologies, utilizing innovative optical technologies to enable high signal integrity and reliability
The DPU can handle networking and storage functions, including RoCE (RDMA over Converged Ethernet), encryption, compression, and erasure-coding offloads. RDMA can move data between systems with less CPU involvement than conventional networking paths. NVIDIA also cites support for GPUDirect Storage and GPU-initiated storage, technologies intended to make data movement between storage and GPU memory more direct or efficient.
NVIDIA describes GPUDirect Storage as a DMA path between storage and GPU memory that can avoid a CPU bounce buffer. That does not bypass every software layer or guarantee a particular application speedup. It depends on a compatible GPU, drivers, operating system and kernel path, storage configuration, and application stack. See NVIDIA’s GPUDirect Storage documentation and its GPU Operator RDMA guidance for configuration details.
Announced specifications—and what the numbers mean
| Specification | Supermicro’s announced detail |
|---|---|
| Chassis | 2U |
| DPUs | Up to four NVIDIA BlueField-3 DPUs |
| Network per DPU | 400-Gb Ethernet or InfiniBand |
| SSD bays | 24 or 36 PCIe Gen5 SSDs |
| Drive formats | E3.S or U.2 |
| Maximum stated capacity | 1.105 PB raw, using 30.71-TB drives |
| Bandwidth claim | More than 250 GB/s of Gen5 SSD bandwidth |
| Availability design | Dual-port, active-active clustering |
These are figures from Supermicro’s announcement, not independent test results. The 1.105-PB figure is raw capacity before formatting, spare drives, overprovisioning, metadata, and protection overhead such as replication or erasure coding. The more-than-250-GB/s figure describes the vendor’s stated SSD bandwidth; it should not be read as guaranteed throughput for an application, a storage software stack, or a GPU cluster. Supermicro also said the system could saturate a 400-Gb/s BlueField-3 link, but end-to-end performance depends on the full configuration.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →“Petascale” in this context describes the product’s scale and raw-capacity positioning; it does not mean every buyer will get a petabyte of usable storage from a deployed system.
Why storage can matter to AI—and when it won’t
AI clusters repeatedly read training data, write checkpoints, preprocess inputs, load embeddings, and serve inference requests. If storage or the path from storage to compute cannot supply data quickly enough, GPUs may wait rather than process. In distributed workloads, traffic also flows between storage and compute nodes and across the cluster’s network fabric.
A high-throughput storage system can help when data delivery is the bottleneck. It will not necessarily shorten training or improve GPU utilization if the limiting factor is data decoding, preprocessing, model code, insufficient parallelism, network congestion, or another part of the pipeline. The relevant question is not simply how fast the SSDs are, but whether the complete path—from media and PCIe through the DPU, network, storage software, and GPU-side software—matches the workload.
Software is part of the system
The JBOF is a building block, not a complete file or object storage service. Supermicro named Hammerspace for data-platform functionality and Cloudian for object storage, as well as Micron and Kioxia as qualified SSD vendors. Buyers still need to choose and validate the storage software, data-protection design, orchestration, and network architecture that make the hardware useful. A parallel file system, an object store, and a data-orchestration platform serve different access patterns and are not interchangeable.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Before committing, confirm that the chosen software supports the exact DPU, Arm configuration, SSD layout, GPU and driver stack, and cluster environment. For GPUDirect Storage in particular, a DPU alone does not activate a compatible end-to-end path.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.A related Supermicro product announced later is different
Supermicro announced a separate Grace-based storage system on March 19, 2025: the 1U ARS-121L-NE316R. It uses an NVIDIA Grace CPU Superchip with 144 Arm Neoverse V2 cores, up to 960 GB of onboard LPDDR5 memory, and 16 hot-swappable E3.S PCIe Gen5 NVMe bays. Supermicro stated a maximum of 983 TB raw capacity with 61.44-TB SSDs and 39.3 PB raw capacity for a 40-system rack; WEKA was identified as a supported storage software partner. This is a CPU-based storage server, not the October 2024 BlueField-3 JBOF. See the 2025 announcement and product page.
What buyers should check
- Workload shape: Identify sequential streaming, random reads, small-file metadata activity, checkpoint writes, retrieval, and mixed training or inference. Peak sequential bandwidth may not predict performance for small files or metadata-heavy jobs.
- Usable capacity: Size for the capacity left after protection, spares, snapshots, metadata, formatting, and operational reserves—not the raw drive total.
- GPU-to-storage ratio: Measure GPU idle time and data-loader waits, then establish how many GPUs each storage node must serve. Avoid buying for a storage bottleneck that has not been demonstrated.
- Network fabric: Match the storage and compute sides. RoCE can require careful lossless-network configuration and operational expertise; InfiniBand may fit an existing HPC environment but entails a specialized ecosystem.
- Protection and failures: Ask how replication or erasure coding, drive replacement, rebuilds, and node failures affect throughput. The launch bandwidth figure does not establish degraded-mode performance.
- Compatibility and operations: Verify software support, drivers, firmware, kernel paths, monitoring, update procedures, and service response. A DPU adds another layer to maintain.
- Power, cooling, and full cost: Include switches, optics, software licenses, support, installation, rack power, and cooling in comparisons. Dense flash and high-speed networking can require substantial infrastructure.
Potential failure points include an oversubscribed network, RoCE congestion or packet loss, a parallel file system poorly tuned for file sizes, metadata overload, slower rebuilds after a failure, or incompatibility among DPU firmware, GPU drivers, GPUDirect Storage components, and storage software. These are reasons to request an application-level proof of concept, not to assume the hardware figures translate directly to workload results.
Where it fits—and where it may not
The BlueField-3 JBOF is aimed at large, software-defined AI and HPC storage deployments that can use dense NVMe flash, fast networking, and DPU offload. It may be excessive for a small AI pilot, routine file serving, or a team seeking a turnkey array with simple per-terabyte pricing. Conventional CPU-based NVMe servers may be easier to integrate and offer broader software flexibility. Enterprise all-flash arrays can prioritize integrated management and support; object storage often suits durable data lakes, while parallel file systems target concurrent high-throughput access. Cloud storage avoids an upfront hardware deployment but brings recurring and data-transfer costs.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsThe October 2024 announcement does not publish a system price or a total-cost comparison. A serious evaluation should request a configuration quote and usable-capacity figure, then compare an application-level benchmark using the intended GPU count, dataset, file-size distribution, storage software, and network. Include the complete bill of materials and support costs rather than comparing SSD bandwidth alone.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

