AI & HPC Storage
Storage for AI Training Clusters
A GPU cluster is only as fast as the data path feeding it. Past a handful of nodes the constraint usually stops being the accelerators and becomes the storage tier — specifically whether data reaches GPU memory without a detour through a CPU bounce buffer. That is what NVIDIA's GPUDirect Storage path exists to remove, and it is why these systems are built differently from a general-purpose array.
The catalogue covers the architectures teams actually shortlist. VAST disaggregates stateless compute nodes from NVMe enclosures. WEKA's WEKApod is a parallel filesystem sized around high-speed adapters per node. Dell PowerScale gives you OneFS scale-out on all-NVMe media, and Pure FlashBlade serves NFS, SMB and S3 concurrently. Published throughput figures are the manufacturers' own and depend heavily on configuration.
We quote the storage with the fabric attached, because that is where these projects go wrong. A parallel filesystem specified without its InfiniBand or high-speed Ethernet is not a working tier — it is a delivery waiting on a second purchase order. Tell us the model sizes, the checkpoint cadence and the GPU count, and the switches, adapters and optics go on the same quote.
11 configurations below · quotes returned within 48 business hours · purchase orders accepted
Storage for AI Training Clusters we configure and supply
Every system below is quoted to your workload — accelerator count, CPU, memory, storage and fabric are specified together rather than sold as a fixed SKU.

FlashBlade//S500 R2 (performance scale-out file & object for AI)
High-throughput unified file and object storage for AI and HPC pipelines.
- Keeps GPUs fed
- Scale without rip-and-replace
- One platform, many protocols

FlashBlade//E (capacity-optimized unstructured data lake)
All-flash capacity at multi-petabyte scale for unstructured data lakes.
- Flash at disk-tier cost
- Grows in steady increments
- Simple at scale

Dell PowerScale F710 1U All-Flash NVMe Node (10x 30.72TB QLC)
High-density 1U scale-out NAS feeding GPU clusters at full bandwidth
- Linear scale-out growth
- Maximum rack density
- GPU-ready throughput

Dell PowerScale F910 2U All-Flash NVMe Node (24x 15.36TB TLC)
Capacity-dense 2U NAS node for enterprise AI data lakes
- Higher per-node capacity
- Single global namespace
- Inline data efficiency

Dell PowerScale F710 Performance Node (10x 7.68TB TLC, 100GbE)
Low-latency 1U NVMe NAS tuned for mixed high-IOPS workloads
- Consistent low latency
- Performance density
- Non-disruptive scaling

NetApp AFF A90 All-Flash Unified Array (200GbE, ONTAP)
High-end unified flash that feeds AI pipelines without starving GPUs.
- GPUs stay fed
- One platform, every protocol
- Six-nines availability

NetApp AFF A800 All-Flash NVMe Array (100GbE, ONTAP AI)
Proven NVMe storage behind NVIDIA DGX deep-learning environments.
- Validated AI architecture
- Throughput at scale
- Data managed, not copied

WEKA WEKApod Nitro (8-Node, ConnectX-8 800Gb/s)
Keep thousands of GPUs saturated with sub-millisecond training data.
- GPUs stay busy
- Scales without rip-and-replace
- Validated for SuperPOD

WEKA WEKApod Prime (2U, 40-Drive Capacity Tier)
Validated AI storage economics for BasePOD-scale GPU deployments.
- Better price-per-petabyte
- Right-sized for BasePOD
- One platform, file and object

VAST Ceres DASE Node (1U BlueField DPU, E1.L NVMe)
Disaggregated all-flash density that scales from terabytes to petabytes.
- Maximum flash per rack U
- Compute and capacity scale apart
- No controller bottleneck

VAST Universal Storage GPUDirect Cluster (Petascale, NVMe-oF)
One petascale namespace feeding large GPU estates over GPUDirect.
- Linear throughput scaling
- One namespace for everything
- Independent compute and capacity
What we need to quote accurately
A configuration quote takes minutes when these are known and days of back-and-forth when they are not. You do not need all of them to start — send what you have.
- The workload: model sizes, training or inference, and expected concurrency
- Node count, and whether this is a single system or a scaling cluster
- Rack power and cooling available per rack, and inlet temperature
- Existing estate — the vendor and management tooling you already run
- Network fabric: InfiniBand, Ethernet, and the speed you are standardised on
- Timeline, and whether the budget is approved or being built
How we quote
- 1
Send the requirement
An email, a bill of materials, or a rough description of the workload. All three work.
- 2
We validate the configuration
Accelerator, chassis, fabric, power and cooling checked against each other before anything is priced.
- 3
Itemised quote within 48 hours
Line-by-line pricing, current availability and lead time, with warranty terms stated.
AI training storage — buyer questions
Do we actually need GPUDirect Storage, or will NFS over 100GbE do?
Often NFS will do. GPUDirect matters when the job is genuinely I/O-bound — large image or video corpora streamed continuously, or heavy checkpoint traffic on a large cluster. If your dataset fits in local NVMe on each node and is read once per epoch, a well-provisioned NFS tier is frequently enough and considerably simpler to run. We will tell you which side of that line your workload sits on rather than defaulting to the more expensive answer.
How much throughput should we buy per GPU?
There is no honest single figure, and any supplier quoting one is guessing. It depends on sample size, whether you stream from object storage or a cached working set, and above all on checkpointing — for most large-model runs the write peak during a checkpoint, not the steady-state read, is what sizes the tier. Send us the model size, node count, checkpoint interval and whether checkpoints are full or sharded, and we will size against the manufacturers' published per-node figures rather than a rule of thumb.
Parallel file system, or an all-flash NAS?
WEKA and VAST are parallel architectures with a single global namespace, built for many clients reading the same dataset at once — which is what a multi-node GPU cluster looks like. PowerScale and FlashBlade are scale-out NAS: simpler to operate, multiprotocol, and a better fit where the same storage also serves engineering or media work. If AI is one workload among several, the NAS platforms usually win on total cost of ownership.
Can you supply the network fabric with it?
Yes, and it belongs on the same quote. These systems are specified around high-speed adapters per node over InfiniBand or 400/800GbE. We specify the switches, adapters, transceivers and cable lengths from your rack elevation, so what arrives can be cabled and brought up rather than half-built.
Compare with other platforms
NVIDIA H100 Systems
NVIDIA H100 Servers
View systemsNVIDIA H200 Systems
NVIDIA H200 Servers
View systemsAI Cluster Fabric
AI Cluster Network Fabric
View systemsAll-Flash Arrays
All-Flash Storage Arrays
View systemsBrowse the full range of GPU servers and AI clusters, or the component catalogue for drives, memory, optics and spares.
