Advertising disclosure: we earn commissions when you shop through the links below.
HPE announced the ProLiant DL585a Gen13 on October 7, 2026 as the most AI-focused system in its AMD-powered ProLiant Gen13 lineup. It is a 10U, 19-inch rack server aimed at AI inference, RAG model fine-tuning, virtualization, analytics and agentic AI. It holds two 6th Gen AMD EPYC 9006 SP7 processors with up to 256 cores each and supports up to eight double-wide PCIe GPUs from Nvidia or AMD over PCIe Gen6. HPE expects it to become available in early 2027 and has not published pricing.
The HPE ProLiant DL585a Gen13 is a high-density, 10U dual-socket enterprise server engineered specifically for sustained AI workloads, retrieval-augmented generation (RAG), and agentic workflows. Built to house two 6th Gen AMD EPYC 9006 "Venice" processors alongside up to eight double-wide PCIe accelerator cards, this system bridges the gap between traditional enterprise virtualization and dedicated scale-up AI supercomputers. It brings data-center compute directly into standard 19-inch rack infrastructure without requiring proprietary OCP rack modifications or specialized liquid-to-air facility retrofits.
Where modern multi-node clusters often force teams into rigid, proprietary interconnect topologies, the DL585a Gen13 prioritizes modular PCIe Gen6 I/O and balanced CPU-GPU resource allocation. Enterprise deployments frequently falter when high-throughput inference engines outrun their host CPUs during pre-processing, context retrieval, and vector database queries. By pairing up to 512 physical CPU cores with massive system memory bandwidth and eight discrete accelerator slots, this chassis targets teams looking for the best hardware for running AI models locally within enterprise-controlled boundaries.
For infrastructure architects sourcing Unknown hardware for AI development across multi-vendor fleets, the platform provides a flexible hardware footprint. Rather than locking administrators into a single accelerator ecosystem, the DL585a Gen13 accepts double-wide PCIe accelerators from both Nvidia and AMD. This makes the HPE ProLiant DL585a Gen13 for AI a versatile host platform for enterprise clusters dedicated to high-volume token generation, parameter-efficient fine-tuning (PEFT), and multi-agent coordination.
The ProLiant DL585a Gen13 is built from the ground up to eliminate host-to-accelerator I/O bottlenecks. Deploying modern local models requires moving massive KV caches and model weights across PCIe lanes with minimal latency.
At the core of the system are two 6th Gen AMD EPYC 9006 SP7 processors delivering up to 256 physical cores each (512 cores and 1024 threads total). This host capacity is critical for running dense vector search indices, pre-tokenization pipelines, and agent runtime orchestration on the CPU side while GPUs remain focused on tensor operations.
Memory throughput is a primary limiter during autoregressive token generation. The DL585a Gen13 features 32 DIMM slots operating across 16 channels, supporting:
Using MRDIMM configurations provides system memory bandwidth that rivals low-tier accelerators. This capability allows high-speed CPU offloading for large parameter weights that exceed discrete accelerator VRAM.
Inter-node scaling and remote direct memory access (RDMA) over converged Ethernet (RoCE) are integrated directly into the chassis:
Delivering power to eight high-TDP accelerators demands extensive electrical headroom:
System administration is handled via HPE iLO 8, which includes automated self-encrypting drive (auto-SED) management and multi-party authorization to secure model weights and proprietary data stored at rest.
When outfitting the system with eight modern GPUs, the HPE ProLiant DL585a Gen13 VRAM for large language models depends on the installed cards. Populating the chassis with eight 80GB or 96GB PCIe accelerators yields 640GB to 768GB of unified VRAM. Populating it with enterprise-tier accelerators sporting larger memory footprints can push total GPU memory past 1TB.
The chassis easily acts as hardware for running 70B parameter models at massive concurrency. Models such as Llama 3.1 70B and Qwen 2.5 72B require roughly 140GB of VRAM in unquantized 16-bit precision:
Hosting giant open-weight models like Llama 3.1 405B requires distributed tensor and pipeline parallelism:
Models such as Mixtral 8x22B and DeepSeek-R1 (671B MoE) benefit significantly from the system's vast memory capacity:
Processing vision-language models like Qwen 2.5-VL or processing 128k-to-1M token contexts requires vast memory allocations for KV cache retention. The high system memory bandwidth supported by MRDIMM (12800 MT/s) makes CPU-side KV cache swapping viable when requests exceed physical GPU VRAM.
The HPE ProLiant DL585a Gen13 is strictly an enterprise-grade and research-tier system. It is engineered for:
This machine is not intended for single-seat developers or hobbyist desk setups. The 10U physical form factor, massive acoustic footprint, and multi-kilowatt power supply demands require standard data-center rack installations, high-voltage PDUs, and dedicated cold-aisle cooling.
When evaluating the HPE ProLiant DL585a Gen13 AI inference performance, architects typically compare it against dedicated SXM-based servers and liquid-cooled OCP hardware.
The Dell PowerEdge XE9680 is a popular 8-GPU server based on dedicated HGX/SXM modules or OAM form factors.
Within HPE's own Gen13 lineup, the XD245 targets ultra-dense hyperscale environments.
For enterprise infrastructure teams prioritizing open accelerator selection, massive CPU compute, and standard rack deployment over closed, proprietary baseboards, the DL585a Gen13 stands as a top-tier platform for production AI workloads.
The top models this device can run at 4-bit, ranked by fit and speed.
Specs not available for scoring. This product is missing VRAM or memory bandwidth data.

Also in Our Store
Docks, fast storage, cables, and a UPS sized for a local AI machine.
