Section10

Hardware Deep-Dive: NVIDIA GPU Ecosystem

NVIDIA is the default enterprise AI hardware choice in 2026, and its lead is widest exactly where a first on-premises build needs it most: the production inference tooling and the OEM (original equipment manufacturer) availability through Dell and Supermicro that let a small team ship without an MLOps department. For Phase 1 across the range this paper addresses (small organizations below 100 employees and medium organizations of 100 to 499), a Dell PowerEdge or Supermicro GPU SuperServer populated with one or two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs covers inference at lower cost and lower facility risk than any eight-GPU SXM alternative, with a clean upgrade path to 8× H200 or B200 SXM once training and high-concurrency multi-tenant workloads arrive in Phase 2. The same platform holds up to 999 employees on inference-dominated workloads, because concurrency growth at that scale is absorbed by additional cards in the same chassis rather than by a different interconnect. That recommendation tracks where deployment risk is lowest for a first build, not a vendor allegiance.

Why CUDA Wins the Tooling Argument

The dominant production LLM inference stack in mid-2026 is CUDA-native. The three engines that cover the field are all developed and tuned first on NVIDIA hardware: vLLM for open-source flexibility, TensorRT-LLM for NVIDIA-specific peak throughput, and SGLang for prefix-cache-heavy serving. NVIDIA reports the B200 sustaining roughly 1,000 tokens per second per user and 60,000 tokens per second per GPU on gpt-oss-120b under its TensorRT-LLM stack, a vendor-reported figure [1]. NVIDIA’s own measurements put TensorRT-LLM draft-model speculative decoding at up to 3.6× faster generation on applicable workloads, measured on Llama 3.1 405B with a 3B draft model across four H200 GPUs [2].

The software gap, not the silicon, is the decisive variable. AMD’s MI300X carries 192 GB of HBM3 at 5.3 TB/s, against 80 GB at 3.35 TB/s for the H100 it competes with, a comparison AMD’s own specification page draws directly [3]. Independent serving benchmarks still rank the NVIDIA systems ahead: Artificial Analysis’s System Load Test runs an 8× MI300X configuration alongside 8× H100, H200, and B200 systems on gpt-oss-120b, and the top result on every headline metric (peak throughput, per-query output speed, and cost per token) belongs to the 8× B200 [4]. The mechanism sits in the serving stack rather than the chips. TensorRT-LLM, the engine behind NVIDIA’s peak numbers, is CUDA-only [5]. vLLM maintains an upstream ROCm backend with pre-built wheels, so AMD serving works without forks on memory-bound workloads, but that upstream path is young: official vLLM Docker images for ROCm only became available on January 20, 2026, and before that date optimized AMD serving ran through AMD’s own downstream prebuilt images [6]. The custom-kernel paths that set the performance frontier, fused attention variants, FP4 inference pipelines, and disaggregated serving, are developed on CUDA first and reach ROCm on a lag. The hardware is comparable; the software stack is not. For a first deployment with no ROCm expertise on staff, that integration risk decides the vendor question.

The NVIDIA developer toolchain spans CUDA, cuDNN, TensorRT, the open-source TensorRT-LLM [5], Triton Inference Server, NVIDIA NIM inference microservices, and the NVIDIA GPU Operator, a vertical stack no competing platform matches end to end. The MXFP4 weight format used by gpt-oss-120b (the Open Compute Project Microscaling FP4 format) fits the full 120B model on a single 80 GB GPU and accelerates natively on Blackwell’s FP4 tensor cores. AMD’s MI350-series datacenter GPUs also accelerate FP4, so the format is not CUDA-exclusive; the hardware tier is. No AMD part answers the RTX PRO 6000 Blackwell at the single-card, 96 GB, air-cooled PCIe class this paper targets for Phase 1, so at that tier the FP4 inference path runs only on NVIDIA silicon.

The Workstation Tier: DGX Spark and DGX Station

The DGX Spark is a developer workstation, not a production inference server. It is a desktop unit roughly the footprint of a Mac Mini, built around the NVIDIA GB10 Grace Blackwell Superchip with 128 GB of unified LPDDR5x memory shared between a 20-core Arm CPU and a Blackwell GPU, in a 240 W appliance delivering up to 1 PFLOP of FP4 compute with sparsity [7]. NVIDIA’s Founders Edition launched at $3,999; partner units from Dell and others run higher, around $4,700–$4,800, with most of the variance in storage [8], [9]. Two units paired over the 200 Gb/s ConnectX-7 NIC form a 256 GB pool that runs models up to 405B parameters at FP4; going beyond two requires a switch [7].

The Spark earns its place here as the cleanest match for a single ML practitioner doing local model development, parameter-efficient fine-tuning, and air-gapped prototyping, with single-unit inference on models up to roughly 200B parameters [7], [8]. It is not a Phase 1 production server for 5–15 concurrent users: the hardware-tier performance matrix in the open-weight model analysis puts a single RTX PRO 6000 at 8,425 tokens per second on a 30B AWQ (activation-aware weight quantization) workload, a figure the Spark cannot approach under concurrent load [10]. Use the Spark for developer workflows and put the production server elsewhere.

The DGX Station occupies the next tier up: a single GB300 Grace Blackwell Ultra Superchip, 748 GB of unified memory (252 GB of HBM3e GPU memory plus 496 GB of LPDDR5X system memory), and up to 20 PFLOP of AI compute in a deskside tower [11]. The shipping configuration uses a salvaged B300 die with seven of its eight HBM3e stacks enabled, which is why it carries 252 GB at 7.1 TB/s rather than the 288 GB at 8 TB/s NVIDIA originally specified in 2025 [12]. That 252 GB of fast HBM3e is the operationally relevant number: it holds a 235B-parameter model at FP8 (about 235 GB of weights) in a single deskside unit, the minimum single-system path to production-quality inference on that model class. NVIDIA markets trillion-parameter support using the full 748 GB pool, but that spills the model into the slower LPDDR5X, and the throughput cost of that path is not yet well characterized [12].

NVIDIA sets no list price and sells no Founders Edition this generation; partners quote pricing individually. MSI’s XpertStation WS300 launched at $85,000 [13], and ServeTheHome reports partner quotes running from roughly $85,000 at the floor, through the $100,000 range for typical configurations, to $120,000–$125,000 at the high end [12]. Orders opened at GTC in March 2026 with partner shipments spread across Q2–Q3 2026 [12]. The system’s power ceiling is 1.6 kW, the practical limit of a North American 120 V outlet, and workstation air cooling handles the single-superchip thermal envelope; that is the architectural break from the rack systems below: no liquid-cooling loop, no facility retrofit, no rack-scale power-delivery problem [12]. For an organization that needs 235B-class quality at one workstation, the DGX Station is the cleanest path; for one that needs higher concurrency or several smaller models in parallel, the rack-mounted options below are the better fit.

The 8-GPU Rack Tier: DGX B200, B300, and OEM Equivalents

The DGX B200 and B300 are full-rack systems for large-model training and high-throughput multi-user inference. The DGX B200 carries eight Blackwell GPUs at 180 GB HBM3e each for 1,440 GB aggregate, NVSwitch fabric delivering 1,800 GB/s of bidirectional NVLink bandwidth per GPU (14.4 TB/s aggregate), a 1,000 W TDP (thermal design power) per GPU, and roughly 14.3 kW of maximum system draw in a 10U air-cooled chassis [14], [15]. Pricing is channel-dependent and not published by NVIDIA; reported figures for eight-GPU Blackwell systems sit in the low-to-mid six figures, so any number is indicative pending a vendor quotation.

The B300 (Blackwell Ultra) raises both capacity and power: eight GPUs at 288 GB HBM3e each for 2,304 GB aggregate, 1,400 W per GPU, and roughly 14 kW of system draw [16]. No OEM ships that full 1,400 W configuration on air. Supermicro’s full-TDP B300 systems are liquid-cooled, with its DLC-2 direct-liquid-cooling line capturing up to 98% of system heat [17], while Dell’s air-cooled PowerEdge XE9780 carries a derated HGX B300 NVL8 configuration at 270 GB and 1,100 W per GPU in a roughly 10U chassis [18], [19].

The production throughput figures matter most for Phase 2 sizing. Eight B200 GPUs running gpt-oss-120b reach about 92,909 tokens per second of peak aggregate throughput in the throughput-optimized configuration, 403 tokens per second per query in the latency-optimized configuration, and the lowest cost in the test at $0.19 per million input plus million output tokens at the 100 tokens-per-second reference speed; DeepSeek-R1 0528 reaches roughly 45,677 tokens per second on the same system. All are point-in-time figures from Artificial Analysis’s continuously updated System Load Test [4]. NVIDIA’s MLPerf Inference v6.0 submission of 288 Blackwell Ultra GPUs across four interconnected GB300 NVL72 racks reached 2.49 million tokens per second on DeepSeek-R1 in the offline scenario (closed division, results published April 2026), far beyond any practical organizational requirement but useful as a ceiling reference [20].

The Dell PowerEdge XE9780 is the operational equivalent for an organization with an existing Dell relationship: eight HGX B300 NVL8 GPUs on an NVSwitch fabric, 2,160 GB aggregate HBM3e, air-cooled, with the XE9780L variant adding direct liquid cooling [18], [19]. Dell states the platform trains LLMs up to 4× faster than the prior-generation 8× HGX H100 system, a vendor-reported generational comparison [21], [22]. For H100 or H200 configurations at lower price points, the predecessor XE9680 remains available; confirm current Dell availability of H100/H200 SKUs against B300 at procurement time, since OEM generation transitions compress the window on older configurations.

Supermicro’s B300 line ships in liquid-cooled and air-cooled configurations [17]. In December 2025 it began volume shipment of 4U and 2-OU (Open Compute Project form factor) DLC-2 liquid-cooled HGX B300 systems, each with eight B300 GPUs, ConnectX-8 SuperNICs at 800 Gb/s, and support for Quantum-X800 InfiniBand or Spectrum-X Ethernet clustering, scaling to 64 GPUs per rack in the 4U design and 144 in the 2-OU design [23]. The 4U DLC-2 HGX B200 system, shipping since August 2025, is the lower-cost entry point [24]. For Phase 1 procurement seeking H200 SXM5 at established pricing, the prior-generation SYS-821GE-TNHR (8× H100/H200 SXM5, dual Intel Xeon, up to 8 TB DDR5, 16 NVMe bays, six 3,000 W redundant power supplies) is the most documented and available H100/H200 system.

The Mid-Range PCIe Tier and the Phase 1 Inflection Point

The Dell PowerEdge XE7745 or a Supermicro 4U GPU SuperServer with one or two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs is the Phase 1 recommendation, and the evidence is counterintuitive enough to need an explicit defense. Akamai’s controlled test, run with NVIDIA’s own NIM profiles on the Llama-3.3-Nemotron-Super-49B reasoning model, measured the RTX PRO 6000 at FP4 delivering 1.63× the throughput of the H100 NVL 96GB at FP8 at 100 concurrent requests, with the advantage narrowing toward parity by 200 [25]. CloudRift’s testing has the RTX PRO 6000 beating the H100 at 28% lower cost per million tokens at single-GPU scale [26]. DatabaseMart, a GPU hosting provider, stress-tested vLLM at 300 concurrent requests and put the card ahead of both the H100 and the A100 80GB on 8B–32B models at production load; the result is vendor-run, but the method is disclosed and it corroborates the two independent tests above [27]. On a 30B AWQ workload, a single card sustains 8,425 tokens per second aggregate [10], roughly 18× the 450-tokens-per-second floor for 15 concurrent users reading at 30 tokens per second each.

Fifteen concurrent requests is a reasonable Phase 1 sizing target, and it reaches headcount only through an assumed simultaneous-use ratio. At 5 percent simultaneous use, fifteen concurrent requests is a 300-seat deployment; at 3 percent it is the 499-employee ceiling of the medium organization size band. Both ratios are this paper’s planning assumptions rather than measured values. Internal AI assistant traffic is bursty enough that workday peaks run above the daily mean. Size the pilot against instrumented peak concurrency and treat headcount as an input to that estimate, not as the estimate.

The mechanism is straightforward. The RTX PRO 6000 Blackwell Server Edition ships with 96 GB of GDDR7 ECC (error-correcting code) memory at 1,792 GB/s and 24,064 CUDA cores on a near-full GB202 die [28], [29]. A 70B model at FP8 needs roughly 70 GB of weights plus about 26 GB of KV-cache headroom at 4K context. That fits inside 96 GB with room for an estimated 13–25 concurrent users, depending on context length and cache precision. At the simultaneous-use ratios assumed here, one card therefore can serve a small organization and potentially a medium one to its 499-employee ceiling; the 70B FP8 case, not the 30B case, is the binding constraint. The same workload on an H100 NVL 96GB runs slower and costs more at the system level. For Phase 1 use cases that top out at 70B-class FP8 or 120B at MXFP4 (confirmed deployable for gpt-oss-120b on the 96 GB card), the RTX PRO 6000 is the better silicon.

The interconnect limit is the honest caveat. SXM GPUs (NVIDIA’s socketed server module form factor) on an HGX board connect over NVLink/NVSwitch at 1,800 GB/s bidirectional per GPU; a PCIe 5.0 x16 slot delivers roughly 64 GB/s in each direction, about 128 GB/s bidirectional, a fourteen-fold gap. When a model is sharded across GPUs (necessary once it exceeds one card’s VRAM) that NVLink advantage is material; when the full model fits on one card, PCIe is equivalent. Multi-GPU PCIe configurations therefore support replica parallelism (independent model copies serving more users) but not efficient tensor parallelism (one larger model split across cards). That distinction is what lets the platform absorb growth past 499 employees and up to 999: added seats need more replicas, not a wider fabric. The XE7745 scales to eight RTX PRO 6000 cards in a 4U air-cooled chassis drawing about 4.8 kW GPU-only at 600 W per card, against 8 kW for an 8× B200 and 11.2 kW for a full-power 8× B300, and well within standard CRAC/CRAH (computer-room air conditioner and air handler) air-cooling envelopes [29].

The Supermicro path for an organization with an existing Supermicro procurement relationship is the SYS-422GL-NR, a 4U NVIDIA-Certified system built on the NVIDIA MGX modular reference design. It accepts up to eight RTX PRO 6000 Blackwell Server Edition cards at 600 W each over PCIe 5.0 x16 in a dual-root topology, the same interconnect class as the XE7745, so its performance profile matches within margin for both replica-parallel inference and single-card workloads that fit inside 96 GB [30]. Dual Intel Xeon 6900-series P-core processors feed the GPUs across 24 DDR5 slots scaling to 6 TB, and four 3,200 W Titanium-level supplies in a 3+1 redundant arrangement carry the roughly 4.8 kW GPU-only draw within a 10°C-to-35°C air-cooled envelope, so Phase 1 needs no facility retrofit [30]. For regulated environments it ships a NIST 800-193-compliant Silicon Root of Trust with cryptographically signed firmware and system lockdown, matching the firmware-integrity posture of the HPE ProLiant option. An organization that wants more local capacity for model weights and datasets can step up to the 5U SYS-522GA-NRT, which trades one rack unit for 24 front hot-swap NVMe bays while supporting the same eight-card configuration [31].

The HPE ProLiant Compute DL380a Gen12 (up to 8× RTX PRO 6000, 4U air-cooled) is the technical equivalent for an organization without a Dell relationship; its performance profile matches the XE7745 within margin because both use the same GPUs over PCIe [32]. HPE’s Gen12 platform adds iLO 7 (integrated Lights-Out management) with a Silicon Root of Trust and quantum-resistant firmware signing, which differentiates it for regulated environments [32]. Absent an existing HPE procurement relationship, it is covered here for completeness rather than as a primary path.

The Phase 1 recommendation rests on three findings working together. The RTX PRO 6000 outperforms the H100 at production inference load on models up to 70B. Its 96 GB ceiling matches the model classes that deliver value at Phase 1 scale: Gemma 4 31B, a 70B-class model at FP8, and gpt-oss-120b at MXFP4. And the XE7745 chassis takes additional RTX PRO 6000 cards as concurrency grows, which carries the recommendation from a fifty-person deployment through the 499-employee ceiling and on to 999 without a platform change, then hands off to a separate 8× H200 or B200 SXM server for Phase 2 training and high-concurrency multi-tenant work. Phase 1 requires no facility retrofit, and that constraint outranks any benchmark margin: rack power and cooling headroom, not silicon, separate a build that ships in Phase 1 from one that waits on a facility assessment.

The Multi-Node Tier: DGX SuperPOD and Vera Rubin

The DGX SuperPOD is hyperscale data center territory only. Current-generation SuperPODs link multiple DGX B200 systems over NVIDIA Quantum InfiniBand for multi-node training; the GB200 NVL72 variant packs 72 GPUs per rack at 132 kW of rack power, fully liquid-cooled under the reference architecture NVIDIA and Vertiv co-developed [33], [34]. The next-generation Vera Rubin DGX SuperPOD, announced at CES 2026, comes in two configurations, eight racks of 72 Rubin GPUs (576 total) or 64 NVL8 nodes (512 total) [35], [36]. NVIDIA announced the Vera Rubin platform’s ramp into full production in June 2026, with partner systems shipping through the second half of the year [37].

The Vera Rubin platform raises per-GPU specifications sharply: 288 GB of HBM4 per GPU at 22 TB/s, an NVIDIA-rated 50 PFLOP of NVFP4 inference (roughly 5× the B200), sixth-generation NVLink at 3.6 TB/s per GPU, and 336 billion transistors on TSMC 3nm [37], [38]. NVIDIA claims a 10× reduction in inference token cost and a 4× reduction in the GPUs needed to train mixture-of-experts (MoE) models, both measured against the GB200 NVL72 [37]. NVIDIA has not published a rack power figure for the Rubin NVL72 at the time of this writing; third-party estimates span roughly 120 kW to more than 250 kW, and the rack is a fully liquid-cooled design regardless of where production power lands, so upgraded facility power and a chilled-water path are prerequisites either way [36], [38]. NVIDIA likewise publishes no NVL72 pricing; on the Blackwell precedent, NVL72-class racks land in the seven figures, and any narrower number is indicative pending a vendor quotation.

These systems mark the upper boundary of the roadmap, kept in view so the Phase 1 architecture never forecloses growth. Phase 1 hardware exceeds Phase 1 throughput requirements by more than an order of magnitude, and an 8× B200 SXM server in Phase 2 covers production agentic workloads across the medium band to 499 employees, and up to 999 where serving rather than pretraining dominates. Serving demand alone does not reach the SuperPOD tier anywhere in that range: NVIDIA’s 288-GPU MLPerf submission produced 2.49 million tokens per second [20], more than three orders of magnitude above the 450-tokens-per-second interactive floor the estimated fifteen concurrent request Phase 1 deployment needs. The tier becomes relevant only if the organization takes on continuous large-scale training of foundation models, which is a change of mission and not a consequence of headcount growth.

System Comparison Across Tiers

SystemGPU ConfigVRAM (Aggregate)InterconnectTDP / System PowerForm FactorPrice TierWorkload Fit
NVIDIA DGX Spark1× GB10 (unified)128 GB unified (CPU+GPU)NVLink-C2C internal240 WDesktop (1.2 kg)$3,999+ (partner ~$4,700–$4,800)Developer workstation; ~1 PFLOP FP4; single-user inference to ~200B params at FP4
NVIDIA DGX Station1× GB300 (unified)748 GB (252 GB HBM3e + 496 GB LPDDR5X)NVLink-C2C internalUp to 1.6 kW systemDeskside tower$85K–$125K (partner-set; verify)Single-superchip training + inference; holds a 235B-class model at FP8
Dell PowerEdge XE77451–8× RTX PRO 6000Up to 768 GB GDDR7 (8×96 GB)PCIe 5.0 x16600 W/GPU; ~4.8 kW GPU-only at 8×4U rack, air-cooledVerify with DellPhase 1 recommendation. 70B-class FP8 on one card; PRO 6000 FP4 ≈ 1.63× H100 NVL FP8; gpt-oss-120b at MXFP4
Supermicro SYS-422GL-NR 4U GPU Server1–8× RTX PRO 6000Up to 768 GB GDDR7 (8×96 GB)PCIe 5.0 x16600 W/GPU; ~4.8 kW GPU-only at 8×4U rack, air-cooled~$75K to ~$167K+ (verify with Supermicro)Phase 1 recommendation. 70B-class FP8 on one card; PRO 6000 FP4 ≈ 1.63× H100 NVL FP8; gpt-oss-120b at MXFP4
Dell PowerEdge XE97808× HGX B300 NVL82,160 GB (8×270 GB HBM3e)NVSwitch1,100 W/GPU (derated)~10U rack, air-cooledVerify with DellPhase 2 training + inference; XE9780L adds liquid cooling
Supermicro 4U DLC-2 B2008× HGX B2001,440 GB (8×180 GB HBM3e)NVSwitch1,000 W/GPU; ~14 kW system4U rack, liquid-cooledVerify with SupermicroPhase 2 training + inference; lower entry vs B300
Supermicro DLC-2 B300 (4U / 2-OU)8× HGX B3002,304 GB (8×288 GB HBM3e)NVSwitch1,400 W/GPU; ~14 kW system4U or 2-OU rack, liquid-cooledVerify with SupermicroLarge-scale training + inference; up to 64 (4U) or 144 (2-OU) GPUs/rack
Supermicro SYS-821GE-TNHR8× H100/H200 SXM5640 GB (8×80 GB H100) or 1,128 GB (8×141 GB H200)NVSwitch700 W/GPU; ~10 kW system8U rack, air-cooled~$33K chassis + GPUs (verify); ~$320K for HGX baseboard with 8× H200Phase 1/2; proven availability
HPE ProLiant DL380a Gen12Up to 8× RTX PRO 6000Up to 768 GB GDDR7PCIe 5.0600 W/GPU4U rack, air-cooledVerify with HPEPhase 1 alternative to XE7745; iLO 7 for regulated environments
NVIDIA DGX B2008× B2001,440 GB (8×180 GB HBM3e)NVSwitch (full mesh)1,000 W/GPU; ~14.3 kW max10U rack, air-cooledIndicative six-figure; verifyProduction training + inference; ~92,909 tok/s on gpt-oss-120b
NVIDIA DGX B3008× B3002,304 GB (8×288 GB HBM3e)NVSwitch (full mesh)1,400 W/GPU; ~14 kW systemRack, liquid-cooledIndicative six-figure; verifyHighest-density training; liquid cooling mandatory
NVIDIA DGX SuperPOD (Rubin NVL72)8 racks × 72 Rubin GPUs (576 total)576 × 288 GB HBM4NVLink 6Reported ~120 kW to 250+ kW per rackRack cluster, liquid-cooledTBD (in production; partner systems H2 2026)Phase 3+ only; hyperscale training + trillion-parameter inference

Prices marked “verify” need a direct quotation from the vendor; NVIDIA does not publish DGX list pricing, and GPU server pricing stayed volatile through 2025–2026 with supply conditions and generation transitions. An H200 NVL 141GB PCIe card (141 GB HBM3e, 600 W) gives organizations H200 capacity in a PCIe form factor without the HGX platform cost, a useful intermediate to request in Dell or Supermicro quotations; the SYS-422GL-NR supports it in the same chassis as the RTX PRO 6000 [30]. Every system here clears the 300–450 tokens-per-second interactive floor for Phase 1 by a wide margin, so the decision among tiers turns on VRAM capacity, capital cost, facility power and cooling headroom, and upgrade-path compatibility, not on baseline concurrency.

Power and Cooling as Hardware Selection Gates

Phase 1 hardware on the RTX PRO 6000 platform draws about 4.8 kW GPU-only at the eight-card maximum, with total system power inside the 15–20 kW per-rack envelope that standard CRAC/CRAH air cooling supports reliably. Phase 2 hardware sits higher but stays air-coolable in most facilities: an 8× H200 SXM server draws roughly 10 kW at the system level, an 8× B200 about 14.3 kW at maximum, near the upper edge of comfortable air-cooling margin. Phase 3+ GB200 NVL72 racks at 132 kW, and Rubin NVL72 racks with reported estimates spanning roughly 120 kW to more than 250 kW, exceed any practical air-cooling ceiling and require direct-to-chip liquid cooling as a deployment prerequisite [33], [34], [36].

The governing standard is ASHRAE (American Society of Heating, Refrigerating and Air-Conditioning Engineers) Technical Committee 9.9’s Thermal Guidelines for Data Processing Environments, 5th Edition, published in March 2021 [39]. That edition defines the air-cooled equipment classes A1–A4 and adds the H1 class for high-density gear that cannot stay within the A1–A4 envelopes. H1 is itself an air-cooled class, carrying a tighter band (recommended 18–22 °C, allowable 15–25 °C) rather than a power-density rating, and ASHRAE leaves it to the equipment maker to certify a product as H1 [40], [41]. Liquid cooling is covered separately, by the W-class envelopes (W17 through W+). Uptime Institute analysis puts perimeter air cooling’s practical ceiling at 20–25 kW per rack, with close-coupled in-row systems extending air’s reach to roughly 50 kW at added capital cost; above 50 kW, liquid cooling becomes the norm [42]. The 5th Edition predates B300- and Rubin-generation hardware and does not address those power densities; ASHRAE has issued a revised and expanded printing of the 5th Edition, but no later edition had superseded it as of mid-2026 [41], [43].

The deployment implication is concrete. Phase 1 on the Dell XE7745 or Supermicro SYS-422GL-NR with RTX PRO 6000 GPUs fits inside any data center that supports a standard 15–20 kW rack, with no chilled-water loop and no CDU (coolant distribution unit) placement to assess. Phase 2 with an 8× H200 or B200 SXM server fits most facilities after a headroom check. Phase 3+ with rack-scale NVL72 hardware requires direct-to-chip liquid cooling, mandatory under NVIDIA’s reference architecture, with a CDU moving heat to the facility chilled-water loop or a dry cooler. Two liquid-cooling options are relevant at this scale: direct-to-chip cold plates with a CDU (used in the DGX SuperPOD, Supermicro DLC-2, and Dell XE9780L), which in Supermicro’s DLC-2 implementation capture up to 98% of system heat [17], and rear-door heat exchangers, which retrofit existing racks but do not reach NVL72-class densities.

The Phase 1 facility-readiness assessment is the gating deliverable. Before locking Phase 1 hardware, document the target rack’s current power capacity in kW, the current cooling method and its maximum rack-kW rating, chilled-water availability and loop temperature, and the physical space for a CDU if liquid cooling is anticipated in Phase 2 or 3+. The RTX PRO 6000 recommendation is in part a recommendation against premature facility commitment: it buys Phase 1 capability while preserving the time and capital to size cooling correctly for the step up.

The Phase 1 Procurement Position

The Phase 1 procurement decision is the Dell PowerEdge XE7745 or the Supermicro SYS-422GL-NR 4U GPU Server with one or two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs, validated against current Dell/Supermicro availability and pricing at order time. One card covers the estimated fifteen concurrent requests with roughly 18× throughput headroom, runs Gemma 4 31B at FP8 and gpt-oss-120b at MXFP4 at production quality plus 70B-class models at FP8, and draws little enough that the eight-card maximum still sits inside any standard air-cooled facility. That single card is sufficient across the range this paper defines (small organizations below 100 employees and medium organizations of 100 to 499) under the simultaneous-use ratios assumed for Phase 1 sizing, with the 13–25 concurrent users of the 70B FP8 case setting the ceiling. The recommendation holds up to 999 employees as well: a second or third card in the same chassis serves the added load as independent replicas, which is the one growth path PCIe handles as well as NVLink. The chassis takes those incremental GPU additions as concurrency grows, which defers the Phase 2 8× H200 or B200 SXM decision until training workloads and multi-tenant concurrency justify the capital and the cooling investment.

The Phase 2 decision is the Dell PowerEdge XE9780 (or the liquid-cooled XE9780L) or the Supermicro DLC-2 HGX B200 system, chosen on current vendor availability, lead time, and the cooling assessment completed during Phase 1. For organizations that prefer proven hardware at established pricing over B300-generation systems on longer lead times, the Supermicro SYS-821GE-TNHR with 8× H200 SXM5 remains the most documented and available option. The Phase 3+ decision waits on Phase 2 operational learning and facility-cooling capacity planning, with an 8x HGX B300 system target or a Vera Rubin DGX SuperPOD configuration on the very high end (in production as of mid-2026, with partner systems shipping through the second half of the year), contingent on a shift into continuous training rather than on headcount growth inside the range this paper covers.

The case against AMD for this first build is not that the silicon falls short on raw specifications: the MI300X carries more memory and bandwidth than the H100 it competes with [3], and Dell lists the 288 GB, 1,400 W MI355X in the same XE9780 chassis family as the B300 [18]. It is that ROCm’s serving stack still imposes integration risk an organization with no ROCm expertise should not absorb on a first deployment: the engine behind the fastest published NVIDIA results has no ROCm build at all [5], upstream vLLM images for ROCm only arrived in January 2026 [6], and the optimizations that define production serving reach CUDA first. For Phase 1, that software risk, not the silicon, settles the vendor question.

References

  1. NVIDIA Corporation, “NVIDIA Blackwell Raises Bar in New InferenceMAX Benchmarks,” NVIDIA Blog, Oct. 9, 2025. [Online]. Available: https://blogs.nvidia.com/blog/blackwell-inferencemax-benchmark-results/. [Accessed: 04-Jun-2026]

    NVGPU-1 Secondary source Back to text

  2. NVIDIA Corporation, “TensorRT-LLM Speculative Decoding Boosts Inference Throughput by up to 3.6x,” NVIDIA Technical Blog, Dec. 2, 2024. [Online]. Available: https://developer.nvidia.com/blog/tensorrt-llm-speculative-decoding-boosts-inference-throughput-by-up-to-3-6x/. [Accessed: 19-Jul-2026]

    NVGPU-2 Secondary source Back to text

  3. Advanced Micro Devices, Inc., “AMD Instinct MI300X Accelerators,” 2026. [Online]. Available: https://www.amd.com/en/products/accelerators/instinct/mi300/mi300x.html. Product page; 192 GB HBM3, 5.3 TB/s peak memory bandwidth, with AMD-published comparison to H100 SXM5 80 GB / 3.35 TB/s. [Accessed: 19-Jul-2026]

    NVGPU-3 Primary source Back to text

  4. Artificial Analysis, “AI Hardware Benchmarking and Performance Analysis,” Artificial Analysis, 2026. [Online]. Available: https://artificialanalysis.ai/benchmarks/hardware. Continuously updated. [Accessed: 04-Jun-2026]

    NVGPU-4 Secondary source Back to text

  5. NVIDIA Corporation, “TensorRT-LLM,” GitHub, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM. GitHub repository. [Accessed: 04-Jun-2026]

    NVGPU-5 Primary source Back to text

  6. vLLM Project, “GPU Installation — AMD ROCm,” 2026. [Online]. Available: https://docs.vllm.ai/en/latest/getting_started/installation/gpu/. Upstream ROCm 6.3+ support, pre-built wheels for ROCm 7.0/7.2.1, official ROCm Docker images available from Jan. 20, 2026, superseding AMD's downstream rocm/vllm images. [Accessed: 19-Jul-2026]

    NVGPU-6 Primary source Back to text

  7. NVIDIA Corporation, “NVIDIA DGX Spark,” 2026. [Online]. Available: https://www.nvidia.com/en-us/products/workstations/dgx-spark/. Product page. [Accessed: 04-Jun-2026]

    NVGPU-7 Primary source Back to text

  8. StorageReview, “NVIDIA DGX Spark Review: The AI Appliance Bringing Datacenter Capabilities to Desktops,” StorageReview.com, Nov. 13, 2025. [Online]. Available: https://www.storagereview.com/review/nvidia-dgx-spark-review-the-ai-appliance-bringing-datacenter-capabilities-to-desktops. [Accessed: 04-Jun-2026]

    NVGPU-8 Secondary source Back to text

  9. Tom's Hardware, “Nvidia DGX Spark review: the GB10 Superchip powers a fast and fun AI toolbox,” Jan. 27, 2026. [Online]. Available: https://www.tomshardware.com/pc-components/gpus/nvidia-dgx-spark-review. [Accessed: 04-Jun-2026]

    NVGPU-9 Secondary source Back to text

  10. D. Trifonov, “RTX 4090 vs 5090 vs PRO 6000: LLM Inference Benchmark,” CloudRift AI, Oct. 9, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx-gpus-for-llm-inference. [Accessed: 04-Jun-2026]

    NVGPU-10 Secondary source Back to text

  11. NVIDIA Corporation, “NVIDIA DGX Station,” 2026. [Online]. Available: https://www.nvidia.com/en-us/products/workstations/dgx-station/. Product page. [Accessed: 04-Jun-2026]

    NVGPU-11 Primary source Back to text

  12. R. Smith, “NVIDIA DGX Station Systems Available at Last: GB300 and GB200 Workstations for Your Desktop,” ServeTheHome, Mar. 20, 2026. [Online]. Available: https://www.servethehome.com/nvidia-dgx-station-systems-available-at-last-gb300-gb200-workstations-for-your-desktop/. [Accessed: 04-Jun-2026]

    NVGPU-12 Contextual source Back to text

  13. TechRadar Pro, “MSI (re)launches $85,000 Nvidia DGX Station workstation with the Nvidia GB300 Ultra, a pair of 400GbE LAN ports, and 768GB of RAM,” 2026. [Online]. Available: https://www.techradar.com/pro/msi-re-launches-usd85-000-nvidia-dgx-station-workstation-with-the-nvidia-gb300-ultra-a-pair-of-400gbe-lan-ports-and-768gb-of-ram. [Accessed: 04-Jun-2026]

    NVGPU-13 Contextual source Back to text

  14. NVIDIA Corporation, “NVIDIA DGX B200,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/dgx-b200/. Product page. Specifications: 1,440 GB total GPU memory, 14.4 TB/s NVLink aggregate, ~14.3 kW max system power. [Accessed: 04-Jun-2026]

    NVGPU-14 Primary source Back to text

  15. Microway, “NVIDIA DGX B200,” 2026. [Online]. Available: https://www.microway.com/product/nvidia-dgx-b200/. Reseller product page. [Accessed: 04-Jun-2026]

    NVGPU-15 Secondary source Back to text

  16. Microway, “NVIDIA DGX B300,” 2026. [Online]. Available: https://www.microway.com/product/nvidia-dgx-b300/. Reseller product page. Specifications: 8× B300 at 288 GB HBM3e each, 1,400 W/GPU, ~14 kW system. [Accessed: 04-Jun-2026]

    NVGPU-16 Secondary source Back to text

  17. Super Micro Computer, Inc., “NVIDIA HGX GPU/AI Server Product Catalog — Blackwell HGX B300, B200, and GB200 NVL72 Solutions,” 2025. [Online]. Available: https://www.supermicro.com/en/accelerators/nvidia. [Accessed: 04-Jun-2026]

    NVGPU-17 Primary source Back to text

  18. Dell Technologies, “PowerEdge XE9780 Technical Guide,” 2026. [Online]. Available: https://www.delltechnologies.com/asset/en-us/products/servers/technical-support/poweredge-xe9780-technical-guide.pdf. Regulatory Model E125S; also the PowerEdge XE-series Spec Sheet. [Accessed: 04-Jun-2026]

    NVGPU-18 Primary source Back to text

  19. Dell Technologies, “PowerEdge XE9780,” 2026. [Online]. Available: https://www.dell.com/en-us/shop/ipovw/poweredge-xe9780. Product page; 8× NVIDIA HGX B300 NVL8 270 GB 1,100 W SXM6 or 8× HGX B200 180 GB 1,000 W SXM6, air-cooled. [Accessed: 19-Jul-2026]

    NVGPU-19 Primary source Back to text

  20. NVIDIA Corporation, “MLPerf AI Benchmarks,” Apr. 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/resources/mlperf-benchmarks/. MLPerf Inference v6.0 results, Closed Division; results retrieved from www.mlcommons.org. [Accessed: 04-Jun-2026]

    NVGPU-20 Primary source Back to text

  21. Dell Technologies, “Dell Technologies Unveils Next-Generation Enterprise AI Solutions with NVIDIA,” Dell Technologies Investor Relations, press release, May 19, 2025. [Online]. Available: https://investors.delltechnologies.com/news-releases/news-release-details/dell-technologies-unveils-next-generation-enterprise-ai. [Accessed: 04-Jun-2026]

    NVGPU-21 Secondary source Back to text

  22. Dell Technologies, “Dell Technologies and NVIDIA Unveil Next-Generation Enterprise AI Solutions,” Dell Newsroom, press release, May 19, 2025. [Online]. Available: https://www.dell.com/en-us/dt/corporate/newsroom/announcements/detailpage.press-releases~usa~2025~05~dell-technologies-and-nvidia-unveil-next-generation-enterprise-ai-solutions.htm. [Accessed: 04-Jun-2026]

    NVGPU-22 Secondary source Back to text

  23. Super Micro Computer, Inc., “Supermicro Expands NVIDIA Blackwell Portfolio with New 4U and 2-OU (OCP) Liquid-Cooled NVIDIA HGX B300 Solutions Ready for High-Volume Shipment,” Supermicro / PR Newswire, press release, Dec. 9, 2025. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-expands-nvidia-blackwell-portfolio-new-4u-and-2-ou-ocp-liquid-cooled. [Accessed: 04-Jun-2026]

    NVGPU-23 Secondary source Back to text

  24. Super Micro Computer, Inc., “Supermicro Expands Its NVIDIA Blackwell System Portfolio with New DLC-2 Systems,” Supermicro, press release, Aug. 11, 2025. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-expands-its-nvidia-blackwell-system-portfolio-new-direct-liquid-cooled-dlc. [Accessed: 04-Jun-2026]

    NVGPU-24 Secondary source Back to text

  25. M. Tabares and C. Lutzer, “Benchmarking NVIDIA RTX Pro 6000 Blackwell on Akamai Cloud,” Akamai Blog, Oct. 30, 2025. [Online]. Available: https://www.akamai.com/blog/cloud/benchmarking-nvidia-rtx-pro-6000-blackwell-akamai-cloud. [Accessed: 04-Jun-2026]

    NVGPU-25 Secondary source Back to text

  26. D. Trifonov, “RTX PRO 6000 vs H100, H200, and L40S: LLM Inference,” CloudRift AI, Nov. 27, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx6000-vs-datacenter-gpus. [Accessed: 04-Jun-2026]

    NVGPU-26 Secondary source Back to text

  27. DatabaseMart, “Pro 6000 vLLM Inference Benchmark: LLM Throughput and Latency Analysis,” DatabaseMart Blog, Jan. 16, 2026. [Online]. Available: https://www.databasemart.com/blog/vllm-gpu-benchmark-pro6000. [Accessed: 04-Jun-2026]

    NVGPU-27 Secondary source Back to text

  28. NVIDIA Corporation, “NVIDIA RTX Blackwell PRO GPU Architecture,” 2025. [Online]. Available: https://www.nvidia.com/content/dam/en-zz/Solutions/design-visualization/quadro-product-literature/NVIDIA-RTX-Blackwell-PRO-GPU-Architecture-v1.0.pdf. Whitepaper v1.0. RTX PRO 6000 Blackwell: 96 GB GDDR7, 1.792 TB/s memory bandwidth, 24,064 CUDA cores. [Accessed: 04-Jun-2026]

    NVGPU-28 Primary source Back to text

  29. Dell Technologies, “PowerEdge XE7745 Spec Sheet,” 2026. [Online]. Available: https://www.delltechnologies.com/asset/en-us/products/servers/technical-support/poweredge-xe7745-spec-sheet.pdf. Up to 8× double-wide 600 W PCIe accelerators incl. RTX PRO 6000 Blackwell Server Edition, 96 GB GDDR7 ECC; 4U air-cooled. [Accessed: 04-Jun-2026]

    NVGPU-29 Primary source Back to text

  30. Super Micro Computer, Inc., “SYS-422GL-NR: DP Intel 4U MGX Dual-Root PCIe GPU System with up to 8 600W GPUs.” [Online]. Available: https://www.supermicro.com/en/products/system/gpu/4u/sys-422gl-nr. [Accessed: 08-Jun-2026]

    NVGPU-30 Primary source Back to text

  31. Super Micro Computer, Inc., “GPU SuperServer SYS-522GA-NRT: DP Intel 5U Dual-Root PCIe GPU System with up to 10 GPUs and extended thermal capacity,” Product Datasheet, 2026. [Online]. Available: https://www.supermicro.com/en/products/system/datasheet/sys-522ga-nrt. [Accessed: 08-Jun-2026]

    NVGPU-31 Primary source Back to text

  32. Hewlett Packard Enterprise, “HPE ProLiant Compute DL380a Gen12,” 2026. [Online]. Available: https://www.hpe.com/us/en/compute/hpe-proliant-compute/dl380a-gen12.html. Product page; up to 8 double-wide RTX PRO 6000 Blackwell Server Edition GPUs, iLO 7, quantum-resistant firmware signing. [Accessed: 19-Jul-2026]

    NVGPU-32 Primary source Back to text

  33. Vertiv, “Vertiv Codevelops with NVIDIA Complete Power and Cooling Blueprint for NVIDIA GB200 NVL72 Platform,” Vertiv Investor Relations, press release, Oct. 17, 2024. [Online]. Available: https://www.vertiv.com/en-us/about/news-and-events/corporate-news/vertiv-codevelops-with-nvidia-complete-power-and-cooling-blueprint-for--nvidia-gb200-nvl72-platform/. [Accessed: 04-Jun-2026]

    NVGPU-33 Secondary source Back to text

  34. NVIDIA Corporation, “NVIDIA Contributes NVIDIA GB200 NVL72 Designs to Open Compute Project,” NVIDIA Technical Blog, Nov. 2024. [Online]. Available: https://developer.nvidia.com/blog/nvidia-contributes-nvidia-gb200-nvl72-designs-to-open-compute-project/. [Accessed: 04-Jun-2026]

    NVGPU-34 Secondary source Back to text

  35. NVIDIA Corporation, “NVIDIA Kicks Off the Next Generation of AI with Rubin — Six New Chips, One Incredible AI Supercomputer,” NVIDIA Investor Relations, press release, Jan. 5, 2026. [Online]. Available: https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Kicks-Off-the-Next-Generation-of-AI-With-Rubin--Six-New-Chips-One-Incredible-AI-Supercomputer/default.aspx. [Accessed: 04-Jun-2026]

    NVGPU-35 Secondary source Back to text

  36. ServeTheHome, “NVIDIA Launches Next-Generation Rubin AI Compute Platform at CES 2026,” Jan. 16, 2026. [Online]. Available: https://www.servethehome.com/nvidia-launches-next-generation-rubin-ai-compute-platform-at-ces-2026/. [Accessed: 04-Jun-2026]

    NVGPU-36 Contextual source Back to text

  37. NVIDIA Corporation, “NVIDIA Vera Rubin NVL72,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/. Product page and specifications, updated Jun. 18, 2026. [Accessed: 19-Jul-2026]

    NVGPU-37 Primary source Back to text

  38. NVIDIA Corporation, “Inside the NVIDIA Vera Rubin Platform: Six New Chips, One AI Supercomputer,” NVIDIA Technical Blog, Jan. 5, 2026. [Online]. Available: https://developer.nvidia.com/blog/inside-the-nvidia-rubin-platform-six-new-chips-one-ai-supercomputer/. [Accessed: 04-Jun-2026]

    NVGPU-38 Secondary source Back to text

  39. ASHRAE Technical Committee 9.9, “Thermal Guidelines for Data Processing Environments,” ASHRAE, Mar. 2021. [Online]. Available: https://www.ashrae.org/technical-resources/bookstore/datacom-series. 5th ed. Atlanta, GA, USA. [Accessed: 14-Jun-2026]

    NVGPU-39 Primary source Back to text

  40. Upsite Technologies, “Major Changes to ASHRAE's Fifth Edition of Thermal Guidelines – Part 2: New Air-Cooled Class for High Density Compute Equipment,” Upsite Technologies Blog, 2024. [Online]. Available: https://www.upsite.com/blog/major-changes-to-ashraes-fifth-edition-of-thermal-guidelines-part-2-new-air-cooled-class-for-high-density-compute-equipment/. [Accessed: 04-Jun-2026]

    NVGPU-40 Contextual source Back to text

  41. Uptime Institute, “New ASHRAE Guidelines Challenge Efficiency Drive,” Uptime Institute Journal, 2021. [Online]. Available: https://journal.uptimeinstitute.com/new-ashrae-guidelines-challenge-efficiency-drive/. [Accessed: 04-Jun-2026]

    NVGPU-41 Secondary source Back to text

  42. Uptime Institute Intelligence, “AI and Cooling: Methods and Capacities,” Uptime Intelligence, Feb. 2026. [Online]. Available: https://intelligence.uptimeinstitute.com/resource/ai-and-cooling-methods-and-capacities. [Accessed: 19-Jul-2026]

    NVGPU-42 Secondary source Back to text

  43. ASHRAE Technical Committee 9.9, “Thermal Guidelines for Data Processing Environments, 5th Edition, Revised and Expanded,” ASHRAE. [Online]. Available: https://www.ashrae.org/file%20library/technical%20resources/bookstore/supplemental%20files/therm-gdlns-5th-r-e-refcard.pdf. Supplemental reference card. [Accessed: 19-Jul-2026]

    NVGPU-43 Primary source Back to text

Contents