Section16
Phased Hardware Deployment Roadmap
For executives. Build the on-premises capability in three phases over five-plus years, never as one Phase-3-scale purchase on day one. The GPU silicon in each phase depreciates; the fine-tuned models, validated agent workflows, and in-house expertise built on it compound. Phase 1 (months 0–12) is a pilot inference server for five to fifteen concurrent users at roughly $60K–$300K depending on configuration. Phase 2 (months 12–24+) adds fine-tuning on the organization’s data, agentic compute for tool-executing agent workflows, and inference scale-out, at roughly $500K–$950K on the H200 training tier or $1.3M–$1.8M on the B300 tier, with the training server the largest single line. Phase 3 (months 24+–60+) adds redundancy, multi-node training fabric, and expanded storage at $215K–$950K depending on scope. Five-plus-year capital expenditure lands near $0.8M–$1.9M on the lean H200 arc and $1.8M–$3.0M on the high B300 arc, directional and subject to vendor quotation. Each phase’s hardware drops into the next phase’s chassis and rack-power budget; expanding the build rather than replacing it. How far to run the arc depends on size: a small organization, fewer than 100 employees in the terms this paper fixes at the outset, often reaches steady state at Phase 1 or Phase 2, while the full three-phase build is sized for a medium organization of 100 to 499 and holds through 999 employees by adding inference nodes rather than redesigning.
Buy the on-premises capability in increments, never as a single Phase-3-scale procurement on day one. A one-shot build over-provisions against capabilities the organization has not yet learned to operate, and it commits capital to a GPU generation that will age out before the workload grows into it. The phased approach wins on a concrete mechanism. Each phase’s hardware fits inside the next phase’s chassis, rack-power envelope, and software stack with no replacement. The Phase 1 Dell PowerEdge XE7745 accepts additional RTX Pro 6000 cards in Phase 2 [1]. The rack-power budget sized for Phase 1 leaves headroom for Phase 2’s training server. The vLLM and Kubernetes stack deployed in Phase 1 runs identically on Phase 2 and Phase 3 hardware. What depreciates over five years is the specific GPU generation. What compounds is the fine-tuned models trained on proprietary data, the validated agent workflows, the indexed institutional knowledge, and the engineering expertise to run the stack.
Concurrent users size the hardware; headcount predicts them only through the share of staff granted access and their duty cycle. Phase 1 measures that ratio under production load, and until it does, any mapping from employees to phases is a planning default rather than a measurement. Phase 1’s five to fifteen concurrent users covers a pilot cohort anywhere in the small to medium organization range this paper addresses, and it is a defensible steady state for a small organization whose AI users are one department. Phase 2’s 25–50 concurrent users is the production tier for a medium organization of 100 to 499 employees. Phase 3’s 50–100+ concurrent users with multi-model serving is where a medium organization near the top of that band lands. A small organization that stops after Phase 1 or Phase 2 has sized the build correctly, and the five-year totals below apply only to organizations that run all three phases.
Cost figures throughout are directional estimates current as of mid-2026. Every line item needs a direct quotation from Dell or Supermicro before commitment, and GPU pricing volatility through 2025–2026 means a quotation even three months old must be refreshed.
Phase 1 (Months 0 to 12): The Pilot
Phase 1 delivers production value to five to fifteen concurrent users from the day it is racked. It is the first increment of the production system rather than a throwaway proof-of-concept, and the inference-server procurement that anchors it sets the trajectory for everything after. Two configurations resolve the choice.
Option A: RTX Pro 6000 (cost-optimal, recommended for most organizations). A Dell PowerEdge XE7745 chassis with one or two NVIDIA RTX Pro 6000 Blackwell Server Edition GPUs (96 GB GDDR7 each), or the Supermicro SYS-422GL-NR 4U air-cooled equivalent. A single RTX Pro 6000 sustains 8,425 tokens per second on a Qwen3-Coder-30B AWQ (activation-aware weight quantization) workload [2], roughly 18× the 450 tokens-per-second floor for fifteen concurrent users reading at thirty tokens per second each. The 96 GB envelope holds a 70B model at FP8 with key–value (KV) cache headroom, and gpt-oss-120B fits in its native MXFP4 format on one card [3]. Akamai’s controlled test on a 49B reasoning model measured the RTX Pro 6000 at FP4 delivering 1.63× the H100 NVL 96 GB’s FP8 throughput at 100 concurrent requests, narrowing toward parity by 200 [4], with independent benchmarking placing it about 28% lower in cost per token than the H100 at single-GPU scale [5]. Total Phase 1 capital expenditure (CapEx) is approximately $60K–$150K+ depending on GPU count and server configuration. Of that, the storage subsystem and 100 GbE switching account for roughly $25K–$50K.
Option B: H200 NVL PCIe (more memory and bandwidth per card, higher cost). Two to four NVIDIA H200 NVL 141 GB PCIe cards (141 GB HBM3e, 4.8 TB/s, 600 W, air-cooled) in the same XE7745 chassis, which Dell qualifies for both parts, or a Supermicro PCIe GPU server [6], [7]. The H200’s 141 GB exceeds the RTX Pro 6000’s 96 GB, and its 4.8 TB/s HBM3e bandwidth is roughly 2.7× the Pro 6000’s 1.79 TB/s GDDR7, which raises decode throughput on memory-bound serving and extends the single-card context and KV-cache ceiling. H200 NVL cards also pair over NVLink bridges, two- or four-way at 900 GB/s per GPU [6], so a model too large for one card shards across the bridged set without the PCIe bottleneck the bridgeless RTX Pro 6000 carries for tensor parallelism. The trade-off is the precision floor: the H200 is Hopper-generation and lacks native FP4 tensor cores, so where MXFP4 or FP4 throughput is the priority (including gpt-oss-120B at its native format), the Blackwell RTX Pro 6000 of Option A is the better-matched part. Choose Option B when per-card memory capacity, NVLink-bridged sharding, and bandwidth for FP8 or FP16 serving of 70B-class-and-larger models outweigh FP4 compute. Total Phase 1 CapEx is approximately $150K–$300K+ depending on GPU count and server configuration.
NVIDIA B200 and B300 systems are Phase 2 training-server sizing, not Phase 1 inference. A single 8× B200 server delivers a peak system output of 92,909 tokens per second on gpt-oss-120B in Artificial Analysis’s continuously updated load test, a point-in-time figure [8], roughly 200× the Phase 1 concurrency threshold; standing one up to serve fifteen users spends capital on capacity the pilot cannot reach.
Phase 1 storage is a high-capacity NVMe pool for model weights and hot retrieval-augmented-generation (RAG) data, a spinning-disk network-attached storage (NAS) tier for cold data, and a separate backup NAS tier covering model weights, vector-database snapshots, and dataset versioning. The backup tier is a named Phase 1 deliverable, not a Phase 3 add-on. The fine-tuned weights, curated datasets, and vector indices are the assets that justify the on-premises investment, and a hardware failure without backups deletes the compounding-value rationale the roadmap rests on. The full server taxonomy specifies the components: SeaweedFS for S3-compatible object storage, which is the maintained open-source path now that MinIO’s community edition has reached end-of-life; Qdrant or pgvector for vectors; and NFS or SMB for model-weight access. Phase 1 networking is 100 GbE switching among the inference server, storage, and management. A single-server Phase 1 deployment neither needs nor benefits from NVIDIA’s InfiniBand.
Phase 1 rack power is the constraint that makes the RTX Pro 6000 recommendation more than a cost argument. Option A draws about 1 kW of total system power with one card (600 W GPU plus roughly 250–500 W of system overhead) and about 4.8 kW GPU-only at the eight-card maximum, inside the air-cooling envelope the hardware analysis established for a standard 15–20 kW rack. Option B’s two-to-four-card H200 NVL configuration draws 1.2–2.4 kW GPU-only, roughly 2–3.2 kW at the system level, within the same envelope. Neither requires a facility cooling retrofit.
What does require attention is the facility cooling assessment, itself a named Phase 1 deliverable that must complete before any hardware order. It records the target rack’s current power capacity in kilowatts, the current cooling method and its maximum rack-kW rating, chilled-water availability and loop temperature should Phase 2 or Phase 3 hardware need direct liquid cooling, and the physical space for a coolant distribution unit (CDU) if liquid cooling enters the plan. Phase 1 hardware needs none of this. The assessment matters because its answer determines whether Phase 2’s training server can be air-cooled in the same facility or whether liquid-cooling infrastructure must be planned and capitalized eighteen months ahead of the hardware that depends on it. The governing standard is ASHRAE TC 9.9’s Thermal Guidelines for Data Processing Environments, 5th Edition, published March 2021 [9]; it predates B200- and B300-generation densities, and no later edition had superseded it as of mid-2026.
Five software capabilities are Phase 1 deliverables, not Phase 2 add-ons. Containerized vLLM inference serving runs under k3s or microk8s with the NVIDIA GPU Operator, supplying the health checks, automated restart, and rolling model updates that raw vLLM lacks and that multi-user production serving requires [10], [11]. OpenTelemetry trace logging captures every request and response with timestamp, user ID, model version, token counts, and latency; it is the operational foundation of the audit control the governance and security sections commit to, and without it that governance claim is aspirational. The RAG ingestion pipeline runs Apache Airflow on a PostgreSQL backend with LlamaIndex or Unstructured.io for document parsing and scheduled re-indexing, because a knowledge-base chatbot whose retrieval corpus is hand-populated goes stale within a month. AI-asset backup schedules weight backups to the separate NAS tier, vector-database snapshots aligned to re-indexing cadence, and dataset versioning replicated to secondary storage, with quarterly restore verification. The Phase 1 security baseline implements the subset of the security architecture that production traffic needs from day one: VLAN segmentation for AI infrastructure traffic, API-key access control on the inference endpoint, and the data-classification policy that governs which inputs the system may process.
Phase 1 capability: model serving for five to fifteen concurrent users and their agents on 30B–120B-class models at FP8 or MXFP4 quality, a RAG-backed internal knowledge chatbot, developer code assistance, and document summarization. The pilot serves real users with real workloads. The operational fluency it builds under production load is what makes the Phase 2 procurement decision evidence-driven rather than speculative. In a small organization the evidence may point to stopping here, since five to fifteen concurrent users and their agents can cover an entire engineering or analyst department in a firm below 100 employees. Two findings justify the Phase 2 order: measured saturation of the Phase 1 server, or a fine-tuning workload the Phase 1 hardware cannot run. Elapsed months are not a third.
Phase 2 (Months 12 to 24+): Expansion
Phase 2 turns the pilot into a production system for 25–50 concurrent users and their agentic workflows, adds custom fine-tuning on proprietary data, and installs the MLOps platform that makes those fine-tunes safely promotable. That user band is the medium organization’s production tier, and it is also where the training server first has a population large enough to keep it busy. Below roughly 100 employees the fine-tuning demand is usually intermittent, and a $400K–$800K training line then buys silicon that idles between runs; where no named workload justifies it, build the agentic compute, embeddings, storage, and MLOps additions and defer the training server. The training-server choice is the largest single capital allocation in the five-year arc.
8× H200 SXM5 (lower-risk, lower-cost). A Supermicro SYS-821GE-TNHR or Dell PowerEdge XE9680 with eight H200 GPUs at 141 GB HBM3e each: 1,128 GB aggregate, well-proven for fine-tuning, about 5.6 kW GPU-only and roughly 10–12 kW total depending on configuration, inside the air-cooling envelope of most enterprise facilities. Training-server CapEx is roughly $400K–$800K. The full Phase 2 build on this path, including the agentic compute, embeddings, and storage additions below, lands near $500K–$950K. This is the recommended path where the Phase 1 cooling assessment confirmed air-cooling capacity but did not provision liquid cooling.
8× HGX B300 (higher-capability, higher-cost, usually liquid-cooled). Eight Blackwell Ultra GPUs at 288 GB HBM3e and 15 PFLOPS of dense NVFP4 each, for 2,304 GB aggregate [12]. At full specification each GPU draws up to 1,400 W and requires liquid cooling. Dell’s air-cooled XE9780 de-rates the HGX B300 to 270 GB and 1,100 W per GPU to fit a roughly 10U air-cooled chassis [13], while Supermicro’s DLC-2 4U runs the full 288 GB, 1,400 W part on direct liquid cooling [14]. NVIDIA does not publish DGX or HGX B300 list pricing; mid-2026 third-party aggregations place an 8-GPU system in the low-to-mid six figures, and a custom Supermicro or Dell HGX build typically runs below a turnkey NVIDIA DGX system. Total Phase 2 on the B300 path is on the order of $1.3M–$1.8M, contingent on quotation. Between the two paths sits the 8× B200 (192 GB per GPU, 1,440 GB aggregate): Supermicro’s 4U DLC-2 HGX B200, shipping since August 2025, is the lower-cost liquid-cooled Blackwell entry [14], and the air-cooled DGX B200 draws about 14.3 kW at maximum [15], near the practical edge of air-cooling margin. It fits organizations that want Blackwell FP4 training economics without the B300’s aggregate capacity.
The decision rule is workload-conditional. The current anchor for the large-model end is NVIDIA’s Nemotron 3 Ultra, a 550B-parameter mixture-of-experts model with 55B active parameters, pretrained in NVFP4 and positioned by NVIDIA for long-running agentic workflows [16]. Its weights occupy roughly 275 GB at the native 4-bit format and about 550 GB at FP8, so an 8× H200 node’s 1,128 GB holds it at either precision with KV-cache headroom, and the same node holds Llama 4 Maverick at BF16. The H200 path therefore covers every model class the multi-user scaling analysis targets for Phase 2 production. The B300 path earns its premium only when the workload demands models past roughly 500B parameters at BF16, or trillion-parameter classes at FP8, without quantization compromise; those exceed the H200 node’s headroom, and only the B300’s 2,304 GB carries them with KV-cache margin. The H200 path is the default. The B300 path warrants justification against a specific workload requirement, not a hedge against future needs.
In May 2026 Dell and AMD added a Phase 2 option that changes the AMD calculus. Dell announced support for the AMD Instinct MI350P (144 GB HBM3e, up to 4,600 peak teraflops at MXFP4, PCIe, air-cooled) in the PowerEdge XE7745 and R7725, with availability starting July 2026 [17], [18]. The MI350P drops into the same XE7745 chassis a Phase 1 Option A deployment already runs: no new server, no liquid-cooling redesign, no separate procurement channel. Dell markets its 144 GB as the highest capacity available in a PCIe accelerator, a claim that checks out against the H200 NVL’s 141 GB [18], [6]. For Phase 2 inference expansion (not the training server), the MI350P in the existing chassis is the lowest-friction AMD adoption path on the table. The re-evaluation criteria from the NVIDIA-versus-AMD comparison still apply: confirm ROCm 7.x stability, obtain independent MI350P-versus-H200 benchmark data, which did not yet exist as of mid-2026, and run a workload-specific cost-per-token comparison before committing. Treat the evaluation as a calendar item for the twelve-month mark.
The dedicated agentic compute server is the Phase 2 addition most often missing from infrastructure plans, and its absence is what converts a working chat deployment into an agentic one that quietly underperforms. Agent tool execution, orchestration, and document parsing are CPU-bound work that starves the inference GPU when it shares the same host. The recommended configuration is an AMD EPYC 9965 (192 cores / 384 threads, Zen 5 Turin, 12-channel DDR5 up to 6400 MT/s, 614 GB/s per socket) or the cost-optimized EPYC 9755 (128 cores), with 512 GB to 1 TB of DDR5-6400 ECC RAM, four to eight NVMe SSDs, and a 25 GbE link to the inference server [19]. CapEx is about $25K–$50K+ fully configured. This server also gives a permanent home to the MLOps platform, data pipeline, and observability backend that ran on temporary infrastructure in Phase 1. For a small organization that defers the training server, this is the Phase 2 purchase that still matters, because agentic workload growth arrives well before fine-tuning demand does.
The embeddings server is the smallest Phase 2 addition: an NVIDIA L40S (48 GB GDDR6, PCIe, 350 W) in a 4U chassis or added to existing inference hardware [20]. The Phase 2 embedding workload, full knowledge-base indexing and periodic re-embedding, does not justify H100 or H200 silicon. CapEx approximately $20K–$30K.
Phase 2 power and cooling follow from the training-server choice. The 8× H200 server fits most enterprise air-cooled racks. The air-cooled de-rated B300 (Dell XE9780, roughly 10U, 13–15 kW maximum total) sits at the threshold of standard air cooling and warrants a dedicated rack with supplemental in-row cooling. The full-spec liquid-cooled B300 (Supermicro DLC-2 4U) requires the CDU and chilled-water loop whose feasibility the Phase 1 assessment established. If that assessment found liquid cooling cannot be provided by Phase 2 procurement, the H200 path is the only viable training server, a constraint to validate against the assessment’s findings rather than discover when hardware arrives.
Phase 2 software adds three named capabilities to the Phase 1 baseline. The MLOps platform is MLflow 3: experiment tracking, a model registry with LoggedModel lineage, and a CI/CD promotion pipeline, self-hosted on the agentic compute server with a PostgreSQL backend and the on-premises object store for artifacts, with no commercial license and no cloud dependency [21]. Without it, fine-tuning produces artifacts of unknown provenance with no audit trail and no rollback; the training hardware exists to produce custom models, and MLflow is the discipline that keeps them from becoming orphan files. Orchestration graduates from k3s to full Kubernetes with the NVIDIA GPU Operator and llm-d for KV-cache-aware routing [22]. Structured quality evaluation adds drift detection and scoring on top of Phase 1’s OpenTelemetry traces; Langfuse is the recommended self-hostable layer, with Arize Phoenix the alternative where stronger statistical drift detection matters.
Once Phase 2 agentic workflows reach production, evaluate prefill/decode disaggregated serving against head-of-line blocking under mixed load. The concurrency-architecture analysis specifies the mechanism; the trigger is specific and observable, namely 95th-percentile (P95) time-to-first-token (TTFT) degrading under mixed agentic load while average GPU utilization stays moderate. That signature is head-of-line blocking, and it is the signal that disaggregation will pay off. Deploying it before the signal appears buys complexity with no return.
Phase 2 capability: custom fine-tuning of domain models on internal data, production agentic workflows for 25–50 users and their agents, a full RAG knowledge base over indexed institutional documentation, and versioned models with CI/CD promotion. The fine-tuned models Phase 2 produces are the assets that justify Phase 3’s redundancy investment; they exist only because Phase 2 hardware was bought, and they need the availability guarantees Phase 3 provides.
Phase 3 (Months 24+ to 60+): Full Deployment
Phase 3 builds the production-grade availability and capacity headroom that separate a working deployment from a resilient one: redundancy, multi-node training fabric, expanded storage, and observability that scales to 50–100+ concurrent users with multi-model serving. Its hardware additions are smaller than Phase 2’s because the foundational decisions were already made. That concurrency band puts Phase 3 at the upper end of the 100 to 499 employee range this paper defines as medium, and the same additions continue to hold for organizations of 500 to 999 employees, a tier IDC still counts inside its small and midsize business category [23]. The second inference server is what carries that extension: it establishes horizontal scale-out, so capacity past it arrives by adding nodes to an existing fabric under an unchanged software stack. Above 999 employees the architecture still functions, but the sizing evidence assembled here does not cover it, and the figures below should be treated as a floor rather than a plan.
The second inference server is the headline Phase 3 addition. Whatever ships at Phase 3 procurement in 2028–2030 will have superseded today’s Vera Rubin generation (288 GB HBM4, 50 PFLOP NVFP4, partner availability from the second half of 2026) [24], but the architectural role holds across generations: horizontal scaling and redundancy against Phase 1’s single point of failure. Phase 1’s server keeps serving inference, Phase 2’s handles training and fine-tuning, and the Phase 3 second inference server enables A/B model testing, zero-downtime model updates, and failover. CapEx approximately $215K–$510K conservative or $350K–$700K mid-range, depending on GPU generation and configuration.
The expanded training cluster adds InfiniBand NDR (400 Gbps) between the Phase 2 training server and additional GPU nodes for multi-node all-reduce. Phase 1 and Phase 2 used 100 GbE because a single training server does not need InfiniBand; a multi-node cluster does. NVIDIA Quantum-2 InfiniBand switches with SHARP run collective operations in the switch fabric, cutting all-reduce latency [25], [26]. Meta’s two 24,576-GPU Llama 3 clusters, one on RoCE Ethernet and one on Quantum-2 InfiniBand, trained large models including Llama 3 itself on both without hitting network bottlenecks [27], [28], which is why Ethernet stays defensible through single-node scale. The InfiniBand premium is real but earns its place only once multi-node training collectives dominate. Most organizations below 500 employees never reach that point, and this line is the first Phase 3 item to cut when they do not. CapEx for the InfiniBand expansion: $50K–$100K.
The storage expansion adds NVMe over Fabrics (NVMe-oF) with GPUDirect Storage for direct GPU-to-storage data paths during training [29], grows the NAS and object-storage tiers to 500 TB to 1 PB for accumulated datasets and checkpoints, and scales out the observability backend that Phase 1’s traces and Phase 2’s quality evaluation generate at production volume. CapEx $50K–$150K+ depending on capacity.
Phase 3 capability: multi-model inference serving several models at once, redundant infrastructure with no single point of failure, multi-agent orchestration at production scale, a continuous fine-tuning pipeline that retrains production models on incremental data, and observability with drift alerting and per-user cost attribution. At the end of Phase 3 the organization holds the on-premises AI capability the strategy argued for: operational, resilient, and scaling against its own demand rather than a vendor’s allocation queue.
The Phase Summary Table
| Phase | Timeline | Hardware Additions | CapEx (directional) | Capability Delivered | User Capacity | Power Increment |
|---|---|---|---|---|---|---|
| Phase 1 | 0 to 12 months | Option A: 1–2× RTX Pro 6000 in Dell XE7745 or Supermicro SYS-422GL-NR; Option B: 2–4× H200 NVL 141 GB PCIe; storage server with backup NAS; 100 GbE switching | Option A: $60K–$150K; Option B: $150K–$300K (storage/networking ≈ $25K–$50K of either) | RAG chatbot, code assistance, document summarization on 30B–120B models | 5–15 users + agents (pilot at any size; steady state for a small organization) | ~2–8 kW (air-cooled) |
| Phase 2 | 12 to 24+ months | 8× H200 (Dell XE9680 / Supermicro SYS-821GE-TNHR), 8× B200, or 8× HGX B300 training server; EPYC 9965/9755 agentic compute; L40S embeddings; expanded storage; optional MI350P in the Phase 1 chassis | H200 path total: ≈$500K–$950K; B300 path total: ≈$1.3M–$1.8M | Custom fine-tuning; production agents; full RAG; MLOps with CI/CD | 25–50 users + agents (medium organization production tier) | +~10 kW (H200) to +13–15 kW (B200 / air-derated B300) |
| Phase 3 | 24+ to 60+ months | Second inference server (post-Rubin class); InfiniBand NDR fabric (+$50K–$100K); NVMe-oF and expanded NAS (+$50K–$150K); observability scale-out | Second server: $215K–$510K conservative; $350K–$700K mid-range | Multi-model serving; redundant infrastructure; continuous fine-tuning pipeline | 50–100+ users + agents (upper medium band; holds through 999 employees) | +7–15 kW |
| Total 5-plus-year | — | — | ≈$0.8M–$1.9M (H200 Phase 2 arc) or ≈$1.8M–$3.0M (B300 Phase 2 arc) | Full multi-node, multi-model, redundant on-premises AI infrastructure | Small through medium, below 500 employees, and to 999 by node addition | — |
The total row spans the Phase 1 option and the Phase 2/3 configuration choices; the Phase 3 InfiniBand fabric and storage expansion ($100K–$250K combined) are incremental for organizations that reach multi-node training, and the lean floor assumes an organization that stops at the conservative Phase 3 second server. An organization below 100 employees that stops at Phase 2 carries the Phase 1 and Phase 2 lines only, and the total row does not describe its commitment. The Phase 1 and Phase 2 software deliverables sit in the master taxonomy and are not duplicated here. Every figure needs a vendor-quotation refresh at procurement. The HGX Blackwell supply environment in mid-2026 remained allocation-constrained, with NVIDIA-confirmed sell-through into mid-year and lead times that varied from weeks to quarters depending on OEM relationship, while H200 and RTX Pro 6000 hardware remained readily available; ordering Phase 1 on immediately available silicon preserves the lead time to bring Phase 2 Blackwell hardware in during the Phase 2 window. Confirm lead times in writing at quotation.
The live decision at the close of this analysis is the Phase 1 order: a Dell PowerEdge XE7745 or Supermicro SYS-422GL-NR with one or two RTX Pro 6000 Blackwell Server Edition cards (Option A), or two to four H200 NVL PCIe cards (Option B), validated against the hardware performance matrix and the facility cooling assessment, with the five Phase 1 software deliverables in place from day one. That order commits the organization to nothing it cannot expand, since every later phase drops into the chassis, rack, and stack it establishes. What remains before installation is the deployment architecture that isolates it: the network segmentation and, where the data classification demands it, the air gap that turn the procurement order into a production system rather than a server in a closet.
References
-
Dell Technologies, “PowerEdge XE7745 Spec Sheet,” 2026. [Online]. Available: https://www.delltechnologies.com/asset/en-us/products/servers/technical-support/poweredge-xe7745-spec-sheet.pdf. Up to 8 double-wide 600 W PCIe accelerators incl. RTX PRO 6000 Blackwell Server Edition, 96 GB GDDR7; 4U air-cooled. [Accessed: 14-Jun-2026]
-
D. Trifonov, “RTX 4090 vs 5090 vs PRO 6000: LLM Inference Benchmark,” CloudRift AI, Oct. 9, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx-gpus-for-llm-inference. [Accessed: 04-Jun-2026]
-
OpenAI, “gpt-oss-120b & gpt-oss-20b Model Card,” arXiv, 2025. [Online]. Available: https://arxiv.org/abs/2508.10925. arXiv:2508.10925. [Accessed: 14-Jun-2026]
-
M. Tabares and C. Lutzer, “Benchmarking NVIDIA RTX PRO 6000 Blackwell on Akamai Cloud,” Akamai Blog, Oct. 30, 2025. [Online]. Available: https://www.akamai.com/blog/cloud/benchmarking-nvidia-rtx-pro-6000-blackwell-akamai-cloud. [Accessed: 14-Jun-2026]
-
D. Trifonov, “RTX PRO 6000 vs H100, H200, and L40S: LLM Inference,” CloudRift AI, Nov. 27, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx6000-vs-datacenter-gpus. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA H200 Tensor Core GPU Datasheet,” 2026. [Online]. Available: https://resources.nvidia.com/en-us-gpu-resources/hpc-datasheet-sc23. H200 NVL: 141 GB HBM3e, 4.8 TB/s, 600 W, PCIe double-wide, 2–4-way NVLink bridge at 900 GB/s, air-cooled. [Accessed: 14-Jun-2026]
-
Dell Technologies, “PowerEdge XE Series Spec Sheet,” 2026. [Online]. Available: https://www.delltechnologies.com/asset/en-us/products/servers/technical-support/poweredge-xe-ai-spec-sheet.pdf. XE7745 qualified GPUs incl. NVIDIA RTX Pro 6000 Blackwell Server Edition 600 W / 96 GB and NVIDIA H200 NVL 600 W / 141 GB. [Accessed: 18-Jul-2026]
-
Artificial Analysis, “AI Hardware Benchmarking and Performance Analysis,” Artificial Analysis, 2026. [Online]. Available: https://artificialanalysis.ai/benchmarks/hardware. Continuously updated; point-in-time figure. [Accessed: 14-Jun-2026]
-
ASHRAE Technical Committee 9.9, “Thermal Guidelines for Data Processing Environments,” ASHRAE, Mar. 2021. [Online]. Available: https://www.ashrae.org/technical-resources/bookstore/datacom-series. 5th ed. Atlanta, GA, USA. [Accessed: 14-Jun-2026]
-
vLLM Project, “vLLM Documentation,” PyTorch Foundation, 2026. [Online]. Available: https://docs.vllm.ai/en/latest/. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA GPU Operator Documentation,” NVIDIA Cloud Native Technologies, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era,” NVIDIA Technical Blog, 2025. [Online]. Available: https://developer.nvidia.com/blog/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai-factory-era/. B300: 288 GB HBM3e, 8.0 TB/s, 15 PFLOPS dense NVFP4, up to 1,400 W. [Accessed: 14-Jun-2026]
-
Dell Technologies, “PowerEdge XE9780 Technical Guide,” 2026. [Online]. Available: https://www.delltechnologies.com/asset/en-us/products/servers/technical-support/poweredge-xe9780-technical-guide.pdf. HGX B300 NVL8, 270 GB/GPU, 1,100 W, ~10U air-cooled; XE9780L direct-liquid-cooled variant. [Accessed: 14-Jun-2026]
-
Super Micro Computer, Inc., “NVIDIA HGX GPU/AI Server Product Catalog — Blackwell HGX B300, B200, and GB200 NVL72 Solutions,” 2026. [Online]. Available: https://www.supermicro.com/en/accelerators/nvidia. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA DGX B200 Datasheet,” 2026. [Online]. Available: https://resources.nvidia.com/en-us-dgx-systems/dgx-b200-datasheet. ~14.3 kW max system power, 10U air-cooled. [Accessed: 18-Jul-2026]
-
NVIDIA Corporation, “NVIDIA Nemotron 3 Ultra,” NVIDIA Research, June 2026. [Online]. Available: https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/. 550B total / 55B active mixture-of-experts, hybrid Mamba-Transformer, pretrained in NVFP4, 1M-token context. [Accessed: 18-Jul-2026]
-
Dell Technologies, “Dell and AMD Are Expanding What's Possible for On-Premises AI,” Dell Technologies Blog, May 7, 2026. [Online]. Available: https://www.dell.com/en-us/blog/dell-and-amd-are-expanding-what-s-possible-for-on-premises-ai/. [Accessed: 14-Jun-2026]
-
AIwire/HPCwire, “Dell, AMD Expand On-Prem AI Platform with Instinct MI350P GPU Support,” May 8, 2026. [Online]. Available: https://www.hpcwire.com/aiwire/2026/05/08/dell-amd-expand-on-prem-ai-platform-with-instinct-mi350p-gpu-support/. [Accessed: 14-Jun-2026]
-
Advanced Micro Devices, Inc., “AMD EPYC 9005 Series Processors,” 2026. [Online]. Available: https://www.amd.com/en/products/processors/server/epyc/9005-series.html. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA L40S GPU Datasheet,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/l40s/. [Accessed: 14-Jun-2026]
-
MLflow Project, “MLflow Documentation,” 2026. [Online]. Available: https://mlflow.org/docs/latest/. MLflow 3. [Accessed: 14-Jun-2026]
-
llm-d Project, “llm-d: Achieve State-of-the-Art Inference Performance with Modern Accelerators on Kubernetes,” GitHub, 2026. [Online]. Available: https://github.com/llm-d/llm-d. CNCF Sandbox. [Accessed: 14-Jun-2026]
-
International Data Corporation, “Worldwide ICT spending to reach $4.3 trillion in 2020 led by investments in devices, applications, and IT services, according to a new IDC spending guide,” Business Wire, press release, Feb. 18, 2020. [Online]. Available: https://www.businesswire.com/news/home/20200218005150/en/Worldwide-ICT-Spending-to-Reach-%244.3-Trillion-in-2020-Led-by-Investments-in-Devices-Applications-and-IT-Services-According-to-a-New-IDC-Spending-Guide. [Accessed: 30-Jul-2026]
-
NVIDIA Corporation, “NVIDIA Vera Rubin NVL72,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/vera-rubin-nvl72/. Product page; 288 GB HBM4, 50 PFLOP NVFP4 per GPU, partner availability H2 2026. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “The NVIDIA Quantum InfiniBand Platform,” 2026. [Online]. Available: https://www.nvidia.com/en-us/networking/products/infiniband/. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA Quantum InfiniBand Switches and Appliances,” 2026. [Online]. Available: https://www.nvidia.com/en-us/networking/infiniband-switching/. [Accessed: 14-Jun-2026]
-
Meta Platforms, Inc., “Building Meta's GenAI Infrastructure,” Engineering at Meta, Mar. 12, 2024. [Online]. Available: https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/. [Accessed: 14-Jun-2026]
-
A. Grattafiori et al., “The Llama 3 Herd of Models,” arXiv, Llama Team, AI @ Meta, 2024. [Online]. Available: https://arxiv.org/abs/2407.21783. arXiv:2407.21783. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA GPUDirect Storage Overview Guide,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/gpudirect-storage/overview-guide/. [Accessed: 18-Jul-2026]