Section13

Full Server Taxonomy: All Systems Required

A production on-premises AI deployment has, at minimum, ten distinct roles. Together, they constitute a small cluster rather than a server purchase. Omitting any one of them either keeps the deployment from reaching production quality or lets it degrade quietly once it gets there. Six are hardware components. The inference server runs the selected model, processes input tokens, and generates output tokens. The training server fine-tunes open models on proprietary data. The agentic compute server executes the tool calls, commands, and orchestration that real agent workflows generate. The storage server holds model weights, chat history, model inputs, model outputs, documents, files, codebases, databases, datasets, configurations, and more. The embeddings server produces the vectors that retrieval depends on, though in Phase 1 it borrows the inference GPU rather than owning hardware. Networking ties those systems together. The remaining four roles are software: the MLOps platform, the data pipeline, the observability layer, and the cluster management layer that provisions, schedules, and monitors the physical fleet. All four are what convert isolated model execution into a maintained, repeatable production capability. Two concerns cut across every hardware line and receive their own treatment below: out-of-band management, which is the recovery path when a node is unreachable, and the facility space, power, and cooling that every purchase order must clear before it is signed. Each component below carries the hardware, software, and dependencies a real deployment needs.

Inference Server

The inference server is the production workload. It runs the model, reads input tokens, generates output tokens, and carries the value the rest of the paper argues for. Phase 1 runs one or two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs in a Dell PowerEdge XE7745 or Supermicro SYS-422GL-NR chassis; a single card sustains 8,425 tokens per second on a 30B AWQ workload and holds a 70B model at FP8 inside its 96 GB VRAM ceiling with ~26 GB left over for the KV cache [1]. One interconnect fact shapes the two-card configuration: the RTX PRO 6000 carries no NVLink, so card-to-card traffic rides PCIe 5.0 [2]. Data parallelism, one model replica per card for added concurrency, sidesteps that limit entirely; tensor parallelism across the pair remains workable for models that exceed a single card’s VRAM, at a bandwidth penalty relative to the NVLink-equipped SXM parts described under the training server. Phase 2 adds a separate 8× H200 or B200 SXM5 server when training and high-concurrency multi-tenant serving arrive.

The serving software is vLLM with the V1 engine as the production default [3]. PagedAttention manages the KV cache in fixed-size pages, the way an operating system pages virtual memory, which drives KV-cache fragmentation close to zero and lets the server batch far more concurrent requests than a contiguous-allocation scheme. The peer-reviewed evaluation reports a 2–4× throughput gain over prior production systems such as FasterTransformer and Orca at matched latency [4]. vLLM is now the default open-source serving engine and runs in production across major model providers and enterprises [3].

Containerization is the second decision most commonly skipped. Raw vLLM as a bare-metal Python process has no health checks, no automated restart, no rolling model updates, and no resource isolation, all of which production multi-user serving requires. Kubernetes supplies them. For a Phase 1 single-node deployment where a full control plane is operational overhead a small team cannot justify, k3s or microk8s are the lightweight equivalents. The NVIDIA GPU Operator runs on either and automates driver installation, CUDA toolkit deployment, device-plugin configuration, and GPU health monitoring [5]. Without it, GPU scheduling in Kubernetes becomes a manual integration project that consumes the time the team needs for model and pipeline work; with it, GPU pods carry the same lifecycle guarantees as any other production service. The Phase 1 baseline is therefore vLLM containerized under k3s with the GPU Operator, not vLLM on bare metal. Phase 2 moves to full Kubernetes as multi-server orchestration and the multi-user scaling architecture come into scope.

A load balancer sits in front of the inference pods only for Phase 2 multi-server deployments; a single-server Phase 1 install defers it. Triton Inference Server is worth considering for ensemble serving and dynamic batching across heterogeneous models, but vLLM’s continuous batching covers the Phase 1 single-model case without the extra component.

GPU Partitioning and Virtualization

A 96 GB GPU is often more GPU than one workload needs. The mechanisms for dividing it are a scheduling decision, not an afterthought. Multi-Instance GPU (MIG) partitions a single physical GPU into fully isolated instances, each with its own memory slice, cache, and compute cores and a guaranteed quality of service (QoS). The RTX PRO 6000 Blackwell Server Edition supports up to four instances per card (24 GB each), and the H100/H200-class data-center GPUs support up to seven [2], [6]. Isolation is the point: a runaway experiment in one instance cannot evict the KV cache of the production model in another. The GPU Operator manages MIG profiles declaratively in Kubernetes, so a partition layout is a configuration file rather than a manual driver procedure [5].

The Phase 1 primary card runs whole. A 70B-class chat model wants the full 96 GB, and MIG stays off. The mechanism earns its place when a second card arrives: one card serves production inference intact while the other is carved into instances for the embedding endpoint, a small utility model, and an isolated development slice. The GPU Operator’s time-slicing mode is the lighter alternative, sharing one GPU across pods without memory isolation or QoS guarantees; it is acceptable for development and wrong for production tenancy.

Hypervisor-based GPU virtualization is the third option (the one to skip by default). NVIDIA vGPU software presents virtual GPUs to virtual machines under a hypervisor and requires a paid per-GPU license [7]. For an organization already standardized on a virtualized estate it integrates cleanly; for everyone else, bare-metal Kubernetes with MIG delivers equivalent isolation without the license cost or the hypervisor layer; it’s the configuration this taxonomy assumes.

Training and Fine-Tuning Server

The training server is the Phase 2 hardware addition that makes custom training on proprietary data operationally real. It runs a periodic load rather than a continuous one: idle for some time and then saturating GPU compute during a lengthy run. The hardware is 8× H200 SXM5 or B200/B300 SXM5 with NVSwitch fabric for all-reduce operations, fed by high-bandwidth storage I/O (NVMe-oF or, at Phase 3, InfiniBand to the SAN), with the software standardized on PyTorch FSDP, DeepSpeed, or Unsloth for distributed training across the GPU domain.

The NVSwitch fabric is what separates an SXM node from a chassis of PCIe cards. Fifth-generation NVLink gives each Blackwell GPU 1.8 TB/s of interconnect bandwidth, up from 900 GB/s per Hopper GPU, and the node’s NVSwitch chips extend those links into an all-to-all fabric in which every GPU reaches every other at full NVLink rate [8]. Gradient all-reduce inside the node therefore never touches PCIe or the external network. That scale-up domain is the reason a single 8-GPU server covers most small-to-mid-organization fine-tuning without any multi-node fabric, and the reason InfiniBand can wait for Phase 3. Checkpoint I/O is the other bottleneck worth engineering. NVIDIA GPUDirect Storage (GDS) opens a direct memory access path between GPU memory and local NVMe or NVMe-oF storage, bypassing the bounce buffer through host CPU memory [9]; enabling it on the training server’s storage path shortens the checkpoint save and restore cycles that long fine-tuning runs repeat constantly.

The MLOps dependency separates a working training server from an expensive science project. A fine-tuning run without experiment tracking is not reproducible. A trained model without a registry has no controlled promotion path and no rollback when it degrades. Sculley et al. named this failure mode “hidden technical debt”: machine-learning components become entangled with their data, hyperparameters, and surrounding infrastructure, so a change in one place propagates unpredictably, and the maintenance cost grows over time and is hard to pay down later [10]. Fine-tuning runs without version tracking incur that debt the moment the first model artifact changes hands. The training server is sound only when paired with the MLOps platform specified later in this section; plan and procure the two together.

Embeddings Server

Embeddings generate the vector representations that drive retrieval-augmented generation (RAG), with a compute profile lighter than primary inference. For Phase 1, the right answer is no dedicated embeddings server. vLLM exposes an embedding endpoint that shares the inference server’s GPU with the chat model. The Sentence Transformers library covers the dominant open-weight embedding models through the same deployment pattern: all-mpnet-base-v2 and all-MiniLM-L6-v2, the E5 family, and newer top-ranked open models such as BGE-M3 [11], [12]. A separate server earns its place once sustained embedding throughput is high enough to contend with chat inference on the primary GPU, a threshold most deployments reach somewhere around a few hundred requests per minute rather than at a fixed number.

When the dedicated server is warranted, the NVIDIA L40S is the right GPU: 48 GB GDDR6 with ECC, 864 GB/s memory bandwidth, up to 1,466 INT8 TOPS with sparsity, PCIe form factor, 350 W TDP, available from Dell and Supermicro in 4U GPU servers [13]. Its throughput trails an H100 or H200, but the embedding workload does not justify that silicon, and the L40S costs substantially less. The NVIDIA A10 (24 GB, 600 GB/s, 150 W) remains an option for very light embedding loads or rack-power-constrained installs [14]. Phase 1 runs embeddings on the inference server’s GPU via vLLM; Phase 2 adds an L40S server if retrieval volume demands it. A MIG instance carved from a second inference GPU, described under GPU partitioning and virtualization, is the middle step between the two.

Agentic Compute Server

The agentic compute server is the component most often left out of AI infrastructure planning. Leaving it out is what turns a working deployment into an agentic one that quietly underperforms. Agents running real workflows generate substantial CPU work: Python subprocess execution, document parsing, code compilation, web-content processing, vector-database queries, tool-result handling, inter-agent message routing, and the orchestration glue between them. With no host of its own, all of it competes for the inference server’s CPU and memory and starves the model of the resources it needs.

Earlier-generation reference CPUs (AMD EPYC 9654 Genoa, Intel Xeon 8592+ Sapphire Rapids) are now superseded. The current AMD recommendation is the EPYC 9965: 192 cores and 384 threads on Zen 5/Zen 5c Turin, 384 MB L3, 12-channel DDR5 up to 6400 MT/s at 614 GB/s per socket, 500 W TDP, and a $14,813 list (1K-unit) price with real-world retail pricing around $9,900–$11,800 as of mid-2026 [15], [16]. Its edge over the EPYC 9654 comes mostly from doubling the core count; AMD reports up to roughly 2× the inference throughput of the prior 96-core generation for a dual-9965 node [16]. The EPYC 9755 (128 cores, Zen 5, 500 W) is the cost-optimized alternative for lighter agentic loads or tighter rack-power budgets. AMD’s successor, the Zen 6 “Venice” series, is expected later in 2026.

On the Intel side, the Xeon 6980P (Granite Rapids) is the current-generation flagship: 128 cores and 256 threads, 504 MB L3, 12-channel DDR5-6400 (MRDIMM-8800 capable), 500 W, at a $13,955 list price [17]. The EPYC 9965 therefore delivers 50% more cores than the Xeon 6980P at a comparable price. At matched 128 cores, Phoronix’s end-of-2025 Linux testing put the EPYC 9755 at 1.63× the 6980P on the geometric mean of nearly 200 benchmarks; Intel took first place only on the AMX-accelerated CPU-inference workloads and the most memory-bandwidth-bound tests utilizing MRDIMM-8800 [18]. For general-purpose Linux server throughput, which is what agentic tool execution and orchestration are, AMD wins on core density and per-dollar performance.

A Phase 1 deployment can use existing server infrastructure for the modest agentic load it carries at that stage. The Phase 2 specification is an AMD EPYC 9965 (or the cost-optimized 9755), 512 GB to 1 TB of DDR5-6000 ECC, four to eight NVMe SSDs for fast scratch storage, and a 25 GbE link to the inference server. The Xeon 6980P is the equivalent for shops with an existing Intel relationship. The same host is the right platform for the CPU-only inference cases (small models, batch jobs, low-volume work where a GPU is not justified), but its primary job is agentic tool execution, orchestration, and the data-pipeline and observability workloads below. It does not replace the GPU inference server. Where no separate management host exists, the same server can also carry the head-node functions and monitoring stack described under cluster management, provisioning, and monitoring.

Data Storage Server

Storage holds the assets the entire deployment compounds value against: model weights from hundreds of gigabytes to several terabytes, training datasets, vector indices, conversation logs, agent memory, and checkpoints. Storage hardware consists of a high-capacity NVMe pool for active weights and hot data, a spinning-disk NAS tier for cold data and backups, 100 GbE or InfiniBand to the compute servers, and redundant controllers. Specify the NVMe pool’s fabric path with NVMe-oF and GPUDirect Storage compatibility from the start, so the Phase 2 training server inherits checkpoint I/O that does not funnel through host CPUs [9].

File storage exposes a directory hierarchy reached by path over protocols like NFS or SMB, which suits a small number of large, sequentially read files such as model weights. Object storage instead keeps data as flat, API-addressed objects (data + metadata + a unique key) in buckets reached over the S3/HTTP interface, which scales horizontally and fits large, growing collections of unstructured items like datasets, conversation logs, checkpoints, and snapshots. For model-weight access, NFS or SMB over the storage fabric is sufficient.

Two software decisions and one operational requirement carry more weight than the hardware.

The first decision is object storage. MinIO, the former leading open-source S3-compatible single-binary object store, wound its community edition down to end-of-life between 2025 and early 2026. Its management console was stripped in mid-2025, community Docker images and binaries stopped publishing in October 2025, the project entered maintenance mode in December 2025, and the GitHub repository was archived read-only in February 2026 and locked again in April 2026 with no further patches [19], [20]. The last community container images predate the October 2025 release that fixed a privilege-escalation CVE, so a default minio/minio pull now ships a known, unpatched vulnerability [20]. The maintained path is the commercial AIStor subscription. Other commercial options are Cloudian (enterprise on-premises object-storage appliances built for petabyte-scale S3 in local data centers), DataCore Swarm (turnkey appliances that keep S3-compatible storage and backups local in resource-limited sites), and Dell ObjectScale (object-storage software for running modern S3 workloads on local all-flash and HDD).

For a Phase 1 deployment the practical open-source choices are the actively maintained alternatives. SeaweedFS is the strongest default and the closest drop-in for MinIO workflows. Written in Go and Apache 2.0-licensed, with more than a decade of development and an O(1)-disk-seek design that stays fast across billions of small objects, it offers full coverage of the core S3 API and is the most production-proven of the open options [21]. Garage is the lightweight, geo-distributed alternative. A single Rust binary with a tiny footprint that runs comfortably on a virtual private server (VPS), with first-class multi-site replication. It is AGPLv3 (the same copyleft posture that made MinIO’s license a concern, fine for purely internal use but a factor if storage is embedded in a distributed product) and its S3 surface is narrower, with replication only and no erasure coding [22]. Ceph’s RADOS Gateway remains the choice for organizations already operating Ceph or needing the widest advanced S3 feature surface, such as object lock and full lifecycle rules, at the cost of materially heavier operations. RustFS, an Apache 2.0 Rust alternative, is emerging but not yet proven enough to deploy in production.

The second decision is the vector database, currently a choice among five credible open-source engines where operational fit is the deciding factor over raw benchmarks. PostgreSQL with the pgvector extension is the path of least resistance for teams already running Postgres. Vectors sit beside relational data on infrastructure the team already backs up and monitors, with no second service to run [23]. It stays comfortable into the millions of vectors, and the PostgreSQL-licensed pgvectorscale extension stretches that ceiling with a disk-resident StreamingDiskANN index that holds the working set on NVMe drives under bounded memory [24]. The cost is tuning: Postgres was not built for vector search, so very large or write-heavy workloads demand more of it than a dedicated engine does. Chroma (Apache 2.0, Rust core) sits at the other low-friction extreme, optimizing for developer experience over scale. A simple install and a few lines of Python stand up an embedded, in-memory or persistent store. It is the fastest route to a working RAG prototype and the default vector store across most LangChain and LlamaIndex examples [25]. Its 1.0 rewrite and the serverless Chroma Cloud added distributed deployment and unified dense, sparse BM25, and full-text search from one codebase, but at this tier Chroma is best read as the prototyping and single-node choice; production multi-tenancy, high availability, and scale past a few million vectors are where it still trails the standalone engines.

The two purpose-built engines that fit this class squarely are Qdrant and Weaviate. Qdrant (Apache 2.0, a single Rust binary) is the strongest default when retrieval is the primary workload and the team does not already run Postgres: it self-hosts simply and gives strong filtered search with low p99 latency under concurrency [26]. Its recent work targets filtered RAG directly, with GPU-accelerated HNSW indexing across NVIDIA, AMD, and Intel hardware, the in-house Gridstore engine that replaced RocksDB for steadier write latency, and a per-query ACORN-1 mode that restores recall when strict metadata filters fragment the graph [27], [28], [29]. Weaviate (BSD-3-Clause, written in Go) is built instead around hybrid search and batteries-included RAG. It fuses BM25 keyword and vector results in one query with a tunable weight, and its module system vectorizes text at ingest through OpenAI, Cohere, Hugging Face, or Google, with reranking and generative search on the same API [30]. That integration removes a layer of glue code, at the cost of more moving parts: a GraphQL-and-REST surface, real Docker or Kubernetes skill to self-host well, and noticeably more RAM than pgvector or Qdrant for an equivalent corpus.

Milvus (Apache 2.0, Go and C++, an LF AI & Data Foundation graduate) is the scale-out heavyweight. A Kubernetes-native architecture separates compute from storage, the index menu is the broadest of the group (HNSW, IVF, DiskANN, SCANN, and GPU variants), and deployments reach into the tens of billions of vectors [31]. It ships three modes: Lite for Python prototyping, Standalone for single-machine production, and Distributed for cluster scale giving teams the option to start small without a later rewrite. For a one-to-few-server deployment, full Distributed Milvus is usually overkill; its advantage appears only once corpus and query volume outgrow what a single Qdrant or Postgres node can carry. For most deployments in this class the choice narrows quickly: pgvector if PostgreSQL is already in production and Qdrant if it is not, Weaviate when hybrid search is the core requirement, Chroma to prototype, and Milvus held in reserve for genuine billion-scale.

The operational requirement that turns this into a Phase 1 deliverable is backup capabilities. Fine-tuned weights, vector indices, and curated datasets are the irreplaceable assets that justify the on-premises investment in the first place. A hardware failure without a backup destroys the compounding value the ROI argument rests on. Three controls belong in Phase 1: weight backups to a separate NAS or tape tier on a defined schedule, vector-database snapshots aligned to re-indexing cadence, and dataset versioning replicated to a secondary pool. All three run on standard tooling: bucket mirroring to secondary storage, the vector store’s snapshot API, and the MLOps artifact store described below. The extra backup tier is cheap. Skipping it forfeits the entire investment thesis.

Networking Infrastructure

Networking is a smaller decision than the GPU-procurement budget suggests, because for Phase 1 and Phase 2 the answer is 100 GbE for everything. A single inference server and a single training server do not need InfiniBand. The 100 GbE link to storage sustains roughly 12.5 GB/s, enough to load a 70B FP16 model (140 GB) from object storage in about eleven seconds, which is the relevant constraint for hot-swap model serving; inter-server coordination at this scale fits comfortably in the same envelope.

“Everything” still divides into planes. NVIDIA’s reference designs run four distinct networks: a compute fabric for east-west traffic (GPU to GPU, server to server), a storage fabric between the compute servers and the storage server, an in-band management network carrying SSH, the Kubernetes API, and monitoring, and a physically separate out-of-band (OOB) management network that reaches each server’s baseboard management controller (BMC) [32]. At Phase 1–2 scale the first three collapse onto the single 100 GbE physical fabric with VLAN separation, which is the segmentation described at the end of this section. The OOB network does not collapse with them; it stays on its own small switch for reasons covered under out-of-band management.

Remote direct memory access (RDMA) is the technology the compute and storage fabrics grow into. RDMA lets a network interface card (NIC) move data directly between the memory of two hosts without traversing either operating system’s network stack, which removes CPU overhead from the data path and cuts latency. InfiniBand carries it natively, and RoCE v2 (RDMA over Converged Ethernet) carries it over routable Ethernet. GPUDirect RDMA extends the same path to the GPU: the NIC reads and writes GPU memory directly, so multi-node gradient exchange and NVMe-oF storage traffic bypass host memory entirely [33]. A Phase 1 single server needs none of this. The procurement decision it drives is forward-looking: buy 100 GbE switches and NICs that support RoCE v2, so that when the Phase 2 training server and NVMe-oF storage arrive, RDMA is a configuration change rather than a hardware swap.

InfiniBand enters at Phase 3, when a multi-node training cluster needs efficient all-reduce operations across two or more training servers. At that point, NDR at 400 Gbps is the current mainstream specification, having displaced 200 Gbps HDR for new builds, and NVIDIA’s Quantum-2 InfiniBand switches with SHARP cut all-reduce latency by running collective operations in the switch fabric [34]. The cost premium over equivalent Ethernet is real but is justified only at multi-node training scale.

Ethernet has a purpose-built answer at the same scale. The NVIDIA Spectrum-X platform pairs Spectrum-4 switches (51.2 Tb/s of switching capacity) with a SuperNIC at the endpoint, either BlueField-3 or ConnectX-8, and adds per-packet adaptive routing and telemetry-driven congestion control tuned for AI traffic patterns; NVIDIA reports 1.6× the AI network performance of off-the-shelf Ethernet fabrics [35]. Spectrum-X is the Ethernet counterpart to a Quantum-2 fabric for a Phase 3 training buildout, attractive where the operations team wants to stay on Ethernet tooling, and it carries a premium over commodity 100 GbE that only multi-node training justifies.

The strongest published evidence that Ethernet suffices at large scale comes from Meta, which built two 24,576-GPU clusters for Llama 3 training (one on RoCE over Arista switches, one on Quantum-2 InfiniBand) and ran large GenAI training on both, including Llama 3 itself, without hitting network bottlenecks [36], [37]. For this paper’s Phase 1–2 scope, the procurement decision is 100 GbE from Arista, Cisco, or NVIDIA/Mellanox, with the InfiniBand call deferred to Phase 3.

The endpoint hardware and cabling deserve the same attention as the switch. NVIDIA’s ConnectX line is the default NIC family for this stack: ConnectX-7 carries 400 Gb/s, and ConnectX-8 reaches 800 Gb/s and integrates a PCIe Gen6 switch, which lets one device supply both GPU-to-NIC connectivity and the network port in eight-GPU RTX PRO server designs [38]. A Phase 1 build at 100 GbE is served by ConnectX-6 Dx or ConnectX-7 class adapters, or equivalent RoCE-v2-capable parts from other vendors. BlueField data processing units (DPUs) put Arm cores and the DOCA software stack on the ConnectX foundation and offload networking, storage, and security processing from the host [39]; NVIDIA’s enterprise reference designs assign them the north-south fabric (traffic entering and leaving the cluster), but, at one-to-few servers, a DPU is hardware to skip, since standard NICs and host-level firewalling cover the same functions. Cabling comes from the LinkX portfolio or its equivalents [40]: direct-attach copper (DAC) inside the rack, where it is the cheapest and lowest-power option, and active optical cables or transceivers with fiber for longer runs. NVIDIA’s Enterprise Reference Architecture documentation flags transceivers and cabling as a common and expensive place to go wrong at cluster scale [41]; the small-scale version of that lesson is to buy DAC wherever reach allows and vendor-validated optics where it does not.

Segmentation is standard practice on top of the chosen fabric: an AI VLAN carries the inference data plane and storage I/O, a separate in-band management VLAN isolates the Kubernetes control plane and monitoring, and a firewall-controlled egress opens only the ports the stack needs (vLLM on 8000, the vector database on 6333–6334, and the object store’s S3 and console ports). BMC traffic belongs on neither VLAN; it gets the physically separate network described next.

Out-of-Band Management

Every server in this taxonomy ships with a baseboard management controller, a small always-on computer with its own network port that exposes power control, remote console, sensor readings, firmware inventory, and virtual media independently of the host operating system. When a GPU node hangs mid-training or an operating-system update goes wrong, the BMC is the recovery path that does not require a drive to the rack. The legacy protocol is the Intelligent Platform Management Interface (IPMI); the current standard is Redfish, the DMTF’s REST- and JSON-based management API that succeeds it [42]. Redfish is the right automation target: scripted firmware audits, power capping, and health polling run against the same endpoints on Dell iDRAC, Supermicro BMCs, and every other current implementation.

Two rules govern the deployment. The OOB network is physically separate: a dedicated 1 GbE switch connecting only BMC ports and the management host, never routed to the AI VLAN and never exposed past the firewall, which is the same isolation NVIDIA’s reference architectures specify [32]. BMCs have a long vulnerability history, and an attacker who controls one controls the server beneath the operating system. The second rule is that BMC credentials rotate off factory defaults on day one. The hardware cost is one small switch and a handful of cables; the failure it prevents, hands-off recovery of a wedged node, is one every deployment eventually meets.

Cluster Management, Provisioning, and Monitoring

Two servers are already a cluster. The deployments that stay healthy are the ones that treat them as a cluster from day one. Hand-configured machines drift; a year in, no one can rebuild the inference server from scratch, and every change becomes an experiment on production. The cluster management role has three parts, provisioning, scheduling, and hardware monitoring, and all three run on a modest management host: a small 1U server or a virtual machine on the agentic compute server.

Provisioning at Phase 1 is deliberately plain: operating-system images plus Ansible (or an equivalent configuration tool) with every playbook in Git, on top of the k3s and GPU Operator deployment already specified for the inference server. NVIDIA Base Command Manager 11 is the step up worth evaluating before Phase 2. It automates bare-metal provisioning, software-image management, firmware updates, and cluster health monitoring from a handful of nodes to very large fleets, and it deploys Slurm or Kubernetes through built-in wizards; a free license option covers small deployments, and entitlement is included with NVIDIA AI Enterprise [43]. It is the same tool underneath NVIDIA’s DGX BasePOD and SuperPOD designs, so adopting it early preserves a growth path. A team fluent in Ansible can defer it; a team without that fluency gets a supported product in place of a tooling project.

Scheduling splits by workload shape. Kubernetes, in the k3s form already deployed, is the right scheduler for serving: long-running services, health checks, rolling updates. Slurm is the batch scheduler of the HPC world, built around job queues, priorities, fair-share accounting, and topology-aware placement with NVIDIA reporting it handles job scheduling for more than 65% of TOP500 systems [44], [45]. The Phase 2 training server is where Slurm earns a place in this taxonomy: once several users compete for eight training GPUs, a queue with fair-share accounting replaces calendar-based coordination, and idle GPU time stops being the default state between hand-scheduled runs. Teams that prefer a single control plane have Kubernetes-native routes to the same outcome: SchedMD’s Slinky project runs Slurm’s daemons as Kubernetes pods, and NVIDIA Run:ai adds GPU-aware quota scheduling on Kubernetes [45], [43]. The recommendation is Kubernetes alone through Phase 1, with Slurm or a Kubernetes-native queueing layer added alongside the training server.

The login-node discipline of HPC applies here scaled down. In a conventional cluster, users never log into compute nodes; they land on a login node and submit work through the scheduler, and the compute nodes carry nothing but scheduled jobs. The management host is this deployment’s login node: the bastion SSH entry point, the scheduler control plane, Base Command Manager where used, and the monitoring stack below. GPU nodes accept no interactive logins. The rule costs nothing and prevents the recurring failure it exists for, stray user processes holding GPU memory that the production serving stack believed was free.

Hardware monitoring is the third function, which is distinct from the LLM observability layer later in this section. Hardware monitoring watches silicon, whereas LLM observability watches output quality. Neither substitutes for the other. Prometheus, a systems monitoring and alerting toolkit, collects the metrics and fires the alerts. Grafana, an analytics and data visualization application, renders the dashboards. Both are open source and the de facto standard pairing [46], [47]. GPU telemetry comes from the NVIDIA Data Center GPU Manager (DCGM) through dcgm-exporter, which the GPU Operator deploys by default, exposing per-GPU utilization, memory occupancy, temperature, power draw, ECC error counts, and XID error events (the driver’s GPU fault codes) as Prometheus metrics [48], [49], [5]. The alerts worth writing first are thermal throttling, ECC error growth, and any XID event, because all three precede hard GPU failures, and the utilization dashboards double as the capacity-planning evidence for the Phase 2 purchase.

One further resource is NVIDIA’s published reference architectures for exactly this class of hardware. The Enterprise Reference Architectures cover clusters of NVIDIA-Certified servers from 4 to 32 nodes, built from four-node scalable units (SUs) and named by a CPU-GPU-NIC-bandwidth pattern [41]. The RTX PRO AI Factory design, pattern 2-8-5-200 (two CPUs, eight RTX PRO 6000 Blackwell Server Edition GPUs, five NICs, 200 Gb/s of east-west bandwidth per GPU), is built on the same inference silicon this taxonomy specifies and documents the switch models, rail-optimized fabric topology, cable counts, and out-of-band design for scaling it out [50], [32]. DGX BasePOD does the same for DGX-based clusters from two nodes to dozens, under Base Command Manager with Slurm or Kubernetes [51]. A Phase 1–2 deployment is smaller than any of these designs and does not need their bill of materials; their value here is as tested checklists for the decisions this section walks through, fabric separation, NIC ratios, cabling, and management design, before those decisions get made by accident.

MLOps Platform

Fine-tuning at Phase 2 needs software that makes training repeatable, versioned, and safely promotable. This is not a hardware addition; the platform runs on the agentic compute server or a small dedicated virtual machine. It is the named software dependency planned alongside the training-server purchase, not bolted on after the first fine-tuned model is already in unmanaged use.

The recommendation is MLflow 3, released in June 2025, with more than 30 million monthly downloads at launch and the de facto standard position in open-source MLOps [52], [53]. MLflow 3 introduced a model-centric architecture with LoggedModel as a first-class entity, which lets you compare model versions across experiments and carries the lineage the Phase 2 promotion workflow depends on [52]. Three capabilities ride on one platform. Experiment tracking logs hyperparameters, dataset versions, training metrics, and evaluation scores for every run, which is the reproducibility control Sculley et al. identified as the core debt failure mode. The model registry stores artifacts with versioning and lineage and exposes aliases such as champion and challenger that CI/CD pipelines target for deployment [54]. The CI/CD integration runs through GitHub Actions, GitLab CI, or Argo Workflows, with environment-separated registries (dev, staging, production) and code promotion across environments rather than manual stage flips on one registry.

The self-hosted Phase 1 configuration is a single command with the S3 URI pointed at the on-premises object store. No external SaaS, no commercial license. Phase 2 swaps the SQLite backend for PostgreSQL when concurrent runs warrant it. Weights & Biases is the cloud-hosted commercial alternative with richer visualization, but it requires either cloud dependency at the free tier or a paid self-hosted license, neither aligned with an on-premises-first architecture. Without this layer the training server produces orphan files of unknown provenance, no audit trail, and no rollback path. The hardware exists to make custom models; MLflow is what keeps them governable.

Data Pipeline Infrastructure

Retrieval pipelines and fine-tuning datasets both need automated, repeatable ingestion, not one-time manual loads. A Phase 1 internal-knowledge chatbot whose retrieval base is hand-populated goes stale within a month and is abandoned within a quarter. Like MLflow, this is a software layer on the agentic compute server, with no additional silicon.

The stack is Apache Airflow for orchestration and scheduling, LlamaIndex or Unstructured.io for document loading and parsing, chunking logic matched to the embedding model’s context window, Sentence Transformers or vLLM’s embedding API for batch embedding, and Qdrant or one of the alternatives for vector storage. Airflow is the center of gravity: Python-native DAGs integrate with every ML framework, parallelize document processing, and bring scheduling, retry logic, logging, and operator integrations for the major vector stores, which makes production ingestion a configured pipeline rather than a coding project [55], [56]. Astronomer draws the operational distinction directly: LangChain and LlamaIndex give you prototype pipelines but no scheduling for freshness, no auditing, no lineage, and no inter-workflow coordination, which is exactly what Airflow adds [57].

One Phase 1 configuration detail matters: Airflow standalone on SQLite is not recommended for production, because SQLite cannot handle concurrent writes. The Phase 1 production configuration is the LocalExecutor on a PostgreSQL backend; Phase 2 adds the CeleryExecutor for distributed task execution when DAG parallelism justifies it. Prefect is the lower-overhead alternative for small teams, one container instead of Airflow’s separate web server, scheduler, and database, and the better choice when Airflow’s footprint outweighs the pipeline volume. The re-indexing schedule lives at the DAG level: a nightly default for most internal sources, supplemented by event-triggered DAGs where a source can fire a webhook on document save. An operational ingestion pipeline is itself a Phase 1 capability milestone, because the chatbot use case that justifies Phase 1 cannot function without it.

LLM Observability Infrastructure

Infrastructure monitoring of the kind the NVIDIA Data Center GPU Manager (DCGM) and Prometheus stack provides tracks GPU utilization, queue depth, and latency, and none of it detects silent quality degradation. A GPU can be healthy and latency well inside its target while the model produces systematically worse output because the input distribution shifted, retrieval quality fell off, or the prompt template drifted from what was evaluated. Catching that requires a layer that captures prompt-response pairs, scores them against quality metrics, and watches for output drift. This is also the operational foundation of the audit trail the governance requirements commit to; without it, the claim of full audit control over inference is aspirational.

The Phase 1 implementation is OpenTelemetry trace logging, which costs nothing and belongs in place from day one. The opentelemetry-sdk and opentelemetry-exporter-otlp packages instrument vLLM-served requests with minimal overhead, turning each request into a span carrying model_id, prompt_tokens, completion_tokens, latency_ms, user_id where API-key routing exists, and trace_id for correlation across an agent workflow. Spans export to a local OpenTelemetry Collector and persist to Jaeger for visualization or to Elasticsearch/OpenSearch for retention [58]. The Phase 1 minimum is that every request and response is logged with timestamp, user ID, model version, and latency.

Phase 2 adds structured quality evaluation and drift detection. The recommendation is Langfuse: MIT-licensed since June 2025, fully self-hostable via Docker Compose or Kubernetes, with nested trace visualization, session tracking, and quality evaluation, all without cloud dependency [59], [60]. One honest footprint note: Langfuse v3 runs on PostgreSQL plus ClickHouse, Redis/Valkey, and S3-compatible blob storage, so it adds two components (ClickHouse and Redis) beyond the Phase 1 stack and is a heavier install than a single container. Arize Phoenix is the open-source alternative with stronger statistical drift detection, better suited to teams with a data-science function. LangSmith offers a self-hosted deployment, but it sits behind the Enterprise plan and its free tier is cloud-only, which rules it out of an on-premises-first architecture [61].

A scoping caveat applies to multi-agent tracing. OpenTelemetry’s GenAI semantic conventions, including the agent and framework spans, remain in development (experimental) status as of mid-2026, with the transition to a stable version not yet finalized [62]. For Phase 1 single-agent and chat logging, OTel’s standard trace and span model is stable and sufficient. For Phase 2 multi-agent workflows where correlating LLM calls across an agent chain matters, Langfuse’s native session tracking is more mature than bare OTel today, which is another reason to plan the Phase 2 move to Langfuse rather than expect OTel alone to cover agentic observability.

Facility Space, Power, and Cooling

Every hardware line in this taxonomy has to clear a facility check before its purchase order is signed. Phase 1 is modest. An inference server carrying two 600 W RTX PRO 6000 Blackwell Server Edition cards, two CPUs, drives, and fans draws on the order of 2.5–3 kW at full load; with the storage server and networking added, the deployment fits one standard 42U rack fed by dedicated 208 V, 30 A circuits, each of which supplies about 5 kW of continuous capacity under the 80% derating rule. Heat is the same number restated: every kilowatt of electrical load becomes a kilowatt of heat, 3,412 BTU per hour, so a 5 kW rack needs on the order of 17,060 BTU/hr of dedicated cooling, which translates to cooling equipment weighing roughly a ton and a half. That belongs in a server room with its own cooling, not on the building’s comfort system. The fan noise of a loaded GPU server settles the location question anyway. The Server Edition card is passively cooled and depends entirely on chassis airflow, which is one more reason the qualified Dell or Supermicro chassis is a requirement rather than a preference. Put the storage, networking, and management components on an uninterruptible power supply (UPS) sized for graceful shutdown; battery runtime for the GPU load itself is rarely worth its cost, and orderly-shutdown automation protects it more cheaply.

Phase 2 changes the electrical class. An 8-GPU SXM system draws up to roughly 10.2 kW for a DGX H200 and roughly 14.3 kW for a DGX B200 at maximum [63], [64], figures beyond what most small-organization server rooms deliver to a single rack. The requirements that follow are a dedicated three-phase feed sized with an electrician before the order, one 8-GPU system per rack, and a cooling assessment against the system’s maximum draw rather than the room’s nameplate. Air cooling remains workable at that one-system-per-rack density with disciplined hot- and cold-aisle airflow. For calibration on where this paper’s scale deliberately stops is NVIDIA’s DGX SuperPOD reference design for DGX B300, which places four systems per rack at roughly 56 kW [65]; a density that assumes purpose-built data-center power and cooling and that the phased roadmap avoids. The facility work for Phase 2 is real but bounded, electrical service and airflow rather than a data-center construction project, and pricing it during Phase 1 keeps it off the critical path when the training server is approved.

The Master Taxonomy Table

The CapEx figures below are directional, for sizing only; every line needs a current vendor quotation before procurement, and GPU pricing stayed volatile enough through 2025–2026 that a quotation more than a few months old should be refreshed.

RolePhaseRecommended HardwareApprox. CapExPrimary SoftwareDepends On
Inference serverPhase 1 (expand Phase 2)Dell XE7745 or Supermicro SYS-422GL-NR, 1–2× RTX PRO 6000 Blackwell Server Edition$60K–$150K (Phase 1)vLLM, k3s or microk8s, NVIDIA GPU OperatorNetworking, Storage
Embeddings serverPhase 1 shared / Phase 2 dedicatedPhase 1: shares inference GPU via vLLM endpoint; Phase 2: NVIDIA L40S (48 GB) PCIe in 4U server$10K–$30K GPU card (Phase 2)Sentence Transformers, vLLM embedding endpointStorage, Inference server
Agentic compute serverPhase 2AMD EPYC 9965 (192c) or 9755 (128c); 512 GB–1 TB DDR5-6000 ECC; 4–8× NVMe; 25 GbE$30K–$80KPython orchestration, Airflow DAGs, tool runtimesInference server, Storage
Training/fine-tuning serverPhase 2Dell or Supermicro, 8× H200/B200/B300 SXM5 with NVSwitch$300K–$1M+PyTorch FSDP or DeepSpeed, MLflowStorage, MLOps platform
Data storage serverPhase 1NVMe primary pool, NAS cold tier, separate backup tier$20K–$100KSeaweedFS or Garage (OSS) / AIStor (commercial); Ceph RGW; NFS/SMB; Qdrant or pgvectorAll compute servers
NetworkingPhase 1 / Phase 3100 GbE (Arista, Cisco, NVIDIA/Mellanox) for Phase 1–2; InfiniBand NDR 400 Gbps for Phase 3$5K–$50K—All servers
Cluster management & monitoringPhase 1 software / Phase 2 BCM optionManagement host: small 1U server or VM on agentic compute server$0–$10K (host, if dedicated)Ansible + Git, k3s (Phase 1); Base Command Manager 11, Slurm or Run:ai (Phase 2); Prometheus, Grafana, DCGM/dcgm-exporterAll servers, OOB network
Out-of-band managementPhase 1Server BMCs, dedicated 1 GbE OOB switch$1K–$3KRedfish API, vendor BMC firmware (iDRAC, Supermicro BMC)None (independent by design)
MLOps platformPhase 2 softwareRuns on agentic compute server$0 (MLflow OSS)MLflow 3, Argo Workflows or GitHub ActionsAgentic compute, Storage
Data pipelinePhase 1 softwareRuns on agentic compute server$0 (Airflow OSS)Apache Airflow on PostgreSQL, LlamaIndex or Unstructured.ioStorage, Embeddings
LLM observabilityPhase 1 softwareRuns on agentic compute server$0 (OTel OSS); Phase 2 Langfuse self-hostedOpenTelemetry Collector (Phase 1); Langfuse or Arize Phoenix (Phase 2, adds ClickHouse + Redis)Inference server
Facility power & coolingPhase 1 / Phase 2208 V, 30 A circuits + dedicated room cooling (Phase 1); three-phase feed, ~14 kW/system cooling (Phase 2)Site-dependent—All hardware

The table makes one pattern plain: four of the ten roles (the MLOps platform, data pipeline, observability, and cluster management) are software layers on the agentic compute server or a small management host. They cost almost nothing in hardware and everything in operational discipline. Skipping them saves no money; it deletes capabilities the rest of the deployment depends on, which is why the agentic compute server is mis-scoped as often as it is omitted. Size that host for the orchestration, pipeline, observability, and management load it actually carries, not for the agent code alone; run the facility numbers before the first purchase order; and the taxonomy holds together as a production system, a small cluster managed as one, rather than a collection of servers.

References

  1. D. Trifonov, “RTX 4090 vs 5090 vs PRO 6000: LLM Inference Benchmark,” CloudRift AI, Oct. 9, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx-gpus-for-llm-inference. [Accessed: 04-Jun-2026]

    TAXO-1 Secondary source Back to text

  2. NVIDIA Corporation, “NVIDIA RTX PRO 6000 Blackwell Server Edition,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/. [Accessed: 15-Jul-2026]

    TAXO-2 Primary source Back to text

  3. vLLM Project, “vLLM Documentation,” PyTorch Foundation, 2026. [Online]. Available: https://docs.vllm.ai/en/latest/. [Accessed: 08-Jun-2026]

    TAXO-3 Primary source Back to text

  4. W. Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” Proc. ACM SIGOPS 29th Symp. on Operating Systems Principles (SOSP '23), 2023. [Online]. Available: https://dl.acm.org/doi/10.1145/3600006.3613165. Koblenz, Germany. [Accessed: 08-Jun-2026]

    TAXO-4 Primary source Back to text

  5. NVIDIA Corporation, “NVIDIA GPU Operator Documentation,” NVIDIA Cloud Native Technologies, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/. [Accessed: 08-Jun-2026]

    TAXO-5 Primary source Back to text

  6. NVIDIA Corporation, “NVIDIA Multi-Instance GPU User Guide,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/tesla/mig-user-guide/. [Accessed: 15-Jul-2026]

    TAXO-6 Primary source Back to text

  7. NVIDIA Corporation, “NVIDIA Virtual GPU (vGPU) Software Documentation,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/vgpu/. [Accessed: 15-Jul-2026]

    TAXO-7 Primary source Back to text

  8. NVIDIA Corporation, “NVIDIA NVLink and NVSwitch,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/nvlink/. [Accessed: 15-Jul-2026]

    TAXO-8 Primary source Back to text

  9. NVIDIA Corporation, “NVIDIA GPUDirect Storage Overview Guide,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/gpudirect-storage/overview-guide/. [Accessed: 15-Jul-2026]

    TAXO-9 Primary source Back to text

  10. D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” Advances in Neural Information Processing Systems, 2015. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html. Vol. 28, pp. 2503–2511. [Accessed: 08-Jun-2026]

    TAXO-10 Primary source Back to text

  11. Sentence Transformers Project, “Sentence Transformers Documentation,” 2026. [Online]. Available: https://sbert.net/. [Accessed: 08-Jun-2026]

    TAXO-11 Primary source Back to text

  12. N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks,” Proc. 2019 Conf. on Empirical Methods in Natural Language Processing, 2019. [Online]. Available: https://arxiv.org/abs/1908.10084. Hong Kong, China, pp. 3982–3992. [Accessed: 08-Jun-2026]

    TAXO-12 Primary source Back to text

  13. NVIDIA Corporation, “NVIDIA L40S GPU Datasheet,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/l40s/. [Accessed: 08-Jun-2026]

    TAXO-13 Primary source Back to text

  14. NVIDIA Corporation, “NVIDIA A10 Tensor Core GPU Datasheet,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/products/a10-gpu/. [Accessed: 08-Jun-2026]

    TAXO-14 Primary source Back to text

  15. Advanced Micro Devices, Inc., “AMD EPYC 9005 Series Processors,” 2026. [Online]. Available: https://www.amd.com/en/products/processors/server/epyc/9005-series.html. [Accessed: 08-Jun-2026]

    TAXO-15 Primary source Back to text

  16. Advanced Micro Devices, Inc., “AMD EPYC 9965 Processor,” 2026. [Online]. Available: https://www.amd.com/en/products/processors/server/epyc/9005-series/amd-epyc-9965.html. [Accessed: 08-Jun-2026]

    TAXO-16 Primary source Back to text

  17. Intel Corporation, “Intel Xeon 6980P Processor (504M Cache, 2.00 GHz) — Product Specifications,” 2026. [Online]. Available: https://www.intel.com/content/www/us/en/products/sku/240777/intel-xeon-6980p-processor-504m-cache-2-00-ghz/specifications.html. [Accessed: 08-Jun-2026]

    TAXO-17 Primary source Back to text

  18. M. Larabel, “Intel Xeon 6980P vs. AMD EPYC 9755 128-Core Showdown with the Latest Linux Software for EOY2025,” Phoronix, Dec. 17, 2025. [Online]. Available: https://www.phoronix.com/review/xeon-6980p-epyc-9755-2025. [Accessed: 08-Jun-2026]

    TAXO-18 Secondary source Back to text

  19. MinIO, Inc., “minio/minio Object Storage,” GitHub, 2026. [Online]. Available: https://github.com/minio/minio. Repository archived read-only Apr. 25, 2026; AGPLv3. [Accessed: 08-Jun-2026]

    TAXO-19 Primary source Back to text

  20. MinIO, Inc., “minio/minio — Releases,” GitHub, Apr. 25, 2026. [Online]. Available: https://github.com/minio/minio/releases. Oct. 2025 privilege-escalation CVE fix; the final community release predates the archive. [Accessed: 08-Jun-2026]

    TAXO-20 Primary source Back to text

  21. SeaweedFS Project, “seaweedfs/seaweedfs — Distributed object storage (S3), file system, and Iceberg tables with O(1) disk access,” GitHub, 2026. [Online]. Available: https://github.com/seaweedfs/seaweedfs. [Accessed: 08-Jun-2026]

    TAXO-21 Primary source Back to text

  22. Deuxfleurs, “Garage: An S3-compatible object store for small, self-hosted, geo-distributed deployments,” 2026. [Online]. Available: https://garagehq.deuxfleurs.fr. [Accessed: 08-Jun-2026]

    TAXO-22 Primary source Back to text

  23. pgvector Project, “pgvector: Open-Source Vector Similarity Search for PostgreSQL,” GitHub, 2026. [Online]. Available: https://github.com/pgvector/pgvector. [Accessed: 08-Jun-2026]

    TAXO-23 Primary source Back to text

  24. Timescale (Tiger Data), “timescale/pgvectorscale — Postgres extension for vector search (StreamingDiskANN), complements pgvector for performance and scale,” GitHub, 2026. [Online]. Available: https://github.com/timescale/pgvectorscale. [Accessed: 08-Jun-2026]

    TAXO-24 Primary source Back to text

  25. Chroma, “chroma-core/chroma — Open-source search and retrieval infrastructure for AI,” GitHub, 2026. [Online]. Available: https://github.com/chroma-core/chroma. [Accessed: 08-Jun-2026]

    TAXO-25 Primary source Back to text

  26. Qdrant, “Qdrant Documentation,” 2026. [Online]. Available: https://qdrant.tech/documentation/. [Accessed: 08-Jun-2026]

    TAXO-26 Primary source Back to text

  27. Qdrant, “Qdrant 2025 Recap: Powering the Agentic Era,” Qdrant Blog, Dec. 17, 2025. [Online]. Available: https://qdrant.tech/blog/2025-recap/. [Accessed: 08-Jun-2026]

    TAXO-27 Secondary source Back to text

  28. Qdrant, “Qdrant 1.13 — GPU Indexing, Strict Mode & New Storage Engine,” Qdrant Blog, Jan. 23, 2025. [Online]. Available: https://qdrant.tech/blog/qdrant-1.13.x/. [Accessed: 08-Jun-2026]

    TAXO-28 Secondary source Back to text

  29. Qdrant, “Qdrant 1.16 — Tiered Multitenancy & Disk-Efficient Vector Search,” Qdrant Blog, Nov. 19, 2025. [Online]. Available: https://qdrant.tech/blog/qdrant-1.16.x/. [Accessed: 08-Jun-2026]

    TAXO-29 Secondary source Back to text

  30. Weaviate, “weaviate/weaviate — An open-source vector database that stores both objects and vectors,” GitHub, 2026. [Online]. Available: https://github.com/weaviate/weaviate. [Accessed: 08-Jun-2026]

    TAXO-30 Primary source Back to text

  31. Milvus / Zilliz, “milvus-io/milvus — A high-performance, cloud-native vector database built for scalable vector ANN search,” GitHub, 2026. [Online]. Available: https://github.com/milvus-io/milvus. [Accessed: 08-Jun-2026]

    TAXO-31 Primary source Back to text

  32. NVIDIA Corporation, “NVIDIA RTX PRO AI Factory Enterprise Reference Architecture — Networking Logical Architecture,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/enterprise-reference-architectures/rtx-pro-ai-factory/latest/network-logical-architecture.html. [Accessed: 15-Jul-2026]

    TAXO-32 Primary source Back to text

  33. NVIDIA Corporation, “GPUDirect RDMA Documentation,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/cuda/gpudirect-rdma/. [Accessed: 15-Jul-2026]

    TAXO-33 Primary source Back to text

  34. NVIDIA Corporation, “NVIDIA InfiniBand Networking,” 2026. [Online]. Available: https://www.nvidia.com/en-us/networking/infiniband/. [Accessed: 08-Jun-2026]

    TAXO-34 Primary source Back to text

  35. NVIDIA Corporation, “NVIDIA Spectrum-X Ethernet Networking Platform for AI,” 2026. [Online]. Available: https://www.nvidia.com/en-us/networking/spectrumx/. [Accessed: 15-Jul-2026]

    TAXO-35 Primary source Back to text

  36. Meta Platforms, Inc., “Building Meta's GenAI Infrastructure,” Engineering at Meta, Mar. 12, 2024. [Online]. Available: https://engineering.fb.com/2024/03/12/data-center-engineering/building-metas-genai-infrastructure/. [Accessed: 08-Jun-2026]

    TAXO-36 Secondary source Back to text

  37. A. Grattafiori et al., “The Llama 3 Herd of Models,” arXiv, Llama Team, AI @ Meta, 2024. [Online]. Available: https://arxiv.org/abs/2407.21783. arXiv:2407.21783. [Accessed: 08-Jun-2026]

    TAXO-37 Primary source Back to text

  38. E. Tweg and N. Dey, “NVIDIA ConnectX-8 SuperNICs Advance AI Platform Architecture with PCIe Gen6 Connectivity,” NVIDIA Technical Blog, May 29, 2025. [Online]. Available: https://developer.nvidia.com/blog/nvidia-connectx-8-supernics-advance-ai-platform-architecture-with-pcie-gen6-connectivity/. [Accessed: 15-Jul-2026]

    TAXO-38 Secondary source Back to text

  39. NVIDIA Corporation, “NVIDIA BlueField Data Processing Units,” 2026. [Online]. Available: https://www.nvidia.com/en-us/networking/products/data-processing-unit/. [Accessed: 15-Jul-2026]

    TAXO-39 Primary source Back to text

  40. NVIDIA Corporation, “NVIDIA LinkX Cables and Transceivers,” 2026. [Online]. Available: https://www.nvidia.com/en-us/networking/interconnect/. [Accessed: 15-Jul-2026]

    TAXO-40 Primary source Back to text

  41. NVIDIA Corporation, “Introducing the NVIDIA Enterprise Reference Architectures,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/enterprise-reference-architectures/white-paper/latest/introduction.html. [Accessed: 15-Jul-2026]

    TAXO-41 Primary source Back to text

  42. DMTF, “Redfish Standard,” 2026. [Online]. Available: https://www.dmtf.org/standards/redfish. [Accessed: 15-Jul-2026]

    TAXO-42 Primary source Back to text

  43. NVIDIA Corporation, “NVIDIA Base Command Manager,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/base-command-manager/. [Accessed: 15-Jul-2026]

    TAXO-43 Primary source Back to text

  44. SchedMD LLC, “Slurm Workload Manager Documentation,” 2026. [Online]. Available: https://slurm.schedmd.com/documentation.html. [Accessed: 15-Jul-2026]

    TAXO-44 Primary source Back to text

  45. NVIDIA Corporation, “Running Large-Scale GPU Workloads on Kubernetes with Slurm,” NVIDIA Technical Blog, May 7, 2026. [Online]. Available: https://developer.nvidia.com/blog/running-large-scale-gpu-workloads-on-kubernetes-with-slurm/. [Accessed: 15-Jul-2026]

    TAXO-45 Secondary source Back to text

  46. Prometheus Authors, “Prometheus Documentation,” 2026. [Online]. Available: https://prometheus.io/docs/introduction/overview/. [Accessed: 15-Jul-2026]

    TAXO-46 Primary source Back to text

  47. Grafana Labs, “Grafana Open Source Documentation,” 2026. [Online]. Available: https://grafana.com/docs/grafana/latest/. [Accessed: 15-Jul-2026]

    TAXO-47 Primary source Back to text

  48. NVIDIA Corporation, “NVIDIA Data Center GPU Manager (DCGM),” NVIDIA Developer, 2026. [Online]. Available: https://developer.nvidia.com/dcgm. [Accessed: 15-Jul-2026]

    TAXO-48 Primary source Back to text

  49. NVIDIA Corporation, “NVIDIA/dcgm-exporter — NVIDIA GPU metrics exporter for Prometheus leveraging DCGM,” GitHub, 2026. [Online]. Available: https://github.com/NVIDIA/dcgm-exporter. [Accessed: 15-Jul-2026]

    TAXO-49 Primary source Back to text

  50. NVIDIA Corporation, “NVIDIA RTX PRO AI Factory Enterprise Reference Architecture — Overview,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/enterprise-reference-architectures/rtx-pro-ai-factory/latest/overview.html. [Accessed: 15-Jul-2026]

    TAXO-50 Primary source Back to text

  51. NVIDIA Corporation, “NVIDIA DGX BasePOD: The Infrastructure Foundation for Enterprise AI — Reference Architecture Featuring NVIDIA DGX B200, H200 and H100 Systems,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/dgx-basepod/reference-architecture-infrastructure-foundation-enterprise-ai/latest/index.html. [Accessed: 15-Jul-2026]

    TAXO-51 Primary source Back to text

  52. MLflow Project, “MLflow Documentation,” 2026. [Online]. Available: https://mlflow.org/docs/latest/. MLflow 3. [Accessed: 08-Jun-2026]

    TAXO-52 Primary source Back to text

  53. Databricks, “MLflow 3.0: Unified AI Experimentation, Observability, and Governance,” Databricks Blog, June 11, 2025. [Online]. Available: https://www.databricks.com/blog/mlflow-30-unified-ai-experimentation-observability-and-governance. [Accessed: 08-Jun-2026]

    TAXO-53 Secondary source Back to text

  54. MLflow Project, “MLflow Model Registry,” 2026. [Online]. Available: https://mlflow.org/docs/latest/ml/model-registry/. [Accessed: 08-Jun-2026]

    TAXO-54 Primary source Back to text

  55. Apache Software Foundation, “Apache Airflow Documentation,” 2026. [Online]. Available: https://airflow.apache.org/docs/. [Accessed: 08-Jun-2026]

    TAXO-55 Primary source Back to text

  56. Apache Software Foundation, “MLOps with Apache Airflow,” 2026. [Online]. Available: https://airflow.apache.org/use-cases/mlops/. [Accessed: 08-Jun-2026]

    TAXO-56 Primary source Back to text

  57. Astronomer, “Best Practices for Orchestrating MLOps Pipelines with Airflow,” Astronomer Documentation, 2026. [Online]. Available: https://www.astronomer.io/docs/learn/airflow-mlops. [Accessed: 08-Jun-2026]

    TAXO-57 Secondary source Back to text

  58. OpenTelemetry Project, “OpenTelemetry Documentation,” Cloud Native Computing Foundation, 2026. [Online]. Available: https://opentelemetry.io/docs/. [Accessed: 08-Jun-2026]

    TAXO-58 Primary source Back to text

  59. Langfuse, “Self-Hosting Langfuse (Open Source LLM Observability),” Langfuse Documentation, 2026. [Online]. Available: https://langfuse.com/self-hosting. [Accessed: 08-Jun-2026]

    TAXO-59 Primary source Back to text

  60. Langfuse, “Doubling Down on Open Source: Open-Sourcing All Product Features Under the MIT License,” Langfuse Blog, June 4, 2025. [Online]. Available: https://langfuse.com/blog/2025-06-04-open-sourcing-langfuse-product. [Accessed: 08-Jun-2026]

    TAXO-60 Secondary source Back to text

  61. LangChain, Inc., “Introducing End-to-End OpenTelemetry Support in LangSmith,” LangChain Blog, Mar. 13, 2026. [Online]. Available: https://blog.langchain.com/end-to-end-opentelemetry-langsmith/. [Accessed: 08-Jun-2026]

    TAXO-61 Secondary source Back to text

  62. OpenTelemetry Project, “Semantic Conventions for Generative AI Systems,” Cloud Native Computing Foundation, 2026. [Online]. Available: https://opentelemetry.io/docs/specs/semconv/gen-ai/. Status: Development. [Accessed: 08-Jun-2026]

    TAXO-62 Primary source Back to text

  63. NVIDIA Corporation, “NVIDIA DGX H200 Datasheet,” 2026. [Online]. Available: https://resources.nvidia.com/en-us-dgx-systems/dgx-h200-datasheet. [Accessed: 15-Jul-2026]

    TAXO-63 Primary source Back to text

  64. NVIDIA Corporation, “NVIDIA DGX B200 Datasheet,” 2026. [Online]. Available: https://resources.nvidia.com/en-us-dgx-systems/dgx-b200-datasheet. [Accessed: 15-Jul-2026]

    TAXO-64 Primary source Back to text

  65. NVIDIA Corporation, “NVIDIA DGX SuperPOD with DGX B300 Systems, NVIDIA Quantum-X800 InfiniBand Switching and AC Power Reference Architecture,” NVIDIA Documentation, 2026. [Online]. Available: https://docs.nvidia.com/dgx-superpod/reference-architecture/scalable-infrastructure-b300-xdr/latest/dgx-superpod-architecture.html. [Accessed: 15-Jul-2026]

    TAXO-65 Primary source Back to text

Contents