Section12

NVIDIA vs. AMD: Structured Comparison and Hardware Recommendation

NVIDIA is the correct vendor for the Phase 1 build; the case holds against the strongest version of the AMD counter-argument. AMD’s MI355X platform delivered 100,282 tokens per second on Llama 2 70B Server at MLPerf Inference 6.0, a result verified by MLCommons and corroborated by independent coverage of the round [1], [2], [3]. Each MI355X carries 288 GB of HBM3E against the H200’s 141 GB [1], [4], and an 8-GPU MI355X node holds Llama 3.1 405B or Llama 4 Maverick at BF16 where an 8× H100 SXM5 node cannot. On hardware, AMD is competitive across most single-node inference scenarios and superior on memory capacity. AMD is not operationally equivalent for the small to medium organization this paper addresses, the range below 500 employees, where one-to-few engineers or a multi-person team operates the deployment and absorbs every failure it produces; that operational gap decides the recommendation. TensorRT-LLM runs only on NVIDIA GPUs and has no ROCm equivalent [5]. FlashAttention 3 has no ROCm port [6]. The community knowledge base an engineer leans on for novel debugging is materially thinner on ROCm. Every system in either vendor’s catalog clears the 300–450 tokens-per-second interactive floor for fifteen concurrent users many times over, so throughput cannot separate the vendors at Phase 1 scale. Operational risk can, and NVIDIA carries far less of it.

The Six Dimensions

Software stack maturity. The largest operational gap is TensorRT-LLM, NVIDIA’s compiled-engine inference runtime, which runs only on NVIDIA GPUs [5]. Argonne National Laboratory’s LLM-Inference-Bench evaluation found TensorRT-LLM both faster and more power-efficient than vLLM across 7B- and 70B-class models on NVIDIA hardware, though the study also recorded model-shape cases where vLLM led, so the advantage is real but workload-dependent rather than uniform [7]. On applicable workloads, TensorRT-LLM’s speculative decoding adds up to 3.6× faster generation [8]. AMD’s nearest answer is AITER, the AI Tensor Engine for ROCm, which supplies AMD-specific fused kernels for mixture-of-experts, attention, and normalization operations in vLLM behind the VLLM_ROCM_USE_AITER flag and ships in vLLM V1 as a first-party feature [9], [10]. AITER closes part of the gap on data center Instinct GPUs; it does not replicate a compiled-engine runtime. FlashAttention 3, an optimized algorithm that accelerates the transformer’s attention mechanism, is likewise unavailable on ROCm because it depends on NVIDIA Hopper- and Blackwell-specific tensor hardware; ROCm falls back to FlashAttention-2-class kernels through Composable Kernel and Triton backends [6]. NVIDIA released CUDA publicly in 2007 and AMD released ROCm in 2016; a nine-year head start. The operational head start runs longer because CUDA reached production-grade AI-framework support years before ROCm did. This dimension goes to NVIDIA.

Model compatibility out of the box. ROCm in 2026 supports every model class this paper targets, a materially better position than its reputation from earlier generations suggests. vLLM lists ROCm as a first-party target with ROCm-specific Python wheels that retire the prior Docker-only friction, named ROCm bug fixes in each release, and expert parallelism, data parallelism, and AITER-optimized multi-head latent attention kernels in V1 [10]. Ollama and llama.cpp run on AMD Instinct GPUs through the HIP/ROCm backend, and SGLang ships first-class ROCm support. The Phase 1 set of Gemma 4 31B, a 70B-class model at FP8, and gpt-oss-120b at MXFP4 has a documented AMD path, as do the larger Phase 2 classes including Llama 4 Maverick, for which AMD publishes a dedicated MI300X deployment guide [11]. MXFP4 serving on the MI355X, CDNA 4’s first FP4-capable generation, is production-demonstrated: AMD’s MLPerf 6.0 submission ran gpt-oss-120b, an MXFP4-native model, at 111% of B200 Offline and 115% of B200 Server single-node performance [1]; the residual step is validating the specific target model at procurement time. Packaged inference microservices now exist on both sides. NVIDIA NIM has shipped since 2024; AMD answered in November 2025 with AMD Inference Microservices (AIM) inside the open-source AMD Enterprise AI Suite, prebuilt vLLM-based containers with OpenAI-compatible APIs whose catalog reached the Llama, Qwen, DeepSeek, and Mistral families with MI350X and MI355X optimization in the suite’s 1.8 release of March 2026 [12]. AIM is a young analog with a narrower catalog and a shorter production history, not a missing capability. The remaining gap is timing: new architectures get tested, benchmarked, and documented on NVIDIA first, and optimization guides arrive on NVIDIA first. NVIDIA wins this dimension narrowly, and the margin shrinks each release cycle.

Community and tooling support. Staffing decides the weight of this dimension, which is where the band below 500 employees is thinnest: the operating assumption throughout this paper is one-to-few engineers, or a multi-person team, running the deployment alongside other duties. The mechanism is mundane. When a novel error surfaces during a short turnaround window, the engineer searches the error; on the CUDA side that search returns nearly two decades of accumulated answers, developer-blog posts, and resolved GitHub issues; the same search on ROCm returns far less. The independent MI300X analysis by Ambati and Diep, an arXiv preprint from October 2025, traces AMD’s residual inference deficit to software maturity, naming collective-communication libraries, compiler optimization for LLM kernels, and stack stability rather than the chip [13]. A community corpus is a function of production-years, and AMD’s documentation investment, real and visible in the per-model optimization guides shipped with ROCm 7.2.3, cannot manufacture that history. For an engineer with no prior ROCm operational experience, novel debugging on AMD takes longer, and that cost compounds across every incident.

Cost per token and memory economics. The answer is workload-conditional, so the blanket “NVIDIA wins” and “AMD wins” framings in vendor marketing both misread the evidence. For the single-GPU serving profile Phase 1 runs, CloudRift’s November 2025 test measured a single RTX PRO 6000 at 3,140 tokens per second against 2,987 for an H100 SXM on GLM-4.5-Air at 4-bit AWQ under vLLM, which works out to roughly 28% lower cost per token ($0.18 versus $0.25 per million tokens) at the study’s rental-rate assumptions and about one third of the H100’s roughly $30,000 hardware cost; the card tested was the Workstation Edition, which shares the GB202 silicon and 96 GB memory of the Server Edition this paper specifies [14]. Akamai’s NIM-profile test on Llama-3.3-Nemotron-Super-49B put the card at 1.63× the H100 NVL 96GB’s FP8 throughput when running FP4 at 100 concurrent requests [15]. Both results carry a boundary: CloudRift’s January 2026 long-context follow-up, at 8K-token input and 8K-token output, measured the B200 at up to 4.9× the RTX PRO 6000’s throughput and found the B200 the cost-per-token leader across every tested model, so the RTX PRO 6000’s cost advantage is a short-to-moderate-context, single-GPU result, the Phase 1 profile, and not a general law [16]. No directly comparable third-party cost-per-token benchmark exists for the MI355X at this model class; MLPerf 6.0 covers Llama 2 70B Server, not 30B-class AWQ workloads. The math reverses at the opposite end. Llama 3.1 405B at BF16 needs roughly 810 GB of GPU memory for weights; a single 8× MI300X node (1.5 TB HBM3) runs it without cross-node tensor parallelism while an 8× H100 SXM5 node (640 GB) cannot. dstack’s benchmark put the 8× MI300X node ahead of the 8× H100 node on both throughput and cost once prompt and batch sizes saturated the H100’s memory [17]. Llama 4 Maverick, at 400 billion total parameters with 17 billion active, follows the same logic at BF16 [18]. AMD wins on system cost for production-precision models that exceed 80 GB per GPU; NVIDIA wins on cost per token for everything that fits inside 80–96 GB.

Memory capacity and bandwidth. AMD’s per-GPU memory lead over the NVIDIA Hopper generation is real, and it becomes decision-relevant only for workloads whose weights and KV cache exceed the NVIDIA part’s capacity at target precision. The MI300X (192 GB HBM3) carries 2.4× the H100 SXM5’s 80 GB, the MI325X (256 GB HBM3E) carries 1.8× the H200’s 141 GB [4], and the MI350X and MI355X (288 GB HBM3E at 8.0 TB/s) match NVIDIA’s B300, which also ships 288 GB of HBM3e at 8.0 TB/s [19]. At the Blackwell Ultra generation, AMD’s flagship single-GPU VRAM advantage is neutralized. Phase 1–2 NVIDIA hardware (RTX PRO 6000 at 96 GB, H100 at 80 GB, H200 at 141 GB) sits below AMD’s current generation on memory; Phase 2–3 NVIDIA hardware (B200 at 180 GB, B300 at 288 GB) closes the gap. The Phase 3+ comparison resets at the Vera Rubin generation (288 GB HBM4, second half of 2026) [20], where AMD’s announced MI455X (432 GB HBM4) would reopen a 50% capacity lead that Phase 2–3 procurement should re-test against then-current NVIDIA specifications [21]. The dimension goes to AMD at the H100/H200 generation, neutral at B300, and back to AMD at the MI455X-versus-Vera-Rubin comparison if the announced MI455X specifications reach production.

OEM availability through Dell and Supermicro. The Dell relationship is asymmetric. Dell’s featured AI server line is primarily NVIDIA-based (XE9780, XE7745, XE9712), with AMD configurations available through the same catalog but not the featured marketing focus [22]. Dell currently offers the PowerEdge XE9680 with 8× MI300X and the PowerEdge XE9785 with 8× MI355X [23], [24]. Supermicro treats both vendors as first-class: its AMD GPU servers include the 10U air-cooled MI355X chassis, the 4U liquid-cooled MI355X system, the 8U air-cooled MI350X system, and the prior-generation AS-8125GS-TNMR2 with 8× MI300X, all through the same channel that serves NVIDIA HGX B200 and B300 systems [25], [26], [27]. Supply favors AMD at the margin for Phase 1 procurement windows. NVIDIA’s November 2025 results confirmed demand running ahead of supply, with CEO Jensen Huang stating that “Blackwell sales are off the charts, and cloud GPUs are sold out” and the company citing half a trillion dollars of Blackwell and Rubin revenue visibility through calendar 2026 [28], so MI355X configurations may carry shorter lead times. Treat that as directional and verify lead times in writing at quotation. The dimension goes to NVIDIA at Dell, neutral at Supermicro, with a slight AMD edge on Phase 1 lead times.

The Comparison Matrix

DimensionNVIDIA (H100 / H200 / B200 / B300)AMD (MI350X / MI355X)Phase 1 WinnerNotes
Software stack maturityTensorRT-LLM (faster and more power-efficient than vLLM per Argonne National Lab; up to 3.6× with speculative decoding); FlashAttention 3; nearly two decades of CUDA corpusROCm 7.2.3; AITER in vLLM V1 (partial mitigation); no TensorRT-LLM; no FlashAttention 3NVIDIALargest operational gap; no AMD compiled-engine runtime
Model compatibilityAll models enabled on NVIDIA first; MXFP4 confirmed on Hopper/Blackwell; NIM shipping since 2024vLLM ROCm first-party; all target classes confirmed; MXFP4 demonstrated at MLPerf 6.0; AIM microservices since Nov. 2025NVIDIA, narrowlyGap is timing, not capability; narrowing each cycle
Community and toolingDominant; nearly two decades of public corpus; NVIDIA Developer programROCm docs strong; smaller community corpus; novel debugging takes longerNVIDIAMost operationally material dimension for a small team
Cost per token (Phase 1 single-GPU)RTX PRO 6000: ~28% lower cost/token than H100 SXM; 1.63× H100 NVL FP8 throughput at FP4No directly comparable third-party single-GPU cost/token benchmark for MI355XNVIDIA (RTX PRO 6000)AMD cost advantage emerges at 405B+, not Phase 1 scale
Cost per token (≥405B BF16)Multi-node H100/H200 required at BF16; single-node only at quantized precisionSingle 8× MI300X (1.5 TB) / MI355X (2.3 TB) node runs 405B BF16 without cross-node parallelismAMDSystem cost and configuration simplicity favor AMD here
Memory capacity & bandwidthRTX PRO 6000: 96 GB; H100: 80 GB; H200: 141 GB; B300: 288 GB / 8.0 TB/sMI300X: 192 GB / 5.3 TB/s; MI350X / MI355X: 288 GB HBM3E / 8.0 TB/sAMD at H100/H200 gen; neutral at B300Advantage applies only to models requiring >80 GB at target precision
Memory capacity (Phase 2/3)B200: 180 GB; B300: 288 GB; Vera Rubin: 288 GB HBM4 (H2 2026)MI355X: 288 GB; MI455X: 432 GB HBM4 (H2 2026)Neutral at B300/MI355X; AMD reopens at MI455XRe-test MI455X vs. Rubin at Phase 3+ procurement
OEM availability, DellFeatured AI line (XE9780, XE7745); Validated DesignsSecondary; XE9680 with MI300X and XE9785 with MI355XNVIDIAAMD configs may require separate channel engagement
OEM availability, SupermicroRTX PRO 6000 and HGX H100/H200/B200/B300 air and liquid available8U MI300X air, 8U MI350X air, 10U MI355X air, 4U MI355X liquid availableNeutralSame procurement channel serves both vendors
Lead times (mid-2026)Cloud GPUs sold out per NVIDIA’s Q3 FY2026 results (Nov. 2025)MI355X may offer shorter lead timesAMD, slightlyDirectional; verify in writing at quotation

Across the six dimensions, NVIDIA takes software stack maturity, model compatibility, and community support outright, and takes Phase 1 cost per token for the model classes a Phase 1 deployment actually runs. AMD takes large-model memory economics at 405B-and-above BF16. The remaining dimensions split by generation and channel: memory capacity favors AMD at the H100/H200 generation and reaches parity at B300, and OEM availability favors NVIDIA at Dell while staying neutral at Supermicro. The verdict does not rest on a vote count. It rests on the two dimensions that govern how a single practitioner or small engineering team absorbs an incident: community and tooling support, which sets how long a novel issue takes to resolve, and software stack maturity, where the absent TensorRT-LLM runtime is the most concrete gap in the comparison.

The Phase 1 Recommendation

The Phase 1 procurement is one or two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs (96 GB GDDR7) in a Dell PowerEdge XE7745 or a Supermicro 4U air-cooled GPU SuperServer. A single card sustains 8,425 tokens per second on Qwen3-Coder-30B AWQ under vLLM [29], roughly 18× the 450 tokens per second a fifteen-user load needs at thirty tokens per second each, and runs 70B models at FP8 inside the 96 GB envelope (about 70 GB for weights plus roughly 26 GB of KV-cache headroom). The same card runs gpt-oss-120b at its native MXFP4 format on one GPU, so the largest Phase 1 model class needs no separate hardware path. Against the H100 NVL 96GB it delivers 1.63× the FP8 throughput when running FP4 at 100 concurrent requests [15], and against the H100 SXM it delivers roughly 28% lower cost per token at single-GPU scale [14].

Headcount does not determine the size of this build; peak concurrent generation does. Dividing the measured 8,425 tokens per second by a thirty-tokens-per-second user rate gives roughly 280 concurrent sessions, which places a single card far above the fifteen concurrent users design point and has the same configuration through the 100 to 499 employee band this paper calls medium. The recommendation continues to hold up to 999 employees on the same arithmetic, subject to two bounds. CloudRift measured aggregate throughput at its own batch settings rather than at 280 concurrent sessions, so that figure is an unvalidated ceiling and not a tested capacity [29]. KV-cache capacity inside the 96 GB envelope, not compute, is what gives way first as context lengths and session counts rise, so an organization near the top of the medium band should measure peak concurrency at its own context lengths and budget the second card against that measurement. Both the Dell XE7745 and the Supermicro 4U GPU SuperServer accept incremental GPU additions up to eight cards in the same chassis, which preserves the Phase 1-to-Phase 2 expansion path without a server swap. The eight-card maximum configuration draws about 4.8 kW GPU-only, which fits inside the 15–20 kW per-rack envelope that standard Computer Room Air Conditioning and Air Handling (CRAC/CRAH) equipment supports with no facility retrofit.

NVIDIA’s community depth and CUDA software maturity are the deciding factors for a first build operated by one-to-few engineers. Those two dimensions drive the recommendation independent of any pre-existing vendor relationship; existing Dell or Supermicro procurement relationships are confirming factors, not the basis of selection. AMD hardware is not recommended for Phase 1. Its throughput exceeds the requirement by orders of magnitude; the disqualifiers are the entry unit and the software stack. The smallest current-generation Instinct deployment is an 8-GPU node [25], [26], a unit that overshoots both the Phase 1 workload and any capital budget scaled to an organization below 500 employees, and the ROCm stack plus the thin community corpus load disproportionate operational risk onto a team with no slack to absorb it.

The Phase 2 Recommendation

Phase 2 stays on NVIDIA for both the inference scale-out and the training server. The path is to add RTX PRO 6000 cards to the Phase 1 server, or stand up a second identical server, for inference capacity, then procure a dedicated 8× H200 or 8× B200 SXM system for post-training and fine-tuning. Two considerations decide the training-server side. The fine-tuning toolchain (PyTorch FSDP, DeepSpeed, Unsloth) develops and validates on CUDA first with differing interconnect ceilings: AMD’s Infinity Fabric supports 8-way tensor parallelism inside a node at 896 GB/s of aggregate peer-to-peer bandwidth per GPU but caps the scale-up domain at eight GPUs through the MI355X generation [25], while NVLink/NVSwitch runs at 1,800 GB/s per GPU and extends to a 72-GPU rack-scale domain in the GB200 and GB300 NVL72 [30]. CloudRift’s long-context benchmark reinforces the B200 as the Phase 2 inference target: at 8K-token input and output, the B200 led every tested model on cost per token and reached up to 4.9× the RTX PRO 6000’s throughput [16]. Phase 2 inference targets including Llama 4 Maverick run on multi-GPU NVIDIA configurations: an 8× H200 node (1,128 GB) holds Maverick at BF16, and an 8× H100 node (640 GB) runs it at FP8, the deployment Meta itself describes on a single H100 DGX host [18]. The RTX PRO 6000 path does not extend to Maverick: the model’s 400 billion total parameters need roughly 200 GB of weights at INT4, above the 192 GB two cards supply, so serving it lands on the SXM node regardless of interconnect [18].

One Phase 2 workload profile warrants serious AMD evaluation: a hard requirement to serve Llama 4 Maverick or Llama 3.1 405B at BF16 with no quantization compromise. INT4 introduces measurable quality degradation that some production agentic workflows cannot accept. If that requirement becomes non-negotiable, a single 8× MI355X node (2.3 TB HBM3E) eliminates the multi-node tensor-parallelism complexity the equivalent NVIDIA configuration would carry, and system cost may favor AMD at that specification. Re-run the decision against then-current NVIDIA Blackwell Ultra (B300, 288 GB) and Vera Rubin (288 GB HBM4) parts, because the VRAM gap that distinguishes AMD at the H100/H200 generation shrinks to parity at B300 and reopens only at the MI455X-versus-Vera-Rubin comparison in the second half of 2026.

When to Revisit This Decision

Phase 2 procurement, twelve to twenty-four months after the Phase 1 pilot, is the first meaningful re-evaluation window. Five conditions should trigger a formal AMD reassessment then, each because it changes a specific term in the analysis above.

  1. The ROCm Core SDK (TheRock) build system reaches production release. The technology-preview stream (7.9 series and later) carries a new build system and ManyLinux-compliant Python-wheel packaging, with a mid-2026 production target [31]. Production release would cut ROCm installation and deployment friction, which is a precondition for taking AMD seriously because deployment friction is part of the operational tax that ruled AMD out at Phase 1.

  2. AMD ships a compiled-engine runtime, or AITER reaches comparable optimization depth on MI455X hardware. This is the single largest software gap in the mid-2026 comparison, and closing it rewrites the first dimension. Verify against independent benchmark data, MLPerf Inference v7.0 or equivalent, rather than vendor claims, because the gap is exactly the kind a vendor blog overstates.

  3. The MI400 series enters production with verified third-party benchmarks. The MI455X (432 GB HBM4, 19.6 TB/s, announced for the second half of 2026) targets Vera Rubin directly [21], [20]. Independent benchmarks against Vera Rubin NVL72 on production workloads at comparable price points are the evidence a Phase 2–3 AMD decision requires, and none exist while both parts are in pre-production or early production.

  4. The VRAM advantage persists against then-current NVIDIA hardware. If the MI455X holds 432 GB against Vera Rubin’s 288 GB, AMD keeps a 50% flagship-generation lead that matters for the largest production-precision models. A higher-VRAM Rubin variant would erase that argument, which is why the capacity case has to be re-checked at procurement rather than assumed forward.

  5. Cost per token is verified on the actual Phase 2 model classes. An independent benchmark comparing AMD accelerators against NVIDIA on the models destined for Phase 2 production must precede any AMD commitment. Vendor benchmarks and MLPerf submissions both inform the picture; neither substitutes for cost analysis on the models actually going into production.

The recommendation is conditional in form and unambiguous in substance for the decision in front of the organization. The order to issue is an NVIDIA RTX PRO 6000 build in a Dell PowerEdge XE7745 or its Supermicro equivalent. The AMD re-evaluation is a dated procurement-planning item one to two years out, carrying the five verification criteria above, to be run against whatever NVIDIA and AMD actually ship by then rather than against what either announced at CES.

References

  1. Advanced Micro Devices, Inc., “MLPerf 6.0: AMD Instinct MI355X GPUs Surpass 1M Tokens/Sec, Power New Workloads and Demonstrate Distributed Inference,” AMD, Apr. 1, 2026. [Online]. Available: https://www.amd.com/en/blogs/2026/amd-delivers-breakthrough-mlperf-inference-6-0-results.html. [Accessed: 08-Jun-2026]

    NVAM-1 Secondary source Back to text

  2. StorageReview, “AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack,” StorageReview.com, Apr. 2, 2026. [Online]. Available: https://www.storagereview.com/news/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack. [Accessed: 08-Jun-2026]

    NVAM-2 Contextual source Back to text

  3. MLCommons, “MLPerf Inference: Datacenter — v6.0 Results,” Apr. 2026. [Online]. Available: https://mlcommons.org/benchmarks/inference-datacenter/. [Accessed: 19-Jul-2026]

    NVAM-3 Primary source Back to text

  4. NVIDIA Corporation, “NVIDIA H200 Tensor Core GPU,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/h200/. Product page; 141 GB HBM3e, 4.8 TB/s memory bandwidth. [Accessed: 19-Jul-2026]

    NVAM-4 Primary source Back to text

  5. NVIDIA Corporation, “TensorRT-LLM,” GitHub, 2026. [Online]. Available: https://github.com/NVIDIA/TensorRT-LLM. GitHub repository; speculative decoding and FP4/INT4 quantization features. [Accessed: 08-Jun-2026]

    NVAM-5 Primary source Back to text

  6. Advanced Micro Devices, Inc., “Model acceleration libraries : Flash Attention 2,” ROCm Documentation, Jan. 22, 2026. [Online]. Available: https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/model-acceleration-libraries.html. Version 7.2.2. [Accessed: 10-Jun-2026]

    NVAM-6 Primary source Back to text

  7. K. T. Chitty-Venkata et al., “LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators,” arXiv, Argonne National Laboratory, Oct. 31, 2024. [Online]. Available: https://arxiv.org/abs/2411.00136. arXiv:2411.00136 (© IEEE 2024). [Accessed: 08-Jun-2026]

    NVAM-7 Primary source Back to text

  8. C. Putterman, L. Vaidya, A. Shah, S. Chetlur, and L. Tewari, “TensorRT-LLM speculative decoding boosts inference throughput by up to 3.6x,” NVIDIA Technical Blog, Dec. 2024. [Online]. Available: https://developer.nvidia.com/blog/tensorrt-llm-speculative-decoding-boosts-inference-throughput-by-up-to-3-6x/. [Accessed: 08-Jun-2026]

    NVAM-8 Secondary source Back to text

  9. Advanced Micro Devices, Inc., “vLLM V1 Performance Optimization — ROCm Documentation,” AMD, 2026. [Online]. Available: https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/vllm-optimization.html. [Accessed: 08-Jun-2026]

    NVAM-9 Primary source Back to text

  10. vLLM Project, “GPU Installation — ROCm,” vLLM, 2026. [Online]. Available: https://docs.vllm.ai/en/stable/getting_started/installation/gpu/. [Accessed: 08-Jun-2026]

    NVAM-10 Primary source Back to text

  11. Advanced Micro Devices, Inc., “Boosting Llama 4 Inference Performance with AMD Instinct MI300X GPUs,” AMD ROCm Blogs, Apr. 28, 2025. [Online]. Available: https://rocm.blogs.amd.com/software-tools-optimization/llama4-performance-b/README.html. [Accessed: 08-Jun-2026]

    NVAM-11 Secondary source Back to text

  12. Advanced Micro Devices, Inc., “AMD Enterprise AI Suite Expands with DeepSeek and Mistral AI models - Adds Support for AMD Instinct MI350X and MI355X GPUs (Version 1.8),” AMD Blogs, Mar. 10, 2026. [Online]. Available: https://www.amd.com/en/blogs/2026/amd-enterprise-ai-suite---version-1-8-.html. [Accessed: 19-Jul-2026]

    NVAM-12 Secondary source Back to text

  13. C. Ambati and T. Diep, “AMD MI300X GPU Performance Analysis,” arXiv, Oct. 31, 2025. [Online]. Available: https://arxiv.org/abs/2510.27583. arXiv:2510.27583. [Accessed: 08-Jun-2026]

    NVAM-13 Secondary source Back to text

  14. D. Trifonov, “RTX PRO 6000 vs H100, H200, and L40S: LLM Inference,” CloudRift AI, Nov. 27, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx6000-vs-datacenter-gpus. [Accessed: 08-Jun-2026]

    NVAM-14 Secondary source Back to text

  15. M. Tabares and C. Lutzer, “Benchmarking NVIDIA RTX PRO 6000 Blackwell on Akamai Cloud,” Akamai Blog, Oct. 30, 2025. [Online]. Available: https://www.akamai.com/blog/cloud/benchmarking-nvidia-rtx-pro-6000-blackwell-akamai-cloud. [Accessed: 08-Jun-2026]

    NVAM-15 Secondary source Back to text

  16. N. Trifonova, “Blackwell Dominates. Benchmarking LLM Inference on NVIDIA B200, H200, H100, and RTX PRO 6000,” CloudRift AI, Jan. 21, 2026. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-b200. [Accessed: 19-Jul-2026]

    NVAM-16 Secondary source Back to text

  17. dstack GmbH, “Exploring Inference Memory Saturation Effect: H100 vs MI300x,” dstack Benchmarks, Dec. 5, 2024. [Online]. Available: https://dstack.ai/blog/h100-mi300x-inference-benchmark/. [Accessed: 08-Jun-2026]

    NVAM-17 Secondary source Back to text

  18. Meta AI, “The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation,” Meta AI Blog, Apr. 5, 2025. [Online]. Available: https://ai.meta.com/blog/llama-4-multimodal-intelligence/. [Accessed: 19-Jul-2026]

    NVAM-18 Secondary source Back to text

  19. NVIDIA Corporation, “Inside NVIDIA Blackwell Ultra: The Chip Powering the AI Factory Era,” NVIDIA Technical Blog, 2025. [Online]. Available: https://developer.nvidia.com/blog/inside-nvidia-blackwell-ultra-the-chip-powering-the-ai-factory-era/. B300: 288 GB HBM3e, 8.0 TB/s, 1,400 W. [Accessed: 08-Jun-2026]

    NVAM-19 Secondary source Back to text

  20. NVIDIA Corporation, “NVIDIA Kicks Off the Next Generation of AI with Rubin — Six New Chips, One Incredible AI Supercomputer,” NVIDIA Investor Relations, press release, Jan. 5, 2026. [Online]. Available: https://investor.nvidia.com/news/press-release-details/2026/NVIDIA-Kicks-Off-the-Next-Generation-of-AI-With-Rubin--Six-New-Chips-One-Incredible-AI-Supercomputer/default.aspx. [Accessed: 08-Jun-2026]

    NVAM-20 Secondary source Back to text

  21. Advanced Micro Devices, Inc., “AMD and its Partners Share their Vision for 'AI Everywhere, for Everyone' at CES 2026,” AMD Newsroom, press release, Jan. 5, 2026. [Online]. Available: https://www.amd.com/en/newsroom/press-releases/2026-1-5-amd-and-its-partners-share-their-vision-for-ai-ev.html. [Accessed: 10-Jun-2026]

    NVAM-21 Secondary source Back to text

  22. Dell Technologies, “Dell Technologies Unveils Next-Generation Enterprise AI Solutions with NVIDIA,” Dell Technologies Investor Relations, press release, May 19, 2025. [Online]. Available: https://investors.delltechnologies.com/news-releases/news-release-details/dell-technologies-unveils-next-generation-enterprise-ai. [Accessed: 08-Jun-2026]

    NVAM-22 Secondary source Back to text

  23. Advanced Micro Devices, Inc., “AMD Delivers Leadership Portfolio of Data Center AI Solutions with AMD Instinct MI300 Series,” AMD Investor Relations, press release, Dec. 6, 2023. [Online]. Available: https://ir.amd.com/news-events/press-releases/detail/1173/amd-delivers-leadership-portfolio-of-data-center-ai-solutions-with-amd-instinct-mi300-series. [Accessed: 08-Jun-2026]

    NVAM-23 Secondary source Back to text

  24. V. Belapurkar and O. Kaven, “Announcing Enhancements to the Dell AI Platform with AMD,” Dell Technologies Blog, May 7, 2026. [Online]. Available: https://www.dell.com/en-us/blog/announcing-enhancements-to-the-dell-ai-platform-with-amd/. [Accessed: 08-Jun-2026]

    NVAM-24 Secondary source Back to text

  25. Super Micro Computer, Inc., “Supermicro Extends AI and GPU Rack Scale Solutions with Support for AMD Instinct MI300 Series Accelerators,” Supermicro, press release, Dec. 6, 2023. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-extends-ai-and-gpu-rack-scale-solutions-support-amd-instinct-mi300-series. [Accessed: 08-Jun-2026]

    NVAM-25 Secondary source Back to text

  26. Super Micro Computer, Inc., “Supermicro Expands Its Portfolio of Performance and Efficiency Driven Air-Cooled AI Solutions Featuring AMD Instinct MI355X GPUs,” Supermicro, press release, Nov. 19, 2025. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-expands-its-portfolio-performance-and-efficiency-driven-air-cooled-ai. [Accessed: 08-Jun-2026]

    NVAM-26 Secondary source Back to text

  27. Super Micro Computer, Inc., “Supermicro delivers performance and efficiency optimized liquid-cooled and air-cooled AI solutions with AMD Instinct™ MI350 series GPUs and platforms,” Supermicro Press Release, press release, June 12, 2025. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-delivers-performance-and-efficiency-optimized-liquid-cooled-and-air-cooled. San Jose, CA, USA. [Accessed: 08-Jun-2026]

    NVAM-27 Secondary source Back to text

  28. NVIDIA Corporation, “NVIDIA Announces Financial Results for Third Quarter Fiscal 2026,” NVIDIA Newsroom, press release, Nov. 19, 2025. [Online]. Available: https://nvidianews.nvidia.com/news/nvidia-announces-financial-results-for-third-quarter-fiscal-2026. [Accessed: 19-Jul-2026]

    NVAM-28 Primary source Back to text

  29. D. Trifonov, “RTX 4090 vs 5090 vs PRO 6000: LLM Inference Benchmark,” CloudRift AI, Oct. 9, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx-gpus-for-llm-inference. [Accessed: 04-Jun-2026]

    NVAM-29 Secondary source Back to text

  30. NVIDIA Corporation, “NVIDIA GB200 NVL72,” 2026. [Online]. Available: https://www.nvidia.com/en-us/data-center/gb200-nvl72/. Product page; 72-GPU NVLink domain at 1.8 TB/s per GPU; GB300 NVL72 with 72 Blackwell Ultra GPUs. [Accessed: 19-Jul-2026]

    NVAM-30 Primary source Back to text

  31. D. Widdows, J. Tseng, S. Todd, C. Sosa, and S. Rahim, “ROCm 7.9 Technology Preview: ROCm Core SDK and TheRock Build System,” ROCm Blogs, Advanced Micro Devices, Inc., Oct. 20, 2025. [Online]. Available: https://rocm.blogs.amd.com/software-tools-optimization/therock/README.html. [Accessed: 10-Jun-2026]

    NVAM-31 Secondary source Back to text

Contents