Section11

Hardware Deep-Dive: AMD GPU Ecosystem

AMD’s Instinct MI355X competes at the top of NVIDIA’s product line on measured inference and leads every NVIDIA part below the Blackwell Ultra B300 on the axis that scales with model size: memory capacity. Each MI355X carries 288 GB of HBM3E at 8.0 TB/s, against 141 GB on the H200 and 180 GB on the B200; NVIDIA’s B300 matches it at 288 GB and 8 TB/s per GPU [1]. An 8-GPU MI350X or MI355X node holds 2.3 TB of HBM3E against 1.1 TB for 8× H200 and 1.4 TB for 8× B200, with an 8× B300 node reaching the same 2.3 TB. At MLPerf Inference 6.0 in April 2026, an 8× MI355X platform delivered 100,282 tokens per second on Llama 2 70B Server, a 3.1× gain over the prior MI325X generation; against the B200 it tied in Offline benchmarking, reached 97% of Server, and reached 119% of Interactive throughput, ratios corroborated by independent coverage of the round [2], [3], [4]. The hardware is competitive and beats every NVIDIA 8-GPU platform on capacity-bound workloads except the B300, which it matches.

Whether that hardware position reaches production depends on software. ROCm has closed most of the distance to CUDA but not all of it, and the gaps that remain decide whether AMD belongs in a first deployment or a later one. For the organizations this paper addresses, small at fewer than 100 employees and medium at 100 to 499, the deciding factor is staffing: whether anyone besides the model owner can provide ROCm support when it breaks. A first on-premises build staffed by one-to-few practitioners with no AMD operational history should buy NVIDIA in Phase 1. That verdict holds through the 500 to 999 employee tier as well, wherever AI infrastructure still rests on a small team.

The Current Generation: MI350X and MI355X

The MI350 Series launched June 12, 2025 and shipped in volume from Q3 2025. It is AMD’s CDNA 4 generation and the right primary reference for any 2026 procurement decision [5]. The MI300X (CDNA 3, 192 GB HBM3, 750 W, December 2023) and MI325X (CDNA 3, 256 GB HBM3E, 1,000 W, late 2024) remain deployed and supported, but the generational gap is large enough that new procurement should target MI350X or MI355X unless an existing fleet drives different economics.

The two SKUs share silicon: 256 compute units on TSMC’s 3 nm process for the compute dies, 288 GB of HBM3E at 8.0 TB/s, and 185 billion transistors across ten chiplets [6], [5]. AMD cut the compute-unit count from the MI300X’s 304 to 256 while roughly doubling per-CU FP8 throughput through a redesigned matrix engine. The net result is 5.0 PFLOPS of dense FP8 (MXFP8/OCP-FP8) on the MI355X and 4.6 PFLOPS on the lower-clocked MI350X, against the MI300X’s 2.6 PFLOPS, a 1.8–1.9× gain at the FLOP level [5]. Both SKUs add the MXFP4 and MXFP6 microscaling formats (OCP block-scaled 4-bit and 6-bit floating point). This is the first AMD generation with FP4 hardware acceleration; dense MXFP4 peaks at 10.1 PFLOPS per GPU.

Thermal envelope separates the two parts. The MI350X runs at 1,000 W total board power (TBP) and qualifies for both air-cooled and liquid-cooled environments. The MI355X runs at 1,400 W TBP and targets liquid cooling, though Supermicro’s 10U air-cooled MI355X chassis began shipping at SC25 in November 2025, which established at the system-design level that the 1,400 W part can run air-cooled, and AMD itself lists the MI355X UBB8 node in both air-cooled and direct-liquid-cooled configurations [7], [2]. For most on-premises sites without existing liquid cooling, the MI350X is the realistic near-term part, and the MI355X arrives when a facility upgrade or new build can justify rear-door heat exchangers or direct liquid cooling.

Both parts keep the OCP Accelerator Module (OAM) and Universal Baseboard (UBB) architecture of the MI300 Series. AMD made this an explicit design constraint so that MI300X-class platform designs carry forward and existing infrastructure can accept MI350X GPUs, subject to cooling and power validation for the higher TBP. The economics are concrete: organizations with prior MI300X investment have an upgrade path that NVIDIA’s SXM transitions between Hopper and Blackwell generally do not match.

On measured performance, AMD’s MLPerf submissions are the most defensible reference; MLCommons published the underlying results [3]. The MI355X delivered 100,282 tokens per second on Llama 2 70B Server at MLPerf Inference 6.0, a 3.1× gain over MI325X on the same benchmark, and AMD demonstrated more than 1 million tokens per second at 11 nodes and 87 GPUs on both Llama 2 70B and GPT-OSS-120B [2], [4]. The picture against NVIDIA’s newer B300 is less favorable than the B200 comparison: on Llama 2 70B the MI355X reached 93% of Server, 92% of Offline, and 104% of Interactive, and on GPT-OSS-120B it landed at 82–91% of B300 single-node performance, a range AMD derives from the official results and independent coverage has not re-verified [2]. AMD’s marketing also cites a 4× generational AI-compute improvement and a 35× inference uplift over MI300X. The 35× figure rests on stacked advantages, not a clean hardware comparison: AMD’s own footnote runs FP4 on the MI355X against FP8 on the MI300X on a Llama 3.1-405B chat workload at 32K input and 1K output, and to hit a 60 ms latency target it serves the MI300X at concurrency 1 and the MI355X at concurrency 64 [5]. The 3.1× MLPerf result is the honest apples-to-apples number; the 35× is a quantization-and-batch-inclusive marketing claim with a long methodology footnote.

The Phase 1 verdict on the hardware side is unambiguous. Divide the MI355X’s 100,282 tokens per second on Llama 2 70B Server by the 20–30 tokens-per-second-per-user floor this paper sets for fluid interactive use, and a single 8-GPU node sustains roughly 3,300 to 5,000 concurrent users/agents, two orders of magnitude beyond Phase 1’s fifteen concurrent requests target [2]. That lower bound is already more than three times above the headcount of a 999-employee organization, the top of the extended range this paper tracks, so employee count alone cannot saturate one node in any organization below 500 employees. Only agent fan-out or a per-request profile heavier than the interactive floor assumes will. Throughput will never be the binding constraint on a Phase 1 AMD deployment. Two other things are. AMD sells the MI350 Series only as OAM modules on 8-GPU baseboards; there is no single-GPU Instinct part comparable to the RTX PRO 6000 Blackwell workstation-class card that anchors the Phase 1 build, so AMD’s smallest current-generation entry point is a datacenter node priced far above the Phase 1 budget. The other constraint is ROCm operational risk, which the software analysis in this section quantifies.

Prior Generations: MI300X and MI325X in Context

The MI300X established AMD’s VRAM-per-dollar position at its December 2023 launch. AMD positioned it as the only single-GPU option then on the market capable of running 70B-parameter inference without sharding, and has kept that capacity argument central ever since [8]. An 8× MI300X node provides 1.5 TB of aggregate HBM3, enough to run Llama 3.1 405B at BF16 on a single node, against an 8× H100 SXM5 node’s 640 GB that forces multi-node deployment for the same model [9].

Independent benchmarking sharpens the picture. On Llama 3.1 405B at FP8, dstack measured an 8× MI300X node as the most cost-efficient and highest-throughput configuration for large prompts and batch sizes, the regime where an 8× H100 node saturates its 80 GB per GPU and its cost per token climbs [10]. The online-serving result cut the other way: against a 4× MI300X replica configuration, the 8× H100 node processed 74% more requests per second and cut time-to-first-token by at least half at every tested request rate, because prefill is compute-bound and eight GPUs distribute it more effectively than four [10]. The Argonne LLM-Inference-Bench study of 7B–72B Llama, Mistral, and Qwen models found the same shape: NVIDIA’s H100 led throughput and throughput-per-watt across the field, while the MI300X ran competitively under vLLM without closing that lead [11]. The MI300X’s edge is memory. It materializes only for models that exceed a single H100’s 80 GB with gains appearing in the decode phase rather than in compute-bound prefill [12].

The MI325X (1,000 W, 256 GB HBM3E, 6.0 TB/s) shipped through 2025 as an incremental CDNA 3 refresh: the same peak FLOP figures as the MI300X with more memory, more bandwidth, and higher TBP. Because it keeps the MI300X’s FP8 compute and adds 64 GB of memory and 0.7 TB/s of bandwidth, the MI325X competes with NVIDIA’s H200 rather than the H100, and its advantage concentrates on memory-bound large-model serving where the bigger HBM pool sustains larger batches without KV-cache eviction [10]. NVIDIA’s counter on its own hardware is TensorRT-LLM, which ROCm cannot match and which the software analysis below takes up. The MI325X persists in deployed fleets and remains supported, but a 2026 buyer faces a component superseded within three quarters of volume shipment; the size of the CDNA 3-to-CDNA 4 jump makes the MI350 Series the rational floor for new AMD procurement.

The MI400 Series and the H2 2026 Roadmap

AMD detailed the full MI400 Series lineup at CES 2026 in January, with production shipments confirmed for the second half of 2026 [13], [14]. The MI400 uses the CDNA 5 architecture, with compute dies on TSMC’s 2 nm (N2) process and I/O dies on 3 nm. AMD previewed the flagship’s headline specifications at its June 2025 launch event: 432 GB of HBM4 at 19.6 TB/s (roughly 2.5× the bandwidth of the MI355X and 1.5× its capacity), and peak compute of 40 PFLOPS FP4 and 20 PFLOPS FP8 per GPU [5].

The lineup spans four parts: the MI455X flagship for rack-scale training and inference, the MI450 volume variant carrying the OpenAI deployment, the MI440X for enterprise platforms in standard 8-GPU servers paired with EPYC Venice CPUs, and the MI430X for HPC and sovereign deployments with full FP64 support [13], [14]. The Helios rack, AMD’s first rack-scale platform, combines 72 MI455X GPUs with 31 TB of HBM4 and 1.4 PB/s of aggregate bandwidth for a claimed 2.9 exaFLOPS FP4 inference and 1.4 exaFLOPS FP8 training per rack, positioned against NVIDIA’s GB300 NVL72 and the forthcoming Vera Rubin rack-scale systems [14], [13]. AMD slates the Helios double-wide rack for the second half of 2026.

The MI400 also moves to UALink, the open Ultra Accelerator Link standard AMD backs as an alternative to NVIDIA’s proprietary NVLink, alongside Infinity Fabric. The MI430X, MI440X, and MI455X are the first accelerators to support it. The intent is to break NVIDIA’s interconnect lock-in, but the lever is contingent: practical UALink fabrics depend on partner switching silicon (from vendors such as Astera Labs, Auradine, Enfabrica, and Xconn) shipping in the second half of 2026, and in its absence Helios systems fall back to UALink-over-Ethernet or traditional mesh topologies [14].

For procurement, the MI400 is a Phase 2 or Phase 3 consideration with a realistic 2027-or-later deployment window for a buyer in this paper’s range. AMD has committed to shipping to partners on launch day, but the OpenAI supply agreement signed October 6, 2025, covering 6 GW of AMD GPUs across multiple generations with a first 1 GW MI450 deployment beginning in the second half of 2026, will absorb substantial early production [15]. An organization ordering one node competes for allocation behind that commitment, which is why the window for a sub-500-employee buyer opens later than the launch date suggests; the 500 to 999 tier gains nothing here either, because order size at that scale still does not command priority allocation. Treat every MI400 figure above as a vendor announcement pending production validation, and verify specifications and pricing at procurement time. The MI500 series teased for 2027, with a claimed 1,000× AI-performance improvement over the MI300X, is something to track rather than plan against [13].

OEM Platforms: Dell, Supermicro, and HPE

Dell, Supermicro, and HPE all ship AMD Instinct configurations, but NVIDIA dominates AI-server volume at each OEM, and AMD-specific sales often require separate engagement from an established NVIDIA buying relationship.

Supermicro’s current AMD Instinct catalog covers four enterprise-relevant configurations. The AS-8125GS-TNMR2 (8U air-cooled, 8× MI300X OAM) is the original MI300X system, shipping since December 2023 [16]. An 8U air-cooled 8× MI350X chassis followed in Q3 2025. The 10U air-cooled 8× MI355X system began shipping at SC25 in November 2025 [7]. For maximum density, Supermicro also ships a 4U liquid-cooled 8× MI355X system. Within an 8-GPU node, Infinity Fabric gives each GPU 896 GB/s of aggregate peer-to-peer bandwidth on the MI300-Series platform, enough for 8-way tensor parallelism without crossing a node boundary [16]. That scale-up domain caps at eight GPUs through the MI355X generation, which is why NVIDIA’s 72-GPU NVLink domain is the rack-scale counter and why AMD’s own answer is Helios with UALink in the MI400 generation.

Dell’s primary AMD offering at MI300X launch was the PowerEdge XE9680 with 8× MI300X, announced alongside the Dell Validated Design for Generative AI with AMD ROCm [8]. For the current generation, Dell announced the air-cooled 10U PowerEdge XE9785 and the direct-liquid-cooled XE9785L in November 2025, each with 8× MI355X accelerators [17]. HPE is a confirmed AMD Instinct integration partner across both the MI300 and MI350 generations [5].

Aggregate per-node VRAM is the procurement headline. An 8-GPU node delivers 1.5 TB on MI300X, 2.0 TB on MI325X, and 2.3 TB on MI350X or MI355X, against 640 GB on 8× H100 SXM5, 1.1 TB on 8× H200 SXM, and 1.4 TB on 8× B200; only an 8× B300 node matches AMD at 2.3 TB [1]. For an organization whose primary workload is large-model inference at production quality without aggressive quantization, per-node memory capacity is the single cleanest argument for AMD over every NVIDIA 8-GPU platform short of the B300. NVIDIA’s larger counter is the GB300 NVL72, which moves the contest to the rack scale rather than the 8-GPU node, a comparison the NVIDIA-versus-AMD analysis that follows takes up directly.

AMD’s 2025–2026 production ramp may yield shorter procurement lead times than allocation-constrained NVIDIA SKUs in some configurations. Allocation dynamics shift quarter to quarter, so verify this with the OEM at quotation time rather than assuming it.

ROCm: Where It Works and Where It Doesn’t

ROCm in 2026 is a working AI software stack for the frameworks that matter most to the use cases this paper targets. It is not at parity with CUDA, and the specific gaps decide whether AMD is a viable Phase 1 choice or a Phase 2 re-evaluation candidate.

The current production release is ROCm 7.2.3, with a technology-preview stream (TheRock / ROCm Core SDK, versions 7.9 and later, currently 7.12.0 as of March 2026) on track to replace the production stream by mid-2026 with a new build system and Python-wheel-and-tarball packaging [18]. MI350X and MI355X are supported in the 7.2.x stream, including SLES 15 SP7 and CDNA 4 coverage. PyTorch ROCm support is upstreamed into the official PyTorch repository and tracks the PyTorch stable cadence, mature for standard training and inference [19].

vLLM is the production inference stack on AMD and the framework AMD contributes to most directly. AMD maintains ROCm-optimized Docker images, publishes ROCm Python wheels that remove the prior Docker-only friction, and ships AITER, the AI Tensor Engine for ROCm, which provides AMD-specific fused kernels for mixture-of-experts, attention, and normalization operations behind the VLLM_ROCM_USE_AITER flag [20], [21]. vLLM V1 on ROCm supports expert parallelism, data parallelism, and AITER-optimized multi-head latent attention (MLA) kernels for the MI300X and MI355X. SGLang carries first-class ROCm support, and llama.cpp runs on AMD GPUs through ROCm/HIP, including consumer RDNA 3 cards.

The decisive gap is the absence of TensorRT-LLM. NVIDIA’s most-optimized inference library runs only on NVIDIA GPUs, and the Argonne LLM-Inference-Bench evaluation found it both faster and more power-efficient than vLLM on NVIDIA hardware, while AMD’s best available paths are vLLM and llama.cpp with no compiled-engine runtime to fall back on [11]. That raises NVIDIA’s per-GPU performance ceiling above AMD’s even where the raw silicon is comparable. AITER within vLLM is AMD’s partial mitigation, not a replacement. Second in consequence is FlashAttention 3, which is built on NVIDIA Hopper tensor hardware (the Tensor Memory Accelerator and warpgroup matrix instructions) and has no ROCm port; ROCm relies on FlashAttention-2-class kernels through Composable Kernel and Triton backends. The remaining gaps are narrower and more situational: bitsandbytes quantization is only partially supported on ROCm, pushing AMD deployments toward GPTQ or AWQ; custom CUDA kernels that assume NVIDIA’s 32-thread warp can break on AMD’s 64-thread wavefront and need kernel-level review; and torch.compile runs on ROCm but compiles more slowly and falls back to eager mode more often through the Inductor backend.

The thinness of ROCm’s community troubleshooting base bites harder than any single missing library, and it bites in proportion to how few people can answer when something breaks. CUDA carries roughly two decades of accumulated public kernels, community answers, and tutorials; ROCm’s equivalent corpus is far smaller, which translates directly into longer debugging cycles when a novel issue surfaces. The measured cost of that thinness is visible in independent testing. AIMultiple’s misconfiguration experiments on the MI300X found that installing stock pip vLLM over AMD’s ROCm container cut throughput 20% and that a wrong dtype flag cost 15%, failure modes with no CUDA equivalent at that severity [22]. A February 2026 technical report on 8× MI325X serving reached the same conclusion from the other direction: on the current ROCm stack, MLA models require a block size of 1 and cannot use KV-cache offloading while grouped-query attention (GQA) models benefit from both, so extracting production throughput demands architecture-by-architecture tuning that CUDA deployments largely amortize through defaults [23]. Organization size predicts exposure to that tuning burden only through staffing. A small organization by this paper’s definition, fewer than 100 employees, that has hired a one-to-few practitioner team absorbs every ROCm regression as added work to the team’s workload; a medium organization of 100 to 499 sits in the same position unless it has staffed additional engineers who own the serving stack rather than the models. The number that governs the decision is how many people can debug a kernel fault without vendor help, which the headcount bands do not report.

The quantified efficiency gap on data center Instinct hardware is narrower than community lore suggests. Most of what remains sits in the stack and in power management rather than in the architecture. The independent analysis by Ambati and Diep is the clearest statement of the pattern: the MI300X sustains on average 45% of its theoretical peak FLOPS across FP8, BF16, and FP16 against up to 93% for the H100 and B200, trails the H100 and H200 most in compute-bound FP8 prefill, and closes much of the gap in memory-bound FP16 decode; the authors attribute the deficit to software stack efficiency (compiler and kernel maturity, and AMD’s ROCm Collective Communication Library trailing NVIDIA’s NCCL) and to dynamic frequency and power throttling under dense tensor load [12]. AIMultiple’s multi-GPU run on the small Llama 3.1 8B model, where AMD’s memory headroom gives no advantage, put single-GPU MI300X throughput at 18,752 tokens per second, about 74% of the H200 under vLLM on both sides [22]. At the other end, on large memory-bound workloads, dstack’s Llama 3.1 405B benchmark put an 8× MI300X node ahead of an 8× H100 node on both throughput and cost once prompt and batch sizes saturated the H100’s memory [10]. The pattern holds across sources: NVIDIA leads on compute-bound and small-model work, AMD closes or wins on memory-bound large-model work, and the deficit that persists is the software stack and its tuning burden more than the chip.

Structured Specification Reference

The table consolidates AMD Instinct specifications across the four product generations relevant to 2026 procurement and Phase 2 planning. FP8 figures are dense (no structured sparsity); all are drawn from primary AMD sources and verified MLPerf submissions.

GPUArchitectureVRAMBandwidthPeak FP8 (dense)TBPForm FactorPrimary Workload Fit
MI300XCDNA 3192 GB HBM35.3 TB/s~2,615 TFLOPS750 WOAM/UBBPrior generation; single-GPU 70B+ inference at BF16; budget deployments
MI325XCDNA 3256 GB HBM3E6.0 TB/s~2,615 TFLOPS1,000 WOAM/UBBIncremental MI300X refresh; large dense-model inference; being superseded
MI350XCDNA 4 (3 nm)288 GB HBM3E8.0 TB/s~4,600 TFLOPS1,000 WOAM/UBB (drop-in)Current shipping part; air or liquid cooled; 70B–400B inference; smallest Instinct deployment unit is the 8-GPU node, above Phase 1 scale
MI355XCDNA 4 (3 nm)288 GB HBM3E8.0 TB/s~5,000 TFLOPS1,400 WOAM/UBB (drop-in)Maximum throughput; primarily liquid-cooled (air-cooled UBB8 available); 100,282 tok/s on Llama 2 70B Server at MLPerf 6.0
MI455X (MI400)CDNA 5 (2 nm/3 nm)432 GB HBM419.6 TB/s~20 PFLOPSTBDTBDH2 2026 target; rack-scale via Helios; frontier training and inference; Phase 2/3 reference

Two cautions govern how the table reads against any NVIDIA comparison. AMD’s 35× MI355X-versus-MI300X inference claim mixes a hardware jump with a quantization step and a 64-fold concurrency difference, as detailed above; the 3.1× MLPerf-validated gain is the defensible apples-to-apples reference. And every AMD throughput figure here reflects vLLM or SGLang because TensorRT-LLM is unavailable; any direct comparison to NVIDIA numbers should state which NVIDIA inference stack it uses, since the NVIDIA figure changes materially between vLLM and TensorRT-LLM.

What This Means for the AMD Decision

The decision reduces to whether AMD’s memory-and-throughput advantage outweighs the operational cost of running ROCm. Across this paper’s range, small at fewer than 100 employees and medium at 100 to 499, it does not, and the entry price compounds the verdict: the smallest current-generation Instinct deployment is an 8-GPU node, while Phase 1 is sized for a single RTX PRO 6000. The MI355X’s 288 GB, its MLPerf parity with the B200, and its near-parity with the B300 are real, and on a capacity-bound workload that exceeds NVIDIA’s per-GPU VRAM below the B300, AMD wins on hardware economics. The absent TensorRT-LLM path, the missing FlashAttention 3, and above all the thin community-troubleshooting base convert that hardware win into longer debugging cycles and more vendor escalations for a team that has no slack to absorb them. The hardware advantage is not large enough to pay for that operational tax at Phase 1. The conclusion extends through the 500 to 999 employee tier for the same reason: that band does not reliably come with additional engineers who own the serving stack, and staffing rather than payroll is what the verdict tracks.

AMD becomes a credible re-evaluation candidate when one of three conditions changes: the team grows past one-to-few practitioners to include an engineer who owns the serving stack, the MI400 and Helios platform with UALink ship and prove out in production, or a capacity-bound workload arrives for which the AMD node’s per-node economics win the head-to-head procurement comparison outright. Each is a Phase 2 trigger, and each is independent of the size band. A 60-person company that hires a second infrastructure engineer clears the bar that a 400-person company running on one practitioner does not.

References

  1. NVIDIA Corporation, “NVIDIA DGX B300.” [Online]. Available: https://www.nvidia.com/en-us/data-center/dgx-b300/. [Accessed: 18-Jul-2026]

    AMDGPU-1 Primary source Back to text

  2. Advanced Micro Devices, Inc., “MLPerf 6.0: AMD Instinct MI355X GPUs Surpass 1M Tokens/Sec, Power New Workloads and Demonstrate Distributed Inference,” AMD, Apr. 1, 2026. [Online]. Available: https://www.amd.com/en/blogs/2026/amd-delivers-breakthrough-mlperf-inference-6-0-results.html. [Accessed: 04-Jun-2026]

    AMDGPU-2 Secondary source Back to text

  3. MLCommons, “MLPerf Inference: Datacenter — v6.0 Results,” Apr. 2026. [Online]. Available: https://mlcommons.org/benchmarks/inference-datacenter/. [Accessed: 18-Jul-2026]

    AMDGPU-3 Primary source Back to text

  4. StorageReview, “AMD Instinct MI355X Achieves MLPerf Inference v6.0 Gains with Over 1 Million Tokens per Second and Supports Scalable ROCm Stack,” StorageReview, Apr. 2, 2026. [Online]. Available: https://www.storagereview.com/news/amd-instinct-mi355x-achieves-mlperf-inference-v6-0-gains-with-over-1-million-tokens-per-second-and-supports-scalable-rocm-stack. [Accessed: 18-Jul-2026]

    AMDGPU-4 Contextual source Back to text

  5. Advanced Micro Devices, Inc., “AMD Instinct MI350 Series and Beyond: Accelerating the Future of AI and HPC,” AMD, June 12, 2025. [Online]. Available: https://www.amd.com/en/blogs/2025/amd-instinct-mi350-series-and-beyond-accelerating-the-future-of-ai-and-hpc.html. [Accessed: 04-Jun-2026]

    AMDGPU-5 Secondary source Back to text

  6. Tom's Hardware, “ISSCC 2026: AMD Discloses How the Instinct MI355X Doubled Per-CU Throughput Despite Lower Compute Unit Count,” Feb. 27, 2026. [Online]. Available: https://www.tomshardware.com/tech-industry/semiconductors/inside-the-instinct-mi355x. [Accessed: 04-Jun-2026]

    AMDGPU-6 Contextual source Back to text

  7. Super Micro Computer, Inc., “Supermicro Expands Its Portfolio of Performance and Efficiency Driven Air-Cooled AI Solutions Featuring AMD Instinct MI355X GPUs,” Supermicro, press release, Nov. 19, 2025. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-expands-its-portfolio-performance-and-efficiency-driven-air-cooled-ai. [Accessed: 04-Jun-2026]

    AMDGPU-7 Secondary source Back to text

  8. Advanced Micro Devices, Inc., “AMD Delivers Leadership Portfolio of Data Center AI Solutions with AMD Instinct MI300 Series,” AMD, press release, Dec. 6, 2023. [Online]. Available: https://ir.amd.com/news-events/press-releases/detail/1173/amd-delivers-leadership-portfolio-of-data-center-ai-solutions-with-amd-instinct-mi300-series. [Accessed: 04-Jun-2026]

    AMDGPU-8 Secondary source Back to text

  9. Advanced Micro Devices, Inc., “AMD Instinct MI300X Platform,” AMD Product Specifications. [Online]. Available: https://www.amd.com/en/products/accelerators/instinct/mi300/platform.html. [Accessed: 18-Jul-2026]

    AMDGPU-9 Primary source Back to text

  10. dstack GmbH, “Exploring Inference Memory Saturation Effect: H100 vs MI300x,” dstack Benchmarks, Dec. 5, 2024. [Online]. Available: https://dstack.ai/blog/h100-mi300x-inference-benchmark/. [Accessed: 04-Jun-2026]

    AMDGPU-10 Secondary source Back to text

  11. K. T. Chitty-Venkata et al., “LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators,” arXiv, Argonne National Laboratory, Oct. 31, 2024. [Online]. Available: https://arxiv.org/abs/2411.00136. arXiv:2411.00136 (© IEEE 2024). [Accessed: 04-Jun-2026]

    AMDGPU-11 Primary source Back to text

  12. C. Ambati and T. Diep, “AMD MI300X GPU Performance Analysis,” arXiv, Oct. 31, 2025. [Online]. Available: https://arxiv.org/abs/2510.27583. arXiv:2510.27583. [Accessed: 04-Jun-2026]

    AMDGPU-12 Secondary source Back to text

  13. Data Center Dynamics, “AMD Unveils Full MI400 Product Lineup, Claims MI500 Chips Will Deliver 1,000× Increase in AI Performance,” Jan. 5, 2026. [Online]. Available: https://www.datacenterdynamics.com/en/news/amd-unveils-full-mi400-product-lineup-claims-mi500-chips-will-deliver-1000x-increase-in-ai-performance/. [Accessed: 04-Jun-2026]

    AMDGPU-13 Contextual source Back to text

  14. A. Shilov, “AMD Touts Instinct MI430X, MI440X, and MI455X AI Accelerators and Helios Rack-Scale AI Architecture at CES,” Tom's Hardware, Jan. 6, 2026. [Online]. Available: https://www.tomshardware.com/tech-industry/artificial-intelligence/amd-touts-instinct-mi430x-mi440x-and-mi455x-ai-accelerators-and-helios-rack-scale-ai-architecture-at-ces-full-mi400-series-family-fulfills-a-broad-range-of-infrastructure-and-customer-requirements. [Accessed: 04-Jun-2026]

    AMDGPU-14 Contextual source Back to text

  15. Advanced Micro Devices, Inc. and OpenAI, “AMD and OpenAI Announce Strategic Partnership to Deploy 6 Gigawatts of AMD GPUs,” AMD Newsroom, press release, Oct. 6, 2025. [Online]. Available: https://www.amd.com/en/newsroom/press-releases/2025-10-6-amd-and-openai-announce-strategic-partnership-to-d.html. [Accessed: 04-Jun-2026]

    AMDGPU-15 Secondary source Back to text

  16. Super Micro Computer, Inc., “Supermicro Extends AI and GPU Rack Scale Solutions with Support for AMD Instinct MI300 Series Accelerators,” Supermicro, press release, Dec. 6, 2023. [Online]. Available: https://www.supermicro.com/en/pressreleases/supermicro-extends-ai-and-gpu-rack-scale-solutions-support-amd-instinct-mi300-series. [Accessed: 04-Jun-2026]

    AMDGPU-16 Secondary source Back to text

  17. Dell Technologies, “PowerEdge XE9785.” [Online]. Available: https://www.dell.com/en-us/shop/ipovw/poweredge-xe9785. [Accessed: 18-Jul-2026]

    AMDGPU-17 Primary source Back to text

  18. Advanced Micro Devices, Inc., “ROCm Release Notes — ROCm 7.2.3 (production) and ROCm Core SDK / TheRock technology preview (7.9.0 and later),” AMD ROCm Documentation. [Online]. Available: https://rocm.docs.amd.com/en/latest/about/release-notes.html. [Accessed: 04-Jun-2026]

    AMDGPU-18 Primary source Back to text

  19. Advanced Micro Devices, Inc., “PyTorch Compatibility — ROCm Documentation,” AMD. [Online]. Available: https://rocm.docs.amd.com/en/latest/compatibility/ml-compatibility/pytorch-compatibility.html. [Accessed: 04-Jun-2026]

    AMDGPU-19 Primary source Back to text

  20. Advanced Micro Devices, Inc., “vLLM V1 Performance Optimization — ROCm Documentation,” AMD. [Online]. Available: https://rocm.docs.amd.com/en/latest/how-to/rocm-for-ai/inference-optimization/vllm-optimization.html. [Accessed: 04-Jun-2026]

    AMDGPU-20 Primary source Back to text

  21. vLLM Project, “GPU Installation — ROCm,” vLLM. [Online]. Available: https://docs.vllm.ai/en/stable/getting_started/installation/gpu/. [Accessed: 04-Jun-2026]

    AMDGPU-21 Primary source Back to text

  22. S. Dogan and E. Sarı, “Multi-GPU Benchmark: B200 vs H200 vs H100 vs MI300X,” AIMultiple, Oct. 6, 2025. [Online]. Available: https://aimultiple.com/multi-gpu. Updated Jun. 30, 2026. [Accessed: 04-Jun-2026]

    AMDGPU-22 Secondary source Back to text

  23. A. Georgiou, “Architecture-Aware LLM Inference Optimization on AMD Instinct GPUs: A Comprehensive Benchmark and Deployment Study,” arXiv, Feb. 23, 2026. [Online]. Available: https://arxiv.org/abs/2603.10031. Technical report, arXiv:2603.10031. [Accessed: 18-Jul-2026]

    AMDGPU-23 Secondary source Back to text

Contents