Section15
Multi-User and Agentic Scaling: Concurrency Architecture
Multi-user and agentic concurrency on GPU inference is a solved architecture as of 2026; what remains unsolved at most organizations is the budget. Production agentic workloads multiply large language model (LLM) inference demand by roughly 3× to 50× per user over direct chat, with the exact factor set by how much planning, tool use, and verification each task involves. Twenty concurrent agentic users therefore load an inference server similar to 60 to 1,000 chat sessions rather than twenty. Which is why the Phase 2 hardware list pairs a dedicated agentic compute server with the GPU inference server for exactly this reason: CPU-bound tool execution stays off the inference path, and the GPU is not starved by work it was never built to run.
The architecture that absorbs that load has five components, each of which defuses a distinct failure mode. A GPU generates one token at a time per inference stream while users and agents submit requests that each need hundreds to thousands of output tokens. Without explicit batching and queue management, one request monopolizes the device while the rest wait. Concurrency past a handful of users produces the tail latency that turns a working system into an unusable one. The five components are continuous batching with paged key–value (KV) cache management, a request scheduler with a queueing discipline matched to the workload, horizontal scaling on Kubernetes with GPU-aware orchestration, dedicated agentic compute that keeps tool execution off the inference server, and prefill/decode disaggregation for when mixed workloads contend for the same GPU. The failure modes arrive in a predictable order as demand grows.
PagedAttention and Continuous Batching
vLLM’s PagedAttention manages the KV cache the way an operating system manages physical RAM: fixed-size pages allocated dynamically across concurrent requests rather than pre-allocated per request. That one decision is what makes multi-user GPU serving tractable. The peer-reviewed SOSP 2023 paper that introduced the mechanism reported a 2–4× throughput gain over the prior state of the art, FasterTransformer and Orca, at equivalent latency, by driving KV-cache memory waste to near zero from the 60–80% that fragmentation and over-reservation cost earlier systems [1], [2]. Pages freed by completed requests return immediately to a global pool and slot into incoming work, so the server sustains batch sizes that static allocation could not fit in memory. An independent 2025 benchmark of vLLM against HuggingFace Text Generation Inference (TGI), a preprint not yet peer reviewed, measured vLLM at up to 24× higher throughput under high-concurrency workloads, with TGI ahead only on tail latency for single-user interactive use [3].
Continuous batching is the other half. Static batching collects a fixed group of requests, runs them together until the last one finishes, and only then admits the next group, so a short request waits behind the longest in its cohort. Continuous batching works at the token-generation iteration: the moment any request in the batch finishes, its slot returns to the scheduler and a waiting request takes it at the next step, with no synchronization barrier. The two mechanisms compound. PagedAttention makes large batches fit in memory; continuous batching keeps them full.
NVIDIA’s TensorRT-LLM implements the equivalent mechanisms as in-flight batching and paged KV caching, with kernel-level tuning for NVIDIA architectures and quantization down to FP8, FP4, INT4 AWQ (activation-aware weight quantization), and INT8 SmoothQuant [4]. For the Phase 1 deployment the server taxonomy section specifies, vLLM containerized under k3s with the NVIDIA GPU Operator, vLLM wins on operational maturity and the breadth of its tooling. TensorRT-LLM earns its added build complexity only at Phase 2, when squeezing the final increment of throughput from H200 or B200 silicon justifies the engineering time; no vendor-neutral benchmark quantifies that increment for this workload, so treat it as a Phase 2 evaluation item rather than a planned gain.
The Request Scheduler
vLLM’s V1 scheduler decides, at every forward pass, which requests run. It tracks a waiting queue of requests that have not gone through prefill and a running queue of requests actively decoding; when a running sequence finishes, its KV blocks return to the free pool and a waiting request takes the slot at the next iteration [5]. That is continuous batching expressed as scheduler behavior. Under memory pressure the scheduler preempts: it evicts lower-priority running sequences, discards their KV blocks, and recomputes them once capacity returns. The swap-to-CPU-memory path belonged to the now-deprecated V0 engine and is not the V1 mechanism [5], [6].
The queueing discipline inside the waiting queue sets tail latency. First-come, first-served (FCFS) is the default and the right call for single-purpose deployments where every user submits the same kind of work: no starvation, no priority inversion, nothing to reason about. For mixed workloads, priority scheduling routes high-priority requests, such as a production agent serving an active user or an interactive coding session, ahead of low-priority ones, such as overnight batch indexing or background re-embedding. vLLM ships priority scheduling but does not enable it by default; an operator turns it on by setting the scheduling policy and assigning each request a priority [5]. When higher-priority work arrives and no KV-cache space is free, the scheduler preempts the lowest-priority running sequences and resumes them once capacity returns.
For a Phase 2 deployment running interactive agents and background pipelines on shared inference hardware, priority scheduling wins. FCFS is simpler and fairer for homogeneous traffic, but it cannot hold an interactive agent’s latency stable while a batch job floods the queue. The configuration cost of priority scheduling is small next to the alternative of standing up separate vLLM instances per priority tier.
The Agentic Workload Multiplier
A single agentic workflow generates many model calls. The peer-reviewed record is clear on the qualitative point: tool-augmented agents run multi-step loops that invoke the model repeatedly to plan, call tools, judge results, and recover from failures [7], and architectures built on tree search or multi-agent debate issue several model calls for a single user query [8]. What the literature does not give is a single multiplier that holds across deployments, because the call count depends on agent architecture, task complexity, and how hard the harness retries. The planning range this paper carries forward from the agentic infrastructure discussion, consistent with the published characterizations and with practitioner experience, is three to ten model calls per user action for simple reason-and-act (ReAct-style) tool-using agents and ten to fifty or more for multi-step planning agents with verification loops.
Twenty concurrent users each running an agentic workflow therefore create a load equivalent to 60 to 1,000 simultaneous chat sessions. The range is wide because agent classes genuinely differ. A code-completion agent that calls the model once per suggestion sits at the floor; a multi-agent orchestrator that decomposes a task across specialists, verifies each output, and revises on failure sits at the ceiling. Size the inference tier against the agent class actually deployed, never against the chatbot baseline.
This is the sizing error the budget turns on. Sizing the inference server to the expected count of concurrent chat users, the natural first instinct, yields a Phase 2 deployment that serves chat well and stalls the moment agentic load arrives. The remedy is architectural: a dedicated agentic compute server, specified in the server taxonomy section, absorbs the CPU-bound tool execution that would otherwise contend with inference, and inference capacity is sized against the agentic multiplier rather than the chat baseline. The hardware roadmap section prices both.
The multiplier has a second consequence. Long-context prefill from an agent’s planning steps, where it ingests a large retrieved document or a long conversation history, competes with latency-sensitive decode on the same GPU, and that contention produces tail-latency spikes that no amount of batch tuning removes. The fix is to stop running prefill and decode on the same device, which is what disaggregation does.
Vertical and Horizontal Scaling
Two scaling dimensions solve two different problems. Vertical scaling means bigger GPUs, more VRAM (GPU memory) per card, and NVLink fabric for multi-GPU coherence inside one chassis. It lowers single-request latency, raises the largest servable model, and is simpler to operate: one server, one set of pods to watch. It also hits a hard ceiling at whatever silicon the current generation ships, which makes it the right Phase 1 answer for a single inference node and nothing more. Horizontal scaling means several inference servers behind intelligent traffic routing, with replicated model weights. It raises total throughput without a ceiling, adds redundancy against node failure, and costs more to operate; for Phase 2 and beyond, that operational cost buys the throughput and redundancy a production deployment cannot do without.
Kubernetes, the container orchestration substrate, is what separates horizontal scaling on paper from horizontal scaling in production. With the NVIDIA GPU Operator it provides GPU isolation between workloads, health-checked pods that restart on failure, and the primitives that traffic splitting and rolling model updates depend on [9]. The shift is industry-wide: the Cloud Native Computing Foundation’s (CNCF) 2025 annual survey, published in January 2026, found that 82% of organizations running containers now run Kubernetes in production, up from 66% two years earlier, and that 66% of organizations hosting generative AI use Kubernetes for at least some of their inference [10]. CNCF’s framing of the result, Kubernetes as the operating system for AI, is promotional; the adoption numbers behind it hold up.
The Phase 1 baseline the server taxonomy section sets is vLLM under k3s or microk8s with NVIDIA’s GPU Operator: single-node Kubernetes that gives a small team production lifecycle management without a full control plane to run. Phase 2 moves to full Kubernetes once a second inference node, the training server, and the agentic compute server all need scheduling against a shared cluster.
The routing layer in front of horizontally scaled inference changed more than anything else in the past year. llm-d, launched in May 2025 by Red Hat with CoreWeave, Google Cloud, IBM Research, and NVIDIA, is a Kubernetes-native distributed serving stack built on vLLM, with KV-cache-aware routing, autoscaling, and prefill/decode disaggregation as first-class capabilities [11], [12]. The project joined the CNCF as a Sandbox project in March 2026, which placed it under vendor-neutral governance, and it has kept a fast release cadence since: v0.7, released in May 2026, moved predicted-latency scheduling to general availability [13], [12]. Its v0.5 benchmarks from February 2026 report roughly 3,100 output tokens per second per B200 decode GPU in a wide expert-parallel configuration, up to 50,000 output tokens per second on a sixteen-prefill-by-sixteen-decode B200 topology, and an order-of-magnitude reduction in time to first token (TTFT) against a round-robin load-balancing baseline [12]. Those are the project’s own results against a naive baseline, but third-party validation now exists: AWS benchmarked llm-d’s disaggregation path on its own B200 instances, with results covered under prefill/decode disaggregation below [14].
KServe complements llm-d rather than competing with it. KServe handles the model-serving API, model-registry integration, and lifecycle; llm-d handles inference-aware, KV-cache-aware routing between vLLM instances beneath it. The Phase 2 recommendation is full Kubernetes with the NVIDIA GPU Operator and llm-d for routing, with KServe added as the serving-API layer if the deployment’s model-management patterns call for it.
Storage and Networking Under Agentic Load
Vector-database query latency lands directly on agent response time, because an agent is sequential; it cannot reason further until retrieval returns. For a Qdrant deployment doing semantic search over millions of well-indexed documents, this paper’s target is sub-50-millisecond query latency for a responsive agent. Past 200 milliseconds per retrieval step, the cost compounds: a ten-step run spends over two seconds in retrieval before a single token is generated. Qdrant on NVMe storage with the Gridstore engine, networked over the 100 GbE storage fabric, meets the target by construction. Holding it there as the corpus grows is the operational job: monitoring 95th-percentile (P95) query latency, tracking index size against the memory budget, and re-sharding before the latency curve bends.
Storage for conversation logs, agent memory, and intermediate state must not throttle the inference path. The PostgreSQL backend that holds LangGraph checkpoints, MLflow tracking, and Langfuse traces lives on NVMe and serves exactly the access pattern it is built for: many small, concurrent reads and writes. The link between the agentic compute server and the inference server should stay well under a millisecond of latency, which the 25 GbE same-rack connection the server taxonomy section specifies delivers without special tuning. The only discipline required is to keep the two servers co-located and not let a well-meaning network design drop a firewall between them.
Monitoring and Capacity Planning
Five metrics define operational visibility for production inference: GPU utilization, queue depth, request latency at the median and the 95th and 99th percentiles (P50/P95/P99), KV-cache hit rate, and storage I/O throughput. vLLM exposes all five through its Prometheus-compatible metrics endpoint, and the NVIDIA GPU Operator surfaces GPU telemetry through DCGM, NVIDIA’s Data Center GPU Manager [15], [9]. KV-cache hit rate is the metric most specific to LLM inference and the one that matters most under agentic load: it measures the share of incoming prompt tokens served from the prefix cache instead of recomputed, and when agents reuse a shared system-prompt prefix across calls, a high hit rate cuts both TTFT and GPU compute directly.
Set baselines and scaling triggers before production traffic arrives, not after. Workable defaults: GPU utilization sustained above 80% for more than five minutes, queue depth above a threshold calibrated to the batch size for more than sixty seconds, and P95 latency more than 20% over its target across a five-minute window. The thresholds are tuning parameters; having them defined before traffic hits is the requirement. A deployment that goes live without them discovers its capacity ceiling by crashing through it.
Prefill/Decode Disaggregation
LLM inference has two computationally distinct phases. Prefill processes the whole input prompt in parallel; it is compute-bound, rewards high FLOP density, and tolerates some batching delay. Decode emits output tokens one at a time, reloading the KV cache from VRAM at each step; it is memory-bandwidth-bound and latency-sensitive. Run both on one GPU and decode wastes the card’s FLOP capacity while a long prefill from one request stalls the decode steps of others; the classic head-of-line blocking pattern. Under chat load that inefficiency is tolerable. Under agentic load, where long-context planning steps throw large prefills against latency-sensitive decode, it becomes the dominant source of tail latency, and adding identical inference servers does not remove it, because each new server carries the same internal contention.
DistServe, at OSDI 2024, made the architectural case [16]. Splitting prefill and decode onto separate worker pools, with the KV cache transferred between them, removed the interference and let the system serve 7.4× more requests at the same latency target, or hold a 12.6× tighter target at the same throughput, while keeping more than 90% of requests inside their service-level objective (SLO). Those are goodput gains, meaning requests served within an SLO rather than raw TTFT percentages. The distinction matters because the size of the improvement depends on the workload. AWS measured up to 70% higher tokens per second from llm-d’s prefill/decode disaggregation path against a standard vLLM deployment, serving OpenAI’s GPT-OSS on B200 hardware as concurrency rose toward 128 [14]. A March 2026 preprint, starting from the observation that prefill/decode disaggregation “has become the standard architecture for modern LLM inference engines,” showed that dynamically routing later-turn prefills to decode nodes, where cached KV state is reused, cuts second-and-later-turn TTFT a further 68% relative to standard disaggregation in multi-turn serving [17]. A single flat percentage would misstate all of these.
Production maturity is settled at large scale. Moonshot AI’s Mooncake platform, which serves the Kimi chatbot on a prefill/decode-disaggregated, KV-cache-centric architecture, won Best Paper at USENIX FAST 2025 and reported operation across thousands of nodes processing over 100 billion tokens daily [18]. What remains open at this paper’s scale is tooling maturity. vLLM V1 supports disaggregated prefilling natively through its KVConnector API, running prefill and decode as separate processes with the KV cache moved across the connector; the documentation still labels the feature experimental as of mid-2026 [19]. That label is substantive: the API is stable enough to validate but not yet frozen, so a Phase 2 deployment adopting it should budget for migration work across vLLM versions. llm-d implements disaggregated serving as a first-class capability and supplies the Kubernetes-native orchestration that makes it tractable to operate.
Disaggregation is a Phase 2 pattern, switched on when agentic workflows reach production and head-of-line blocking shows up under real load. Phase 1 single-server inference gains nothing from it, since there is no inter-server KV transfer to optimize, and deploying it before the contention appears buys operational complexity with no return. The trigger is specific and observable: P95 TTFT degrading under mixed agentic load while average GPU utilization stays moderate. That signature is head-of-line blocking, and it is the cue to disaggregate.
The Phase 2 Scaling Architecture
The five components assemble into one deployable system: vLLM with PagedAttention and continuous batching as the inference engine, the V1 scheduler with priority scheduling for mixed workloads, horizontal scaling on full Kubernetes with the NVIDIA GPU Operator and llm-d for KV-cache-aware routing, dedicated agentic compute that keeps tool execution off the inference path, and prefill/decode disaggregation switched on when agentic load makes head-of-line blocking the binding constraint. Each component resolves a failure mode the one before it cannot with failure modes arriving in that order as load increases. Together they produce a system in which a second, third, and fourth inference node buy additional throughput rather than additional contention. The serving hardware implied by these calculations is the input the financial plan has to fund.
References
-
W. Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” Proc. ACM SIGOPS 29th Symposium on Operating Systems Principles (SOSP '23), 2023. [Online]. Available: https://dl.acm.org/doi/10.1145/3600006.3613165. Koblenz, Germany, pp. 611–626. [Accessed: 14-Jun-2026]
-
vLLM Project, “vLLM: Easy, Fast, and Cheap LLM Serving with PagedAttention,” vLLM Blog, June 20, 2023. [Online]. Available: https://blog.vllm.ai/2023/06/20/vllm.html. [Accessed: 14-Jun-2026]
-
S. Kolluru, “Comparative Analysis of Large Language Model Inference Serving Systems: A Performance Study of vLLM and HuggingFace TGI,” arXiv, Nov. 17, 2025. [Online]. Available: https://arxiv.org/abs/2511.17593. arXiv:2511.17593 [cs.DC]. Preprint, not peer-reviewed. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “TensorRT-LLM Overview,” NVIDIA Technical Documentation, 2026. [Online]. Available: https://nvidia.github.io/TensorRT-LLM/overview.html. Covers in-flight batching and paged KV caching. [Accessed: 14-Jun-2026]
-
A. Gordic, “Inside vLLM: Anatomy of a High-Throughput LLM Inference System,” vLLM Blog, Sept. 5, 2025. [Online]. Available: https://blog.vllm.ai/2025/09/05/anatomy-of-vllm.html. [Accessed: 14-Jun-2026]
-
vLLM Project, “Optimization and Tuning — vLLM Documentation,” 2026. [Online]. Available: https://docs.vllm.ai/en/latest/configuration/optimization/. V1 default preemption mode: RECOMPUTE. [Accessed: 14-Jun-2026]
-
H. Go and S. Park, “A Study on Classification Based Concurrent API Calls and Optimal Model Combination for Tool Augmented LLMs for AI Agent,” Scientific Reports, July 1, 2025. [Online]. Available: https://doi.org/10.1038/s41598-025-06469-w. Vol. 15, art. 20579. [Accessed: 14-Jun-2026]
-
Arunkumar V, G. R. Gangadharan, and R. Buyya, “Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents,” arXiv, Jan. 18, 2026. [Online]. Available: https://arxiv.org/abs/2601.12560. arXiv:2601.12560 [cs.AI]. Preprint. [Accessed: 14-Jun-2026]
-
NVIDIA Corporation, “NVIDIA GPU Operator Documentation,” NVIDIA Cloud Native Technologies, 2026. [Online]. Available: https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/. [Accessed: 14-Jun-2026]
-
Cloud Native Computing Foundation, “The CNCF Annual Cloud Native Survey: The Infrastructure of AI's Future,” CNCF, Jan. 20, 2026. [Online]. Available: https://www.cncf.io/reports/the-cncf-annual-cloud-native-survey/. [Accessed: 14-Jun-2026]
-
Red Hat, “Red Hat Launches the llm-d Community, Powering Distributed Gen AI Inference at Scale,” press release, May 20, 2025. [Online]. Available: https://llm-d.ai/blog/llm-d-press-release. [Accessed: 14-Jun-2026]
-
llm-d Project, “llm-d: Achieve State-of-the-Art Inference Performance with Modern Accelerators on Kubernetes,” GitHub, 2026. [Online]. Available: https://github.com/llm-d/llm-d. v0.5 benchmarks, Feb. 2026; v0.7 release, May 2026; CNCF Sandbox. [Accessed: 14-Jun-2026]
-
C. Costa, C. Coleman, and R. Shaw, “Welcome llm-d to the CNCF: Evolving Kubernetes into SOTA AI Infrastructure,” CNCF Blog, Mar. 24, 2026. [Online]. Available: https://www.cncf.io/blog/2026/03/24/welcome-llm-d-to-the-cncf-evolving-kubernetes-into-sota-ai-infrastructure/. IBM Research, 'Donating llm-d to the Cloud Native Computing Foundation': https://research.ibm.com/blog/donating-llm-d-to-the-cloud-native-computing-foundation. [Accessed: 14-Jun-2026]
-
V. Gangasani, A. Smith, and G. Annem, “Introducing Disaggregated Inference on AWS Powered by llm-d,” AWS Machine Learning Blog, Mar. 16, 2026. [Online]. Available: https://aws.amazon.com/blogs/machine-learning/introducing-disaggregated-inference-on-aws-powered-by-llm-d/. [Accessed: 16-Jul-2026]
-
vLLM Project, “vLLM Documentation,” 2026. [Online]. Available: https://docs.vllm.ai/en/latest/. [Accessed: 14-Jun-2026]
-
Y. Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-Optimized Large Language Model Serving,” Proc. 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI '24), 2024. [Online]. Available: https://www.usenix.org/conference/osdi24/presentation/zhong-yinmin. Santa Clara, CA, USA, pp. 193–210. [Accessed: 14-Jun-2026]
-
Z. Li, J. Liu, Z. Xu, Y. Zhang, T. Rabbani, and C. Zhang, “Not All Prefills Are Equal: PPD Disaggregation for Multi-Turn LLM Serving,” arXiv, Mar. 9, 2026. [Online]. Available: https://arxiv.org/abs/2603.13358. arXiv:2603.13358 [cs.NI]. Preprint. [Accessed: 14-Jun-2026]
-
R. Qin et al., “Mooncake: Trading More Storage for Less Computation — A KVCache-Centric Architecture for Serving LLM Chatbot,” Proc. 23rd USENIX Conference on File and Storage Technologies (FAST '25), 2025. [Online]. Available: https://www.usenix.org/conference/fast25/presentation/qin. Santa Clara, CA, USA, pp. 155–170. Best Paper Award. [Accessed: 16-Jul-2026]
-
vLLM Project, “Disaggregated Prefilling (Experimental),” vLLM Feature Documentation, 2026. [Online]. Available: https://docs.vllm.ai/en/latest/features/disagg_prefill/. [Accessed: 14-Jun-2026]