Section9
CPU-Only Model Deployments: Capabilities, Limitations, and Risks
CPU-only inference is a legitimate starting point for a single engineer running experiments on existing hardware. It is not a foundation for production deployments, and treating it as one creates migration costs that exceed the cost of buying the right hardware in the first place.
The appeal is real and worth naming. Inference looks free on hardware already in the rack no procurement cycle, no lead times, and no GPU server to stand up and maintain. A practitioner installs Ollama or builds llama.cpp from source, pulls a quantized 7B model, and within an hour has a working local assistant on a server already sitting in the rack. Ollama’s documentation makes that setup nearly frictionless [1], and llama.cpp is built for minimal-setup inference across a wide range of hardware, locally or in the cloud [2]. For an individual engineer testing whether on-premises AI is worth the larger conversation, this is the right first move. The trouble begins when the proof-of-concept becomes the production plan.
What CPU Inference Actually Delivers
The throughput numbers are unambiguous. A single AMD EPYC 9654 socket carries 96 Zen 4 cores and a theoretical 460.8 GB/s of DDR5 memory bandwidth across twelve channels [3]; with every channel populated it generates roughly 4 tokens per second on a 34B Q4_0 model [4]. A tuned single-socket EPYC 9554 running DDR5-5600 across all twelve channels reaches about 7 tokens per second on a 70B Q4_K_M model, measured with full llama-bench output and disclosed thread settings [5]. Push to full F16 precision on a dual-socket Turin platform and single-socket generation on Llama-3.1-70B falls to a few tokens per second; the second socket buys almost nothing, because token generation scales poorly once cross-NUMA memory traffic dominates [6]. A Georgia Tech characterization found that spreading inference across two sockets (96 cores) ran slower than a single 48-core socket, and Puget Systems saw decode fall roughly 40% when a dual-EPYC box was pushed to use both sockets, because the inter-socket interconnect throttles the memory bandwidth decode depends on [7], [8]. Smaller models run faster: server hardware with AVX2/AVX512 instruction sets and fully populated memory channels reaches roughly 5–15 tokens per second on 7B INT4 and 1–7 tokens per second on 70B INT4, the spread driven by memory speed, bandwidth, and thread tuning. For CPU inference, memory bandwidth is the biggest limiting factor [9], [10], [11].
The GPU comparison varies by tier and quantization, but the gap holds. A single RTX 5090 GPU with 32 GB of GDDR7 memory at 1,792 GB/s generates roughly 145–186 tokens per second on 8B models at Q4_K_M [12], [13], about 1.5–1.8× the RTX 4090 on the same bandwidth-bound decode, since the 5090’s VRAM bandwidth is close to 1.8× the 4090’s bandwidth. Absolute figures swing hard with the harness: LocalScore’s methodology-controlled run reports 66.3 generated tokens per second on Llama 3.1 8B and 45.5 tok/s on Qwen2.5 14B [14], while less-constrained community runs report rates two to three times higher [13]. Red Hat’s emerging-tech group makes the same caution the basis for a standardized CPU benchmarking framework, noting that vendors publish isolated best-case throughput without reproducible methodology, which leaves infrastructure teams unable to compare claims [15]. Single-user throughput compares cleanly only within one methodology, not across suites.
The RTX PRO 6000 Blackwell Server Edition carries 96 GB of GDDR7 memory at 1.6 TB/s [16]. Its single-user rate on small models lands within roughly 10% of the RTX 5090’s. The two cards differ by about that much in bandwidth on decode operations where bandwidth governs. The 96 GB memory capacity is the operationally meaningful difference. A 70B Q4_K_M model occupies about 40–44 GB of memory [5], leaving roughly 50 GB for KV cache and runtime overhead on a single card [17]. Production-serving throughput matters more than single-user speed and is where the two architectures diverge most. Continuous batching on a GPU lifts aggregate throughput into the thousands of tokens per second across concurrent requests by amortizing each weight read over many in-flight tokens [18]. GPU-rental operators report a single PRO 6000 in the high thousands of tokens per second on a 30B AWQ model under high-concurrency vLLM serving; read those as vendor-reported rather than independent [19]. The order of magnitude is the point: GPU serving sustains thousands of tokens per second per card under load, against single digits on a CPU socket.
The CPU-to-GPU ratio depends on quantization and model size. At 7B Q4_K_M the single-user gap runs roughly 4–13×; at 7B Q8_0, where more bytes move per token, it widens further. The 10–30× range commonly cited for GPUs holds as a representative middle when the model fits in GPU VRAM, provided the quantization level travels with the number rather than being treated as fixed. The mechanism splits by phase. Prefill, which processes the whole prompt at once, is compute-bound, and recent server CPUs narrow the gap here with dedicated matrix units (Intel’s Advanced Matrix Extensions, or AMX, on Sapphire Rapids and later Xeons; IBM’s Matrix Multiply Assist; Arm’s Matrix Extension) that the Georgia Tech study measured at 6.3–9.1× the prefill throughput of the prior CPU generation [7]. Decode, which generates one token at a time, is memory-bandwidth-bound, because each token requires reading the full set of weights and a single-request batch leaves almost no arithmetic to overlap that read [20].Decode is where CPUs lose. The EPYC 9654’s 460 GB/s of theoretical DDR5 bandwidth is between a quarter and a third of what the RTX 5090 (1,792 GB/s) or the PRO 6000 Server Edition (1,600 GB/s) moves from GDDR7 [3], [12], [16], [7].
The ratio inverts once a model no longer fits in GPU VRAM. When the GPU must offload weights to system RAM and pull them back across PCIe for every token, the same study measured an AMX-equipped Xeon beating an A100 by 12.7× in throughput on a 30B model and an H100 by 5× on a 66B model, because the GPUs spent 59–95% of their time moving weights over the bus rather than computing [7]. That finding is the argument for sizing VRAM to the model rather than against it. The Phase 1 choice of 96 GB on the PRO 6000 exists precisely to keep the target 70B-class models resident in VRAM, on the fast side of this crossover, instead of in the offload regime where a CPU would win by default.
The Multi-User Problem
The standard framing says three concurrent users each get a third of single-user throughput. That undersells the problem, because the mechanism that enables multi-user concurrency on GPUs is the one a CPU cannot run.
Out of the box, the llama.cpp inference engine runs a single slot, an isolated execution context and KV cache allocation, and processes requests one at a time. A user who submits behind two others waits for both to complete before seeing a first token. At 5 tokens per second on a 500-token response, the third user in a queue waits roughly 200 seconds for that first token; ten deep means nine full request cycles of waiting. Enabling parallel slots and continuous batching removes the full-queue wait. It does not remove the ceiling. On a GPU, continuous batching raises aggregate throughput because batching many requests lifts arithmetic intensity into the regime where the GPU’s idle parallel compute finally gets used; vLLM’s PagedAttention and continuous batching deliver 2–4× the throughput of earlier optimized serving systems at equal latency, and over an order of magnitude more than a naive serving loop, by managing the KV cache so the batch can grow [18]. A CPU holds no comparable reserve of parallel compute to activate. The same memory-bandwidth limit that caps single-user decode also caps the batch, so added users mostly divide a fixed throughput while tail latency climbs and the KV cache splits across slots, shrinking each request’s usable context window.
The user-experience thresholds show how thin the margin was to begin with. Interactive reading is equivalent to 5 tokens per second where anything below that rate reads as unusably slow. 10–30 tok/s keeps up with the average reading pace and anything above 30 tok/s outruns the average. A single user at 4 tokens per second on a 70B model sits at the edge of tolerable when nothing else runs. Add one concurrent request and it falls off the edge.
Memory Footprint and the Server That Cannot Spare It
A 70B INT4 model in GGUF format needs roughly 40–44 GB of RAM for holding the model parameters (weights), reflecting the format’s roughly 4.8 effective bits per parameter after block-metadata overhead [21], [22]. The common “40 GB” shorthand sits within rounding error; the weights alone run 40–44 GB, with another 2–4 GB per request of KV cache at 4K context and 2–4 GB of operating-system overhead on top. KV cache is the part that scales with load. It grows linearly with both sequence length and the number of concurrent requests, and at long context with a full batch it can exceed the weights themselves [7]. Red Hat’s CPU benchmarking work budgets it as the per-request footprint × planned concurrency × a 1.25 safety margin, which is the discipline that keeps a shared server off an out-of-memory kill during a traffic spike [15]. Smaller models scale down proportionally. A 7B INT4 model needs about 4.5–5 GB and a 13B INT4 model about 7–8 GB; INT8 doubles those figures and FP16 quadruples them [23].
A modern dual-socket server with 256–512 GB of DDR5 RAM holds even the largest GGUF model without strain. Where issues start to present themselves is when AI workloads are added to CPU-based servers on top of their existing workloads. A general-purpose application server with 64–128 GB of RAM cannot serve a 70B INT4 model alongside its existing database queries, web services, and build jobs without risking out-of-memory conditions or swap thrashing under concurrent load. That is operational risk, not theoretical risk, and it is the most common way CPU inference fails in production. Not because inference is slow, but because inference destabilizes whatever else the server was supposed to run.
Quality Degradation: A More Precise Account
INT4 models do show measurable degradation on complex reasoning, math, and coding. The conventional framing is wrong about why, and the why drives the deployment decision.
On GPUs, the most comprehensive study to date evaluated FP8, INT8, and INT4 across the Llama-3.1 family over more than 500,000 evaluations and found FP8 effectively lossless at every model scale, well-tuned INT8 within 1–3% of the FP16 baseline, and INT4 weight-only competitive with INT8 [24]. Four-bit weights, calibrated well, are not the problem. The specific quantization scheme, and where it runs, is. Degradation concentrates by task rather than spreading evenly. An independent run of Qwen3-32B at GPTQ-INT4 quantization lost under two points on MMLU-Pro yet dropped about eight points on HumanEval code generation [25], and a systematic study of quantized reasoning models found that bit-widths below 8 introduce real accuracy risk concentrated on mathematical and multi-step reasoning, with model size and task difficulty as the deciding factors [26]. The average retention figure hides the part that matters.
CPU inference under llama.cpp uses the GGUF format, which applies simpler per-block quantization (unless the model’s build supplies an importance matrix) and lack the activation-aware calibration that GPTQ and AWQ quantization techniques on GPUs apply by default. Many distributed GGUF quants are uncalibrated, and their losses run higher and more task-sensitive than GPU-side GPTQ or AWQ. Benchmarks on aggressive GGUF quantization show degradation surfacing earlier on mathematical reasoning (GSM8K) than on general language tasks [27]. Compounding matters more than per-token loss. A 2–3% error rate per reasoning step reads as tolerable in isolation; chained across a 10–20 step agentic workflow, it becomes the dominant cause of task failure. For knowledge-base question answering over short context, where one retrieval-grounded answer settles the question, GGUF INT4 on CPU is genuinely fine. For agentic planning, tool-use chains, and code generation with multi-step verification, it is not, and the failure mode is not a slightly worse answer but a chain that breaks halfway through while the agent loops on a malformed tool call.
A second myth cuts the other way and deserves the same correction. Lower-bit quantization is widely assumed to mean faster inference everywhere, on the logic that fewer bytes per weight means less bandwidth per token. That holds on bandwidth-bound CPU decode. It does not hold universally: a University of Virginia characterization across Apple Silicon and NVIDIA GPUs found that the cost of dequantizing packed weights back to compute precision can offset the bandwidth saving, so a 4-bit model is not reliably faster than an 8-bit one on every platform [28]. Quantization is a quality-and-speed trade whose sign depends on the hardware, not a free speedup.
The Migration Trap
Sculley and colleagues named the dynamic that governs what happens next [29]. In machine-learning systems, components entangle with their surroundings until changing one forces a re-audit of everything connected to it. Their “changing anything changes everything” principle was about input data and features, and the same logic applies to serving infrastructure. CPU-tuned workflows encode the infrastructure’s limits as design assumptions: long timeouts because generation is slow, single-request queues because batching buys little, no token streaming because the latency profile makes it pointless, output-length caps because longer responses mean unacceptable waits, and prompt designs squeezed to fit the model classes that fit in available RAM. None of these choices is wrong for the environment that produced it. All of them turn wrong the moment the serving stack changes.
Rebuilding those workflows on GPU-accelerated inference means re-architecting for streaming, redesigning queue behavior, retuning prompts for the larger models the GPU now permits, and replacing the request-handling layer to use continuous batching. That work consistently costs more than buying GPU hardware first and designing against its capabilities. The teams that pay it are the ones that treated the proof-of-concept as a production foundation.
Where CPU Inference Belongs
None of this means CPU inference has no place. Its place is specific, and the specificity is the point.
Air-gapped and edge deployments live within the single-user, low-throughput envelope where CPU performance is enough. At the smallest model sizes, CPUs can be the faster choice outright. A 2025 study running a 1-billion-parameter model on an iPhone 15 Pro measured the CPU at 17 tokens per second against 12.8 on the device’s GPU, because for a model that small the overhead of staging data to the GPU outweighs its compute advantage [30]. A technician’s lightweight assistant on an air-gapped resource-constrained machine and an analyst’s local reference model on a workstation inside a SCIF (Sensitive Compartmented Information Facility) both sit in this envelope, and both are often where GPU procurement is genuinely infeasible for reasons of procurement, supply, regulation, or facility power and cooling. Embedded inference on industrial hardware without the power, thermal, and procurement budget for a GPU has no alternative; the CPU is the only processor in the box. Secondary classification and routing models that read short inputs and feed a larger pipeline run comfortably at 5–15 tokens per second, because they classify rather than generate, and classification does not need interactive-chat speed.
A second niche sits at the opposite extreme: models too large for any affordable GPU. A Q4 quantization of DeepSeek-R1 or Llama-3.1-405B can run through hybrid CPU/GPU offload or on a large unified-memory system, trading speed for the ability to run at all. On a dual-EPYC workstation with the KTransformers hybrid-offload library, Puget Systems clocked DeepSeek-R1 at roughly 10–14 tokens per second in decode, with a long prompt taking minutes to answer [8]; Apple’s unified-memory M-series reaches similar territory for ultra-large models at lower cost than a multi-GPU NVIDIA build [28]. That is useful for batch and non-interactive work and unusable for anything a person waits on live.
CPU compute also earns a place as a component of a GPU deployment rather than a substitute for one. An agentic compute server uses CPUs as the processing layer for tool execution, document parsing, retrieval orchestration, and the dozens of non-inference operations that surround an agentic workflow. That is the correct architectural use of CPU in modern AI infrastructure: handling everything that is not matrix math, while the GPUs handle what is.
What remains is not whether to put GPUs in the rack but which ones, how many, and in what configuration. That decision turns on two points the constraints above make concrete: the model classes the organization needs to serve, and the concurrency it must sustain before throughput collapses. Both are sizing questions, and both rest on procurement limits.
References
-
Ollama, “Ollama Documentation,” 2026. [Online]. Available: https://docs.ollama.com. [Accessed: 02-Jun-2026]
-
G. Gerganov et al., “llama.cpp: LLM Inference in C/C++,” GitHub, ggml-org. [Online]. Available: https://github.com/ggml-org/llama.cpp. Software repository, 2023–2026. [Accessed: 02-Jun-2026]
-
Advanced Micro Devices, “AMD EPYC 9654,” AMD product specifications. [Online]. Available: https://www.amd.com/en/products/processors/server/epyc/4th-generation-9004-and-8004-series/amd-epyc-9654.html. 2022–2026. Per-socket memory bandwidth 460.8 GB/s; 96 cores, 192 threads; 12 channels DDR5-4800. [Accessed: 02-Jun-2026]
-
ggml-org, “CPU Performance,” ggml-org/llama.cpp, GitHub discussion #3167. [Online]. Available: https://github.com/ggml-org/llama.cpp/discussions/3167. 2023–2025. First-party community measurement (EPYC 9654, 34B Q4_0). [Accessed: 02-Jun-2026]
-
ahelpme.com, “LLM Inference Benchmarks with llama.cpp and AMD EPYC 9554 CPU,” Mar. 2025. [Online]. Available: https://ahelpme.com/ai/llm-inference-benchmarks-with-llamacpp-with-amd-epyc-9554-cpu/. First-party measurement with full llama-bench output (70B Q4_K_M, single-socket EPYC 9554, 12-channel DDR5-5600). [Accessed: 02-Jun-2026]
-
ggml-org, “Dual Epyc Genoa/Turin Token Generation Performance Bottlenecks,” ggml-org/llama.cpp, GitHub discussion #11733, 2025. [Online]. Available: https://github.com/ggml-org/llama.cpp/discussions/11733. First-party community measurement of NUMA-scaling behavior. [Accessed: 02-Jun-2026]
-
S. Na, G. Jeong, B. H. Ahn, J. Young, T. Krishna, and H. Kim, “Understanding Performance Implications of LLM Inference on CPUs,” 2024 IEEE International Symposium on Workload Characterization (IISWC), 2024. [Online]. Available: https://seonjinna.github.io/assets/pdf/iiswc24_CPULLM.pdf. [Accessed: 12-Jun-2026]
-
J. Allman, “Exploring Hybrid CPU/GPU LLM Inference,” Puget Systems, Mar. 20, 2025. [Online]. Available: https://www.pugetsystems.com/labs/hpc/exploring-hybrid-cpu-gpu-llm-inference/. [Accessed: 12-Jun-2026]
-
Clarifai, “llama.cpp: Fast Local LLM Inference, Hardware Choices & Tuning,” Clarifai, Inc., Mar. 17, 2026. [Online]. Available: https://clarifai.com/blog/ilama.cpp. [Accessed: 02-Jun-2026]
-
M. Larabel, “llama.cpp Benchmark,” OpenBenchmarking.org / Phoronix Test Suite. [Online]. Available: https://openbenchmarking.org/test/pts/llama-cpp. Test profile pts/llama-cpp, 2024–2026. [Accessed: 02-Jun-2026]
-
M. Larabel, “Llama.cpp AI Performance with the GeForce RTX 5090,” Phoronix, Jan. 27, 2025. [Online]. Available: https://www.phoronix.com/review/nvidia-rtx5090-llama-cpp/2. Review page 2 of 3. [Accessed: 18-May-2026]
-
NVIDIA Corporation, “New GeForce RTX 50 Series Graphics Cards & Laptops Powered By NVIDIA Blackwell Bring Game-Changing AI and Neural Rendering Capabilities To Gamers and Creators,” NVIDIA GeForce, Jan. 6, 2025. [Online]. Available: https://www.nvidia.com/en-us/geforce/news/rtx-50-series-graphics-cards-gpu-laptop-announcements/. 32 GB GDDR7, 1,792 GB/s memory bandwidth. [Accessed: 02-Jun-2026]
-
InsiderLLM, “RTX 5090 vs DGX Spark vs AMD: The Ultimate Local LLM Benchmark (2026),” InsiderLLM, Mar. 25, 2026. [Online]. Available: https://insiderllm.com/guides/rtx-5090-local-ai-benchmarks/. Community benchmark, uncontrolled methodology. Updated 2026-10-02. [Accessed: 02-Jun-2026]
-
LocalScore / Mozilla Builders, “NVIDIA GeForce RTX 5090 — Accelerator Results,” LocalScore benchmark. [Online]. Available: https://www.localscore.ai/accelerator/155. 2025–2026. [Accessed: 02-Jun-2026]
-
M. Tahhan, J. Harrigan, A. Ivanov, P. Power, and L. M. Zuccarelli, “Benchmarking AI Inference on CPUs: A Transparent Blueprint for the Enterprise,” Red Hat Emerging Technologies, May 28, 2026. [Online]. Available: https://next.redhat.com/2026/05/28/benchmarking-ai-inference-on-cpus-a-transparent-blueprint-for-the-enterprise/. [Accessed: 12-Jun-2026]
-
NVIDIA Corporation, “NVIDIA RTX PRO 6000 Blackwell Server Edition,” NVIDIA Cloud and Data Center. [Online]. Available: https://www.nvidia.com/en-us/data-center/rtx-pro-6000-blackwell-server-edition/. 2025–2026. 96 GB GDDR7, 1.6 TB/s bandwidth. [Accessed: 02-Jun-2026]
-
ggml-org, “Performance of llama.cpp on NVIDIA CUDA,” ggml-org/llama.cpp, GitHub discussion #15013. [Online]. Available: https://github.com/ggml-org/llama.cpp/discussions/15013. 2024–2026. Community reference for KV-cache headroom. [Accessed: 02-Jun-2026]
-
W. Kwon et al., “Efficient Memory Management for Large Language Model Serving with PagedAttention,” Proc. 29th ACM Symp. Operating Systems Principles (SOSP '23), Association for Computing Machinery, 2023. [Online]. Available: https://arxiv.org/abs/2309.06180. Koblenz, Germany, pp. 611–626. doi:10.1145/3600006.3613165. [Accessed: 02-Jun-2026]
-
Spheron Network, “RTX PRO 6000 Benchmarks: 30B AWQ, 70B FP8, and Cost per Million Tokens,” Mar. 9, 2026. [Online]. Available: https://www.spheron.network/blog/rent-nvidia-rtx-pro-6000/. Vendor-reported (GPU-rental operator). [Accessed: 02-Jun-2026]
-
Y. Fu, P. Bailis, I. Stoica, and H. Zhang, “Break the Sequential Dependency of LLM Inference Using Lookahead Decoding,” Proc. 41st Int. Conf. Machine Learning (ICML), 2024. [Online]. Available: https://arxiv.org/pdf/2402.02057. arXiv:2402.02057. [Accessed: 02-Jun-2026]
-
SitePoint, “VRAM for 70B Models: Why 16GB GPU Is the Minimum in 2026,” Feb. 15, 2026. [Online]. Available: https://www.sitepoint.com/vram-requirements-70b-models-16gb-gpu-minimum-2026/. [Accessed: 18-May-2026]
-
Spheron Network, “GPU Memory Requirements for LLMs: VRAM Calculator,” Mar. 20, 2026. [Online]. Available: https://www.spheron.network/blog/gpu-memory-requirements-llm/. [Accessed: 18-May-2026]
-
Latitude, “We Tested Quantized LLMs: Cost and Performance Results,” Latitude.so, Oct. 2024. [Online]. Available: https://latitude.so/blog/quantized-llms-cost-performance-results. Updated Mar. 2026. [Accessed: 02-Jun-2026]
-
E. Kurtic, A. N. Marques, S. Pandit, M. Kurtz, and D. Alistarh, “'Give Me BF16 or Give Me Death'? Accuracy-Performance Trade-Offs in LLM Quantization,” Proc. 63rd Annu. Meeting Assoc. Computational Linguistics (ACL), Vol. 1: Long Papers, 2025. [Online]. Available: https://arxiv.org/abs/2411.02355. Vienna, Austria, pp. 26872–26886. [Accessed: 02-Jun-2026]
-
AIMultiple Research, “LLM Quantization: BF16 vs FP8 vs INT4,” AIMultiple, Apr. 15, 2026. [Online]. Available: https://aimultiple.com/llm-quantization. Updated Apr. 15, 2026. First-party measurement (Qwen3-32B, GPTQ-INT4 on H100, lm-evaluation-harness). [Accessed: 02-Jun-2026]
-
R. Liu et al., “Quantization Hurts Reasoning? An Empirical Study on Quantized Reasoning Models,” Proc. Conf. Language Modeling (COLM), 2025. [Online]. Available: https://arxiv.org/abs/2504.04823. arXiv:2504.04823. [Accessed: 02-Jun-2026]
-
Ionio, “Benchmarking Quantized LLMs: What Works Best for Real Tasks?” Ionio AI, 2025. [Online]. Available: https://www.ionio.ai/blog/llm-quantize-analysis. [Accessed: 18-May-2026]
-
A. Benazir and F. X. Lin, “Benchmarking and Characterization of Large Language Model Inference on Apple Silicon,” Proc. ACM Meas. Anal. Comput. Syst., 2026. [Online]. Available: https://dl.acm.org/doi/10.1145/3771563. doi:10.1145/3771563. [Accessed: 12-Jun-2026]
-
D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” Advances in Neural Information Processing Systems, 2015. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html. Vol. 28, pp. 2503–2511. [Accessed: 02-Jun-2026]
-
H. Zhang and J. Huang, “Challenging GPU Dominance: When CPUs Outperform for On-Device LLM Inference,” arXiv, May 2025. [Online]. Available: https://arxiv.org/abs/2505.06461v1. arXiv:2505.06461v1. [Accessed: 12-Jun-2026]