Section7

Generative AI Fundamentals for Decision Makers

Every strategic argument in this paper reduces to one distinction: whether the AI capability runs on infrastructure the organization controls or on infrastructure a vendor controls. The technical vocabulary below builds toward that distinction. Without it, vendor marketing collapses categorically different things into the same word. “Model” can mean a piece of software the organization runs under its own license, on its own hardware, governed by its own policy. It can also mean a remote API call to a service that retains both the data and the capability. The terms in this section make that difference legible and provide the foundation required to evaluate vendor claims on their merits.

Neural Networks

A neural network is a layered mathematical function that learns patterns from examples rather than from hand-written rules. During training, the network’s internal parameters (its weights) adjust across millions of examples until its outputs match the desired results. No engineer specifies what a “cat” looks like or what a “valid invoice” contains. The network discovers the pattern by exposure to data, sometimes billions or trillions of examples worth.

The contrast with traditional software matters for procurement and governance discussions. A conventional system encodes an expert’s rules and produces predictable, auditable behavior. A neural network encodes a statistical pattern discovered from data, and its behavior is bounded by what it has seen, not by what an engineer wrote. Training data quality is therefore a first-order infrastructure concern, not a downstream detail to be sorted out after the hardware arrives.

Large Language Models

A Large Language Model (LLM, or simply “model”) is a neural network trained on text at a scale that would have been infeasible a decade ago. Many current LLMs also train on images, audio, and structured data. The “large” refers to parameter count: the number of learned weights, ranging from roughly one billion at the small end to over a trillion at the frontier. The “language model” refers to the predominantly human-language training data and the model’s ability to capture linguistic patterns within it.

Parameter count correlates with capability but does not determine it. A well-architected 30-billion-parameter model fine-tuned on high-quality domain data routinely outperforms a generic 200-billion-parameter model on the specific task that matters. Architecture, training data quality, post-training methodology, and fine-tuning each carry weight comparable to raw size. Treat parameter count as one input to capability assessment, not the headline.

Transformers

The Transformer is the architecture underlying every major LLM in production today. Engineers at Google Brain and the University of Toronto introduced it in their 2017 paper “Attention Is All You Need,” which replaced the sequential processing of earlier recurrent networks with an attention mechanism that weighs the relationships between all tokens in a context window simultaneously [1]. That parallelism is what makes modern LLMs trainable at scale and what allows them to handle long, structured instructions coherently.

The original paper demonstrated the architecture on machine translation, improving over the prior state of the art by more than two BLEU points while requiring substantially less training time [1]. It is now the single most-cited work in modern AI. Every LLM discussed in this paper is a Transformer variant.

Graphics Processing Units

A Graphics Processing Unit (GPU) is a specialized processor, also called an accelerator, originally designed to render the thousands of triangles and pixels that make up a video game frame, a task that requires performing the same mathematical operation on enormous numbers of values simultaneously. That architectural characteristic, massive parallelism across thousands of small compute cores, turns out to be exactly what neural network training and inference require. Multiplying a matrix of billions of weights against an input vector is structurally identical to shading millions of pixels: the same operation, repeated across a vast array of independent values, all at once. A modern CPU (Central Processing Unit) handles this kind of workload poorly because CPUs are designed around a small number of powerful cores optimized for sequential, branching logic. A GPU designed for AI workloads contains thousands to tens of thousands of smaller cores purpose-built to execute the same instruction across massive datasets in parallel, completing in milliseconds what a CPU would require minutes to process.

The GPU has become the foundational compute unit of modern AI for this reason. Every benchmark, throughput figure, and hardware recommendation in this paper refers to GPU-based infrastructure. The specifications that matter for AI workloads differ from those for gaming or traditional visualization. Three figures govern AI performance: GPU memory capacity, which determines how much of the model fits on-device; memory bandwidth, which determines how quickly weights can be moved from memory to the compute cores during each token generation step; and compute throughput in the floating-point and integer precisions relevant to the target quantization level. A GPU with 80 GB of high-bandwidth memory and a GPU with 24 GB of slower memory can both run the same model at the right quantization level, but they will exhibit materially different performance and throughput. GPU memory, memory bandwidth, and compute throughput are three independent constraints, not a single “power” number. Understanding that is the prerequisite for evaluating every hardware recommendation that follows.

Tokens

A token is the unit of text a model processes. It is not a word and not a character. English text averages roughly 0.75 words per token, so a 1,000-word document occupies around 1,300 tokens. Code, non-Latin scripts, and unusual proper nouns consume tokens at different rates.

Tokens are the unit in which everything is measured: context window size, throughput, pricing, and VRAM consumption. When a vendor advertises “200K context,” it means 200,000 tokens, which roughly equates to 150,000 English words or about 450 single-spaced pages. When a GPU specification lists tokens per second, that number describes output speed for a specific model running on specific hardware under specific operational configurations. The number alone does not tell the full performance story, but it is a useful indicator of throughput when run on a known environment.

Context Windows and the KV Cache

The context window is the maximum number of tokens a model can consider at once. Everything outside that window is invisible to the model: earlier conversation turns, prior documents, previous tool calls. This is a hard engineering constraint, not a soft preference, and it has direct hardware implications. Context windows vary by model from a few thousand tokens to as much as one million, with recent developments pushing beyond that limit.

The mechanism that makes context windows expensive is the KV (key-value) cache. As a Transformer processes a context, it computes intermediate attention values for every token and stores them so each successive token can be generated without recomputing the whole context from scratch. The cache lives in GPU memory and grows linearly with context length multiplied by batch size (requests/sequences processed concurrently), and roughly with model size. The required KV cache is different for each model since it is dependent on the model’s architecture and optimizations. A 70-billion-parameter model serving a 128,000-token context can consume tens of gigabytes of VRAM in cache alone, on top of the tens to hundreds of gigabytes required for the model weights themselves. A model’s KV cache size can be roughly calculated with the following formula:

KV = 2 × Batch_Size × Sequence_Length × Num_Layers × Num_KV_Heads × Head_Dimension × Bytes_Per_Element

  • “2” accounts for both the Key and Value matrices
  • Batch_Size is the number of sequences (requests) processed at the same time
  • Sequence_Length is the number of tokens in the context
  • Num_Layers is the total number of transformer attention layers in the model
  • Num_KV_Heads is the number of Key-Value attention heads
  • Head_Dimension is the dimensionality of each individual attention head (e.g. 128)
  • Bytes_Per_Element is the memory footprint per numerical value (e.g. 2 bytes for FP16/BF16, 1 byte for FP8/INT8, or 0.5 bytes for packed 4-bit formats)

For example, a simple calculation of the Qwen3-8B model at FP8 with 32K context and batch size of one, this results in 2 × 1 × 32,768 × 36 × 8 × 128 × 1 = 2,415,919,104 bytes = 2.25 GB required memory. With ten requests processed concurrently (batch size), this turns into 22.5 GB of required memory on top of the memory required to load the entire model. For Google’s Gemma 4 31B model at FP8 with 32K context and batch of one, the required KV cache is 2.8 GB. For OpenAI’s gpt-oss-120B it is 0.5 GB.

This is the single most important fact to understand before sizing inference hardware. A GPU that holds the model weights comfortably can still run out of memory the moment users start sending long documents or running multi-turn agent conversations. The server taxonomy section develops this in detail.

Prompts

A prompt is the input text sent to a model to elicit a response. In its simplest form, a prompt is a user’s question. In production systems, a prompt is a structured assembly of several components concatenated before the model ever sees them: a system prompt that establishes the model’s role, constraints, and output format; the conversation history of prior user and model turns; any retrieved context injected by a RAG layer; the tool definitions available to an agent; and the user message itself. Everything in that assembly consumes tokens from the context window.

Prompt construction is where most production AI quality is won or lost. Two deployments running the same model on the same hardware against the same question will produce materially different results based on how the prompt is assembled. A well-constructed system prompt encodes the organization’s terminology, escalation rules, refusal behavior, and output structure. A poorly constructed one leaves the model to guess. This is the discipline of prompt engineering: the systematic design, testing, and versioning of prompts as production artifacts, treated with the same rigor as application code.

The practical implication for infrastructure is that prompts are not free. A 4,000-token system prompt sent on every request to every user multiplies directly into KV cache pressure, throughput cost, and latency. Sizing inference hardware requires an estimate of the full prompt size in production, not just the user’s typed message.

Context Engineering

Context engineering extends prompt engineering from the wording of a single instruction to the management of everything the model sees. Anthropic, which formalized the term in a September 2025 engineering publication, defines it as the practice of deciding which tokens, drawn from every source that can reach the model, should occupy the context during inference, and of maintaining that selection as a session evolves [2]. The prompt assembly described above is the starting point. Context engineering governs how that assembly changes across a long session: which conversation turns are retained verbatim, which are compacted into summaries, which documents are loaded up front versus retrieved when needed, and which tool outputs are truncated before they enter the window.

The discipline exists because a larger context is not automatically a better one. Model recall degrades as the window fills, a failure mode the practitioner literature calls context rot, so tokens compete for a finite attention budget and every marginal token carries a cost in the model’s ability to use the rest [2]. The core techniques respond to that constraint: compaction summarizes older conversation history to reclaim window space, structured note-taking moves durable state into external files the model re-reads on demand, and just-in-time retrieval loads references at the moment they are needed rather than preloading an entire corpus [2].

For infrastructure planning, context engineering is where the KV cache mathematics developed earlier meets operational policy. Every token admitted to the window consumes cache memory for the life of the sequence and prefill compute on every request that carries it, so the compaction thresholds, retrieval policies, and tool output limits an application enforces translate directly into the concurrent capacity of a given GPU. Context assembly discipline also determines how much benefit the serving layer can extract from prefix reuse: the RadixAttention mechanism described in the inference engine comparison below reuses cached computation only when requests share stable prefixes, and an application that reshuffles its context on every turn forfeits that reuse. Hardware sizing should therefore assume the context management policy the application will actually enforce, not the maximum window the model supports.

Inference vs. Training

Two operations share the same model weights but differ in almost every way that matters for infrastructure. Training produces a model by repeatedly running data through it, measuring how wrong its outputs are, and adjusting its weights to reduce that error. Each step is a forward pass followed by a backward pass that computes gradients and an optimizer update that applies them. Building a model from scratch repeats this across trillions of tokens, typically over weeks to months, on hundreds or thousands of GPUs networked together. Inference runs the finished model in the forward direction only: it takes a new input and produces an output, with no gradients and no weight changes. Training is how a model learns; inference is how it works. Training happens occasionally, when a model is built, updated, or tuned. Inference happens at every user request, every agent step, every embedding lookup, every retrieval query. For almost every organization, the dominant workload is inference.

The mechanical difference produces a memory difference that directly shapes hardware requirements. Inference must hold the model weights plus the KV cache in GPU memory. Training must hold considerably more: the weights, a full set of gradients the same size as the weights, the optimizer states that the update rule maintains, and the intermediate activations from the forward pass that the backward pass needs to compute gradients. Under the standard accounting for mixed-precision training with the Adam optimizer, model state alone consumes roughly 16 bytes per parameter (two bytes for the half-precision weights, two for the gradients, and twelve for the optimizer’s full-precision master weights, momentum, and variance), against the two bytes per parameter that half-precision inference weights require [3]. The original ZeRO analysis makes the gap concrete: a 1.5-billion-parameter model needs at least 24 GB of GPU memory to train but only about 3 GB to hold for inference, before activations are even counted [3]. A GPU that comfortably serves a given model can be an order of magnitude short of the memory needed to train the same model.

The workloads also differ in character, which changes how the hardware is procured and operated. Training is a batch job. It can be scheduled, checkpointed, paused, and resumed, and at scale it is bound less by any single GPU than by the interconnect that synchronizes gradients across many GPUs, which is why training clusters invest heavily in high-bandwidth fabrics such as NVLink and InfiniBand. Because it is episodic, training capacity is a natural candidate for rented or burst infrastructure used for the duration of a job and released afterward. Inference is the opposite: a continuously available service held to latency targets, where requests arrive unpredictably and a single well-specified node can often serve the workload without cross-GPU synchronization. Inference capacity is therefore the part of the stack an organization most benefits from owning and keeping online, which is the central case this paper develops.

This asymmetry compounds over a deployed model’s lifetime. Training is a large one-time expenditure; inference is a recurring one that accumulates with every request, and, for a model in sustained production use, the cumulative inference load overtakes the one-time cost of training it. Industry disclosures point the same way: Google attributed roughly 60% of its machine-learning energy use over 2019 to 2021 to inference rather than training [4], and Meta has reported that inference, not training, draws the largest share of its AI infrastructure capacity [5]. An organization that sizes its hardware around the occasional training run while underestimating sustained inference will find the recurring workload, not the episodic one, defines its actual requirements.

The practical sequencing follows directly. Full pretraining of a foundation model is out of reach for any organization without the significant resources to do so, as the foundation models subsection establishes, so the relevant form of training is not building a model but adapting one. Parameter-efficient fine-tuning with the LoRA and QLoRA methods, covered in the fine-tuning subsection, brings training capability within reach of the same server-class GPUs provisioned for inference, collapsing the training-versus-inference hardware distinction for small and medium organizations. Size the infrastructure for sustained inference throughput across the expected user population first. Add fine-tuning capability second, on the same or incrementally expanded hardware, once the use case is proven and the data is curated. The phased hardware roadmap section formalizes this sequencing and specifies the configurations at each step.

Prefill and Decode

Inference is not a single operation. Every request a model serves passes through two computationally distinct phases which govern most of the hardware sizing and inference engine decisions that follow.

The prefill phase processes the entire input prompt in a single parallel forward pass. The model reads every token of the system prompt, conversation history, retrieved context, and user message at once, computes the attention values across the whole sequence, and writes the resulting key-value pairs into the KV cache. Prefill ends when the model emits the first output token. Because it performs large matrix multiplications across the full sequence length simultaneously, prefill saturates the GPU’s compute cores and is compute-bound: its speed is limited by floating-point throughput, one of the three constraints introduced in the GPU subsection [6]. Its governing metric is time to first token (TTFT), the delay before any response appears. For a short prompt, prefill is nearly instantaneous. For a retrieval-augmented request that prepends thousands of context tokens, prefill can run to hundreds of milliseconds and dominate the perceived latency of the entire request.

The decode phase generates the response one token at a time. Each step reads the previous token, the full accumulated KV cache, and the model weights, computes a single new token, appends its key-value pair to the cache, and repeats until the response is complete. This step is a small matrix-vector multiplication rather than a large matrix-matrix one, so it exposes almost no parallelism and leaves the compute cores largely idle while the memory bus runs at capacity. Decode is therefore memory-bandwidth-bound: its speed is set by how fast weights and KV cache stream out of GPU memory, not by how fast the cores can multiply, mapping directly onto memory bandwidth, the second of the GPU subsection’s three constraints [6]. Its governing metric is time per output token (TPOT), also called inter-token latency, which determines how smoothly the response streams.

The two phases place opposite demands on identical hardware: prefill wants raw compute, decode wants memory bandwidth. When both run on the same GPU, as in a standard single-server deployment, they interfere. A long prefill stalls the decode steps of other in-flight requests, inflating both TTFT and TPOT under concurrent load [6]. This is why production serving systems schedule the two phases deliberately, through techniques such as the continuous batching discussed in the inference engine subsection and, in larger clusters, prefill-decode disaggregation, which assigns each phase to a separate pool of GPUs so each runs on hardware matched to its profile [7]. For an organization sizing its first inference server, the practical consequence is that a single tokens-per-second figure conceals two different bottlenecks. A workload dominated by long inputs, such as RAG or document analysis, stresses prefill compute, while a workload dominated by long generated outputs, such as drafting, summarization, or agent loops, stresses decode bandwidth. Sizing hardware requires an estimate of which phase the expected workload emphasizes, not a single throughput number.

Inference Engines

An inference engine (also called a runtime or serving framework) is the software layer that loads a model’s weights, manages GPU memory, schedules incoming requests, and produces output tokens. The engine sits between the model file and the application. Its design decisions, including how it allocates KV cache memory, how it batches concurrent requests, and what quantization formats it supports, have a larger impact on throughput, latency, and hardware utilization than almost any hardware variable. Selecting the wrong engine for a workload can make a correctly specified GPU cluster perform as if it were half its size.

The throughput hierarchy on NVIDIA hardware is broadly: TensorRT-LLM leads, followed by SGLang, then vLLM, then llama.cpp, then Ollama. Hardware flexibility runs in roughly the reverse order: llama.cpp and Ollama support the widest range of hardware, while TensorRT-LLM is locked to NVIDIA infrastructure. Neither ranking is the full story. The right engine is the one that matches the operational constraints (deployment timeline, team expertise, concurrency requirements, hardware standardization), not the one that achieves peak tokens-per-second on a benchmark.

vLLM is the de facto standard for production multi-user inference on NVIDIA hardware. Its core innovation is PagedAttention, developed by the UC Berkeley Sky Computing Lab [8]. PagedAttention applies the core idea of operating system virtual memory to KV cache management. Just as a process addresses memory through a logical page table that maps to non-contiguous physical frames, each active sequence in vLLM addresses its KV cache through a logical block table that maps to non-contiguous physical blocks in GPU memory. The practical effect is that KV cache memory is allocated on demand rather than pre-reserved, eliminating the GPU memory fragmentation endemic to naive implementations. PagedAttention is paired with continuous batching, which keeps the GPU busy at every decoding step. Research has shown this combination can yield nearly a 5× increase in throughput over monolithic setups, comfortably handling many parallel requests on the same hardware. vLLM exposes an OpenAI-compatible API, which means existing applications can point at a vLLM endpoint without code changes [9].

SGLang is an open-source inference engine developed at UC Berkeley that has emerged as a direct competitor to vLLM for production multi-user workloads, and in several benchmarks surpasses it [10]. Its core innovation is RadixAttention, a KV cache management approach that automatically identifies and reuses shared prefixes across requests. Where vLLM’s PagedAttention solves the problem of where to store KV cache blocks in GPU memory, RadixAttention solves the problem of whether to compute them at all. If two requests share an identical system prompt, the KV cache for that prompt is computed once and shared across both, rather than duplicated. SGLang consistently outperforms vLLM on prefix-sharing workloads, with throughput gains ranging from modest improvements on simple tasks to 6.4× on highly repetitive prompt structures [10]. The practical consequence for on-premises deployments is significant. Any architecture where many users share a common system prompt (which describes nearly every production deployment) benefits directly from RadixAttention’s prefix reuse without any application-level changes. For random, diverse queries with no shared context, the advantage over vLLM shrinks. SGLang exposes an OpenAI-compatible API, carries comparable setup complexity to vLLM, and should be evaluated alongside vLLM as a candidate production engine for any deployment where multi-turn conversation, RAG retrieval, or agentic loops constitute the primary workload pattern [11].

llama.cpp is a C/C++ inference engine written by Georgi Gerganov, the same author as the GGUF format. GGUF is highly optimized to run well even on a standard CPU, making it accessible to users without dedicated GPU hardware. Tools like Ollama and LM Studio abstract away much of the complexity, often relying on GGUF models behind the scenes. llama.cpp is the engine that made serious LLM inference on consumer hardware practical, and it remains the reference implementation for CPU-first and memory-constrained deployments. It does not support continuous batching in the same manner as vLLM, which limits its throughput under concurrent load. On CPU inference, generation speed for a quantized 13-billion-parameter model on a capable workstation-class processor typically runs to roughly 8–15 tokens per second for a single request, with server throughput below one request per second under concurrent load. llama.cpp is appropriate for development workstations, air-gapped environments, and single-user deployments where operational simplicity outweighs throughput requirements [12].

Ollama wraps llama.cpp behind a container-friendly management layer with a simple command-line interface and local API. Ollama is the right choice when a model needs to be running locally in under five minutes. Developer experience is unmatched, but Ollama does not scale past single-user workloads. On time to first deployment, Ollama wins by hours compared to vLLM’s minutes-to-hours and TensorRT-LLM’s days-to-weeks. Ollama is the appropriate engine for developer laptops, proof-of-concept demonstrations, and environments where infrastructure expertise is limited. It is not a production serving engine for multi-user workloads [13].

TensorRT-LLM is NVIDIA’s open-source inference library for accelerating LLM inference on NVIDIA GPUs. It delivers its strongest performance gains through low-precision quantization which can double throughput and halve memory consumption with minimal accuracy impact. TensorRT-LLM is appropriate for organizations with mature MLOps practices and stable model selections at workloads where quantization-driven throughput advantages translate into material infrastructure cost savings; it is not appropriate for early-phase deployments where model selection remains in flux [14].

The engine selection decision should be made concurrently with hardware sizing, not after. An inference server specified for vLLM’s continuous batching profile cannot be naively reused for TensorRT-LLM’s compiled-plan workflow without operational process changes. And a deployment budgeted for Ollama’s single-user throughput will require re-architecture, not just reconfiguration, if user load grows to production scale.

Speculative Decoding

Speculative decoding is an inference-time optimization that reduces token-generation latency without altering model weights, model architecture, or output quality. The standard Transformer inference process is inherently sequential: generating K tokens requires K separate forward passes through the full model, one token at a time. Speculative decoding addresses this by computing several tokens in parallel. A smaller, faster approximation model (the draft model) proposes multiple candidate tokens ahead, and the larger target model verifies those candidates in a single parallel forward pass, accepting or rejecting them without changing the output distribution. The technique was introduced independently and concurrently by researchers at Google Research and DeepMind in 2022–2023 [15], [16].

The infrastructure implication is that speculative decoding extracts more throughput and lower latency from hardware that is already provisioned, rather than requiring additional GPUs. The gains are workload-dependent: the technique performs best when the draft model’s proposals are frequently accepted by the target model, which is typical of interactive workloads such as chat, document Q&A, and agent loops where context is coherent and predictable. Both vLLM and SGLang provide built-in support for speculative decoding as a configurable option [17]. For organizations deploying interactive AI applications where user-facing latency is a primary concern, speculative decoding is a standard configuration consideration at the inference engine layer, not an advanced optimization deferred to a later phase.

Transformer Architecture Optimizations

The attention mechanism that makes Transformers powerful is also their primary bottleneck: the time and memory cost of self-attention grow quadratically with sequence length [18]. A family of optimizations has made long-context inference practical on commodity GPU hardware. Because these optimizations live inside the models and inference engines rather than in operator configuration, their effects are easy to overlook when sizing infrastructure, yet they determine the real memory footprint and throughput of a deployment. Two categories matter for planning.

The first reorganizes how attention is computed without changing its mathematical result. FlashAttention, introduced by researchers at Stanford in 2022, is the foundational example [18]. A naive implementation writes the full N×N attention matrix to the GPU’s high-bandwidth memory (HBM) and reads it back, and on modern GPUs that memory traffic, not the arithmetic, is the limiting factor. FlashAttention is IO-aware: it uses tiling to keep blocks of the computation in the GPU’s small but fast on-chip SRAM, computes attention block by block, and never materializes the full matrix in HBM. The result is an exact attention computation, with no approximation and no quality loss, that reduces memory traffic to linear in sequence length and delivers a reported 7.6× speedup on the attention computation itself [18]. Two successors refined the approach for newer hardware. FlashAttention-2 improved work partitioning across GPU threads for roughly a 2× gain over the original [19], and FlashAttention-3 exploits the asynchrony and FP8 low-precision support of NVIDIA’s Hopper GPU generation to reach 1.5 to 2.0× over FlashAttention-2 and up to 75 percent of the H100’s theoretical maximum throughput [20]. FlashAttention is now the default attention kernel in the production inference engines discussed above, including vLLM and SGLang. An operator does not configure it directly, but it is the reason a given GPU can serve far longer contexts than a naive implementation would permit.

The second category changes the attention architecture itself to shrink the KV cache. Standard multi-head attention gives every query head its own key and value projections, so the cache stores a separate set of keys and values per head. Multi-Query Attention (MQA), proposed in 2019, takes the opposite extreme: all query heads share a single key-value head, which shrinks the cache sharply but can degrade quality [21]. Grouped-Query Attention (GQA), introduced in 2023, interpolates between the two by dividing the query heads into groups, with each group sharing one key-value head [22]. By tuning the number of groups, GQA captures most of MQA’s memory savings while holding quality close to full multi-head attention, and it has become the standard attention design in modern open-weight model families, including the Llama, Mistral, and Qwen lines [22].

The infrastructure consequence ties directly back to KV cache sizing. Because GQA reduces the number of key-value heads, it reduces the per-token KV cache footprint by the same proportion, which eases exactly the memory-bandwidth pressure that bounds the decode phase. A model built with GQA serving a long context consumes materially less cache memory than an equivalent multi-head model, and that difference changes how many concurrent requests a given GPU can support. When sizing inference hardware, confirm which attention scheme a candidate model uses, because the KV cache sizing considerations depend on the key-value head count the model actually ships with, not a generic worst case.

Fine-Tuning

Fine-tuning continues training on a pre-trained model using a smaller, specialized dataset. The model retains its general capabilities and acquires domain-specific behavior: vocabulary, formatting conventions, reasoning patterns, internal terminology. A fine-tuned model on proprietary internal data routinely outperforms a general model on domain-specific tasks. It also produces a private asset that no competitor and no vendor can replicate.

Fine-tuning is one of the primary strategic justifications for on-premises infrastructure. Sending proprietary training data to a third-party fine-tuning API surrenders the very advantage that fine-tuning creates. The resulting model and the differentiated capability it represents live on the vendor’s hardware, governed by the vendor’s terms. The on-premises path produces the same artifact under the organization’s own control.

The computational barrier to fine-tuning has fallen dramatically due to two techniques that make the process feasible on the same inference hardware an organization already operates. The first is LoRA (Low-Rank Adaptation), introduced by researchers at Microsoft in 2021 [23]. LoRA freezes the pre-trained model weights and injects trainable rank decomposition matrices into each layer of the Transformer architecture, greatly reducing the number of trainable parameters for downstream tasks. Rather than retraining billions of weights, LoRA trains only a small set of adapter matrices that are added on top of the frozen base model. Compared to full fine-tuning of a 175-billion-parameter model, LoRA can reduce the number of trainable parameters by 10,000 times and the GPU memory requirement by 3 times, while performing on par with or better than full fine-tuning on downstream tasks — and, unlike other adapter approaches, it introduces no additional inference latency.

The second technique is QLoRA, which extends LoRA by combining it with quantization [24]. QLoRA backpropagates gradients through a frozen, 4-bit quantized pretrained language model into LoRA adapters, reducing memory usage enough to fine-tune a 65-billion-parameter model on a single 48-gigabyte GPU while preserving full 16-bit fine-tuning task performance. The practical consequence is that fine-tuning a production-scale model no longer requires a dedicated cluster of high-end GPUs running for weeks. A single server-class GPU — the same hardware provisioned for inference — can produce a domain-adapted model in hours or days. For organizations whose strategic case for on-premises infrastructure rests on owning a fine-tuned model trained on proprietary data, LoRA and QLoRA are the techniques that make that case economically viable at the hardware scales this paper recommends.

Foundation Models

A foundation model is a large neural network trained on broad, general-purpose data at massive scale, designed to be adapted to a wide range of downstream tasks rather than built for any single one. The term was coined by researchers at Stanford’s Center for Research on Foundation Models in 2021 and has since become the organizing concept behind every major AI product in production today. LLMs are one category of foundation model. Multimodal models that process images, audio, and video alongside text are another. The unifying characteristic is the training strategy: expose the model to enormous quantities of diverse data, let it learn general representations of language, reasoning, and structure, then adapt those representations to specific tasks through fine-tuning, prompting, or retrieval.

The strategic importance of this concept for infrastructure planning is that foundation models are not built; they are procured or adopted. No organization under a few thousand employees will train a foundation model from scratch. The compute required runs to tens of millions of dollars and months of GPU-hours at scale. The practical decision is which foundation model to adopt as the base for fine-tuning and deployment, and whether that base model is open-weight (self-hostable, fine-tunable on internal data, privately governed) or closed-source (API-only, vendor-governed, data leaving the organization on every request). Every capability argument in this paper is an argument about which foundation model sits at the center of the deployment and under what terms the organization controls it.

Multimodal Models

A multimodal model is a neural network that processes and generates across more than one data type within a single architecture. Text-only LLMs accept text and produce text. Multimodal models accept combinations of text, images, audio, video, and structured data, and can generate across those modalities as well. The current frontier models (GPT, Claude, Gemini, and Llama 4’s multimodal variants) are all multimodal to varying degrees, accepting text and images as inputs as standard capability.

The infrastructure implications differ from text-only deployment in two ways. First, image and video inputs consume significantly more tokens than equivalent text, accelerating KV cache growth and context window consumption. A single high-resolution image can occupy hundreds to thousands of tokens depending on the model’s visual encoding approach. Second, the EU licensing restriction on Llama 4’s multimodal variants, discussed in the open-weight subsection below, does not apply to its text-only variants. The distinction between multimodal and text-only is therefore not only a capability question but a legal one for organizations operating under EU jurisdiction. For organizations whose use cases center on document analysis, image understanding, or mixed-media workflows, multimodal capability is a first-tier selection criterion, not an upgrade path.

Mixture of Experts

A Mixture of Experts (MoE) model is a neural network architecture in which the standard dense feed-forward layer inside each Transformer block is replaced by a collection of parallel sub-networks (the experts) alongside a lightweight router (or gating network) that selects which experts process each token. MoE models dynamically route input tokens to the most relevant experts based on their characteristics, rather than passing every token through the entire set of parameters. This design leads to sub-linear scaling of the floating-point operations required as model size increases, thereby significantly reducing the computational cost relative to a dense model of the same total parameter count. The practical consequence of this design is a separation between two figures that are often conflated: total parameters and active parameters. Total parameters describes everything stored in GPU memory; all experts combined. Active parameters describes only the subset of the network that executes during any given token’s forward pass. Mixtral 8x7B illustrates this concretely: each token has access to 46.7 billion total parameters, but only 12.9 billion active parameters are used during inference, because the router selects only two of eight available expert blocks at each layer. The model outperforms or matches Llama 2 70B and GPT-3.5 across evaluated benchmarks despite this selective activation [25]. The MoE architecture in its modern form was established by Google’s Switch Transformers work in 2022, which demonstrated that routing each token to a single expert was sufficient to achieve strong scaling properties, enabling models with over a trillion total parameters to be trained at a fraction of the compute required by an equivalent dense model [26].

The hardware implications of MoE architecture are distinct from those of dense models and must be understood before any procurement decision involving MoE models. The full set of expert weights must reside in GPU memory simultaneously, because the router cannot predict at serving time which experts any given token will require. A MoE model with 400 billion total parameters therefore demands GPU memory provisioned for 400 billion parameters, comparable to a dense model of that size, even though the compute per token is closer to that of a 50-billion-parameter dense model. This means that while MoE reduces the per-token compute cost, it does not reduce the memory footprint: all expert weights must be loaded and available for routing, even if only a small fraction are active for any given token. The result is that MoE models offer a favorable throughput-per-compute profile but an unfavorable memory-per-capability profile compared to dense models of equivalent quality. For on-premises deployments, this distinction matters directly: an organization evaluating a MoE model must size GPU memory for the model’s total parameter count, then budget inference compute against the model’s active parameter count. Confusing the two figures produces hardware specifications that are either over-provisioned on compute or, more dangerously, under-provisioned on memory, rendering the deployment inoperable before the first user request arrives.

Small Language Models

A small language model (SLM) is a language model compact enough to load and serve inference on a single consumer-class device rather than a datacenter cluster. No parameter count fixes the boundary. One widely cited survey scopes SLMs as decoder-only Transformers between 100 million and 5 billion parameters [27]. NVIDIA researchers instead anchor the definition to deployment, calling a model small when it fits on a common consumer device and serves a single user at practical latency; noting that as of 2025 they would treat most models under 10 billion parameters as small [28]. The deployment-based definition is the one that matters for infrastructure planning, because the property an organization actually cares about is whether the model runs on the hardware in front of it.

SLMs earn a decision-maker’s attention because the capability-per-parameter curve has moved sharply in their favor. Microsoft’s Phi line established that training data quality, not raw scale, sets much of a small model’s ceiling: the Phi-3 technical report presents Phi-3-mini, a 3.8-billion-parameter model, as reaching quality comparable to GPT-3.5 while running on a phone [29], and its successor Phi-4-mini, also 3.8 billion parameters and built with grouped-query attention, matches the performance of models twice its size on the math and coding tasks that require multi-step reasoning [30]. The pattern holds across vendors. Alibaba reports that its 4-billion-parameter Qwen3 model, released under the Apache 2.0 license, performs comparably to the prior generation’s 72-billion-parameter Qwen2.5-Instruct model [31]. Google’s Gemma line spans a 1-billion-parameter text model built for on-device use up to larger variants designed to run on a single consumer GPU or TPU host [32], and the current Gemma 4 generation, released under Apache 2.0 in April 2026, adds effective-2-billion and effective-4-billion variants aimed at mobile, edge, and browser deployment [33]. A well-chosen small model in 2026 does work that a mid-tier LLM did in 2024.

The strategic argument for SLMs is sharpest in agentic systems. NVIDIA researchers argue, in a 2025 position paper, that small models are not merely adequate but preferable for the majority of agentic work, on three grounds: they are already capable enough for the narrow, repetitive, format-constrained tasks that dominate agent workloads; they are more operationally flexible, because a specialized model can be fine-tuned overnight on a few GPU-hours rather than over weeks; and they are cheaper to serve, which the authors estimate at roughly 10 to 30 times lower inference cost than a 70-to-175-billion-parameter model for a 7-billion SLM, presented as an order-of-magnitude figure drawn from mixed sources rather than a controlled benchmark [28]. The paper stops short of claiming SLMs displace LLMs outright. It advocates heterogeneous systems in which small specialized models handle the bulk of invocations and a larger model is called only when a task genuinely needs open-ended reasoning [28]. This is a live position, not settled consensus, and the same paper states the standing counter-argument fairly: scaling laws give a larger model of the same generation a durable edge in general language understanding, an edge that matters whenever a task cannot be cleanly decomposed into narrow steps [28].

For on-premises deployment the consequences are direct and favorable. An SLM at 4-bit quantization occupies a few gigabytes of memory, which places it within reach of a single consumer GPU, an integrated neural processing unit (NPU), or CPU-only inference through the llama.cpp path described in the inference engine comparison. The KV cache that bounds concurrency shrinks with the model, and the grouped-query attention standard in these families shrinks it further, so a modest GPU serves more concurrent SLM sessions than the same card could serve of a frontier model. The fine-tuning economics compound the effect: the LoRA and QLoRA methods covered in the fine-tuning subsection bring specialization of a small model within a few GPU-hours on the same hardware provisioned for inference, so an organization can maintain a library of task-specialized models rather than paying to run one large generalist for every request. That pattern maps onto the multi-user scaling and agentic infrastructure this paper develops. Route the narrow, high-volume steps to on-premises SLMs and reserve a larger open-weight model, or a metered API call, for the fraction of requests that need it.

The boundary is real and worth stating plainly. A small model fine-tuned for a defined task can match a much larger model on that task, but it does not carry the broad world knowledge or the open-domain reasoning of a frontier model; pushing an SLM past the task it was shaped for exposes that gap quickly. The deployment question is therefore not whether an SLM is as capable as an LLM in general, which it is not, but whether the specific job in front of it is narrow enough that a small specialized model clears the bar at a fraction of the hardware and operating cost. For a large share of the classification, extraction, routing, and structured-generation work inside a typical organization, it is, and the model-selection chapter treats capable open-weight SLMs as Phase 1 candidates for exactly that reason.

Reasoning Models and Adjacent Architectures

Three developments sit alongside the size and architecture categories above and change how a deployment is sized, even though none has displaced the standard Transformer as the default.

Reasoning models trade inference-time computation for accuracy. Instead of emitting an answer directly, the model generates an extended internal chain of intermediate steps before its final response, a technique OpenAI introduced in its o1 model in 2024 [34] and DeepSeek showed could be induced through reinforcement learning alone, without human-annotated reasoning traces, in its R1 model in 2025 [35]. The capability gain on math, coding, and multi-step logic is real. The infrastructure cost is equally real and easy to under-budget: a reasoning model can emit thousands of hidden reasoning tokens per query, each of which is a decode step that consumes memory bandwidth and KV cache. A workload that shifts from direct-answer to reasoning-mode generation can multiply its per-query decode load severalfold, which lands squarely on the decode-bound bottleneck identified in the prefill and decode discussion. Several open-weight families now expose reasoning as a switchable mode, so one model serves both profiles and the operator, not the vendor, decides when the extra tokens are worth their cost.

Distillation is the training technique behind most capable small models. A large, high-quality teacher model generates outputs that a smaller student model is trained to reproduce; this transfers behavior the student could not have learned as efficiently from raw data [36]. Its importance here is that distillation is how reasoning capability reaches deployable sizes: DeepSeek released a family of distilled models from 1.5 to 70 billion parameters, built on Qwen and Llama backbones and trained on reasoning traces generated by R1, several of which carry much of the parent model’s reasoning ability at a fraction of its footprint [35]. For an organization, distillation is the reason a task-specialized small model can inherit the strengths of a frontier model it could never afford to run in production.

State-space models are the most credible current challenger to attention itself. The Mamba architecture, introduced in 2023, replaces the quadratic-cost attention mechanism with a selective state-space recurrence whose cost grows linearly with sequence length, and reports up to 5 times the inference throughput of a comparable Transformer on long sequences [37]. Pure state-space models trail Transformers on tasks that demand precise recall from long context, so the production pattern to date is hybrid: architectures that interleave state-space layers with a smaller number of attention layers, as in NVIDIA’s Nemotron-H and Hymba small models, keep most of the efficiency while restoring the recall [28]. These architectures remain the exception rather than the default. For long-context, high-throughput workloads on constrained hardware they are worth evaluating, because the memory-bandwidth pressure that bounds decode on a Transformer is exactly what the linear-time formulation relieves.

Embeddings

An embedding is a dense numerical representation (a vector) of a piece of text, image, or other input. Two pieces of content with similar meaning produce embeddings that are close together in vector space. This is what allows a system to find “documents about contract termination” without those exact words appearing in the documents, and to surface relevant content whose phrasing differs from the query.

Embeddings are computationally cheaper than LLM inference but benefit from GPU acceleration at scale, particularly when indexing large corpora or supporting high query volume. They are the connective tissue beneath Retrieval-Augmented Generation.

Retrieval-Augmented Generation

Retrieval-Augmented Generation (RAG) is the pattern in which a model’s response is informed by content retrieved from an external knowledge base at query time. The user asks a question, the system retrieves the most relevant passages from a vector store of embedded documents, and the model generates its response with those passages in its context window. The model never has to “learn” the proprietary content. It reads it at inference time.

RAG and fine-tuning are complementary, not competitive. Fine-tuning teaches the model how to behave in the domain: terminology, format, reasoning style. RAG provides the current facts the model needs to answer a specific question. A mature on-premises deployment uses both: a fine-tuned base model that knows the organization’s language, paired with a retrieval layer that grounds every answer in the current document set. The combination of private model, private data, and private retrieval is the architectural endpoint the rest of this paper is engineered to produce.

Hallucination

A hallucination is a model output that is factually wrong, fabricated, or unsupported by evidence, presented with the same confidence as a correct answer. The model does not know it is wrong. It has no internal fact-checking mechanism; it generates the most statistically plausible continuation of the text, and sometimes the most plausible continuation is false. The term is widely used in the industry but is technically imprecise. What the model produces is better understood as a confident confabulation, not a perceptual error.

Hallucination rates vary significantly by model, task, and measurement method. Independent benchmarks show that even top-tier frontier models hallucinate on specific task types while appearing reliable on others. A model that scores near-zero on summarization faithfulness can simultaneously hallucinate on a substantial portion of open-domain knowledge questions. No current model has eliminated hallucination. The question for deployment planning is not whether a model will hallucinate, but how often, on which task types, and what controls are in place to catch it.

The infrastructure implications are direct. RAG reduces hallucination on questions about specific documents by grounding the model’s response in retrieved text it can actually cite. Fine-tuning on domain data reduces hallucination on domain-specific terminology and conventions. Guardrail layers (software that intercepts model output before it reaches the user and checks for confidence thresholds, out-of-scope responses, or policy violations) add a detection layer independent of the model itself. None of these eliminates the risk; all of them are standard practice in production deployments handling consequential decisions. Organizations deploying AI for any workflow with downstream legal, financial, or operational consequences need an explicit hallucination mitigation strategy before the system goes live, not after the first failure.

Quantization

Quantization is the process of reducing the numerical precision used to store a model’s weights (its internal parameters). A model trained at full precision stores each weight as a 32-bit floating-point number (FP32) within four 8-bit bytes. Serving that model in production at 16-bit (FP16 or BF16) cuts memory consumption roughly in half with negligible quality loss. Pushing to 8-bit (INT8) or 4-bit (INT4) can reduce memory by 4× to 8× relative to full precision, at progressively greater quality trade-offs.

The practical consequence is that quantization determines which models fit on which hardware. A 70-billion-parameter model at FP16 requires roughly 140 GB of GPU memory (70B parameters × 16-bits-per-param ÷ 8-bits-per-byte = 140 GB), beyond the capacity of most single GPUs currently available. The same model at 4-bit quantization fits in approximately 35–40 GB, within the reach of a single high-end server GPU. Quantization is therefore not a performance optimization applied after hardware procurement; it is a variable that shapes the procurement decision itself. A server specified to run a quantized model cannot necessarily run the same model at full precision, which matters when fine-tuned model quality depends on precision level. Infrastructure planning must specify both the target model and the target quantization level together, not independently.

A distinct quantization methodology worth understanding is Activation-aware Weight Quantization (AWQ), developed at MIT and awarded Best Paper at MLSys 2024. Where GGUF-based quantization schemes apply uniform or block-wise compression across all weight tensors, AWQ operates on a different principle: not all weights in a neural network contribute equally to output quality, and compressing them equally is therefore a precision-allocation mistake. AWQ is based on the observation that weights are not equally important: protecting only 1% of salient weights can greatly reduce quantization error. The method searches for the optimal per-channel scaling that protects these salient weights by observing the activations produced during inference, not the weights themselves — identifying which weight channels consistently drive large-magnitude activations and therefore carry disproportionate influence over the model’s output. The activation-guided approach allows AWQ to reach 3-bit and 4-bit compression with quality preservation that standard uniform quantization at the same bit-width cannot match [38]. Critically, AWQ does not rely on any backpropagation or reconstruction during the quantization process, so it preserves the model’s generalization ability across different domains and modalities without overfitting to the calibration dataset. The practical distinction from GGUF-based quantization is deployment context: AWQ produces GPU-native quantized weights optimized for CUDA tensor core execution, and is the format of choice when serving quantized models at production throughput on NVIDIA hardware through engines such as vLLM, SGLang, and TensorRT-LLM. GGUF remains appropriate for CPU-capable and consumer-hardware deployments. For organizations running GPU-based inference servers at scale, AWQ-quantized model weights — which are pre-computed and available for most major open-weight model families on Hugging Face — represent a path to 4-bit memory efficiency without the quality penalty that equivalent GGUF compression at aggressive quantization levels can introduce.

Model File Formats

A model file format is the binary container in which a trained model’s weights, architecture metadata, and configuration are packaged for distribution and deployment. The format determines which inference engines can load the model, which hardware it can run on, and what security posture the deployment carries. Two formats dominate the current open-weight landscape and any infrastructure decision will touch at least one of them.

GGUF (GPT-Generated Unified Format) is the most widely encountered format for locally deployed, quantized models. GGUF is a binary format designed for fast loading and saving, ease of reading, and self-contained model description. It is the successor to earlier formats and is designed to be unambiguous, containing all the information needed to load a model. It is also extensible, so that new information can be added to models without breaking compatibility. The format was developed by Georgi Gerganov, the creator of llama.cpp, and has become the standard exchange format for the open-weight model ecosystem. Unlike tensor-only file formats such as safetensors, GGUF encodes both the model’s weights and a standardized set of metadata [39].

The internal structure of a GGUF file is sequential and self-describing. A GGUF file is organized into four major logical regions: a fixed-length header containing a magic number, format version, tensor count, and metadata key-value pair count; a variable-length metadata section containing key-value pairs encoding model configuration; a tensor info section describing each weight tensor’s name, shape, data type, and byte offset; and a bulk tensor data section containing the quantized weight values themselves, aligned to a 32-byte boundary [39]. This self-contained design means a single .gguf file carries everything an inference engine needs to load and run the model without external configuration files.

GGUF defines various quantization types directly within its specification, facilitating efficient memory mapping and loading. The practical consequence is that a quantization level (Q4_K_M, Q5_K_M, Q8_0, and so on) is encoded into the file itself, eliminating ambiguity about what precision the weights were stored at. When a vendor or repository lists a model in GGUF format, the quantization suffix in the filename describes exactly the precision configuration inside.

Safetensors is the predominant format for full-precision models distributed through the Hugging Face ecosystem and consumed directly by production training and serving frameworks. PyTorch model weights have historically been saved using Python’s pickle utility into .bin files, but pickle is not secure: pickled files may contain malicious code that executes on load. Safetensors is a secure alternative purpose-built for sharing model weights. The security distinction is not academic. Organizations downloading models from public repositories face a supply-chain risk with any pickle-based format, because loading the file executes arbitrary Python code embedded by whoever created it. Safetensors eliminates that risk through safe deserialization with no hidden side effects [40].

Beyond security, safetensors offers a performance advantage during distributed loading. Lazy loading is supported, which is useful in distributed settings where only some of the tensors need to be loaded. This capability allowed the BLOOM model to load in 45 seconds on 8 GPUs instead of 10 minutes with regular PyTorch weights [41].

The practical distinction for procurement and deployment planning is this: models downloaded from Hugging Face for use in a Python-based serving stack will typically arrive as safetensors files, requiring GPU memory sufficient to hold the full-precision or half-precision weights. The same model converted to GGUF can run on a broader range of hardware, including CPU-only servers, at the cost of the conversion step and potential quality degradation at aggressive quantization levels. Neither format is universally superior; each serves a different operational context.

Agents, Tools, Harnesses, and Orchestration

An agent is an LLM given access to tools (web search, code execution, file access, internal APIs) and controlled by an orchestration loop that lets it plan, act, observe the result, and revise. The OECD’s February 2026 formal paper distinguishes an AI agent (a system exhibiting autonomy, goal-pursuit, perception, and action) from agentic AI as a paradigm (systems designed to operate in open-ended, less predictable environments with greater autonomy than traditional AI) [42]. Industry uses the terms interchangeably. For governance purposes, they should not be.

The supporting vocabulary maps to discrete pieces of infrastructure. A tool is a callable function the agent invokes: a database query, a code interpreter, an API call, a system command. A harness is the software wrapper around the model that provides the system prompt, manages memory across turns, routes tool calls, and handles failure modes. Orchestration coordinates multiple agents working on decomposed subtasks, typically with one orchestrator agent dispatching specialized worker agents in parallel and synthesizing their results.

This is not a future paradigm. Agentic systems are in production at every major technology vendor as of 2026. Anthropic’s Model Context Protocol, the open standard for connecting agents to tools, crossed 97 million installs by March 2026, moving from an experimental specification to foundational infrastructure [43]. Multi-agent architectures are the current production pattern for complex software engineering workflows. The agentic infrastructure and multi-user scaling sections discuss the deployment implications in full.

Harness Engineering

Harness engineering treats the harness defined in the agent vocabulary above as a designed artifact rather than incidental glue code. The discipline covers the instructions, tools, permissions, memory mechanisms, and verification steps that surround a fixed model, on the premise that agent reliability is determined as much by that surrounding structure as by the model itself. Anthropic’s November 2025 engineering report on long-running agents documented the motivating failure mode: an agent working across many context windows operates in discrete sessions, each new session begins with no memory of the one before, and a capable model without structural support loses coherence on any task that outlives a single window [44].

The published pattern resembles a shift handoff between engineers. An initializer establishes the project structure and a machine-readable statement of the goal. The working agent then records its state in artifacts that persist outside the context window, including a progress log written for the next session and a feature list with explicit pass and fail states; in Anthropic’s demonstration, an agent building a full chat application worked from more than 200 such feature descriptions, all initially marked failing, and advanced them only as verification confirmed the behavior [44]. Verification gates keep each session from building on progress a previous session merely claimed.

Tool design is the other half of the discipline. Anthropic’s companion guidance treats the tool interface as the agent’s user experience: parameter names must be unambiguous, error responses should state the corrective action rather than return an opaque code, and tool outputs need pagination or truncation with sensible defaults so a single call cannot flood the window, a constraint Claude Code enforces with a 25,000-token default cap on tool responses [45].

The harness itself runs on ordinary CPU infrastructure, but its design decisions set the load profile the GPU inference tier must absorb. Session restart frequency, compaction policy, tool response caps, and verification retries all shape token throughput, context length distribution, and KV cache pressure. A poorly engineered harness does not surface in monitoring as a software defect; it surfaces as inflated inference cost and degraded task completion, which is why the agentic infrastructure section treats harness design as part of capacity planning rather than an application-layer afterthought.

Loop Engineering

Loop engineering is the outermost layer in the progression that runs from prompt engineering through context engineering and harness engineering, and the first that removes the practitioner from the position of issuing requests. Its subject is the agentic loop: the cycle in which an agent reasons about the next step, acts, observes the result, and repeats until a termination condition is met. The pattern descends from the ReAct work published by researchers at Princeton University and Google, which demonstrated that interleaving reasoning traces with actions and environmental observations outperforms either reasoning or acting alone [46]. Every orchestration framework in the agent vocabulary above runs a variant of this cycle. Loop engineering is the practice of designing that cycle as a system instead of supervising it turn by turn.

The term entered common usage in June 2026 through an essay by Addy Osmani, an engineering director at Google, who framed the discipline as replacing oneself as the person who prompts the agent and designing the system that does the prompting instead [47]. The design surface consists of the elements the model does not supply: the trigger that starts a run, whether a schedule, an event, or a queue of work items; the goal specification with a completion condition that can actually be checked; the verification step that decides whether the condition is met, through test suites, deterministic assertions, or a second evaluator agent; the persistence layer that carries state between runs; and the stop rules that bound cost and terminate a loop pursuing an unachievable goal [47]. The completion condition dominates the rest. A goal verifiable by mechanical means, such as a passing test suite, produces a governable loop. A goal that requires judgment to assess produces an expensive one, and writing the specification that makes such a goal checkable is where the engineering effort concentrates.

The infrastructure consequence is that unattended loops decouple token consumption from human attention. An interactive session generates load bounded by how fast its operator reads and responds; a scheduled loop generates load bounded only by its stop rules. Under metered API consumption that load is an uncapped variable cost that scales with agent activity rather than headcount. Under owned infrastructure it is utilization of capacity already paid for, which strengthens the fixed-cost argument the ROI section quantifies. Loops also shift the workload profile in a direction the serving layer can exploit: runs that repeat against a stable goal specification and system prompt are heavily prefix-shared, precisely the pattern that the prefix reuse and continuous batching techniques in the inference engine comparison reward. The stop rules and verification gates that make loops economical are also the control points that make them auditable, and the governance section returns to them as the accountability boundary for autonomous workloads.

Open Source, Open Weight, Closed Source, and Proprietary Fine-Tuned

This is the most consequential distinction in this section, and the one most frequently mishandled in vendor conversations. The four categories are not interchangeable. The differences determine what an organization can actually do with a model.

Open source. The model weights, training code, training data information, and license terms together satisfy a formal definition of open-source AI. The Open Source Initiative published Version 1.0 of its Open Source AI Definition (OSAID) in October 2024, requiring four freedoms (use, study, modify, share) along with disclosure of training data information, complete training code, and parameters under OSI-approved terms [48]. Very few capable models meet this standard. OLMo from the Allen Institute for AI and Pythia from EleutherAI pass. The OSI explicitly found that Llama 2, Mixtral, Phi-2, and Grok do not [49]. Llama 4, by the same criteria, does not qualify either. Truly open-source models exist but lag the frontier in capability.

Open weight. Model weights are released publicly under a license that permits self-hosting and modification, but the release fails the formal open-source definition. Typically the training data is not disclosed, the training code is not released, or the license imposes specific commercial restrictions. This is the operative category for almost every model an organization will actually deploy on-premises. Meta’s Llama 4 family is the canonical example. The Llama 4 Community License Agreement, effective April 5, 2025, grants royalty-free worldwide rights to use, modify, and distribute the models commercially, subject to a 700-million-monthly-active-user threshold, a “Built with Llama” attribution requirement, and compliance with Meta’s Acceptable Use Policy [50]. One material new term not present in earlier Llama versions: the multimodal Llama 4 models are not licensed to individuals domiciled in, or companies with principal place of business in, the European Union [50]. Organizations with an EU footprint need to confirm this restriction does not block their intended use before standardizing on Llama 4 multimodal variants. Industry terminology continues to diverge here. Meta still describes Llama as “open source,” while Google and Microsoft have agreed to drop the term for models that do not meet the OSI definition. This paper uses “open weight” for technical precision.

Closed source. The model architecture, weights, and training process are proprietary. Access is provided through an Application Programming Interface (API). The current frontier closed-source models (OpenAI’s GPT family, Anthropic’s Claude family, and Google’s Gemini family) are delivered exclusively through the vendor’s API and a small set of cloud partners [51], [52]. The output is usable. The model is not inspectable, not modifiable, not self-hostable. Every request sends data to the vendor. Any fine-tuning capability offered happens on the vendor’s hardware under the vendor’s terms. The capability is rented, not owned.

Proprietary fine-tuning. An open-weight base model further trained on internal data, hosted on internal infrastructure, governed by internal policy. The resulting model is a private asset. No vendor sees the training data. No competitor can replicate the capability without the same data and the same infrastructure. This is the highest-value endpoint of an on-premises AI investment.

The distinction between closed-source and open-weight is the one that decides strategic posture. A closed-source-only strategy means every interaction with the AI system involves sending data to a third party governed by that party’s commercial terms, retention policy, and jurisdictional exposure. An open-weight strategy paired with on-premises infrastructure means the data never leaves, the model can be fine-tuned on proprietary corpora without surrendering them, and the capability becomes a durable asset rather than a recurring expense. Vendor marketing routinely collapses these categories. When a hyperscaler offers “your private model” running on their hardware under their license, that is closed-source delivery with a private endpoint. It is not a proprietary fine-tuned asset in any sense that survives a vendor relationship change.

The rest of this paper treats this vocabulary as operative. The model selection chapter works through the deployable open-weight families and identifies Phase 1 candidates. The closed-source chapter takes apart the vendor delivery model and the structural ceiling it imposes. Hardware, software, and operational chapters then specify what running open-weight models well actually costs. The strategic argument returns to one point at every layer: the choice of where the model runs is also the choice of who owns the capability.

References

  1. A. Vaswani et al., “Attention is all you need,” Advances in Neural Information Processing Systems, Dec. 2017. [Online]. Available: https://papers.nips.cc/paper_files/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa-Paper.pdf. NeurIPS Vol. 30, pp. 5998–6008.; arXiv published 2017-06-12. [Accessed: 10-May-2026]

    WAIA-1 Primary source Back to text

  2. Anthropic, “Effective context engineering for AI agents,” Anthropic Engineering Blog, Sept. 29, 2025. [Online]. Available: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents. [Accessed: 22-Jul-2026]

    WAIA-2 Secondary source Back to text

  3. S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He, “ZeRO: Memory Optimizations Toward Training Trillion Parameter Models,” Proc. Int. Conf. for High Performance Computing, Networking, Storage and Analysis (SC), Nov. 2020. [Online]. Available: https://arxiv.org/abs/1910.02054. Atlanta, GA, USA, pp. 1–16. [Accessed: 22-Jul-2026]

    WAIA-3 Primary source Back to text

  4. D. Patterson et al., “The Carbon Footprint of Machine Learning Training Will Plateau, Then Shrink,” Computer, July 2022. [Online]. Available: https://arxiv.org/abs/2204.05149. Vol. 55, no. 7, pp. 18–28. [Accessed: 22-Jul-2026]

    WAIA-4 Primary source Back to text

  5. C.-J. Wu et al., “Sustainable AI: Environmental Implications, Challenges and Opportunities,” Proc. Machine Learning and Systems (MLSys), 2022. [Online]. Available: https://arxiv.org/abs/2111.00364. Vol. 4, pp. 795–813. [Accessed: 22-Jul-2026]

    WAIA-5 Primary source Back to text

  6. P. Patel et al., “Splitwise: Efficient Generative LLM Inference Using Phase Splitting,” Proc. 51st ACM/IEEE Annu. Int. Symp. Computer Architecture (ISCA), 2024. [Online]. Available: https://arxiv.org/abs/2311.18677. Buenos Aires, Argentina, Jun.–Jul. 2024, pp. 118–132. [Accessed: 22-Jul-2026]

    WAIA-6 Primary source Back to text

  7. Y. Zhong et al., “DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving,” Proc. 18th USENIX Symp. Operating Systems Design and Implementation (OSDI), July 2024. [Online]. Available: https://arxiv.org/abs/2401.09670. Santa Clara, CA, USA, pp. 193–210. [Accessed: 22-Jul-2026]

    WAIA-7 Primary source Back to text

  8. W. Kwon et al., “Efficient memory management for large language model serving with PagedAttention,” Proceedings of the 29th Symposium on Operating Systems Principles (SOSP '23), Association for Computing Machinery, 2023. [Online]. Available: https://arxiv.org/pdf/2309.06180. [Accessed: 10-May-2026]

    WAIA-8 Primary source Back to text

  9. vLLM Project, “vLLM: The high-throughput and memory-efficient inference and serving engine for LLMs,” 2026. [Online]. Available: https://vllm.ai/. [Accessed: 10-May-2026]

    WAIA-9 Primary source Back to text

  10. L. Zheng et al., “SGLang: Efficient execution of structured language model programs,” Proceedings of the 38th International Conference on Neural Information Processing Systems (NeurIPS '24), 2024. [Online]. Available: https://arxiv.org/pdf/2312.07104. [Accessed: 10-May-2026]

    WAIA-10 Primary source Back to text

  11. SGLang Project, “SGLang documentation,” 2026. [Online]. Available: https://docs.sglang.io/. [Accessed: 10-May-2026]

    WAIA-11 Primary source Back to text

  12. ggml-org, “llama.cpp: LLM inference in C/C++,” GitHub. [Online]. Available: https://github.com/ggml-org/llama.cpp. Software repository. [Accessed: 10-May-2026]

    WAIA-12 Primary source Back to text

  13. Ollama, Inc., “Ollama.” [Online]. Available: https://ollama.com/. [Accessed: 10-May-2026]

    WAIA-13 Primary source Back to text

  14. NVIDIA Corporation, “TensorRT-LLM Overview,” NVIDIA Technical Documentation, 2026. [Online]. Available: https://nvidia.github.io/TensorRT-LLM/overview.html. [Accessed: 10-May-2026]

    WAIA-14 Primary source Back to text

  15. Y. Leviathan, M. Kalman, and Y. Matias, “Fast inference from transformers via speculative decoding,” Proceedings of the 40th International Conference on Machine Learning (ICML '23), 2023. [Online]. Available: https://arxiv.org/pdf/2211.17192. pp. 19274–19286. [Accessed: 22-Jul-2026]

    WAIA-15 Primary source Back to text

  16. C. Chen, S. Borgeaud, G. Irving, J.-B. Lespiau, L. Sifre, and J. Jumper, “Accelerating large language model decoding with speculative sampling,” arXiv, Feb. 2023. [Online]. Available: https://arxiv.org/abs/2302.01318. arXiv:2302.01318. [Accessed: 22-Jul-2026]

    WAIA-16 Secondary source Back to text

  17. H. Xia et al., “Unlocking efficiency in large language model inference: A comprehensive survey of speculative decoding,” Findings of the Association for Computational Linguistics: ACL 2024, Association for Computational Linguistics, Aug. 2024. [Online]. Available: https://doi.org/10.18653/v1/2024.findings-acl.456. pp. 7655–7671. [Accessed: 22-Jul-2026]

    WAIA-17 Primary source Back to text

  18. T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré, “FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness,” Advances in Neural Information Processing Systems (NeurIPS), 2022. [Online]. Available: https://arxiv.org/abs/2205.14135. Vol. 35. [Accessed: 22-Jul-2026]

    WAIA-18 Primary source Back to text

  19. T. Dao, “FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning,” Proc. 12th Int. Conf. on Learning Representations (ICLR), 2024. [Online]. Available: https://arxiv.org/abs/2307.08691. [Accessed: 22-Jul-2026]

    WAIA-19 Primary source Back to text

  20. J. Shah, G. Bikshandi, Y. Zhang, V. Thakkar, P. Ramani, and T. Dao, “FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision,” Advances in Neural Information Processing Systems (NeurIPS), 2024. [Online]. Available: https://arxiv.org/abs/2407.08608. Vol. 37. [Accessed: 22-Jul-2026]

    WAIA-20 Primary source Back to text

  21. N. Shazeer, “Fast Transformer Decoding: One Write-Head is All You Need,” arXiv, Nov. 2019. [Online]. Available: https://arxiv.org/abs/1911.02150. arXiv:1911.02150. [Accessed: 22-Jul-2026]

    WAIA-21 Secondary source Back to text

  22. J. Ainslie, J. Lee-Thorp, M. de Jong, Y. Zemlyanskiy, F. Lebrón, and S. Sanghai, “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints,” Proc. 2023 Conf. on Empirical Methods in Natural Language Processing (EMNLP), Dec. 2023. [Online]. Available: https://arxiv.org/abs/2305.13245. Singapore, pp. 4895–4901. [Accessed: 22-Jul-2026]

    WAIA-22 Primary source Back to text

  23. E. J. Hu et al., “LoRA: Low-Rank Adaptation of Large Language Models,” Proc. 10th Int. Conf. on Learning Representations (ICLR), 2022. [Online]. Available: https://arxiv.org/abs/2106.09685. [Accessed: 10-May-2026]

    WAIA-23 Primary source Back to text

  24. T. Dettmers, A. Pagnoni, A. Holtzman, and L. Zettlemoyer, “QLoRA: Efficient Finetuning of Quantized LLMs,” Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https://arxiv.org/abs/2305.14314. Vol. 36. [Accessed: 10-May-2026]

    WAIA-24 Primary source Back to text

  25. A. Q. Jiang et al., “Mixtral of Experts,” arXiv, Jan. 2024. [Online]. Available: https://arxiv.org/abs/2401.04088. arXiv:2401.04088. [Accessed: 10-May-2026]

    WAIA-25 Primary source Back to text

  26. W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,” Journal of Machine Learning Research, 2022. [Online]. Available: https://www.jmlr.org/papers/v23/21-0998.html. Vol. 23, no. 120, pp. 1–39. [Accessed: 10-May-2026]

    WAIA-26 Primary source Back to text

  27. Z. Lu et al., “Small Language Models: Survey, Measurements, and Insights,” arXiv, 2024. [Online]. Available: https://arxiv.org/abs/2409.15790. arXiv:2409.15790. [Accessed: 22-Jul-2026]

    WAIA-27 Secondary source Back to text

  28. P. Belcak et al., “Small Language Models are the Future of Agentic AI,” arXiv, NVIDIA Research, June 2, 2025. [Online]. Available: https://arxiv.org/abs/2506.02153. arXiv:2506.02153v2 [cs.AI], revised Sep. 15, 2025. [Accessed: 22-Jul-2026]

    WAIA-28 Secondary source Back to text

  29. M. Abdin, J. Aneja, H. Awadalla, A. Awadallah, A. A. Awan, and N. Bach et al., “Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone,” arXiv, 2024. [Online]. Available: https://arxiv.org/abs/2404.14219. arXiv:2404.14219. [Accessed: 22-Jul-2026]

    WAIA-29 Primary source Back to text

  30. Microsoft, “Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs,” arXiv, 2025. [Online]. Available: https://arxiv.org/abs/2503.01743. arXiv:2503.01743. [Accessed: 22-Jul-2026]

    WAIA-30 Primary source Back to text

  31. Qwen Team, Alibaba, “Qwen3 Technical Report,” arXiv, 2025. [Online]. Available: https://arxiv.org/abs/2505.09388. arXiv:2505.09388. [Accessed: 22-Jul-2026]

    WAIA-31 Primary source Back to text

  32. Gemma Team, Google DeepMind, “Gemma 3 Technical Report,” arXiv, 2025. [Online]. Available: https://arxiv.org/abs/2503.19786. arXiv:2503.19786. [Accessed: 22-Jul-2026]

    WAIA-32 Primary source Back to text

  33. Google, “Gemma 4 model overview,” Google AI for Developers, 2026. [Online]. Available: https://ai.google.dev/gemma/docs/core. [Accessed: 22-Jul-2026]

    WAIA-33 Primary source Back to text

  34. A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, and A. Low et al., “OpenAI o1 System Card,” arXiv, 2024. [Online]. Available: https://arxiv.org/abs/2412.16720. arXiv:2412.16720. [Accessed: 22-Jul-2026]

    WAIA-34 Primary source Back to text

  35. DeepSeek-AI, D. Guo, D. Yang, H. Zhang, and J. Song et al., “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning,” Nature, 2025. [Online]. Available: https://arxiv.org/abs/2501.12948. Vol. 645, pp. 633–638. doi:10.1038/s41586-025-09422-z. [Accessed: 22-Jul-2026]

    WAIA-35 Primary source Back to text

  36. G. Hinton, O. Vinyals, and J. Dean, “Distilling the Knowledge in a Neural Network,” arXiv, 2015. [Online]. Available: https://arxiv.org/abs/1503.02531. arXiv:1503.02531. [Accessed: 22-Jul-2026]

    WAIA-36 Secondary source Back to text

  37. A. Gu and T. Dao, “Mamba: Linear-Time Sequence Modeling with Selective State Spaces,” arXiv, 2023. [Online]. Available: https://arxiv.org/abs/2312.00752. arXiv:2312.00752. [Accessed: 22-Jul-2026]

    WAIA-37 Secondary source Back to text

  38. J. Lin et al., “AWQ: Activation-Aware Weight Quantization for LLM Compression and Acceleration,” Proc. 7th Conf. Machine Learning and Systems (MLSys 2024), May 2024. [Online]. Available: https://arxiv.org/abs/2306.00978. Santa Clara, CA, USA. [Accessed: 10-May-2026]

    WAIA-38 Primary source Back to text

  39. G. Gerganov and llama.cpp contributors, “GGUF file format specification,” GitHub, 2024. [Online]. Available: https://github.com/ggml-org/ggml/blob/master/docs/gguf.md. [Accessed: 10-May-2026]

    WAIA-39 Primary source Back to text

  40. Hugging Face, “Safetensors,” 2024. [Online]. Available: https://huggingface.co/docs/safetensors/index. [Accessed: 10-May-2026]

    WAIA-40 Primary source Back to text

  41. Hugging Face, “Load safetensors,” 2024. [Online]. Available: https://huggingface.co/docs/diffusers/main/en/using-diffusers/using_safetensors. [Accessed: 10-May-2026]

    WAIA-41 Primary source Back to text

  42. OECD, “The agentic AI landscape and its conceptual foundations,” OECD Artificial Intelligence Papers, OECD Publishing, Feb. 13, 2026. [Online]. Available: https://doi.org/10.1787/396cf758-en. No. 56. [Accessed: 10-May-2026]

    WAIA-42 Primary source Back to text

  43. Anthropic, “2026 Agentic Coding Trends Report,” Jan. 21, 2026. [Online]. Available: https://resources.anthropic.com/hubfs/2026%20Agentic%20Coding%20Trends%20Report.pdf. [Accessed: 10-May-2026]

    WAIA-43 Secondary source Back to text

  44. Anthropic, “Effective harnesses for long-running agents,” Anthropic Engineering Blog, Nov. 26, 2025. [Online]. Available: https://www.anthropic.com/engineering/effective-harnesses-for-long-running-agents. [Accessed: 22-Jul-2026]

    WAIA-44 Secondary source Back to text

  45. Anthropic, “Writing effective tools for agents — with agents,” Anthropic Engineering Blog, Sept. 11, 2025. [Online]. Available: https://www.anthropic.com/engineering/writing-tools-for-agents. [Accessed: 22-Jul-2026]

    WAIA-45 Secondary source Back to text

  46. S. Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models,” Proc. 11th Int. Conf. on Learning Representations (ICLR), 2023. [Online]. Available: https://arxiv.org/abs/2210.03629. [Accessed: 22-Jul-2026]

    WAIA-46 Primary source Back to text

  47. A. Osmani, “Loop engineering,” addyosmani.com, June 7, 2026. [Online]. Available: https://addyosmani.com/blog/loop-engineering/. [Accessed: 22-Jul-2026]

    WAIA-47 Contextual source Back to text

  48. Open Source Initiative, “The open source AI definition — Version 1.0,” Oct. 28, 2024. [Online]. Available: https://opensource.org/ai/open-source-ai-definition. [Accessed: 10-May-2026]

    WAIA-48 Primary source Back to text

  49. Open Source Initiative, “Open source AI — Validation results,” 2025. [Online]. Available: https://opensource.org/ai. [Accessed: 10-May-2026]

    WAIA-49 Primary source Back to text

  50. Meta Platforms, Inc., “Llama 4 community license agreement,” Apr. 5, 2025. [Online]. Available: https://www.llama.com/llama4/license/. [Accessed: 10-May-2026]

    WAIA-50 Primary source Back to text

  51. Anthropic, “System Card: Claude Opus 4 & Claude Sonnet 4,” May 2025. [Online]. Available: https://www.anthropic.com/claude-4-system-card. [Accessed: 10-May-2026]

    WAIA-51 Primary source Back to text

  52. Anthropic, “Models overview,” Claude API documentation, 2026. [Online]. Available: https://platform.claude.com/docs/en/about-claude/models/overview. [Accessed: 10-May-2026]

    WAIA-52 Primary source Back to text

Contents