Section14

AI Agents and Agentic Infrastructure

Agentic systems are a production capability in 2026, deployable on the hardware this paper specifies, and they impose demands that chat-style inference planning does not capture: each user task multiplies into several model calls, the tool work surrounding those calls runs on CPU rather than GPU, and every tool an agent can invoke is an attack surface. A large language model (LLM) on its own is an engine that processes input tokens (prefill), generates output tokens (decode), and returns to idle. An agent is a layer on top of the LLM that takes an objective, plans a sequence of steps, executes them against external systems, observes the results, and revises when something fails. That leap requires three capabilities at once: context management, tool access, and an orchestration loop. Remove any one and the system reverts to an elaborate chatbot. With all three, it is an actor that pursues multi-step objectives across the systems an organization already runs: writing and submitting code, retrieving and citing internal knowledge, parsing logs and proposing fixes, decomposing work and delegating it.

What an Agent Is, Precisely

Context management is the first requirement: the system prompt, retrieved memory, conversation history, and tool-call results assembled into the model’s working window at each step. Tool access is the second: the set of callable functions the model invokes to read or change the outside world, including database queries, code-execution sandboxes, application programming interface (API) calls, file operations, and web retrieval. The orchestration loop is the third, and the one early prototypes most often lack: the control logic that lets the model plan, act, observe, and revise. The Organisation for Economic Co-operation and Development (OECD) formalized the distinction at the policy level in its February 2026 paper, defining an AI agent as a system that perceives, pursues goals, and acts with a degree of autonomy, and agentic AI as the broader paradigm of systems built to operate in open-ended, less predictable environments [1]. The agent literature treats the plan–act–observe loop as the canonical architecture [2]. The boundary is operational: an LLM wired to a single function is a tool-using chatbot; an LLM inside a loop that can call several tools, judge the outcomes, and try again is an agent.

The orchestration loop is also where the infrastructure cost lives. Every plan–act–observe iteration is another LLM call with published architectures stacking them: staged pipelines add separate model calls for planning, tool invocation, and answer synthesis [3]. Designs built on tree search or multi-agent debate multiply the count further [2]. No published study yields a single defensible multiplier, so this paper does not cite one. The planning range consistent with the documented architectures is roughly three to ten model calls per user action for simple tool-using agents that interleave reasoning and action in the ReAct pattern, and ten to fifty or more for multi-step planning agents with verification steps. The multi-user scaling discussion inherits this range for its capacity calculations. The point here is that agentic capacity cannot be extrapolated from single-turn chatbot benchmarks.

Agent Harnesses and Frameworks

An agent harness is the software wrapper around the model that supplies the system prompt, manages memory across turns, routes tool calls, enforces guardrails, and handles failure modes. The model supplies the reasoning; the harness supplies everything that turns reasoning into supervised execution. An agent framework provides the code-level building blocks for constructing that harness: tool calling, memory structures, and routing logic. Two open-source frameworks dominate the self-hosted production niche this paper targets, with opposite design philosophies. A third path, the custom harness, is one experienced practitioners take more often than framework marketing suggests.

LangGraph reached 1.0 general availability on October 22, 2025 and has issued point releases since under a no-breaking-changes-until-2.0 commitment [4]. LangChain, its developer, reports production deployments at Uber, LinkedIn, and Klarna, the three with published case studies, with Replit and Elastic among the customers named in less public detail [5], [6], [4]. It models an agent as a directed cyclic graph: typed state flows through nodes (agent steps) joined by edges (conditional transitions), with state checkpointed at each node for durable execution. The capabilities that justify its boilerplate are the ones production needs: recovery from the exact node where execution halted, human-in-the-loop interrupts that pause for review, streaming output, multi-agent coordination, and integration with the Model Context Protocol (MCP) tool standard described below. An equivalent simple multi-agent flow takes meaningfully more code in LangGraph than in CrewAI; that extra code is the cost of the control surface. For stateful workflows with non-trivial branching and recovery requirements, LangGraph is the appropriate choice.

CrewAI is a standalone framework with its own runtime, tools, and memory model, built independently of LangChain [7], [8]. It pairs Crews (role-defined agent teams) with Flows (event-driven orchestration), and its own documentation steers production users toward Flows. Its strengths are prototyping speed and ergonomics. Its weaknesses for production are the absence of built-in checkpointing for long-running work, coarser error handling than LangGraph, and limited control over conditional branching between agents. One architectural difference carries a direct cost on owned hardware: CrewAI’s role-based design places each agent’s role, goal, and backstory, plus prior conversation context, into every turn’s prompt [7], where LangGraph passes a node only the state it explicitly reads. The consequence is per-task token overhead that grows with agent count and tool fan-out; with on-premises inference every token is GPU time. The common advice is to prototype in CrewAI and migrate stateful workflows to LangGraph later; the better plan for anything expected to reach production is to skip the round trip and start in LangGraph.

The field expanded quickly around these two. The OpenAI Agents SDK shipped in March 2025 [9], and Google’s Agent Development Kit followed in April 2025 [10]. Microsoft’s Agent Framework reached 1.0 on April 3, 2026; it consolidates AutoGen and Semantic Kernel into a single supported software development kit (SDK) with MCP and agent-to-agent interoperability [11]. None displaces LangGraph with selective CrewAI prototyping in this paper’s stack, but each is a credible default for an organization already standardized on the corresponding vendor’s platform.

Custom harnesses are the third path and, for a first production agent, often the right one. Framework abstractions carry costs that compound in production: measurable token overhead, failure traces that thread through layers the team did not write, and organization-specific integration patterns that fall outside what a framework anticipates. A purpose-built harness (a Python loop, a small state object, direct client calls to the vLLM serving endpoint the server taxonomy specifies, and tool functions written against the organization’s own APIs) gives predictable behavior and transparent failure modes with no abstraction tax. Framework leaderboard positions offer little guidance for this decision. A November 2025 preprint evaluating six agents on 300 enterprise tasks found that agents optimized for accuracy alone cost 4.4 to 10.8 times more than cost-aware alternatives with comparable performance, that task success fell from 60% on a single run to 25% when consistency across eight consecutive runs was required, and, in a 15-expert evaluation, that ranking agents on cost, latency, and reliability alongside accuracy predicted production success far better (ρ = 0.83) than accuracy-only ranking (ρ = 0.41) [12]. The study is a single-author preprint with a small expert panel, so its numbers are directional; the direction matches operational experience. Choose a framework for its control surface and operational fit rather than its benchmark position. The recommended sequence is a custom harness for the first production agent, LangGraph when a second agent justifies shared infrastructure, and CrewAI for prototyping rather than production.

Model Context Protocol

MCP is the standard that made tool-connected agents practical to build and operate. Anthropic released it on November 25, 2024 as an open client–server protocol: an MCP server exposes capabilities (tools, resources, prompts) under a defined schema, and any MCP client, whether an integrated development environment (IDE), a desktop assistant, or an agent harness, consumes them through the same interface [13]. It dissolves the M×N integration problem: before MCP, every one of M model clients needed a bespoke integration with every one of N tools; with MCP, one server for a database, an internal API, or a document store works with every compatible client.

Governance changed in a way that matters for procurement. On December 9, 2025 Anthropic donated MCP to the Agentic AI Foundation, a directed fund under the Linux Foundation co-founded with Block and OpenAI and backed by Google, Microsoft, AWS, Cloudflare, and Bloomberg, which places the protocol under vendor-neutral stewardship [14], [15]. The current specification is version 2025-11-25 [16]. Cross-vendor adoption is broad: OpenAI committed in March 2025 and Google DeepMind in April 2025, official SDKs cover Python, TypeScript, C#, Java, and a half-dozen other languages, and Anthropic reported more than 10,000 active public MCP servers at the time of the donation [14]. Reviewing the year through its Technology Radar, Thoughtworks credited MCP with bringing agentic AI into the mainstream faster than the industry expected [17].

For the on-premises architecture this paper specifies, MCP is the integration layer between the agentic compute server and the rest of the organization’s systems. Each capability the organization wants agents to use (the ticketing system, the document store, the monitoring API, the code repository) becomes an MCP server, and those servers run on the agentic compute server defined in the server taxonomy section. Harnesses reach them through the protocol rather than through bespoke wrappers, so adding a tool to the agent fleet is a server deployment, not a code change in every harness.

MCP’s own specification is blunt about the risk these capabilities carry: it treats tool invocations as arbitrary code execution, instructs clients to consider tool annotations untrusted unless they come from a trusted server, and requires explicit user consent before a tool runs [16]. The protocol states these commitments; enforcing them is the deployment’s job. The controls under the security constraints heading below apply from the first installation, and the security section develops the full enforcement architecture.

Agent Types Relevant to Technical Organizations

Six categories cover the deployable use cases for small to mid-sized technical organizations, each with its own infrastructure dependencies. Each is deployable in 2026, though the public reference base is thicker for some than others.

Coding agents take a task description, write code, run tests, iterate on failures, and create pull requests. GitHub Copilot, Cursor, Continue.dev, and Replit all run production agentic coding systems; Replit built its agent on LangGraph [5]. The dependencies are a code-execution sandbox isolated from production, repository access through an MCP server or direct Git, and test-runner integration. For software development organizations, the high-value case is internal tooling and drafted infrastructure code that engineers review, rather than unsupervised generation of customer-facing production code.

Knowledge and Q&A agents are the most mature agentic pattern, built on retrieval-augmented generation (RAG) pipelines that retrieve from internal documentation and generate cited answers. They depend on the embedding endpoint, the vector database, the RAG pipeline, and access controls that honor document-level permissions. The Phase 1 chatbot that justifies the initial hardware is exactly this pattern.

Document drafting agents generate structured reports, proposals, and specifications from prompts and retrieved context. Published case studies are thinner here than for coding or Q&A agents, so the pattern carries more integration risk than its apparent simplicity suggests. Dependencies are file-system access through MCP, template management, and the same retrieval layer the knowledge agents use.

Sysadmin and troubleshooting agents parse logs, query monitoring systems, and propose or execute remediation. Log analysis and hypothesis generation are reliably deployable today; automated remediation requires a human-in-the-loop gate for anything irreversible, such as restarting a service, rolling back a deployment, or changing a configuration. Dependencies are MCP access to monitoring APIs (Prometheus, Grafana, and the observability backends from the server taxonomy section), infrastructure-management endpoints, and a mandatory approval gate on execution. This is where a small infrastructure team stands to gain the most, and where the security discipline must be strictest.

Data analysis agents run queries against structured sources, generate visualizations, and interpret results. They depend on database access through MCP servers, a Python or SQL execution environment (the same sandbox the coding agents need), and rendering for charts and tables. The production pattern keeps that environment isolated from production-data writes by default.

Orchestration and meta-agents decompose complex tasks, delegate to specialists, and aggregate results. This is the most architecturally demanding category and the heaviest load on the agentic compute server; Microsoft’s Agent Framework 1.0, LangGraph, and CrewAI all support the hierarchical coordination it needs [11]. Its dependencies are all of the above plus an inter-agent message-routing layer and a coordination state store. Orchestration is what turns the individual categories from one-off tools into a connected system.

Infrastructure Requirements Specific to Agentic Workloads

The fact that drives the server taxonomy is that agents generate substantial work outside LLM inference. Tool execution, document parsing, code compilation, web-content retrieval, vector-database queries, batch embedding, JSON (JavaScript Object Notation) serialization for MCP messages, inter-agent routing, and state persistence are all CPU-bound. Anthropic’s engineering work on MCP makes the operational problem concrete: as the number of connected tools grows, loading every tool definition up front and pushing intermediate results back through the model’s context window slows agents and runs up cost. The fix is to load tools on demand, filter data before it reaches the model, and execute logic in code, all of which runs on CPU [18].

Co-locating that work with GPU inference is the failure mode. A GPU server’s CPU complement is sized to feed and coordinate the GPUs. Schedule parallel document parsing and Python subprocesses onto the same cores and they contend with vLLM’s batch coordination and request routing; inference latency spikes under load, and agent-step latency becomes unpredictable. The server taxonomy section resolves this with the dedicated agentic compute server specified for Phase 2: an AMD EPYC 9965 (192 cores) or EPYC 9755 (128 cores), 512 GB to 1 TB of DDR5-6000 ECC memory, four to eight NVMe SSDs, and a 25 GbE link to the inference server. The core count is deliberate. Parsing, retrieval, and orchestration parallelize cleanly across cores, and they draw on the processor’s twelve-channel DDR5 memory system, rated at up to roughly 614 GB/s per socket, for document and vector operations.

State management runs on this server too. LangGraph’s checkpointed graph execution, conversation-history persistence, agent memory across turns, and the orchestration state for multi-agent work all sit in CPU-side storage backed by the storage server’s PostgreSQL and S3-compatible object store. That state grows roughly with active-conversation count and agent-step depth, which is modest at single-agent scale and material at the multi-tenant Phase 2 scale the scaling section budgets for.

Agentic workflows also live with a latency floor that chat does not impose. Model calls within a task are sequential because each step depends on the previous step’s output, so run time stacks call by call, and a non-trivial agent run completes in minutes to hours rather than at the seconds-per-token pace of a chat reply [2]. The floor is architectural; sequential dependencies between calls cannot be parallelized away. Capacity planning for agents therefore targets sustained throughput across the fleet of running tasks with the scaling section treating the inference and agentic workloads as separate budgets.

Security Constraints for Agentic Systems

An agent with tool access can be tricked into taking actions with potential security consequences. Prompt injection, adversarial instructions embedded in a retrieved document, a web page, or a tool response that hijack the agent’s own instructions heads the Open Worldwide Application Security Project (OWASP) 2025 Top 10 for LLM Applications as LLM01:2025. The list’s expanded excessive-agency entry (LLM06:2025) names the agent-specific failure: too many tools, too-broad permissions, or high-impact actions taken without human approval [19]. A January 2026 peer-reviewed review in MDPI’s Information concluded that prompt injection is a fundamental architectural vulnerability of systems that read instructions and data through a single channel, which demands defense in depth rather than any single fix [20]. A concurrent systematization of attacks on agentic coding assistants catalogs 42 distinct techniques across input manipulation, tool poisoning, and protocol exploitation, including lookalike tools that quietly displace trusted ones, and reports that adaptive attacks defeat most published defenses, with success rates above 85% against state-of-the-art protections [21].

CVE-2025-53773, a published Common Vulnerabilities and Exposures entry, ends the theoretical-risk argument. A prompt injection in GitHub Copilot, a production agentic coding assistant, let attacker-controlled text in a source file or README rewrite the workspace settings to enable an auto-approve mode, which disabled the confirmation step gating the agent’s shell commands and opened a path to remote code execution. Microsoft rates it 7.8 (High) on the Common Vulnerability Scoring System and patched it in the August 2025 release [22]. The mechanism is the cautionary detail: the exploit worked by switching off the human-in-the-loop gate.

The mitigations apply from Phase 1 forward. Keep retrieved content and instructions in separate channels so the model does not read a fetched document as a command, and sanitize content before it enters agent context. Constrain behavior with deterministic system prompts and explicit refusal rules. Require human approval before any irreversible action: deleting a file, sending email, deploying code, changing a configuration, moving money. Scope tool permissions to least privilege, with agents acting under caller-scoped identities rather than admin accounts. Require explicit consent before invoking any tool that touches an external system. No single control stops prompt injection; the published consensus is that none can on its own, which is why the controls are layered [20], [21].

The human-in-the-loop gate is the most important of these and the first to erode under deployment pressure. Once an agent handles a class of action competently, the pull to drop the approval step is constant, and CVE-2025-53773 shows what happens when the gate comes off, whether an attacker removes it or an operator does. The gate stays for any action that cannot be reversed, regardless of how many times the agent has performed it correctly. That is the operational form of the argument the security section makes in full.

Held to that discipline, the architecture compounds. Each MCP server added to the fleet is reusable by every agent, each fine-tuned model sharpens every workflow that calls it, and each action class brought under the approval gate widens what agents can safely do. The binding question then stops being whether an agent can be built and becomes how much concurrent agentic throughput the specified hardware sustains once real users and real workflows arrive. Because each agent task spends three to fifty model calls rather than one, that budget is a different calculation from the chatbot case, which is where the multi-user scaling analysis picks up.

References

  1. OECD, “The Agentic AI Landscape and Its Conceptual Foundations,” OECD Artificial Intelligence Papers, OECD Publishing, Feb. 13, 2026. [Online]. Available: https://doi.org/10.1787/396cf758-en. No. 56. [Accessed: 13-Jun-2026]

    AGEN-1 Primary source Back to text

  2. Arunkumar V., G. R. Gangadharan, and R. Buyya, “Agentic Artificial Intelligence (AI): Architectures, Taxonomies, and Evaluation of Large Language Model Agents,” arXiv, Jan. 18, 2026. [Online]. Available: https://arxiv.org/abs/2601.12560. arXiv:2601.12560. [Accessed: 13-Jun-2026]

    AGEN-2 Secondary source Back to text

  3. H. Go and S. Park, “A Study on Classification Based Concurrent API Calls and Optimal Model Combination for Tool Augmented LLMs for AI Agent,” Scientific Reports, July 2025. [Online]. Available: https://doi.org/10.1038/s41598-025-06469-w. Vol. 15, art. 20579. [Accessed: 13-Jun-2026]

    AGEN-3 Primary source Back to text

  4. LangChain, “LangGraph 1.0 Is Now Generally Available,” LangChain Changelog, Oct. 22, 2025. [Online]. Available: https://changelog.langchain.com/announcements/langgraph-1-0-is-now-generally-available. [Accessed: 13-Jun-2026]

    AGEN-4 Primary source Back to text

  5. LangChain AI, “LangGraph Documentation: Agent Orchestration Framework,” LangChain, Inc.. [Online]. Available: https://langchain-ai.github.io/langgraph/. [Accessed: 13-Jun-2026]

    AGEN-5 Primary source Back to text

  6. LangChain AI, “langchain-ai/langgraph,” GitHub. [Online]. Available: https://github.com/langchain-ai/langgraph. GitHub repository. [Accessed: 13-Jun-2026]

    AGEN-6 Primary source Back to text

  7. CrewAI, Inc., “CrewAI Documentation: Introduction and Architecture.” [Online]. Available: https://docs.crewai.com/en/introduction. [Accessed: 13-Jun-2026]

    AGEN-7 Primary source Back to text

  8. CrewAI, Inc., “crewAIInc/crewAI,” GitHub. [Online]. Available: https://github.com/crewaiinc/crewai. GitHub repository. [Accessed: 13-Jun-2026]

    AGEN-8 Primary source Back to text

  9. OpenAI, “New Tools for Building Agents (Agents SDK),” OpenAI, Mar. 2025. [Online]. Available: https://openai.com/index/new-tools-for-building-agents/. [Accessed: 13-Jun-2026]

    AGEN-9 Secondary source Back to text

  10. Google, “Agent Development Kit (ADK),” Google Cloud, Apr. 2025. [Online]. Available: https://google.github.io/adk-docs/. [Accessed: 13-Jun-2026]

    AGEN-10 Primary source Back to text

  11. Microsoft, “Microsoft Agent Framework Version 1.0,” Microsoft DevBlogs, Apr. 3, 2026. [Online]. Available: https://devblogs.microsoft.com/agent-framework/microsoft-agent-framework-version-1-0/. [Accessed: 13-Jun-2026]

    AGEN-11 Secondary source Back to text

  12. S. Mehta, “Beyond Accuracy: A Multi-Dimensional Framework for Evaluating Enterprise Agentic AI Systems,” arXiv, Nov. 18, 2025. [Online]. Available: https://arxiv.org/abs/2511.14136. arXiv:2511.14136. [Accessed: 13-Jun-2026]

    AGEN-12 Secondary source Back to text

  13. Anthropic, “Introducing the Model Context Protocol,” Anthropic, Nov. 25, 2024. [Online]. Available: https://www.anthropic.com/news/model-context-protocol. [Accessed: 13-Jun-2026]

    AGEN-13 Secondary source Back to text

  14. Anthropic, “Donating the Model Context Protocol and Establishing the Agentic AI Foundation,” Anthropic, Dec. 9, 2025. [Online]. Available: https://www.anthropic.com/news/donating-the-model-context-protocol-and-establishing-of-the-agentic-ai-foundation. [Accessed: 13-Jun-2026]

    AGEN-14 Secondary source Back to text

  15. The Linux Foundation, “Linux Foundation Announces the Formation of the Agentic AI Foundation (AAIF), Anchored by New Project Contributions Including Model Context Protocol (MCP), goose and AGENTS.md,” press release, Dec. 9, 2025. [Online]. Available: https://www.linuxfoundation.org/press/linux-foundation-announces-the-formation-of-the-agentic-ai-foundation. [Accessed: 17-Jul-2026]

    AGEN-15 Secondary source Back to text

  16. Model Context Protocol, “Specification, Version 2025-11-25,” Agentic AI Foundation / Linux Foundation, Nov. 25, 2025. [Online]. Available: https://modelcontextprotocol.io/specification/2025-11-25. [Accessed: 13-Jun-2026]

    AGEN-16 Primary source Back to text

  17. Thoughtworks, “The Model Context Protocol's Impact on 2025,” Dec. 11, 2025. [Online]. Available: https://www.thoughtworks.com/en-us/insights/blog/generative-ai/model-context-protocol-mcp-impact-2025. [Accessed: 13-Jun-2026]

    AGEN-17 Contextual source Back to text

  18. Anthropic, “Code Execution with MCP: Building More Efficient AI Agents,” Anthropic Engineering. [Online]. Available: https://www.anthropic.com/engineering/code-execution-with-mcp. [Accessed: 13-Jun-2026]

    AGEN-18 Secondary source Back to text

  19. OWASP Foundation, “OWASP Top 10 for Large Language Model Applications 2025,” OWASP Foundation. [Online]. Available: https://owasp.org/www-project-top-10-for-large-language-model-applications/. [Accessed: 13-Jun-2026]

    AGEN-19 Primary source Back to text

  20. S. Gulyamov et al., “Prompt Injection Attacks in Large Language Models and AI Agent Systems: A Comprehensive Review of Vulnerabilities, Attack Vectors, and Defense Mechanisms,” Information, Jan. 7, 2026. [Online]. Available: https://doi.org/10.3390/info17010054. Vol. 17, no. 1, art. 54. [Accessed: 13-Jun-2026]

    AGEN-20 Primary source Back to text

  21. N. Maloyan and D. Namiot, “Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems,” arXiv, Jan. 24, 2026. [Online]. Available: https://arxiv.org/abs/2601.17548. arXiv:2601.17548. [Accessed: 13-Jun-2026]

    AGEN-21 Secondary source Back to text

  22. National Vulnerability Database and Microsoft Security Response Center, “CVE-2025-53773,” NIST/MITRE, Aug. 12, 2025. [Online]. Available: https://nvd.nist.gov/vuln/detail/CVE-2025-53773. CVSS 3.1 base score 7.8 (High). See also the Microsoft Security Response Center advisory 'CVE-2025-53773: GitHub Copilot and Visual Studio Remote Code Execution Vulnerability.'. [Accessed: 13-Jun-2026]

    AGEN-22 Primary source Back to text

Contents