Section22

Security Architecture for On-Premises AI

Deploying on-premises AI inference moves the security function onto the organization. For most organizations the largest exposure is not the inference server behind the firewall. It’s the online consumer AI tools employees are already feeding confidential data into, known as shadow AI. IBM’s 2025 breach study, conducted by Ponemon Institute across 600 breached organizations, traced one breach in five to exactly that [1]. An on-premises deployment gives the organization control over shadow AI exposure, because a credible internal substitute is the best control that provides unlimited use of secure AI capabilities without the quarterly increasing cloud AI costs from an increase in token consumption.

There are several continually evolving documents and standards that cover various security architecture aspects for AI deployments. The Open Worldwide Application Security Project (OWASP) reissued its Top 10 for LLM Applications on August 4, 2026, as an update to their 2025 edition, where only the top two entries maintained their rank [2], [3]. Its agentic companion list, OWASP Top 10 for Agentic Applications, dates to December 2025 [4]. The National Institute of Standards and Technology (NIST) AI Risk Management Framework (AI RMF) 1.0 from January 2023 names “Secure and Resilient” as a characteristic of trustworthy AI and defers implementation to the NIST Cybersecurity Framework without specifying a topology or a rule set, with NIST stating that 1.0 is itself under revision [5], [6]. NIST AI 100-2e2025 provides the federal attack taxonomy for generative systems and mentions that current mitigations cannot fully prevent the techniques it catalogs [7]. The NIST Special Publication (SP) 800-53 Control Overlays for Securing AI Systems (COSAiS) project provides a series of guidelines for securing AI systems, bridging taxonomy and implementation, and covers generative AI, predictive AI, single-agent systems, multi-agent systems, and security controls for AI developers. As of mid-2026, the project remains in development with only the predictive-AI outline released for comment [8]. The specifications below provide an engineering position consistent with those documents.

Network Segmentation

The inference and training servers sit on an isolated virtual local area network (VLAN) with no direct route to the open internet. External model downloads pass through a controlled proxy that allowlists specific repositories: Hugging Face’s official organization repositories, NVIDIA’s NGC registry, and any vendor sources procurement has approved. In-band management traffic, including the Kubernetes control plane and monitoring telemetry, runs on its own VLAN, separate from the inference data plane. Baseboard management controller (BMC) traffic runs on the physically separate out-of-band network specified in the server taxonomy section. Firewall rules open only the ports the production stack needs: vLLM on 8000 [9], the vector database on 6333–6334 [10], MLflow on 5000 [11], and the object store’s S3 and console ports.

Two important rules increase the deployment’s network security. The first is the proxy allowlist. An organization that permits unrestricted access to online repositories from the inference VLAN has implicitly approved every model and dataset on those repositories, including the ones known to contain malicious payloads [12]. The allowlist is where the supply-chain verification described below is actually enforced. The second is the absence of general-purpose data egress from the inference VLAN, which usually falls under perimeter security, but also provides protections against prompt injection. An agent talked into exfiltrating data has to reach an external destination. Having a data plane with no route out reduces a successful injection to a local failure. User workstations reach GPU servers through a bastion host or VPN-gated jump access, never by direct SSH.

Access Controls for Model APIs

Every inference endpoint requires authenticated access: application programming interface (API) keys at minimum, OAuth, or single sign-on where the identity infrastructure supports it. Role-based access needs to be implemented because not all models deserve the same policy. A general-purpose Gemma 4 31B instance answering routine questions is one tier; a model fine-tuned on customer records, financial models, or HR documents is a separate tier with stricter access and audit requirements. OWASP’s 2026 edition moved Excessive Agency from sixth place to third as LLM03:2026 on the evidence of what has gone wrong in production agentic deployments [2]. The agentic list sharpens the same idea into least agency, which bounds not only what a system can reach but how far it may act before checking back [4]. The logic that governs an agent’s access to tools governs a user’s access to models.

Each agent instance gets its own unique agent identity with minimal privileges and access needed to do the job it is designed to do. An agent that can log in as an administrator, for example, inherits every permission the administrator account holds, resulting in the audit trail capturing actions taken by the account instead of actions taken by the agent. Scoped, per-agent credentials are easy to implement at Phase 1 and are the precondition for attributing any future incident.

Rate limiting prevents resource exhaustion. A user who runs a script that fires ten thousand inference requests degrades service for everyone else. OWASP’s 2026 edition raises Unbounded Consumption to LLM06:2026 and reframes it around cost asymmetry, where a single crafted input can trigger a computation workload out of proportion through extended reasoning, multimodal content, or a chain of tool calls compared to what it cost the sender to craft the input [2]. Token-aware limiting in this case offers better control over request-aware limiting: a token ceiling per user per window, a cap on reasoning effort for untrusted callers, and a circuit breaker at the agent level that halts a run exceeding its budget. Configure them at the API gateway or in the vLLM serving layer. Thresholds are organization-specific; the control is not optional at any multi-tenant scale.

Data Classification and AI Interaction Policy

The defensible policy has three tiers, each mapped to a deployment surface. Public and internal-general data, such as marketing material, general operational documents, and public filings, can be processed by any approved AI system, including cloud AI platforms from vendors the organization has a commercial agreement with. Confidential business data, such as pricing models, competitive analysis, internal financial projections, and source code, should be processed by on-premises AI systems and never shared with non-confidential cloud platforms, whatever the vendor’s terms say. Highly sensitive data, such as HR records, legal-privileged content, customer personally identifiable information (PII), and data regulated under HIPAA or GDPR, should only be processed by restricted-access AI systems with audit logging on every interaction and human review before any output is acted on.

Documentation alone accomplishes very little. Technical enforcement comes from binding the tiers to the access controls at the model API layer, so that the classification decides which endpoint a caller can access. The NIST Generative AI Profile (NIST AI 600-1) treats data-sensitivity classification for AI interactions as a governance control and provides the federal-standard authority for tiering. The specific three-tier policy mentioned above provides general guidelines consistent with the NIST GenAI profile [13].

Shadow AI Policy

Employees are already using public consumer AI tools with data that may be confidential. Netskope’s Cloud and Threat Report from January 2026, drawn from millions of users between October 2024 and October 2025, found 47% of generative AI users still reach these tools through personal, unmanaged accounts, either exclusively or alongside approved ones, with the average organization recording 223 data-policy violations per month, double the prior year [14]. Harmonic Security’s analysis, also from January 2026, looked at 22,458,240 enterprise prompts sent during calendar year 2025 and identified 579,113 sensitive-data exposures across 665 distinct AI tools, of which 16.9%, or 98,034 instances, occurred on personal free-tier accounts where no audit trail exists [15]. IBM and Ponemon put a cost on these incidents: organizations with high shadow AI use averaged $670,000 more per breach, with shadow AI incidents compromising personally identifiable information in 65% of cases against a 53% global average and intellectual property in 40% against 33% [1]. Gartner projects that more than 40% of enterprises will experience a security or compliance incident tied to unauthorized shadow AI by 2030 [16]. Netskope and Harmonic both sell shadow-AI discovery products, so each has an interest in the conclusion its research supports; the figures are worth mentioning because the methods and observation windows are fully disclosed and because the IBM breach study and the Gartner projection arrive at the same conclusion without either depending on the Netskope and Harmonic research. CISA’s December 2025 joint guidance with eight partner agencies treats this as a present problem, noting that generative AI tools are already in use inside pipeline, power, and utility environments and that prohibition of these tools tends to move the behavior out of sight instead of securely protecting against it [17].

The best protection against shadow AI is to provide the organization with secure and capable alternative AI tools the organization controls. Netskope measured personal-account use falling from 78% to 47% in a single year while use of organization-managed accounts rose from 25% to 62% [14]. Provisioning a credible alternative demonstrably shifts behavior, and nothing else in the published record shifts it that far. That is the empirical basis for the policy design below, and it is also the strongest security argument for the on-premises build that the rest of this paper specifies.

The policy needs at least four components. First, an approved-tools list specifying which tools are permitted for which data classifications, with each entry justified by the contract the organization actually holds rather than by the vendor’s marketing posture. A commercial agreement with enterprise data-protection terms may clear confidential business data, while any consumer free tier clears nothing above public and internal-general data. Harmonic’s concentration finding makes this tractable, since six applications accounted for 92.6% of all sensitive-data exposure in its dataset, so a list that governs those six covers most of the risk without attempting to enumerate hundreds of tools [15]. Second, prohibited uses: no source code, customer data, financial data, or anything classified confidential or higher goes to a non-approved non-confidential cloud AI tool, enforced through data loss prevention (DLP) systems, network monitoring, and browser extensions that flag pasting into known AI endpoints. Third, a reporting channel for cases the policy does not cover, whether a Slack channel, an email alias, or a ticket queue. The mechanism is trivial to build, and its absence is a common reason these policies fail, because without it people default to whatever tool is already open. Fourth, enforcement that runs primarily through education and through approved tooling that makes the compliant choice the easy one.

That fourth component is the one most often skipped and the one the evidence most directly supports. An on-premises inference server running Gemma 4 31B that answers most internal-document questions is the operational answer to the consumer chatbot, because the approved alternative covers the work. Education explains why the policy exists; the substitute is what makes following it the path of least resistance. IBM found that only 37% of the breached organizations it studied had any policy to manage AI or detect shadow AI at all, so for most organizations the first move is having a policy in place [1].

Acceptable Use

The acceptable-use policy governs the on-premises infrastructure, where technical enforcement is direct, and approved cloud AI platforms. Workloads that consume AI resources with no institutional value should be prohibited: bulk synthetic-content generation for non-business purposes, speculative or redundant inference, and use of tools to produce content that would breach the organization’s ethics or responsible-use policies. The resource-allocation argument made earlier in the paper provides the rationale, that GPU capacity is finite and was bought with capital, and the acceptable-use policy turns it into an enforceable rule. LLM03:2026 Excessive Agency and the NIST Generative AI Profile both set least privilege and purposeful use as governance principles; these prohibitions are that principle applied to the human side of the inference infrastructure [2], [13].

Model Output Auditing

Log every inference request and response for sensitive model interactions; retaining the logs for a period the organization can defend. The retention window is not an engineering preference: it follows from the records-retention schedule and whatever sector rule applies to the data class involved, and set by whoever owns those obligations rather than copied from a default. Establish a review process that decides which interactions need human review, which get batch sampling, and which are fed into quality-evaluation pipelines.

One finding in the 2026 OWASP edition argues for this layer more directly than anything in the 2025 list did. Misinformation rose to LLM07:2026 from LLM09:2025 due to the newly available incident data. Practitioners voted AI misinformation near the bottom, but the incident record placed it near the top. It is the widest gap in the document of the vote understating a risk [2]. The reason is the shift in what model output now does. A confident wrong answer that stops in a chat window is a quality problem; the same answer feeding a tool call, a generated commit, or a downstream agent becomes a wrong action in a production system [2]. Silent quality degradation is therefore a security concern that can be easily missed by infrastructure monitoring.

The audit commitment and the infrastructure that implements it have to be designed together. The observability layer specified in the server taxonomy section, OpenTelemetry trace logging in Phase 1 and self-hosted Langfuse or Arize Phoenix in Phase 2, is what makes this policy operationally real. Phase 1’s OpenTelemetry instrumentation captures each request with a timestamp, user ID, model version, token counts, and latency, persisted to a local Jaeger or Elasticsearch backend with no cloud dependency. Phase 2 adds structured quality evaluation against locally defined criteria, drift detection across output distributions over time, and per-user cost attribution. An audit commitment without the observability layer is aspirational; observability without an audit policy is an unused capability. OWASP’s agentic list pairs least agency with observability for this reason, on the argument that constraining an agent without seeing what it does is risk reduction no one can verify, and watching an unconstrained agent is surveillance without a control [4]. The stored prompt-and-response record is also the raw material for the observable-behavior audit that the supply-chain review below requires, which is a second reason to implement it before the first model reaches any user.

Retrieval Corpus and Vector Store

An organization’s corpus made available to an AI for retrieval-augmented generation (RAG) is a security asset with the same standing as the model weights and enables most of the value derived from the Phase 1 chatbot. OWASP keeps Vector and Embedding Weaknesses on the 2026 list as LLM09:2026, covering poisoned embeddings, cross-tenant leakage, and inversion attacks that reconstruct source text from stored vectors [2]. To secure the vector database containing the retrieval corpus, Qdrant, pgvector, or whichever engine the deployment selects from the server taxonomy’s shortlist needs its own authentication, its own tenant separation, and its own place in the backup schedule.

Two controls provide most of the protection. The first is enforcing document-level permissions at query time rather than at ingestion time. A retriever that ignores the caller’s entitlements silently collapses the three-tier data classification above into a single tier, because the generation model answers fluently from whatever it was handed and has no way of knowing if the caller was entitled to the retrieved data. The second is provenance on everything ingested. The Airflow ingestion pipeline specified in the server taxonomy section is the enforcement point: each document enters with a recorded source, timestamp, and permission set, with material from sources the organization does not control marked as untrusted before it reaches an embedding.

Corpus poisoning is the reason these controls belong in Phase 1. Chen et al. demonstrated at NeurIPS 2024 that an optimized trigger placed in a retrieval-augmented generation (RAG) knowledge base or within an agent’s long-term memory can reach an 82% retrieval success rate and a 63% end-to-end attack success rate [18]. Their attack needs no model training and no access to the serving infrastructure, only write access to the corpus. The remediation asymmetry is what makes this different from a session-scoped attack: OWASP’s 2026 prompt-injection entry now explicitly covers instructions that persist in memory or a retrieval store, and an injection that survives in the index taints every future query that retrieves it until somebody rebuilds the index [2]. Ending the session fixes nothing. Re-indexing on a defined cadence, with the ingestion checks reapplied on each pass, is the control that does.

Prompt Injection in Agentic Systems

OWASP’s 2026 Top 10 for LLM Applications has prompt injection in first place as LLM01:2026. Along with the typical prompt injection techniques, it also covers cross-modal attacks that hide instructions in images or audio, injections that persist from session to session, and the impact an agent adds once the injected instruction reaches a tool [2], [3]. The split between direct and indirect injection is what matters operationally. Direct injection needs adversarial access to the agent interface, which is real but bounded. Indirect injection just needs an agent to process attacker-controlled external content and is the dominant threat in any agentic system that retrieves documents, fetches web pages, reads email, or consumes tool responses. Greshake et al. established the indirect category in peer-reviewed work at ACM AISec in 2023 [19]. NIST’s finalized adversarial machine learning taxonomy classifies supply chain attacks, direct prompt injection, and indirect prompt injection as the three generative-AI attack classes, stating that the available mitigations do not fully prevent them [7]. A January 2026 peer-reviewed paper in MDPI’s Information reaches the same conclusion from the architecture: a system that reads instructions and data through one channel has a structural vulnerability requiring defense-in-depth approaches rather than singular solutions [20].

Weighed on the incident record alone, prompt injection would have dropped out of the top ten entirely; the practitioner vote is what kept it at number one. OWASP reads the sparse incident record as evidence of defensive effort rather than of shrinking risk, on the argument that a well-known vulnerability class which mature teams work on continuously yields fewer clean exploits in public databases [2]. Incident counts are a poor signal for this threat class and should not lead to fewer resources protecting against prompt injection.

A 2026 systematization-of-knowledge preprint synthesizing 78 studies published between 2021 and 2026 catalogs 42 distinct attack techniques spanning input manipulation, tool poisoning, protocol exploitation, multimodal injection, and cross-origin context poisoning reports that attack success against state-of-the-art defenses exceeds 85% once the attacker adapts to the defense. Of the 18 published defenses it analyzes, most achieve under 50% mitigation against adaptive attacks [21]. The AgentDojo benchmark (NeurIPS 2024), with 97 realistic tasks and 629 security test cases across email, banking, travel, and workspace environments, provides another version of the same picture: existing injection attacks break some of the security properties but not all of them, with current models failing many of the tasks even with no attacker present [22].

What changed between 2024 and 2026 is where the defenses place their enforcement. Filtering and detection layers that sit inside the model’s own reasoning path quickly become outdated against adaptive attackers. The designs that work move enforcement outside the model entirely. CaMeL, from Google DeepMind and ETH Zürich, extracts control flow and data flow from the trusted user query, runs untrusted data through a quarantined model with no tool access, and uses a custom interpreter that tracks provenance and checks capability-based policies before every tool call. CaMeL completes 77% of AgentDojo tasks with provable security against the attack class it defines, against 84% for the same undefended system [23]. The security guarantee costs 7 points of task completion and, by the authors’ own account, requires that somebody write and maintain the policies, which is where the burden lands in practice. StruQ formalizes the narrower version of the same idea at the prompt boundary, treating the query as a structured object so that retrieved content cannot be read as instruction [24]. OWASP’s 2026 edition states the same conclusion as design guidance: the objective is not a model that cannot be fooled but a surrounding system in which a fooled model breaks nothing that matters [2].

The controls that follow are layered because no single one closes the gap. Sanitize retrieved content before it enters agent context, stripping control characters, normalizing encodings, and flagging suspicious instruction patterns. Bracket document content as untrusted data the model is told not to execute. Deny general egress from the inference VLAN, so a successful injection has nowhere to send what it takes. Scope every tool to least privilege under a caller-scoped identity.

Two of those four need a caveat the 2025 OWASP list did not cover. Sanitization and bracketing both depend on instructions the model reads, leading OWASP to rename its System Prompt Leakage entry to Hidden Context Exposure as LLM08:2026 and extending it to cover everything assembled into the context window that the user never sees: tool and function schemas, retrieved policy text, and developer instructions [2], [3]. Its guidance is to design on the assumption that hidden context is discoverable and that nothing placed in the context window is a secret [2]. Bracketing is hardening against attacks and is worth configuring because it raises the effort required by an attacker. However, it is not a boundary, and a design that treats it as one has no boundary at all, which is precisely why the egress rule and the tool scoping sit beside it.

Tools connected through the Model Context Protocol (MCP), which the agentic infrastructure section makes the integration layer between agents and internal systems, earn the same discipline. A tool response is external content and deserves the same skepticism a web page gets, which is why the 2026 agentic risk list treats tool misuse and agentic supply-chain compromise as distinct entries [4]. OWASP’s February 2026 guide to secure MCP server development names what makes these servers different from ordinary APIs: they operate with delegated user permissions across dynamic tool-based architectures and chained tool calls, which increases the potential impact of a single vulnerability. The guide provides best practices for secure architecture, authentication and authorization at the server, strict input validation, session isolation between callers, and a hardened deployment [25]. Session isolation is the one most often missed on a single-node Phase 1 build, where every agent shares a host and nothing separates their tool sessions by default.

The human-in-the-loop is the control most often sacrificed under deployment pressure. Any agent action that cannot be reversed, such as deleting a file, sending email, deploying code, changing a configuration, or moving money, requires explicit human approval before it runs. The published Copilot vulnerability mentioned in the agentic infrastructure section is the cautionary case, because the exploit allowed an attacker to use prompt injection to enable the workspace’s auto-approve setting, effectively disabling the human-in-the-loop verification gate. The defensible policy is to ensure agent actions are reversible.

Supply Chain Security for Open-Weight Models

Organizations obtain open-weight models by downloading the model weight files from a trusted repository. OWASP LLM04:2026 goes over LLM supply chain risks, vulnerabilities, and mitigations, which includes looking at whether the downloaded model artifacts on disk are the files the publisher released [2]. Current practice has three tiers, each addressing a distinct risk.

Integrity verification is the baseline. Every model weight file pulled from Hugging Face or any other repository should be checked against the hash value the original developer published on their official site. Hugging Face’s hf_hub_download() performs this automatically when checksums are published, but it does not re-verify cached files. Air-gapped deployments require explicit manual verification as described in the air-gapped deployment section. Two gaps sit beside the hash. The first is the serialization format. The peer-reviewed CACM analysis of malicious model uploads shows that Python pickle files execute embedded code during deserialization. For this reason, Hugging Face developed safetensors as a secure format to store and load machine learning model weights without the security risks of older formats like Python’s pickle. [12]. The second is the chat template. Work by Pillar Security and Fujitsu Research of Europe showed that the industry standard Jinja2 template bundled inside a GGUF file (a single, self-contained binary file format designed to store and run large language models) is an executable program that runs on every inference call, is something an attacker can modify without touching a single weight, and the resulting backdoor bypasses weight-based integrity checks; evading the automated scans on the public model hub [26]. Organizations running GGUF models for CPU-based inference, such as the air-gapped and CPU-only deployment case, should pin a known-good template at the serving layer instead of trusting the one shipped in the file.

Provenance documentation is the second tier. For every deployed model, the organization records the source URL, download timestamp, hash and verification method, originating organization and any disclosed government or military research affiliations, license terms, and any published behavior research on the specific model. This sits in the procurement record, not a separate compliance file. NIST SP 800-218A, the federal community profile for secure development of generative AI and dual-use foundation models, addresses model-weight integrity verification and provenance documentation under its Protect-the-Software practices and is written for acquirers as well as producers. The framework frames these as recommended practice rather than mandate, but, for anyone selling into government, they function as expected practice [27].

Behavioral review is the third tier; separating real supply-chain security from checkbox compliance. Hubinger et al. showed that backdoors can be trained into large language models and survive the standard safety-training techniques they tested, including supervised fine-tuning, reinforcement learning, and adversarial training. A counterintuitive result in the tested configurations showed that adversarial training taught models to recognize and better conceal their triggers instead of removing them [28]. Follow-on work constructed cryptographically protected backdoors that, under standard cryptographic assumptions, no polynomial-time red-teaming method can detect or trigger even with full white-box access [29]. Both results bound what any review can promise. A behavioral audit raises the cost of hiding something, but it does not certify that nothing is hidden.

The same discipline extends to model weights the organization produces. OWASP’s Data and Model Poisoning entry, LLM05:2026, includes fine-tuning subversion in the 2026 revision, which places the Phase 2 training server inside the supply chain review [2]. A fine-tuned model inherits whatever was included in its training set, and a training set assembled from logged production traffic inherits whatever an attacker managed to put into that traffic. The MLflow experiment tracking and model registry specified in the server taxonomy section provide the controls: dataset versions recorded per run, lineage from artifact back to training data, and a promotion path that can be reversed when a model starts behaving differently after a retraining run.

The Chinese-origin model families that the open-weight model section covers, such as Qwen, DeepSeek, GLM, Kimi, MiniMax, and others, are the prominent current examples that warrant this review, on provenance grounds rather than on any documented backdoor. As of mid-2026 no peer-reviewed publication documents a cryptographic backdoor in a specific model from any of them, and this paper does not imply otherwise. Risks specific to the cloud product do not necessarily apply to self-hosted weights. DeepSeek’s hosted API exposed a ClickHouse database in January 2025 with chat history, API tokens, and backend logs reachable without authentication along with DeepSeek’s terms allowing the storage of user data on Chinese servers under Chinese law [30]. However, a self-hosted, MIT-licensed deployment behind the organization’s firewall isn’t able to transmit anything to the originating vendor. The model’s behavioral risk, however, does remain within the model’s weights, and is why the review has to account for each unique model even if they are all from the same originating country and/or vendor.

Over the last couple years, there have been more studies indicating the increased importance of reviewing model behavior over model origin. In February 2025, MITRE’s OCCULT evaluation found DeepSeek-R1, released in January 2025, to be the first model in its panel to answer more than 90% of the TACTL offensive-cyber knowledge benchmark correctly [31]. An independent May 2026 evaluation of the DeepSeek V4 Pro preview from that April, conducted from public weights and cloud API with no developer cooperation, put the model at 79.5% on the Cybench agentic capture-the-flag benchmark. Compared against 80% for GPT-5.2 released on December 2025, this places offensive-cyber capability three to six months behind the Western frontier [32]. The same evaluation found several methods that push certain models to exhibit harmful behavior. A publicly circulated roleplay jailbreak template from 2023, sent to the model as a single user message with no system-prompt access and no model tuning, drove V4 Pro and V4 Flash’s StrongREJECT jailbreak rate from 0.6% to 77.8% and dropped its refusal rate on the AgentHarm agentic-misuse suite to 0.95%. FAR.AI independently reached 98% to 100% attack success across CBRN, cyber, and terrorism content using a similar approach. Kimi K2.6 and Qwen3.6 Max, both released within the same six months, resisted the same attacks, indicating DeepSeek V4’s weak safeguards [32]. Origin did not predict the outcome nor did it predict behavior under pressure. On Anthropic’s Agentic Misalignment evaluations that are designed to identify a model’s harmful insider-threat-style behaviors, the same evaluation measured deliberate harmful actions in 51.5% of Gemini 3.1 Pro samples and 35.0% of V4 Pro samples, with Claude Opus 4.6 and GPT-5.4 at zero [32]. The provenance review is important because it forces a per-model measurement that guides organizations to make informed decision on the models they should use and the ones they should avoid.

The review for any model in this category documents model origin and corporate structure, government or military research affiliations of the originating institution, published behavior research on the specific model, an observable-behavior audit covering content restrictions and geopolitical framing, and legal counsel’s guidance on applicable procurement regulations. Along with those, there is one additional methodological detail that should be included in the audit design. The same independent evaluation mentioned above found that when using language models to judge pieces of information, judge origin shifted scores on open ordinal rubrics, with Western judges rating identical transcripts at 4.08 on average on a ten-point concern scale against 2.97 for Chinese judges, while binary rubrics showed no significant gap across nine judges [32]. An internal audit that grades on a scale inherits the model’s judging behavior, while one that asks yes-or-no questions does not.

Until the review is signed off, initial production use belongs in non-sensitive, air-gap-isolated contexts. For organizations whose posture does not require this scrutiny, the open-weight models section’s Chinese-origin set remains competitive on capability, with one condition: any deployment where untrusted users can craft prompts should treat DeepSeek V4, in both the Pro and Flash tiers, as effectively unguarded at the model layer and enforce refusal policy outside it. For organizations with government contracts, defense adjacency, or sensitive data classifications, US-origin models such as Google’s Gemma 4 family, NVIDIA’s Nemotron 3 family, and gpt-oss-120b under Apache 2.0 or similar licensing covers a similar capability range without the procurement overhead.

The Architecture as a Whole

Each of these nine security controls depends on the others and should be incorporated into the on-premises design from the very beginning. Access controls without audit logging produce policy compliance no one can verify. Audit logging without an observability layer that captures prompt-and-response pairs records who called the API but not what they asked or what was returned. Supply-chain verification without a network architecture that controls download paths may produce valid checksums on weights pulled through unsanctioned channels. Retrieval-corpus provenance without query-time permission enforcement produces a well-documented index that answers questions the caller had no right to ask with information the caller was not approved to have. A shadow-AI policy without an approved alternative that covers real work produces the personal-account use the policy was meant to stop. Acceptable-use rules without audit infrastructure produce governance that exists only on paper. Prompt-injection defenses without least-privilege tool scoping and an egress-denied data plane leave the model as the only thing standing between an attacker’s text and an irreversible action, which is the arrangement every current source says will fail.

The architecture works when all components are designed together and continues to work through the operational discipline that keeps them current: a versioned policy, a named owner for each control, and a scheduled review that treats security as a standing function. Schedule the reviews to keep up with the changing industry standards and best practices; the OWASP lists reissue annually and the NIST framework and its overlays are currently mid-revision. Logging and observability must be implemented from the very beginning; enabling the audit trail before the first user session. An organization that switches logging on after an incident has no record of what its models did while nobody was watching, the systems that were affected, and any potential damage that was done.

References

  1. IBM Corporation, “IBM Report: 13% of Organizations Reported Breaches of AI Models or Applications, 97% of Which Reported Lacking Proper AI Access Controls,” IBM Newsroom, press release, July 30, 2025. [Online]. Available: https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls. [Accessed: 29-Jul-2026]

    SECAR-1 Secondary source Back to text

  2. OWASP GenAI Security Project, “OWASP Top 10 for LLM Applications 2026,” OWASP Foundation, Aug. 4, 2026. [Online]. Available: https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/. v1.0. [Accessed: 14-Aug-2026]

    SECAR-2 Primary source Back to text

  3. OWASP GenAI Security Project, “OWASP Top 10 for LLM Applications 2025,” OWASP Foundation, Nov. 18, 2024. [Online]. Available: https://genai.owasp.org/resource/owasp-top-10-for-llm-applications-2025/. Superseded by the 2026 edition; kept as the baseline for the ranking and scope changes. [Accessed: 14-Aug-2026]

    SECAR-3 Primary source Back to text

  4. OWASP GenAI Security Project, “OWASP Top 10 for Agentic Applications for 2026,” OWASP Foundation, Dec. 9, 2025. [Online]. Available: https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/. [Accessed: 29-Jul-2026]

    SECAR-4 Primary source Back to text

  5. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, U.S. Department of Commerce, Jan. 26, 2023. [Online]. Available: https://doi.org/10.6028/NIST.AI.100-1. [Accessed: 26-Jun-2026]

    SECAR-5 Primary source Back to text

  6. National Institute of Standards and Technology, “AI Risk Management Framework,” NIST AI Resource Center, 2026. [Online]. Available: https://airc.nist.gov/airmf-resources/airmf/. The page states that AI RMF 1.0 is being updated and a revised version is in progress. [Accessed: 14-Aug-2026]

    SECAR-6 Primary source Back to text

  7. A. Vassilev, A. Oprea, A. Fordyce, H. Anderson, X. Davies, and M. Hamin, “Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations,” NIST AI 100-2e2025, National Institute of Standards and Technology, Mar. 24, 2025. [Online]. Available: https://csrc.nist.gov/pubs/ai/100/2/e2025/final. doi:10.6028/NIST.AI.100-2e2025. [Accessed: 29-Jul-2026]

    SECAR-7 Primary source Back to text

  8. National Institute of Standards and Technology, “SP 800-53 Control Overlays for Securing AI Systems (COSAiS),” NIST Computer Security Resource Center. [Online]. Available: https://csrc.nist.gov/projects/cosais. Concept paper Aug. 14, 2025; predictive-AI annotated outline Jan. 8, 2026. [Accessed: 29-Jul-2026]

    SECAR-8 Primary source Back to text

  9. vLLM Project, “vLLM Documentation,” PyTorch Foundation, 2026. [Online]. Available: https://docs.vllm.ai/en/latest/. [Accessed: 26-Jun-2026]

    SECAR-9 Primary source Back to text

  10. Qdrant, “Qdrant Documentation,” 2026. [Online]. Available: https://qdrant.tech/documentation/. [Accessed: 26-Jun-2026]

    SECAR-10 Primary source Back to text

  11. MLflow Project, “MLflow Documentation,” 2026. [Online]. Available: https://mlflow.org/docs/latest/. [Accessed: 26-Jun-2026]

    SECAR-11 Primary source Back to text

  12. A. K. Sood and S. Zeadally, “Malicious AI Models Undermine Software Supply-Chain Security,” Communications of the ACM, 2025. [Online]. Available: https://cacm.acm.org/research/malicious-ai-models-undermine-software-supply-chain-security/. doi:10.1145/3704724. [Accessed: 29-Jul-2026]

    SECAR-12 Primary source Back to text

  13. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile,” NIST AI 600-1, U.S. Department of Commerce, July 2024. [Online]. Available: https://doi.org/10.6028/NIST.AI.600-1. [Accessed: 26-Jun-2026]

    SECAR-13 Primary source Back to text

  14. Netskope Threat Labs, “Cloud and Threat Report: 2026,” Netskope, Jan. 6, 2026. [Online]. Available: https://www.netskope.com/resources/cloud-and-threat-reports/cloud-and-threat-report-2026. Vendor telemetry, Oct. 2024 to Oct. 2025. [Accessed: 29-Jul-2026]

    SECAR-14 Secondary source Back to text

  15. Harmonic Security, “What 22 Million Enterprise AI Prompts Reveal About Shadow AI in 2025,” Jan. 2026. [Online]. Available: https://www.harmonic.security/resources/what-22-million-enterprise-ai-prompts-reveal-about-shadow-ai-in-2025. AI Usage Index dataset, 22,458,240 prompts, Jan. 1 to Dec. 31, 2025. [Accessed: 29-Jul-2026]

    SECAR-15 Secondary source Back to text

  16. Gartner, Inc., “Gartner Identifies Critical GenAI Blind Spots That CIOs Must Urgently Address,” Gartner Newsroom, press release, Nov. 19, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-11-19-gartner-identifies-critical-genai-blind-spots-that-cios-must-urgently-address0. [Accessed: 29-Jul-2026]

    SECAR-16 Secondary source Back to text

  17. Cybersecurity and Infrastructure Security Agency et al., “Principles for the Secure Integration of Artificial Intelligence in Operational Technology,” CISA, Dec. 3, 2025. [Online]. Available: https://www.cisa.gov/resources-tools/resources/principles-secure-integration-artificial-intelligence-operational-technology. [Accessed: 26-Jun-2026]

    SECAR-17 Primary source Back to text

  18. Z. Chen, Z. Xiang, C. Xiao, D. Song, and B. Li, “AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases,” Advances in Neural Information Processing Systems 37 (NeurIPS 2024), 2024. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2024/hash/eb113910e9c3f6242541c1652e30dfd6-Abstract-Conference.html. [Accessed: 14-Aug-2026]

    SECAR-18 Primary source Back to text

  19. K. Greshake, S. Abdelnabi, S. Mishra, C. Endres, T. Holz, and M. Fritz, “Not What You've Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” Proc. 16th ACM Workshop on Artificial Intelligence and Security (AISec '23), 2023. [Online]. Available: https://dl.acm.org/doi/10.1145/3605764.3623985. pp. 79–90. doi:10.1145/3605764.3623985. [Accessed: 26-Jun-2026]

    SECAR-19 Primary source Back to text

  20. S. Gulyamov et al., “Prompt Injection Attacks in Large Language Models and AI Agent Systems: A Comprehensive Review of Vulnerabilities, Attack Vectors, and Defense Mechanisms,” Information, Jan. 7, 2026. [Online]. Available: https://www.mdpi.com/2078-2489/17/1/54. Vol. 17, no. 1, art. 54. doi:10.3390/info17010054. [Accessed: 29-Jul-2026]

    SECAR-20 Primary source Back to text

  21. N. Maloyan and D. Namiot, “Prompt Injection Attacks on Agentic Coding Assistants: A Systematic Analysis of Vulnerabilities in Skills, Tools, and Protocol Ecosystems,” arXiv, Jan. 24, 2026. [Online]. Available: https://arxiv.org/abs/2601.17548. arXiv:2601.17548 [cs.CR]. Preprint, not peer reviewed. [Accessed: 29-Jul-2026]

    SECAR-21 Secondary source Back to text

  22. E. Debenedetti, J. Zhang, M. Balunović, L. Beurer-Kellner, M. Fischer, and F. Tramèr, “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,” Advances in Neural Information Processing Systems 37 (NeurIPS 2024) Datasets and Benchmarks Track, 2024. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2024/hash/97091a5177d8dc64b1da8bf3e1f6fb54-Abstract-Datasets_and_Benchmarks_Track.html. doi:10.52202/079017-2636. [Accessed: 29-Jul-2026]

    SECAR-22 Primary source Back to text

  23. E. Debenedetti et al., “Defeating Prompt Injections by Design,” arXiv, Google DeepMind and ETH Zürich, June 24, 2025. [Online]. Available: https://arxiv.org/abs/2503.18813. arXiv:2503.18813v2 [cs.CR]. [Accessed: 29-Jul-2026]

    SECAR-23 Secondary source Back to text

  24. S. Chen, J. Piet, C. Sitawarin, and D. Wagner, “StruQ: Defending Against Prompt Injection with Structured Queries,” Proc. 34th USENIX Security Symposium (USENIX Security 25), Aug. 2025. [Online]. Available: https://www.usenix.org/conference/usenixsecurity25/presentation/chen-sizhe. Seattle, WA, USA, pp. 2383–2400. [Accessed: 26-Jun-2026]

    SECAR-24 Primary source Back to text

  25. OWASP GenAI Security Project, “A Practical Guide for Secure MCP Server Development,” OWASP Foundation, Feb. 16, 2026. [Online]. Available: https://genai.owasp.org/resource/a-practical-guide-for-secure-mcp-server-development/. [Accessed: 14-Aug-2026]

    SECAR-25 Primary source Back to text

  26. A. Fogel, O. Hofman, E. Cohen, and R. Vainshtein, “Inference-Time Backdoors via Hidden Instructions in LLM Chat Templates,” arXiv, Feb. 2026. [Online]. Available: https://arxiv.org/abs/2602.04653. arXiv:2602.04653 [cs.CR]. Submitted to the ICLR 2026 Workshop on Trustworthy AI. [Accessed: 29-Jul-2026]

    SECAR-26 Secondary source Back to text

  27. H. Booth, M. Souppaya, A. Vassilev, M. Ogata, M. Stanley, and K. Scarfone, “Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile,” NIST SP 800-218A, National Institute of Standards and Technology, July 26, 2024. [Online]. Available: https://csrc.nist.gov/pubs/sp/800/218/a/final. doi:10.6028/NIST.SP.800-218A. [Accessed: 29-Jul-2026]

    SECAR-27 Primary source Back to text

  28. E. Hubinger, C. Denison, J. Mu, M. Lambert, M. Tong, and M. MacDiarmid et al., “Sleeper Agents: Training Deceptive LLMs That Persist Through Safety Training,” arXiv, 2024. [Online]. Available: https://arxiv.org/abs/2401.05566. arXiv:2401.05566. [Accessed: 26-Jun-2026]

    SECAR-28 Secondary source Back to text

  29. A. Draguns, A. Gritsevskiy, S. R. Motwani, C. Rogers-Smith, J. Ladish, and C. Schroeder de Witt, “Unelicitable Backdoors in Language Models via Cryptographic Transformer Circuits,” arXiv, 2024. [Online]. Available: https://arxiv.org/abs/2406.02619. arXiv:2406.02619, revised Feb. 2025. [Accessed: 26-Jun-2026]

    SECAR-29 Secondary source Back to text

  30. G. Nagli, “Wiz Research Uncovers Exposed DeepSeek Database Leaking Sensitive Information, Including Chat History,” Wiz Blog, Jan. 29, 2025. [Online]. Available: https://www.wiz.io/blog/wiz-research-uncovers-exposed-deepseek-database-leak. [Accessed: 26-Jun-2026]

    SECAR-30 Secondary source Back to text

  31. M. Kouremetis et al., “OCCULT: Evaluating Large Language Models for Offensive Cyber Operation Capabilities,” arXiv, Feb. 18, 2025. [Online]. Available: https://arxiv.org/abs/2502.15797. arXiv:2502.15797 [cs.CR]. MITRE Public Release Case No. 25-0076. [Accessed: 29-Jul-2026]

    SECAR-31 Secondary source Back to text

  32. Neo Research, “Evaluating DeepSeek v4 Pro for Frontier Risks,” Neo Research (Singapore AI Safety Hub), May 29, 2026. [Online]. Available: https://neoresearch.ai/research/deepseek-v4-pro-safety-evaluation/. [Accessed: 29-Jul-2026]

    SECAR-32 Secondary source Back to text

Contents