Section23
Air-Gapped and Critical Infrastructure AI Deployment
Air-gapped AI is a solved problem at the inference layer and an unsolved one at the operational layer. A true air gap permits zero external network connectivity at inference time: no internet, no cloud API calls, no external update feeds. That constraint rules out every cloud-hosted product, including the closed frontier models that lead the Artificial Analysis Intelligence Index [1]. It does not rule out modern AI capability. NVIDIA ships official air-gap deployment workflows for its NIM inference microservices [2]; Los Alamos National Laboratory researchers ran vLLM inference on the laboratory’s DGX pod to evaluate open-weight models for legacy FORTRAN-to-C++ translation, a study cleared for unlimited release [3]; HPE sells a turnkey air-gapped Private Cloud AI product to governments and regulated industries [4] and won a ten-year, $931 million contract to deliver an air-gapped-managed private cloud to the Defense Information Systems Agency [5]. The models run offline today. What does not come for free is the operational model around them, and that is where air-gapped deployments succeed or fail.
Where Air-Gapped AI Earns Its Place
The use cases that justify the architectural overhead are the ones where the data itself cannot leave the facility. CISA’s December 2025 joint guidance, co-authored with eight partner agencies across five countries, is premised on machine learning, large language models, and AI agents already being integrated into operational technology (OT) environments such as pipelines, power plants, and utilities, not merely projected for future use [6]. The practical risk is unsanctioned AI use: operators reach for cloud chat tools for OT-adjacent work whether or not policy permits, and the durable control is a sanctioned alternative inside the air gap rather than a prohibition that staff route around.
The strongest single use case is predictive maintenance on equipment telemetry, because it has the longest operational track record and the best-quantified returns. Two peer-reviewed papers establish that machine-learning methods now drive the discipline: supervised models such as support vector machines and neural networks for fault classification and remaining-useful-life prediction, and deep architectures such as convolutional and long short-term memory networks for complex sensor and time-series data [7], [8]. The program-level economics come from the U.S. Department of Energy, whose Federal Energy Management Program guide reports roughly tenfold return on investment, a 25–30% reduction in maintenance cost, a 70–75% reduction in breakdowns, and a 35–45% reduction in downtime when predictive maintenance replaces reactive maintenance [9]. Those figures predate the current model generation and describe predictive maintenance as a whole rather than LLM-specific deployments; read them as the discipline’s documented ceiling, not an AI-attributable delta. The reason this work belongs inside the air gap is the data. Flow rates, pressure profiles, production volumes, and vibration signatures are process-sensitive telemetry that competitors and adversaries would pay to obtain, and that sensitivity requires local inference.
The same logic extends to anomaly detection on network traffic and sensor streams, to log analysis at the volumes critical-infrastructure systems produce, and to incident-response support grounded in locally stored playbooks. NIST’s April 2026 concept note for a Critical Infrastructure Profile names “AI-enhanced deterministic diagnostic assistants” with traceable, auditable rationales as an expected use case for critical-infrastructure operators [10], which is the category an air-gapped, runbook-grounded incident assistant falls into.
Two further patterns push AI into internal capability development. Operational runbook automation converts documented procedures into agent-executable workflows with tool access scoped to internal systems. Documentation assistants and policy-compliance checking against locally stored regulatory documents share the same value: institutional knowledge applied at scale, with the knowledge base never leaving the facility. Two more are organizational extensions of that pattern rather than externally validated critical-infrastructure use cases, and they belong in the analysis on those terms. The first is facility-specific software development, where air-gapped engineers use a fully on-premises code assistant such as Tabnine, whose inference runs entirely inside the customer network with no cloud fallback [11]. The second is institutional-knowledge capture, where the system becomes the durable record of facility expertise that would otherwise retire with the staff who hold it.
What Changes When the Network Disappears
Four supply-chain constraints reshape every downstream decision, and the order reflects dependency: model selection, agent design, and lifecycle management all assume the supply-chain problem is already solved.
Model weights and updates. Every model deployed in the air gap is downloaded on a connected system, cryptographically verified, and physically transferred to the isolated environment. NVIDIA’s NIM workflow formalizes this in two phases. On a connected host, download-to-cache or create-model-store pre-stages model weights and container assets. The operator transfers the cache by approved physical media or scp to the air-gapped host and runs the NIM container with NIM_MODEL_PATH pointed at local assets and no NVIDIA NGC API key or Hugging Face token set, which blocks outbound calls entirely [2]. Each subsequent update re-runs the procedure. The cost of a model update is therefore measured not in seconds but in the physical-transfer cycle: preparation on a connected system, hash verification, approved-media transit, staging, validation, deployment. For a facility with strict media-handling rules, that cycle runs days to weeks, and it is the binding constraint on model-lifecycle planning.
Local retrieval and knowledge bases. Retrieval-augmented generation (RAG) runs entirely against locally curated knowledge bases, with no real-time web search, no cloud vector databases, and no external knowledge graphs. Execution is straightforward, since open-source vector stores such as Chroma, Qdrant, Milvus, and Weaviate run as ordinary on-premises services with no internet dependency. The procurement implication is the harder part: someone must curate the corpus and keep it current through the same physical-transfer discipline that governs model updates. A regulatory document set that refreshes quarterly is manageable; a real-time threat-intelligence feed is not.
Dependency management. Python packages, CUDA libraries, OS updates, and NVIDIA driver releases do not reach the air-gapped host through a package manager’s default path. Containerized deployment cuts this burden sharply by packaging Python dependencies, CUDA libraries, and inference-framework components inside the image, so transferring the container alongside model weights makes runtime dependency resolution disappear. That is the reason NVIDIA NIM suits air-gapped operation: the container is the dependency-management solution. Non-containerized environments fall back to manual package dependency downloads on a connected machine and offline installation. OS package mirrors and driver updates need a local mirror, typically an internal Ubuntu or RHEL mirror refreshed on a scheduled physical transfer. Driver management is operationally significant because the CUDA driver version must stay compatible with the toolkit version the inference framework requires, and an uncoordinated driver update can break inference until the matching toolkit catches up.
Model provenance and weight verification. Cryptographic verification of model weights is the supply-chain control that matters most, and current practice runs three tiers. The baseline is SHA-256 file-hash verification of every weight file against hashes published by the original developer, with one discipline that decides whether it works: the hashes must come from the original source, such as the developer’s model card or official repository, before the system is air-gapped, never from a third-party mirror afterward. NVIDIA NIM adds a second tier, the cryptographic profile hash (a string such as 09e2f8e68f78ce94bf79d15b40a21333cea5d09dbe01ede63f6c957f4fcfab7b) used both in the download-to-cache command and the NIM_MODEL_PROFILE runtime variable, carrying verification through the deployment lifecycle [2]. The third tier is chain-of-custody documentation: source URL, download timestamp, SHA-256 hash, verification method, physical media, handling personnel, and destination system, recorded in writing. This is established defense-context practice and aligns with the supply-chain controls in NIST SP 800-82 Revision 3 [12]. Encryption applied directly to the weights themselves, such as the tensor-level scheme proposed in the December 2025 CryptoTensors work, remains active research rather than production standard and should not appear in procurement documentation as a current control [13].
Model Selection Inside the Air Gap
Open-weight models are the only option at the weights layer, but the commercial serving infrastructure around them is not similarly constrained. NVIDIA NIM under NVIDIA AI Enterprise licensing, and tools such as Tabnine Enterprise, can operate air-gapped provided the license terms permit offline activation. Verify that eligibility with the vendor before standardizing on any commercial stack: NVIDIA NIM explicitly supports offline activation [2], whereas some commercial software performs periodic online license validation and fails silently in a true air gap [11].
Within the open-weight category, the Phase 1 selection logic from the Open Weight Models section applies directly. For the security posture typical of critical infrastructure and regulated industries, the strongest US-origin options on a single RTX PRO 6000 96 GB are Gemma 4 31B (Apache 2.0) for general reasoning and coding, gpt-oss-120b (Apache 2.0, MXFP4) for the 120B tier with configurable chain-of-thought, and Nemotron 3 Super (NVFP4) where the NIM path is already in the stack. All deploy via vLLM, and none carries the Chinese-origin provenance-review overhead that Qwen and DeepSeek require; discussed in the Security Architecture section. Where auditability is itself the procurement driver, for government contractors and regulated operators that must trace model behavior to specific training data, the Allen Institute’s Olmo 3 is the differentiated choice: its OlmoTrace tooling links outputs to training-data decisions in a way no other open-weight model matches.
A hardware qualifier matters here. In air-gapped environments where GPU infrastructure is genuinely infeasible, whether from supply constraints, facility power and thermal limits, or procurement channels restricted under the International Traffic in Arms Regulations (ITAR), CPU inference on small quantized models is a valid choice for single-user or very-low-concurrency deployments. The CPU-Only Model Deployments section sets the envelope: a modern server CPU with fully populated DDR5 memory channels reaches roughly 5–15 tokens per second on 7B INT4 and 1–7 tokens per second on 70B INT4. That serves a single technician’s lightweight assistant on an air-gapped resource-constrained machine or an analyst’s reference model on a workstation inside a SCIF (Sensitive Compartmented Information Facility); it does not serve multi-user workloads, and the hard performance ceiling means the deployment must be scoped to that limit.
Agents Without the Open Internet
Agent tool access in an air-gapped environment is limited to locally available resources: internal databases, file systems, locally hosted APIs, and code-execution environments. The Model Context Protocol (MCP) supports this without modification because it is transport-agnostic. MCP servers can run on the inference host over stdio, on the local network over the specification’s HTTP transport (Streamable HTTP, which superseded the earlier HTTP-plus-SSE transport), or in a local container, with no external connectivity required [14]. Any locally hosted capability can be exposed as an MCP tool.
The discipline air-gapped agents require is not a protocol feature but an operational control: an explicit allowed-tool list at the harness level. MCP clients may, by default, attempt to discover every server in their environment, so an air-gapped deployment registers only locally hosted, pre-approved servers and constrains the agent to that set through the system prompt. This matters because NIST SP 800-82 Revision 3 cautions that active network scanning can disrupt live ICS (industrial control system) environments [12], a constraint that applies directly to any agent with network-adjacent tool access in an OT context. The OWASP prompt-injection guidance discussed in the Security Architecture section supplies the security rationale; the implementation is a closed allowlist, audited and version-controlled like any other production access control.
Operational Model Management as Release Engineering
The model lifecycle in an air-gapped environment is controlled software release, not ad hoc update. Sculley and colleagues framed the underlying principle: model components accumulate hidden technical debt when versioning, reproducibility, and rollback are not enforced as engineering discipline [15]. The air gap amplifies the cost of skipping it. A model update that misbehaves in production cannot be patched by re-pulling from a vendor registry the way a connected deployment can. Remediation runs the full physical-transfer cycle of preparation, verification, approved-media transit, staging, validation, and deployment, which is days to weeks for facilities with strict media procedures.
The answer has three parts. NVIDIA NIM anchors versions through the nim_runtime_manifest.yaml cache file, generated on first deployment and reused on restart, with deletion forcing regeneration against the next available assets [2]. Self-hosted MLflow runs entirely on-premises and supplies model registry, versioning, and rollback with no external dependency. Pre-production validation becomes a documented procedure: deploy the candidate to a non-production instance, run a defined evaluation suite against locally stored prompts and expected outputs, compare against threshold, then promote or reject, with the suite itself under version control. Rollback procedures get tested before they are needed. A rollback executed for the first time during an incident is not a rollback; it is an experiment.
What the Regulators Have Said and Have Not
No single published government document addresses AI deployment in air-gapped environments specifically as of mid-2026. The applicable governance stack is a combination. NIST AI RMF 1.0 [16] supplies the base risk-management framework. NIST AI 600-1, the Generative AI Profile [17], covers foundation-model risks including supply chain and autonomous-agent behavior. CISA’s December 2025 joint OT guidance [6] addresses AI in operational technology, and NIST SP 800-82 Revision 3 [12] supplies the OT security baseline that governs most environments where air-gapped AI lives. The NIST AI RMF Critical Infrastructure Profile, released as a concept note on April 7, 2026 [10], is the forthcoming guidance that will address AI governance across IT, OT, and ICS for critical-infrastructure operators. The full profile is not yet published, so it functions as forthcoming direction rather than a completed framework.
Organizations deploying AI in isolated environments should adopt the existing stack now and position for the Critical Infrastructure Profile when it lands. CISA’s four principles structure the work directly: understand AI and its OT-specific risks and educate personnel; assess the business case for AI in a given OT use before deploying, and manage the data-security risks at the boundary; establish AI governance; and embed oversight, safety, and security into deployed systems, monitoring, validating, and refining them continuously [6]. SP 800-82’s consequence-based risk approach, which ranks human safety above environmental harm and both above operational and equipment losses, provides the lens for weighing any specific AI use against the OT availability priority that inverts standard IT security ordering [12].
The Honest Frame
Air-gapped AI is architecturally straightforward with open-weight models, local serving, and the release discipline any well-run software environment already practices. The evidence is concrete: a national laboratory evaluating open-weight LLMs on local DGX hardware through vLLM [3], a defense agency standing up an air-gapped-managed private cloud to deliver AI services [5], and a commercial market for fully air-gapped AI developer tooling [11]. The defense-context provenance of these cases matters less than the pattern they validate, and the same stack of open-weight models, NIM or vLLM serving, local RAG, and allowlisted MCP tools scales down to the non-defense enterprise without the classified-environment compliance load.
The hard work is not the inference. It is three disciplines. The supply-chain discipline keeps weights, dependencies, and knowledge bases current through controlled physical transfer. The release-engineering discipline treats every model update as a versioned, validated, rollback-capable release. The agent-harness discipline holds tool access to a closed allowlist audited against the OT baseline. Organizations that staff and process those three deploy AI inside the air gap. Organizations that treat the inference server as the deliverable meet the rest of the iceberg in production. The Phase 1 and Phase 2 hardware decisions in the Phased Hardware Deployment Roadmap section carry over unchanged, and whether that hardware returns value depends entirely on the three disciplines above.
References
-
Artificial Analysis, “Comparison of AI Models Across Intelligence, Performance, and Price,” 2026. [Online]. Available: https://artificialanalysis.ai/models. [Accessed: 12-Jun-2026]
-
NVIDIA Corporation, “Air Gap Deployment for NVIDIA NIM for Large Language Models,” NVIDIA NIM for LLMs Documentation. [Online]. Available: https://docs.nvidia.com/nim/large-language-models/latest/deploy-air-gap.html. [Accessed: 12-Jun-2026]
-
N. R. Ranasinghe et al., “LLM-Assisted Translation of Legacy FORTRAN Codes to C++: A Cross-Platform Study,” Proc. 1st Workshop on AI and Scientific Discovery: Directions and Opportunities (AISD 2025), May 2025. [Online]. Available: https://aclanthology.org/2025.aisd-main.6/. Albuquerque, NM, USA, pp. 58–69. Also arXiv:2504.15424; LANL release LA-UR-25-22376. [Accessed: 12-Jun-2026]
-
Hewlett Packard Enterprise, “HPE Private Cloud Air-gapped Solutions.” [Online]. Available: https://www.hpe.com/us/en/solutions/air-gapped-private-cloud-solutions.html. [Accessed: 12-Jun-2026]
-
Hewlett Packard Enterprise, “HPE Awarded $931M Other Transaction Agreement to Modernize DISA Datacenter,” Business Wire, press release, Nov. 25, 2025. [Online]. Available: https://www.businesswire.com/news/home/20251125602131/en/HPE-Awarded-%24931M-Other-Transaction-Agreement-to-Modernize-DISA-Datacenter. [Accessed: 12-Jun-2026]
-
Cybersecurity and Infrastructure Security Agency et al., “Principles for the Secure Integration of Artificial Intelligence in Operational Technology,” CISA, Dec. 3, 2025. [Online]. Available: https://www.cisa.gov/resources-tools/resources/principles-secure-integration-artificial-intelligence-operational-technology. [Accessed: 12-Jun-2026]
-
R. Muvunzi et al., “Artificial Intelligence and Robotics in Predictive Maintenance: A Comprehensive Review,” Frontiers in Mechanical Engineering, Jan. 2026. [Online]. Available: https://www.frontiersin.org/journals/mechanical-engineering/articles/10.3389/fmech.2025.1722114/full. Vol. 11, art. 1722114. doi:10.3389/fmech.2025.1722114. [Accessed: 12-Jun-2026]
-
T. Bitam, A. Yahiaoui, D. E. Boubiche, R. Martínez-Peláez, H. Toral-Cruz, and P. Velarde-Alvarado, “Artificial Intelligence of Things for Next-Generation Predictive Maintenance,” Sensors, Dec. 2025. [Online]. Available: https://www.mdpi.com/1424-8220/25/24/7636. Vol. 25, no. 24, art. 7636. doi:10.3390/s25247636. [Accessed: 12-Jun-2026]
-
G. P. Sullivan, R. Pugh, A. P. Melendez, and W. D. Hunt, “Operations & Maintenance Best Practices: A Guide to Achieving Operational Efficiency (Release 3.0),” Federal Energy Management Program, U.S. Department of Energy, Aug. 2010. [Online]. Available: https://www1.eere.energy.gov/femp/pdfs/om_5.pdf. PNNL-19634. [Accessed: 12-Jun-2026]
-
National Institute of Standards and Technology, “Concept Note: AI RMF Profile on Trustworthy AI in Critical Infrastructure,” NIST Information Technology Laboratory, U.S. Department of Commerce, Apr. 6, 2026. [Online]. Available: https://www.nist.gov/programs-projects/concept-note-ai-rmf-profile-trustworthy-ai-critical-infrastructure. Updated Jul. 17, 2026. Project status: ongoing; staff: M. Stanley, R. Sheh. [Accessed: 12-Jun-2026]
-
Tabnine, “What It Really Takes to Be Air-Gapped: Inside the Architecture of Secure AI Development,” Tabnine Blog, June 11, 2025. [Online]. Available: https://www.tabnine.com/blog/what-it-really-takes-to-be-air-gapped/. [Accessed: 12-Jun-2026]
-
K. Stouffer et al., “Guide to Operational Technology (OT) Security,” NIST Special Publication 800-82, Revision 3, National Institute of Standards and Technology, Sept. 2023. [Online]. Available: https://nvlpubs.nist.gov/nistpubs/SpecialPublications/NIST.SP.800-82r3.pdf. doi:10.6028/NIST.SP.800-82r3. [Accessed: 12-Jun-2026]
-
H. Zhu, S. Li, Q. Li, and Y. Jin, “CryptoTensors: A Light-Weight Large Language Model File Format for Highly-Secure Model Distribution,” arXiv, Dec. 2025. [Online]. Available: https://arxiv.org/abs/2512.04580. arXiv:2512.04580 [cs.CR]. [Accessed: 12-Jun-2026]
-
Anthropic and contributors, “Model Context Protocol Specification.” [Online]. Available: https://modelcontextprotocol.io. 2024–2026. [Accessed: 12-Jun-2026]
-
D. Sculley et al., “Hidden Technical Debt in Machine Learning Systems,” Advances in Neural Information Processing Systems, 2015. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2015/hash/86df7dcfd896fcaf2674f757a2463eba-Abstract.html. Vol. 28, pp. 2503–2511. [Accessed: 12-Jun-2026]
-
National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI 100-1, U.S. Department of Commerce, Jan. 26, 2023. [Online]. Available: https://nvlpubs.nist.gov/nistpubs/ai/nist.ai.100-1.pdf. doi:10.6028/NIST.AI.100-1. [Accessed: 12-Jun-2026]
-
C. Autio et al., “Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile,” NIST AI 600-1, National Institute of Standards and Technology, July 26, 2024. [Online]. Available: https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.600-1.pdf. doi:10.6028/NIST.AI.600-1. [Accessed: 12-Jun-2026]