Section3

Owning the Stack: Geopolitical and Economic Arguments for On-Premises AI

An organization choosing where to put its AI workloads is also choosing how much geopolitical, tariff, and resource-allocation turbulence to price into its operating costs in perpetuity. Three governments restrict who can buy which chips. One foundry in Taiwan fabricates nearly every leading-edge accelerator. Tariff schedules change quarterly. Water tables drop near data center clusters. Labor markets reorganize around who can direct AI systems and who cannot. On-premises deployment is the choice to absorb these forces once, as capital, rather than indefinitely, as someone else’s pricing decision.

Executive reading guide: Points one and two (supply chain timing and tariff exposure) carry the directly actionable procurement argument. Read those plus the closing position. The remaining points provide depth for governance framing, stakeholder communication, and the longer resilience case.

The geopolitical fracture in AI hardware supply

Leading-edge AI silicon is concentrated to a degree without precedent in modern infrastructure. TSMC (Taiwan Semiconductor Manufacturing Company) fabricates nearly every advanced GPU in commercial use, including the NVIDIA H100, H200, and Blackwell generations (B200 and B300). A single fab campus in Hsinchu sits roughly one hundred miles from a strategic competitor with declared territorial claims. This is the substrate of the global AI buildout.

U.S. export controls have been reshaping access to that substrate since October 2022. The Bureau of Industry and Security (BIS) restricted exports of advanced computing chips and semiconductor manufacturing items to China and related entities through Interim Final Rule 87 FR 62186, naming NVIDIA’s A100 and H100 explicitly [1]. NVIDIA responded with the A800 and H800, lower-spec China variants engineered specifically to fall under the performance thresholds. The October 2023 update closed that workaround by adding a performance density threshold; per NVIDIA’s own securities disclosure, the government accelerated the restriction’s effective date to October 23, 2023 [2]. The December 2, 2024 rule added controls on high-bandwidth memory, twenty-four additional categories of semiconductor manufacturing equipment, and 140 Entity List entries [3]. BIS has stated its intention to update these controls at least annually, and the December 2024 Government Accountability Office review documents ongoing compliance challenges with implementation [4]. The controls have since begun cutting in both directions: in January 2026, BIS relaxed licensing for H200-class exports to China to case-by-case review, paired with a new import tariff discussed in the next subsection, an arrangement that changed the economics of the same chips twice in fourteen months [5].

The downstream effect on Western buyers was not a quiet bureaucratic story. Oracle’s Larry Ellison described a dinner at Nobu Palo Alto where he and Elon Musk spent the evening “begging” NVIDIA CEO Jensen Huang for GPU allocation [6]. That is two of the largest buyers on the planet describing their own procurement posture. Allocation friction has not fully unwound; Instrumental’s supply chain analysis reports that many AI components still carry lead times of 36 to 52 weeks, forcing procurement commitments long before pricing or tariff liability can be known [7]. The binding constraint remains TSMC fab capacity meeting concurrent demand from every hyperscaler at once, with export restrictions adding allocation friction on top.

The supposed counterweight to this concentration, domestic and EU fab investment, operates on a timeline that does not help anyone procuring hardware in 2026 or 2027. The CHIPS and Science Act, signed in August 2022, authorized $52.7 billion in manufacturing incentives. By December 2024, the Commerce Department had allocated more than $32 billion of the $39 billion in direct incentives, including major commitments to TSMC’s Phoenix expansion [8]. A joint Semiconductor Industry Association and Boston Consulting Group study projects these investments will grow U.S. share of sub-10nm logic manufacturing from zero percent in 2022 to twenty-eight percent by 2032 [9]. Fabs take three to five years to build and commission. The first meaningful output from CHIPS-funded advanced logic capacity arrives in 2027 at earliest, with the twenty-eight percent figure attached to 2032.

Europe is in a similar position with worse prospects. The EU Chips Act (Regulation EU 2023/1781) entered force September 21, 2023, targeting a doubling of the EU’s global semiconductor market share from ten percent to twenty percent by 2030 [10]. The European Court of Auditors, the EU’s own independent audit body, concluded in April 2025 that the Act is very unlikely to be sufficient to hit the 2030 target, citing insufficient financing and slow implementation [11]. The geographic source of supply is shifting in the official capacity plans. The timing of that shift is not.

This is where the procurement timing argument becomes operative. Hardware purchased and installed in 2026 carries a known acquisition cost, has already converted supply-chain risk from procurement risk to operational risk, and delivers capacity the moment it is racked. Hardware ordered in 2027 or 2028 is exposed to whatever the next BIS rule, next Chinese counter-restriction, or next allocation crunch produces. Owning hardware does not eliminate supply chain dependency; drivers, CUDA updates, firmware patches, and eventual replacement still flow through the same constrained channels. But it converts a recurring exposure into a one-time exposure. This paper also addresses the AMD alternative as a vendor-diversification path; the open-weight model strategy covered later in the paper partially insulates the software layer from single-vendor lock-in. The procurement timing argument stands on its own regardless of how those longer-horizon decisions resolve.

Tariffs and the cost-pass-through asymmetry

Tariff policy adds a second layer of cost volatility on top of supply chain concentration, and the structure of who absorbs that cost is what makes the procurement argument economic rather than merely strategic.

Section 301 tariffs from the first Trump administration imposed seven and a half to twenty-five percent duties on Chinese-sourced electronic components and server assemblies [12]; both the Biden and second Trump administrations kept them in force. Section 232 metals tariffs have escalated repeatedly: the original 2018 rates of twenty-five percent on steel and ten percent on aluminum were raised to fifty percent on both in June 2025, copper was added at fifty percent in July 2025, and an April 2026 proclamation restructured the regime into tiered rates of fifty, twenty-five, and fifteen percent applied to the full customs value of covered products, with metal-intensive electrical grid equipment such as transformers and switchgear at the fifteen percent transitional rate through 2027 [13]. Server chassis, rack frames, power supply housings, and the thermal and power infrastructure of every data center build sit inside that scope. The April 2, 2025 tariff package added a ten percent baseline on all imports, with semiconductors initially excluded [14]. That exclusion did not hold. On January 14, 2026, Proclamation 11002 imposed the first product-specific Section 232 tariff on chips: twenty-five percent on a narrow category of advanced AI accelerators, naming the NVIDIA H200 and AMD MI325X as covered examples, with broad exemptions for chips used in U.S. data centers, research, startups, and public-sector applications, and with a second, broader tariff phase and a Commerce report on the data center chip market explicitly reserved for mid-2026 [5]. The current exemptions protect domestic deployment; the two-phase structure is the administration telling buyers, in writing, that the schedule can change again.

The financial scale is now documented rather than speculative. U.S. AI data centers collectively paid more than $6 billion in tariffs in 2025, according to Instrumental’s industry analysis [7]. CBRE estimates commercial project construction costs rose three to five percent from tariff-related materials increases alone [15]. A 2020 joint study by the Semiconductor Industry Association (SIA) and Boston Consulting Group put the ten-year total cost of ownership for a U.S.-based wafer fabrication facility at roughly 30 to 50 percent above an equivalent facility in Taiwan, South Korea, or Singapore, and 37 to 50 percent above China [16]; separately, Cushman & Wakefield projected in late 2025 that Section 232 tariffs on steel, aluminum, and copper, metals foundational to data center construction, would drive construction materials costs up by approximately 9 percent and total data center project costs up by 4.6 percent [17], [18].

The component vulnerability is structural. Semiconductors represent more than half of total AI server cost, per SemiAnalysis data cited by CSIS, with the GPU as the dominant cost item, manufactured by NVIDIA at TSMC in Taiwan [19]. Memory from SK Hynix and Samsung adds Korean exposure. In 2024, the United States imported approximately $77.5 billion in computer and electronic products from Taiwan, $142.4 billion from Mexico, and $64.5 billion from China, according to U.S. Department of Commerce trade data [20], [21], [22]. Non-semiconductor components (cooling, power, networking, chassis) represent at least a quarter to a third of data center cost and are broadly exposed through steel, aluminum, copper, and electronics duties.

The asymmetry that matters for the deployment decision is who eats these costs over time. CSIS notes that hyperscalers may absorb tariff increases in the short term to maintain AI development roadmaps, but pass-through to customer pricing is a matter of timing rather than principle [19]. Cloud API pricing reflects input costs eventually, and the customer has no contractual mechanism to lock yesterday’s price for tomorrow’s inference. An organization that purchases hardware in 2026 pays the 2026 tariff schedule once, on a depreciable capital asset. An organization that buys cloud inference pays whatever pricing schedule the provider sets next quarter, then the quarter after that, then the quarter after that, with cost increases driven by tariff rounds the customer did not vote for, on hardware sourced from countries the customer did not select. The cloud customer has elected an indefinite tariff exposure. The hardware buyer has elected a one-time exposure.

The financing structure behind cloud AI pricing

Central institutional analyses now describe the current AI investment cycle in directly comparable historical terms. The International Monetary Fund’s chief economist Pierre-Olivier Gourinchas told reporters at the October 2025 World Economic Outlook briefing that the current tech investment surge carries echoes of the dot-com boom of the late 1990s: it was the internet then, it is AI now [23]. The Bank of England’s Financial Policy Committee, in its December 2025 Financial Stability Report, judged that risky asset valuations “remain materially stretched, particularly for technology companies focused on Artificial Intelligence,” with U.S. equity valuations close to the most stretched since the dot-com bubble, heightening the risk of a sharp correction whose lending losses could raise financial stability risks [24]. The Bank for International Settlements documented the financing side twice in 2026: its March Quarterly Review reported hyperscaler corporate bond issuance topping $100 billion in 2025, mostly at maturities over five years, alongside off-balance-sheet joint ventures funded by private credit that the BIS calls “shadow borrowing,” obligations economically akin to debt that largely reside outside corporate balance sheets [25]. Its June 2026 Annual Economic Report went further, warning that the five largest hyperscalers will spend over $1 trillion on AI capital expenditure across 2025 and 2026 combined, that these commitments are outpacing their earnings and free cash flow, and that intense winner-take-most competition raises the risk of collective overcommitment to projects with uncertain returns [26]. None of these institutions has predicted timing or outcome. They documented the pattern. For an organization choosing between cloud AI consumption and on-premises infrastructure, the documented pattern matters more than any forecast of how it resolves, because the cloud customer is exposed to the resolution either way.

The numbers anchor the pattern. CreditSights estimated the five largest U.S. hyperscalers (Amazon, Microsoft, Alphabet, Meta, and Oracle) spent approximately $443 billion on capital projects in 2025 and projected roughly $602 billion for 2026, with about 75 percent directed at AI-specific infrastructure [27]; the IEA independently puts the 2025 figure for five large technology companies above $400 billion, with a further 75 percent increase expected in 2026 [28]. Moody’s raised its 2026 forecast to $785 billion in May 2026 for a six-company basket that adds CoreWeave, and projects nearly $1 trillion in 2027 [29]. Goldman Sachs Global Institute’s baseline model, which Goldman is careful to frame as an assumption-sensitive estimate rather than a forecast, implies approximately $7.6 trillion in cumulative AI capital expenditure between 2026 and 2031 across compute, data centers, and power [30]. Two changes inside those figures matter for the customer side of the relationship. First, the hyperscalers are now spending more on AI infrastructure than they generate in earnings and free cash flow, a financial-model change the BIS dates to this cycle [26]. Second, they are closing the gap with bond issuance and with off-balance-sheet vehicles whose obligations are economically equivalent to debt [25]. Bonds carry interest. Interest gets paid out of revenue. Cloud AI customers are the revenue. The provider’s debt service becomes a floor under the price the customer is offered, independent of what the inference itself costs to deliver.

The historical comparators show what happens to customers, not only to equity markets, in cycles of this kind. The NASDAQ Composite peaked at 5,048.62 on March 10, 2000 and fell about 78 percent to 1,114.11 by October 2002 [31]. Running alongside that equity collapse was a less-discussed infrastructure event. Between 1996 and 2001, the U.S. telecommunications industry poured more than $500 billion into laying fiber-optic cable, adding switches, and building wireless networks, deploying tens of millions of miles of fiber across the country [32], [33]. After the bust, much of that installed fiber remained unused, earning the nickname “dark fiber,” while many of the telecommunications firms themselves collapsed [32], [34]. The Federal Reserve Bank of Richmond’s 2003 Economic Quarterly documents how firms such as Global Crossing and WorldCom, which had constructed long-haul fiber networks anticipating surging bandwidth demand, became emblematic casualties of the bust; both ultimately filed for bankruptcy as overcapacity in long-haul fiber and regulatory uncertainty destroyed the investment thesis underpinning the boom [35]. Their customers were the second-order story. Organizations that had signed multi-year capacity contracts on the assumption of vendor continuity had to renegotiate terms with bankruptcy trustees, migrate to surviving providers, or absorb service degradation as failing operators cut staff. The IEEE Communications Society’s September 2025 retrospective points to the 19th-century U.K. railway mania as the older instance of the same arc: a near five-fold expansion of the network between 1840 and 1852 that ultimately returned about one-quarter of projected revenue, with most original operators absorbed or liquidated [33]. The physical infrastructure outlasted the operators in both cases. The customer contracts written against those operators did not.

One reported difference between the present AI cycle and the 1995–2000 cycle deserves a precise statement rather than a confident takeaway. The hyperscalers are profitable and unlikely to fail at the corporate level. What the institutional sources document is a narrower question: whether the AI revenue model being built will generate enough cash to justify the spending and service the new debt at the rate it is being incurred. If it does, AI compute remains tight and cloud pricing reflects that scarcity. If it does not, the historical pattern at the customer level is operator consolidation, contract disruption, and pricing volatility even when the largest providers themselves remain solvent.

The implication for an AI strategy and roadmap is structural rather than predictive. Cloud AI consumption commits an organization to the resolution of the AI capex cycle on whatever schedule that resolution arrives, with re-pricing risk on a quarterly billing cadence the customer does not control. Hardware ownership commits the organization to a known capital cost on a depreciable asset and a five-to-seven-year hardware-refresh cycle the organization does control. The first ties an organization’s AI capability to the AI industry’s financing structure. The second ties it to the institution’s own work. On-premises deployment is not insulated from the AI industry; drivers and hardware-refresh cycles continue to flow through the same supply chain. What it does not depend on is the financing structure underneath cloud pricing holding together through the cycle the IMF, the Bank of England, and the Bank for International Settlements are now documenting in print.

Energy, water, and the resource footprint of AI at scale

The physical resources AI consumes are no longer a secondary story. The International Energy Agency’s April 2025 Energy and AI report, the first intergovernmental analysis of the energy-AI relationship at sector scale, places 2024 global data center electricity consumption at approximately 415 TWh, or roughly 1.5 percent of global electricity demand [36]. U.S. data centers alone consumed 183 TWh in 2024, more than four percent of national electricity use. Data center electricity consumption has grown twelve percent annually since 2017, more than four times faster than total electricity consumption. The IEA’s April 2026 update reports 2025 global data center consumption of roughly 485 TWh, a seventeen percent increase from 2024, against three percent growth in overall global electricity demand [28].

The projection deserves careful reading by anyone planning infrastructure on a five-year horizon. The IEA base case projects global data center electricity demand will roughly double by 2030 to 945–950 TWh, approximately three percent of global electricity [36], [28]. A typical AI-focused data center consumes as much electricity annually as 100,000 households [36]; the largest facilities under construction are designed to consume twenty times that figure. AI is the fastest-growing driver of data center electricity demand, with AI-focused facilities expanding their consumption by fifty percent in 2025 alone and projected to triple their electricity use between 2025 and 2030 [28].

Water consumption follows a similar curve with worse local consequences, because water cannot be transmitted long distances the way electricity can. Global data centers consumed roughly 560 billion liters of water in 2023 [37]. Li and colleagues document that training GPT-3 in Microsoft’s U.S. data centers can consume a total of 5.4 million liters of water, including 700,000 liters of scope-1 on-site water consumption as a one-time training cost, and that GPT-3 inference consumes roughly 500 milliliters of water for every ten to fifty medium-length responses depending on when and where it is deployed [38]. The same authors project global AI demand is projected to require 4.2 to 6.6 billion cubic meters of water withdrawal annually by 2027, exceeding the total annual withdrawal of four to six countries the size of Denmark, or half the United Kingdom’s [38].

The supply chain underneath this physical footprint has been tightening in parallel. The IEA’s 2025 Global Critical Minerals Outlook reports that the average market share of the top three refining nations for copper, lithium, nickel, cobalt, graphite, and rare earth elements rose to eighty-six percent in 2024, up from roughly eighty-two percent in 2020 [39]. China dominates refining for most of these. The escalation of 2024 and 2025 is the relevant story for AI hardware. China restricted gallium, germanium, and antimony exports to the U.S. in December 2024, added tungsten, tellurium, bismuth, indium, and molybdenum controls in February 2025, and placed seven heavy and medium rare earth elements under control in April 2025. Gallium prices outside China doubled within five months of the December 2024 restriction [39]. Gallium is essential for compound semiconductors used in some AI inference paths; copper sits in every interconnect and power delivery system in the rack, and has carried its own fifty percent Section 232 tariff since July 2025 [13].

The on-premises position on this is not that local inference is more efficient per query than hyperscale inference. It is not. Hyperscalers run more efficient cooling, denser packing, and better power utilization effectiveness than any small facility will match. The on-premises position is that the workload is bounded. A 6 to 10 kW inference server supporting a few hundred internal users runs only the models that organization needs, only when its staff need them, with no obligation to maintain idle capacity for anonymous public traffic. The marginal energy and water cost of each inference is absorbed by purposeful institutional work rather than distributed across millions of zero-marginal-value requests. The efficiency gap per query is real and goes to the hyperscaler. The total resource consumption gap, per institution served, goes the other direction.

Resource allocation under mispriced infrastructure

The previous point depends on a structural observation about how hyperscale AI is priced. Cloud inference pricing systematically underweights the environmental and supply chain costs documented above. Electricity, water, and the embedded carbon of silicon disappear into per-token rates that do not vary with grid carbon intensity, local water stress, or refinery concentration. When infrastructure is priced below its true resource cost, workloads that would not survive accurate pricing proliferate.

The categories that absorb disproportionate hyperscale capacity relative to institutional value (bulk advertising content generation, engagement-optimized media, speculative inference at consumer scale) consume the same physical resources as workloads solving material problems. This is not an indictment of individual users making rational decisions inside a pricing system that does not penalize waste. It is an observation about what mispriced infrastructure produces at the population level.

Granular workload-type attribution at the level the argument would benefit from (what fraction of global inference serves research, what fraction enterprise, what fraction consumer entertainment) is not published by the IEA, Gartner, IDC, or any institutional source as of this writing. The data does not exist at that resolution, and the absence is itself part of the transparency gap the governance argument later in this paper addresses. What the IEA does document is that efficiency gains per AI task are running behind volume growth. Per-task energy consumption is dropping by at least an order of magnitude annually, faster than at any prior point in energy history, but total demand is growing faster still, with AI agents flagged as an emerging high-intensity category [28]. Mispriced infrastructure plus accelerating volume produces aggregate resource pressure regardless of the workload mix underneath.

A related thread is the share of AI capacity directed at active harm. The IEA’s April 2025 report notes that cyberattacks on utilities and critical infrastructure, increasingly AI-enabled, tripled over the preceding four years [36]. CISA’s AI security guidance documents the broader category of AI-generated phishing, synthetic media, and disinformation campaigns now consuming production-scale inference capacity [40]. The infrastructure cost of these workloads is not separately quantified in current institutional literature, but the directional evidence is clear: a non-trivial share of global AI compute is being used to attack the same organizations paying for it.

On-premises deployment prices these costs accurately. Electricity shows up on the utility bill. Cooling shows up in the facilities budget. Hardware depreciation shows up on the balance sheet. There is no abstraction layer that converts kilowatt-hours into tokens and tokens into a flat per-call price. Pricing accuracy disciplines workload selection. Every inference is traceable to an internal user or agent serving a documented business or research purpose, because the institution pays the metered cost of every inference on its own infrastructure. The argument is structural rather than statistical: the on-premises deployment model is bounded, accountable, and purpose-priced by design.

Labor market disruption and the talent recruiting signal

The labor market evidence on AI converged in 2025 in ways that change the talent calculation for organizations without internal AI capability. McKinsey Global Institute’s 2023 U.S. analysis projected that by 2030, activities accounting for up to thirty percent of currently worked hours could be automated, with approximately twelve million occupational transitions required in the U.S. alone [41]. The 2024 European extension reached similar conclusions for the EU and UK, with demand for STEM and healthcare professionals projected to grow seventeen to thirty percent between 2022 and 2030, and demand for office, production, and customer service work projected to decline [42].

The 2025 data made the divergence concrete. Brynjolfsson, Chandar, and Chen, analyzing high-frequency ADP payroll records covering millions of workers, document a sixteen percent relative employment decline, as of September 2025, for early-career workers aged 22 to 25 in the most AI-exposed occupations since late 2022, after controlling for firm-level shocks, while employment for experienced workers in the same occupations remained stable or grew; the declines concentrate where AI automates rather than augments the work [43]. Fifty-one percent of organizations surveyed by McKinsey reported that generative AI was reducing their need for entry-level roles, and Bureau of Labor Statistics data cited in the same analysis shows unemployment among college graduates aged 23 to 27 rising from 3.25 percent in 2019 to 4.59 percent in 2025 [44]. The 2025 World Economic Forum Future of Jobs Report, drawn from over 1,000 employers across 22 industries and 55 economies, projects 170 million new roles and 92 million displaced by 2030, a net 78 million job increase that disguises 22 percent disruption to current employment composition. AI and information processing technologies alone are projected to create 11 million jobs and displace 9 million [45]. Sixty-three percent of employers cite the skills gap as the primary barrier to business transformation.

The talent retention implication is the operative point for a small organization building AI capability. As AI fluency becomes a professional baseline expectation in engineering, research, and knowledge work, organizations without internal AI infrastructure become less attractive to the workers most able to direct AI systems productively. Engineers who want to work with on-premises model fine-tuning, agent orchestration, and structured retrieval systems will not stay long in environments where their AI work is rate-limited by a vendor’s API quota and shaped by a vendor’s content policies. Building on-premises infrastructure is a recruiting and retention signal as much as a technical decision.

Regulatory exposure runs in the same direction. The EU AI Act (Regulation EU 2024/1689) entered force August 1, 2024 [46]. The prohibited-use provisions took effect February 2, 2025. General-purpose AI model obligations, including documentation, transparency, and systemic-risk reporting requirements covering every major LLM, became applicable August 2, 2025. Full compliance for high-risk AI systems is currently scheduled for August 2, 2026, with possible extension to December 2027 under the Digital Omnibus package in trialogue as of April 2026 [47]. The extraterritorial reach is the part U.S.-headquartered organizations frequently underestimate: Article 2 brings any provider or deployer whose AI system produces output used in the EU into scope, regardless of where the organization or its infrastructure sits [48]. Maximum fines reach €35 million (approximately $40 million at the time of this writing) or seven percent of global group annual revenues for prohibited practices, exceeding GDPR ceilings.

Cloud AI providers serving EU users carry their own General-Purpose AI obligations. Organizations using those APIs become deployers, with high-risk obligations if the use case falls in a regulated category and audit obligations attached to the deployer regardless. On-premises deployment of open-weight models, serving only internal users, is structurally easier to bring into compliance because the deploying organization controls the model behavior, the audit logs, the data handling, and the documentation; every element the Act actually requires. The frameworks for demonstrating responsible use exist: Partnership on AI’s ABOUT ML transparency framework and the AI Now Institute’s 2025 Artificial Power landscape report both provide governance scaffolding [49], [50]. Both presuppose the deploying organization can answer questions about what its models do, how they were trained or adapted, and which inferences they produced. Those answers are tractable only when the infrastructure is owned.

On-premises AI as resilience infrastructure

The arguments above synthesize into a single position. Organizations that own their AI infrastructure are insulated from supply chain shocks affecting future hardware procurement, from tariff escalation affecting cloud cost-pass-through, from cloud provider outages and policy changes, and from regulatory shifts affecting third-party data handling. The cloud-dependent organization’s AI capability is a function of vendor uptime, vendor pricing, vendor content policy, vendor data residency decisions, and geopolitical constraints on the vendor’s data center footprint, none of which the customer controls. The on-premises organization’s AI capability is a function of its own infrastructure, its own policies, and its own procurement decisions.

The right analogy for executive framing is generator infrastructure, not software licensing. An on-site generator does not produce cheaper electricity than the grid. It produces independent electricity during grid failure. On-premises AI does not produce cheaper inference per query than a hyperscaler does. It produces independent capability during cloud disruptions, vendor policy shifts, market shocks, and regulatory dislocations. The financial argument lives at the institutional level, not the per-query level: a generator pays for itself the first time the grid drops during a critical production window, and on-premises AI pays for itself the first time a vendor’s policy change or rate increase would have foreclosed a capability the institution had built its workflows around.

The capability compounds. Each year of operation accumulates fine-tuned models adapted to institutional vocabulary and tasks, indexed institutional knowledge that no external vendor sees, validated agent workflows shaped by actual internal use, and operational expertise in a small team that knows the stack end to end. A cloud-dependent peer at the same five-year mark has a stack of API invoices and whatever lock-in their vendor has built into its tooling. The resilience gap between the two compounds in adaptability as well as capability: the on-premises organization responds to AI advances on its own schedule, with its own data, on its own hardware, rather than waiting for a vendor to release the relevant feature in the relevant tier at the relevant price.

Research and technology for the benefit of all

The defensive case for on-premises deployment is now established. The positive case is more important and rarely articulated. The decision to deploy AI on-premises is not only a choice about avoiding risk. It is a choice about how the organization’s share of finite global AI resources (silicon, electricity, cooling water, refined gallium, engineer-hours) gets allocated.

Internal AI infrastructure directed at engineering, scientific work, technical research, documentation, knowledge management, and operational efficiency contributes directly to institutional capability. That capability compounds into better products, faster problem-solving, more resilient operations, and more skilled staff. Those outcomes benefit the organization’s people, its customers and partners, and, through the work the organization actually does in the world, the broader communities it serves. The infrastructure is being built so that real people inside this organization can do their real work better. That is a defensible answer to the question of what this share of global AI capacity is for.

“Research and technology for the benefit of all” is the phrase to anchor this position. It belongs in internal AI governance documentation, in the acceptable use policy, in the framing of the AI strategy for any stakeholder audience that asks why the organization is doing this. It is not a marketing tagline. It is an operational constraint that shapes what the infrastructure is for and, equally, what it is not for. Later sections of this paper cover the security and access controls that enforce the constraint along with the governance scaffolding around it.

The procurement window for hardware that can sustain this position is open now and tightening. The hardware decisions that follow are where the position becomes a budget line.

References

  1. U.S. Department of Commerce, Bureau of Industry and Security, “Implementation of Additional Export Controls: Certain Advanced Computing and Semiconductor Manufacturing Items; Supercomputer and Semiconductor End Use; Entity List Modification,” Federal Register, Oct. 2022. [Online]. Available: https://www.federalregister.gov/documents/2022/10/13/2022-21658/implementation-of-additional-export-controls-certain-advanced-computing-and-semiconductor. Vol. 87, no. 197, pp. 62186–62220. [Accessed: 18-May-2026]

    FRAC-1 Primary source Back to text

  2. U.S. Department of Commerce, Bureau of Industry and Security, “Implementation of Additional Export Controls: Certain Advanced Computing Items; Supercomputer and Semiconductor End Use; Updates and Corrections,” Federal Register, Oct. 2023. [Online]. Available: https://www.federalregister.gov/documents/2023/10/25/2023-23055/implementation-of-additional-export-controls-certain-advanced-computing-items-supercomputer-and. Vol. 88, no. 205, pp. 73424–73519. [Accessed: 18-May-2026]

    FRAC-2 Primary source Back to text

  3. U.S. Department of Commerce, Bureau of Industry and Security, “Commerce Strengthens Export Controls to Restrict China's Capability to Produce Advanced Semiconductors for Military Applications,” press release, Dec. 2024. [Online]. Available: https://www.bis.gov/press-release/commerce-strengthens-export-controls-restrict-chinas-capability-produce-advanced-semiconductors-military. 89 FR 96790. [Accessed: 18-May-2026]

    FRAC-3 Primary source Back to text

  4. U.S. Government Accountability Office, “Export Controls: Commerce Implemented Advanced Semiconductor Rules and Took Steps to Address Compliance Challenges,” Dec. 2024. [Online]. Available: https://www.gao.gov/products/gao-25-107386. GAO-25-107386. [Accessed: 18-May-2026]

    FRAC-4 Primary source Back to text

  5. Executive Office of the President, “Adjusting Imports of Semiconductors, Semiconductor Manufacturing Equipment, and Their Derivative Products into the United States,” Jan. 14, 2026. [Online]. Available: https://www.whitehouse.gov/presidential-actions/2026/01/adjusting-imports-of-semiconductors-semiconductor-manufacturing-equipment-and-their-derivative-products-into-the-united-states/. Proclamation 11002. White & Case LLP analysis: 'President Trump Orders Narrowly Targeted 25% Section 232 Tariff on Certain Advanced Semiconductor Articles,' Jan. 2026: https://www.whitecase.com/insight-alert/president-trump-orders-narrowly-targeted-25-section-232-tariff-certain-advanced. [Accessed: 02-Jul-2026]

    FRAC-5 Primary source Back to text

  6. Fortune, “Larry Ellison and Elon Musk 'Begged' Nvidia's Jensen Huang for More GPUs Over a Fancy Sushi Dinner,” Sept. 16, 2024. [Online]. Available: https://fortune.com/2024/09/16/larry-ellison-elon-musk-begged-nvidias-jensen-huang-more-gpus-fancy-sushi-dinner/. [Accessed: 02-Jul-2026]

    FRAC-6 Contextual source Back to text

  7. Instrumental Inc., “AI Data Centers Paid $6B+ in Tariffs in 2025 — A Cost to U.S. AI Competitiveness?” Mar. 17, 2026. [Online]. Available: https://instrumental.com/resources/build-better-news/ai-data-centers-paid-6b-in-tariffs-in-2025-a-cost-to-u-s-ai-competitiveness/. [Accessed: 18-May-2026]

    FRAC-7 Contextual source Back to text

  8. The Conference Board, “Policy Backgrounder: The Future of the CHIPS and Science Act,” Mar. 13, 2025. [Online]. Available: https://www.conference-board.org/research/ced-policy-backgrounders/the-future-of-the-CHIPS-and-Science-Act. [Accessed: 18-May-2026]

    FRAC-8 Secondary source Back to text

  9. Semiconductor Industry Association and Boston Consulting Group, “Emerging Resilience in the Semiconductor Supply Chain,” SIA, May 2024. [Online]. Available: https://www.semiconductors.org/wp-content/uploads/2024/05/Report_Emerging-Resilience-in-the-Semiconductor-Supply-Chain.pdf. [Accessed: 18-May-2026]

    FRAC-9 Secondary source Back to text

  10. European Parliament and Council of the European Union, “Regulation (EU) 2023/1781 of the European Parliament and of the Council of 13 September 2023 Establishing a Framework of Measures for Strengthening Europe's Semiconductor Ecosystem and Amending Regulation (EU) 2021/694 (Chips Act),” Official Journal of the European Union, 2023. [Online]. Available: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32023R1781. Vol. L 229, pp. 1–53. [Accessed: 18-May-2026]

    FRAC-10 Primary source Back to text

  11. European Court of Auditors, “The EU's Strategy for Microchips,” Apr. 2025. [Online]. Available: https://www.eca.europa.eu/ECAPublications/SR-2025-12/SR-2025-12_EN.pdf. Special Report SR-2025-12. [Accessed: 18-May-2026]

    FRAC-11 Primary source Back to text

  12. Sandler, Travis & Rosenberg, P.A., “Section 301 Tariffs on China.” [Online]. Available: https://www.strtrade.com/trade-news-resources/tariff-actions-resources/section-301-tariffs-on-china. [Accessed: 18-May-2026]

    FRAC-12 Secondary source Back to text

  13. Congressional Research Service, “Section 232 Tariffs on Steel and Aluminum,” May 2026. [Online]. Available: https://www.congress.gov/crs-product/IN12519. CRS Insight IN12519. [Accessed: 02-Jul-2026]

    FRAC-13 Primary source Back to text

  14. Executive Office of the President, “Regulating Imports With a Reciprocal Tariff to Rectify Trade Practices That Contribute to Large and Persistent Annual United States Goods Trade Deficits,” Federal Register, Apr. 7, 2025. [Online]. Available: https://www.federalregister.gov/documents/2025/04/07/2025-06063/regulating-imports-with-a-reciprocal-tariff-to-rectify-trade-practices-that-contribute-to-large-and. Vol. 90, no. 65, pp. 15041–15109. [Accessed: 18-May-2026]

    FRAC-14 Primary source Back to text

  15. CBRE, “On Again, Off Again: Tariffs & Commercial Real Estate,” CBRE Insights, Mar. 19, 2025. [Online]. Available: https://www.cbre.com/insights/briefs/on-again-off-again-tariffs-and-commercial-real-estate. [Accessed: 18-May-2026]

    FRAC-15 Secondary source Back to text

  16. Semiconductor Industry Association and Boston Consulting Group, “Government Incentives and U.S. Competitiveness in Semiconductor Manufacturing,” SIA, Sept. 2020. [Online]. Available: https://www.semiconductors.org/wp-content/uploads/2020/09/Government-Incentives-and-US-Competitiveness-in-Semiconductor-Manufacturing-Sep-2020.pdf. [Accessed: 18-May-2026]

    FRAC-16 Secondary source Back to text

  17. Cushman & Wakefield, “Tariffs Add New Pressure to Commercial Construction Budgets,” Business Wire, press release, Oct. 7, 2025. [Online]. Available: https://www.businesswire.com/news/home/20251007478380/en. [Accessed: 18-May-2026]

    FRAC-17 Secondary source Back to text

  18. Center for Strategic and International Studies, “The Impact of Tariffs on the AI Data Center Buildout: Balancing Supply Chain Security and AI Infrastructure Leadership,” CSIS, May 2026. [Online]. Available: https://www.csis.org/analysis/impact-tariffs-ai-data-center-buildout-balancing-supply-chain-security-and-ai. [Accessed: 18-May-2026]

    FRAC-18 Secondary source Back to text

  19. Center for Strategic and International Studies, “How Tariffs Could Derail the United States' $3 Trillion AI Buildout,” CSIS, Aug. 8, 2025. [Online]. Available: https://www.csis.org/analysis/how-tariffs-could-derail-united-states-3-trillion-ai-buildout. [Accessed: 18-May-2026]

    FRAC-19 Secondary source Back to text

  20. California Chamber of Commerce, “Taiwan Trading Partner Portal,” Advocacy.CalChamber.com, 2025. [Online]. Available: https://advocacy.calchamber.com/international/portals/taiwan/. [Accessed: 20-May-2026]

    FRAC-20 Secondary source Back to text

  21. California Chamber of Commerce, “Mexico Trading Partner Portal,” Advocacy.CalChamber.com, 2026. [Online]. Available: https://advocacy.calchamber.com/international/portals/mexico/. [Accessed: 20-May-2026]

    FRAC-21 Secondary source Back to text

  22. California Chamber of Commerce, “China Trading Partner Portal,” Advocacy.CalChamber.com, 2026. [Online]. Available: https://advocacy.calchamber.com/international/portals/china/. [Accessed: 20-May-2026]

    FRAC-22 Secondary source Back to text

  23. International Monetary Fund, “Press Briefing Transcript: World Economic Outlook, Annual Meetings 2025,” IMF, Oct. 14, 2025. [Online]. Available: https://www.imf.org/en/news/articles/2025/10/14/tr-10-14-25-press-briefing-transcript-world-economic-outlook-annual-meetings-2025. Press briefing remarks by P.-O. Gourinchas. [Accessed: 22-May-2026]

    FRAC-23 Primary source Back to text

  24. Bank of England, “Financial Stability Report — December 2025,” Financial Policy Committee, Bank of England, Dec. 2, 2025. [Online]. Available: https://www.bankofengland.co.uk/-/media/boe/files/financial-stability-report/2025/opening-remarks-december-2025. Opening remarks by A. Bailey, Governor. [Accessed: 22-May-2026]

    FRAC-24 Primary source Back to text

  25. E. Eren, I. Krohn, and K. Todorov, “Financing the AI Infrastructure Boom: On- and Off-Balance Sheet Borrowing,” Bank for International Settlements Quarterly Review, Bank for International Settlements, Mar. 16, 2026. [Online]. Available: https://www.bis.org/publ/qtrpdf/r_qt2603u.htm. [Accessed: 22-May-2026]

    FRAC-25 Primary source Back to text

  26. Bank for International Settlements, “Progress and Peril,” Annual Economic Report 2026, Ch. I, BIS, June 28, 2026. [Online]. Available: https://www.bis.org/publ/arpdf/ar2026e1.htm. [Accessed: 02-Jul-2026]

    FRAC-26 Primary source Back to text

  27. CreditSights, “Technology: Hyperscaler Capex 2026 Estimates,” CreditSights Research, Nov. 25, 2025. [Online]. Available: https://know.creditsights.com/insights/technology-hyperscaler-capex-2026-estimates/. [Accessed: 22-May-2026]

    FRAC-27 Secondary source Back to text

  28. International Energy Agency, “Key Questions on Energy and AI,” IEA, Apr. 2026. [Online]. Available: https://www.iea.org/reports/key-questions-on-energy-and-ai. [Accessed: 18-May-2026]

    FRAC-28 Primary source Back to text

  29. Moody's Research, “AI Hyperscalers' 2027 Capex Will Near $1 Trillion, Fueling AI Growth and Memory Shortage,” Moody's Investors Service, May 11, 2026. [Online]. Available: https://www.moodys.com/research/Artificial-Intelligence-Data-Centers-US-Hyperscaler-capex-to-near-Sector-In-Depth--PBC_1483702. Data Center Dynamics report: Moody's: Hyperscaler Capex Forecasts Marked Up by $85bn, to Close in on $1trn by 2027: https://www.datacenterdynamics.com/en/news/moodys-hyperscaler-capex-forecasts-marked-up-by-85bn-to-close-in-on-1trn-by-2027/. [Accessed: 22-May-2026]

    FRAC-29 Secondary source Back to text

  30. Goldman Sachs Global Institute and Goldman Sachs Global Investment Research, “Tracking Trillions: The Assumptions Shaping the Scale of the AI Build-Out,” Goldman Sachs, May 2026. [Online]. Available: https://www.goldmansachs.com/insights/articles/tracking-trillions-the-assumptions-shaping-scale-of-the-ai-build-out. [Accessed: 22-May-2026]

    FRAC-30 Secondary source Back to text

  31. Nasdaq, Inc., “NASDAQ Composite (NASDAQCOM),” Federal Reserve Bank of St. Louis FRED. [Online]. Available: https://fred.stlouisfed.org/series/NASDAQCOM. [Accessed: 22-May-2026]

    FRAC-31 Primary source Back to text

  32. R. E. Litan, “The Telecommunications Crash: What To Do Now?” Policy Brief #112, The Brookings Institution, Dec. 1, 2002. [Online]. Available: https://www.brookings.edu/articles/the-telecommunications-crash-what-to-do-now/. [Accessed: 22-May-2026]

    FRAC-32 Secondary source Back to text

  33. A. Weissberger, “Big Tech Spending on AI Data Centers and Infrastructure vs the Fiber Optic Buildout During the Dot-Com Boom (& Bust),” IEEE Communications Society Technology Blog, Sept. 27, 2025. [Online]. Available: https://techblog.comsoc.org/2025/09/27/big-tech-spending-on-ai-data-centers-and-infrastructure-vs-the-fiber-optic-buildout-during-the-dot-com-boom-bust/. [Accessed: 22-May-2026]

    FRAC-33 Contextual source Back to text

  34. S Cherry, “A Telecom Diet Rich in Fiber,” IEEE Spectrum, July 1, 2009. [Online]. Available: https://spectrum.ieee.org/a-telecom-diet-rich-in-fiber. [Accessed: 22-May-2026]

    FRAC-34 Contextual source Back to text

  35. E. A. Couper, J. P. Hejkal, and A. L. Wolman, “Boom and Bust in Telecommunications,” Federal Reserve Bank of Richmond Economic Quarterly, 2003. [Online]. Available: https://www.richmondfed.org/-/media/RichmondFedOrg/publications/research/economic_quarterly/2003/fall/pdf/wolman.pdf. Fall 2003. [Accessed: 22-May-2026]

    FRAC-35 Primary source Back to text

  36. International Energy Agency, “Energy and AI,” IEA, Apr. 2025. [Online]. Available: https://www.iea.org/reports/energy-and-ai. [Accessed: 18-May-2026]

    FRAC-36 Primary source Back to text

  37. A. de Vries-Gao, “The Carbon and Water Footprints of Data Centers and What This Could Mean for Artificial Intelligence,” Patterns, 2025. [Online]. Available: https://doi.org/10.1016/j.patter.2025.101430. Vol. 7, no. 1, p. 101430. [Accessed: 18-May-2026]

    FRAC-37 Primary source Back to text

  38. P. Li, J. Yang, M. A. Islam, and S. Ren, “Making AI Less 'Thirsty': Uncovering and Addressing the Secret Water Footprint of AI Models,” arXiv, 2023. [Online]. Available: https://arxiv.org/abs/2304.03271. arXiv:2304.03271, revised 2025. [Accessed: 18-May-2026]

    FRAC-38 Secondary source Back to text

  39. International Energy Agency, “Global Critical Minerals Outlook 2025,” IEA, 2025. [Online]. Available: https://www.iea.org/reports/global-critical-minerals-outlook-2025. [Accessed: 18-May-2026]

    FRAC-39 Primary source Back to text

  40. Cybersecurity and Infrastructure Security Agency, “Artificial Intelligence,” CISA. [Online]. Available: https://www.cisa.gov/ai. [Accessed: 18-May-2026]

    FRAC-40 Primary source Back to text

  41. McKinsey Global Institute, “Generative AI and the Future of Work in America,” McKinsey & Company, July 2023. [Online]. Available: https://www.mckinsey.com/mgi/our-research/generative-ai-and-the-future-of-work-in-america. [Accessed: 18-May-2026]

    FRAC-41 Secondary source Back to text

  42. McKinsey Global Institute, “A New Future of Work: The Race to Deploy AI and Raise Skills in Europe and Beyond,” McKinsey & Company, May 2024. [Online]. Available: https://www.mckinsey.com/mgi/our-research/a-new-future-of-work-the-race-to-deploy-ai-and-raise-skills-in-europe-and-beyond. [Accessed: 18-May-2026]

    FRAC-42 Secondary source Back to text

  43. E. Brynjolfsson, B. Chandar, and R. Chen, “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” Stanford Digital Economy Lab working paper, Nov. 2025. [Online]. Available: https://digitaleconomy.stanford.edu/app/uploads/2025/11/CanariesintheCoalMine_Nov25.pdf. Originally Aug. 2025. [Accessed: 02-Jul-2026]

    FRAC-43 Secondary source Back to text

  44. McKinsey & Company, “How AI Is — and Isn't — Changing the Future of Work,” Apr. 6, 2026. [Online]. Available: https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/the-organization-blog/how-ai-is-and-isnt-changing-the-future-of-work. [Accessed: 18-May-2026]

    FRAC-44 Secondary source Back to text

  45. World Economic Forum, “The Future of Jobs Report 2025,” WEF, Jan. 7, 2025. [Online]. Available: https://www.weforum.org/publications/the-future-of-jobs-report-2025/. [Accessed: 18-May-2026]

    FRAC-45 Secondary source Back to text

  46. European Parliament and Council of the European Union, “Regulation (EU) 2024/1689 of 13 June 2024 Laying Down Harmonised Rules on Artificial Intelligence (Artificial Intelligence Act),” Official Journal of the European Union, July 12, 2024. [Online]. Available: https://eur-lex.europa.eu/legal-content/EN/TXT/?uri=CELEX:32024R1689. [Accessed: 18-May-2026]

    FRAC-46 Primary source Back to text

  47. Modulos AI, “Does the EU AI Act Apply to US Companies? Extraterritorial Scope Explained,” Apr. 2026. [Online]. Available: https://www.modulos.ai/blog/eu-ai-act-us-companies/. [Accessed: 18-May-2026]

    FRAC-47 Contextual source Back to text

  48. Morgan Lewis, “The EU AI Act Is Here — With Extraterritorial Reach,” July 26, 2024. [Online]. Available: https://www.morganlewis.com/pubs/2024/07/the-eu-artificial-intelligence-act-is-here-with-extraterritorial-reach. [Accessed: 18-May-2026]

    FRAC-48 Secondary source Back to text

  49. Partnership on AI, “ABOUT ML: Annotation and Benchmarking on Understanding and Transparency of Machine Learning Lifecycles,” 2024. [Online]. Available: https://partnershiponai.org/workstreams/about-ml/. [Accessed: 18-May-2026]

    FRAC-49 Secondary source Back to text

  50. AI Now Institute, “Artificial Power: 2025 Landscape Report,” 2025. [Online]. Available: https://ainowinstitute.org/publications/research/ai-now-2025-landscape-report. [Accessed: 18-May-2026]

    FRAC-50 Secondary source Back to text

Contents