Section1
Executive Summary
An organization should own the artificial intelligence (AI) capability it competes on and heavily depend on while renting out the rest through confidential cloud platform agreements. As of mid-2026, on-premises AI infrastructure built on OEM hardware, running open-weight models, and using open-source tooling is the only arrangement in which the organization maintains cost-effective control over their entire AI stack while keeping sensitive proprietary data in-house with the ability to run as much inference as the power budget allows. The financial justification starts with avoided cloud cost, which is derived from published vendor pricing and is independent of productivity assumption. Once realistic seat and agentic spend is credited, the recommended on-premises configuration carries a five-year premium near $500K over an all-cloud baseline with productivity gains held at zero. A conservative 10% productivity gain on well-suited tasks turns the investment net positive. The premium buys secure on-premises AI capabilities and assets that compound in value year-over-year, and are fully customizable to meet the organization’s detailed specifications with no token limits and no per-token costs. The organization will be able to develop, grow, and use their on-premises AI infrastructure within a stable financial environment where costs are more predictable and less susceptible to market and economic fluctuations.
The decision presented to leadership is smaller than the length of this paper suggests. It is a Phase 1 procurement order for one inference server sized to handle five to fifteen concurrent users within a small (<100 employees) to medium (100 to 499 employees) size organization, priced at $60K–$150K on the recommended configuration. The five to fifteen concurrent users figure captures the estimated number of active inference requests happening at the same time within a population of 100 to 500 employees where most employees are either not currently using the system, crafting a prompt, or reviewing the output, which does not require inference processing. The Phase 1 deployment requires a passing facility assessment before procuring any hardware, and completing several security and governance tasks before the first production inference request is processed. The cloud subscriptions the organization already pays for stay in place for commodity work. Every purchase after Phase 1 depends on what the pilot measures.
Why Adopt AI at All
Budget pressure and constrained hiring are the usual opening for an AI proposal. That opening assumes its conclusion, because it treats AI as a proven multiplier of output. An organization whose people and processes work today is entitled to ask what justifies the cost and the organizational change.
The skeptic has evidence. In a representative survey of roughly 6,000 senior executives in the United States, the United Kingdom, Germany, and Australia, 89% reported no effect of AI on their firm’s labor productivity (sales per employee) over the past three years [1]. The same body of evidence explains the result. AI raises the output of an organization that already functions, and it does so only on work that fits the model’s strengths. Among 5,172 customer-support agents at a Fortune 500 software firm, an AI assistant raised issues resolved per hour by 15% on average and by about 30% for novice and lower-skilled agents [2]. Vendor telemetry covering about 10,000 software developers shows the result when nothing around the AI step changes. Heavy-adoption teams merged 98% more pull requests, but review time rose 91% and delivery metrics stayed flat [3]. Generation got faster; the queue moved to review.
The evidence supports specific gains: code generated against established internal patterns, questions answered from the organization’s own documentation, troubleshooting that compresses expert diagnosis time, and documentation written instead of deferred. The domain-by-domain assessment in the Where AI Delivers Real Value section sets out that evidence. It is the right starting point for an executive who first needs to be persuaded that AI belongs in the organization at all. The AI Productivity Paradox section is its counterweight: it documents where deployments fail and why.
For a stable organization, the most valuable of these gains is knowledge continuity. Ramp-to-productivity for knowledge-based roles runs from a six-month floor to twelve to fifteen months; organizations lose between a third and two-thirds of new hires within the first year [4]. An indexed knowledge base built from project records and incident reports does not change attrition rates. It lowers what each departure costs, since what a departing engineer knew stays queryable by the successor.
The more pressing reason is trajectory. Credentialed observers disagree on whether AI has reached the macroeconomic data. Apollo Global Management’s chief economist finds AI absent from employment and productivity data and, outside the largest technology firms, from profit margins and earnings expectations [5]. Erik Brynjolfsson reads the 2025 data as the start of a productivity harvest that several more quarters must confirm [6]. The strategy here holds under either reading. Workflow designs and curated fine-tuning data accumulate only through use. A competitor that began Phase 1 in 2025–2026 will therefore hold two to four years of them by 2028–2030, and no later budget can buy that interval back.
Measurement has to be instrumented rather than surveyed. In a randomized trial by Model Evaluation & Threat Research (METR), experienced software developers estimated afterward that AI had made them 20% faster, when it had in fact made them 19% slower [7]. The entry point for an organization still deciding is modest: one inference server on-premises, a handful of use cases whose output validates quickly, and a twelve-month pilot measured against instrumented baselines. The five-year roadmap is for organizations ready to scale once the pilot results indicate so.
What Cloud AI Covers and Where It Stops
Cloud AI subscriptions are the correct tool for commodity work. Microsoft 365 Copilot, GitHub Copilot, ChatGPT Business, Claude, and Google Workspace AI handle drafting, summarization, and routine code well against low-sensitivity material. The organization should keep the ones it uses.
The limit appears at the work where AI becomes a durable advantage. As of May 2026, none of Microsoft, OpenAI, Anthropic, Google, or Amazon Web Services (AWS) let an organization fine-tune a model on arbitrary internal sources, export the resulting fine-tuned model weights, and run them on its own hardware. Microsoft 365 Copilot Tuning limits its preview to tenants holding at least 5,000 Copilot licenses [8]. No organization this paper addresses has that many employees. In May 2026, OpenAI stopped organizations without prior fine-tuning history from starting new jobs, and it will accept no new fine-tuning job from anyone after January 6, 2027 [9]. Anthropic’s only public fine-tuning path is Claude 3 Haiku on Amazon Bedrock, which is text-only and capped at 32K context [10]. Google offers the widest range of cloud tuning options, but the tuned Gemini model remains reachable only through Google’s application programming interfaces (APIs) [11]. In each case the organization leases a tuned capability and never holds the asset.
Agents follow the same pattern. Microsoft Copilot Studio runs multi-agent orchestration on Microsoft’s infrastructure and offers no independent audit of agent memory or state. It also adds capacity packs, from $200 per tenant per month for 25,000 messages, on top of per-seat licensing [12]. An agent that works over internal data, with auditable memory and a per-inference audit trail, needs infrastructure the organization runs.
The resulting architecture sorts every AI-touching task by data sensitivity, customization need, and token volume. Commodity work stays in the cloud. Proprietary corpora, custom-trained models, agents with persistent memory, audited inference, and long-running high-token workloads move on premises. The cloud line item stays in the budget at a smaller size.
Open-weight models closed most of the capability gap in 2026. On the Artificial Analysis Intelligence Index (v4.1), the strongest open-weight model in mid-2026, Kimi K3, scored 57 against 60 for Claude Fable 5 [13]. Epoch AI measured the lag between the best open-weight and closed models at four to six months [14]. That comparison overstates what Phase 1 hardware can run. Kimi K3 needs roughly 1.4 TB of memory for its weights at 4-bit quantization, far beyond a single server-class card. The U.S.-origin models this paper recommends for one 96 GB card, Gemma 4 31B and gpt-oss-120b, score 29 and 24 on the same index. Phase 1 earns its return by utilizing open-weight models to process proprietary corpora through audited, long-running, high-token inference workloads that best fit generation-heavy and easy-to-validate tasks. The serving stack is built from the first day to route requests to larger models as the infrastructure grows.
Control of the Data and of the Record
Where inference runs determines who controls the organization’s data and the record of what its models did with it. Commercial-tier contracts from Anthropic, OpenAI, and AWS exclude customer data from model training [15], [16], [17]. For organizations on those contracts, the familiar worry that proprietary material will train a vendor’s next model is largely settled. Sovereignty is a separate matter, because the data still passes through vendor infrastructure. Vendor terms also change: Anthropic’s 2025 consumer-terms revision brought consumer conversations into model training unless users opted out [18]. Employees on personal accounts create exposure that no enterprise contract reaches.
That last exposure has a measured price. In IBM’s 2025 breach study with the Ponemon Institute, one in five organizations studied had suffered a breach traceable to unsanctioned AI use. High levels of that shadow AI added about $670,000 to the average breach cost [19], [20]. Substitution is what moves that number. Netskope measured personal-account use of generative AI falling from 78% to 47% in one year, as organization-managed use rose from 25% to 62% [21]. An approved internal system that matches the speed and capability of consumer tools is the control that changes behavior. A policy that only prohibits use pushes the same use out of sight.
Ownership also produces the record. From the first production query, request logging captures timestamp, user, model version, token counts, and latency. Each output falls under a defined review tier with a named approver. Governance reviews and incident response draw on that record as evidence. The European Union (EU) AI Act also applies to organizations outside the EU whose AI output is used there [22], [23]. A policy can be written after the fact. The record of what the models did in production cannot be reconstructed later.
Availability is part of the same question of control. On June 12, 2026, three days after their launch, a U.S. Department of Commerce export-control order restricted foreign-national access to Claude Fable 5 and Claude Mythos 5. Anthropic could not verify nationality in real time, so it suspended both models for every user worldwide [24]. The Department lifted the controls on June 30. Fable 5 returned to users globally on July 1 [24]. The 19-day suspension was neither an outage nor a contract dispute, and no service-level agreement covered it. An open-weight model already running on owned hardware keeps serving through that kind of event.
The Financial Case
The financial case depends on where the organization starts. For an organization already paying for cloud AI, it rests on avoided cost that reconciles against current invoices. Take a 75-person reference organization over 60 months at mid-2026 published pricing. GitHub Copilot Business at $19 per seat per month comes to $85,500 [25]. The Microsoft 365 Copilot add-on at $30 comes to $135,000. It rises to $310,500 all-in for an organization that must also buy the $39 E3 base license Copilot requires [26], [27]. ChatGPT Business and Claude Team Standard at $20 per seat come to $90,000 apiece [28], [29].
Seats are the floor. Agentic workloads bill per token at API rates, and retrieved documents and tool results make up most of those tokens input. Consider ten production workflows against Claude Sonnet 4.6, priced at $3 per million input and $15 per million output tokens; ignoring prompt caching to simplify the calculation. If each processes four million tokens a day at roughly 80% input, the ten together cost about $6,600 a month, or about $396,000 over five years [30]. Pricing through 2025–2026 moved in one direction. Anthropic shifted Enterprise token consumption to full API rates with no built-in discount between November 2025 and February 2026 [31]; GitHub put every Copilot plan on usage-based billing on June 1, 2026 [32]. For an organization doing serious agentic or long-running work, realistic five-year cloud exposure runs $500K to $1M+ on top of the per seat costs.
For an organization not yet using AI, the same figures price the cloud path at the moment of adoption. The comparison is between building capability now and paying a higher entry price later. That later price includes vendor fees and price adjustments that trend upward as the four largest hyperscalers recover roughly $725 billion in planned 2026 capital expenditure [33]. It also includes the operating experience a later start gives up.
On premises, five-year total cost of ownership (TCO) runs about $1.3 to $1.45 million on the recommended NVIDIA H200 path. The higher-capability NVIDIA B300 path runs about $2.8 to $2.95 million. On both paths, enterprise support contracts make up most of the operating expense, with electricity being a smaller share. With productivity held at zero, the H200 path costs roughly $1.2 million more than an all-cloud baseline counting base seats alone. That premium narrows toward $500K once realistic seat and per token API spend is credited. The B300 path carries a net premium of roughly $1.5 to $2.0 million. It earns that premium only when an identified workload requires larger models that cannot run on the H200 node.
Productivity gains sit on top of that floor as upside. The productivity model assumes a 10% gain, set below the 15% field average [2]. It applies that gain to the 40% of working hours where AI is deployed, across 50 knowledge workers at a fully loaded $90 an hour. The result is $374,400 a year and $1.87 million over five years ($90/h × 80h/wk × 26wk/yr × 50 workers × 40% use × 10% gain × 5yrs). That is enough to make the H200 path net positive and to cover most of the B300 premium. The gain materializes only under three conditions. The deployment must target well-suited tasks. The downstream workflow must be redesigned to absorb the added volume. Recovered hours must be turned into output. The avoided-cost floor does not depend on any of those conditions.
Two further returns stay outside the quantitative model by design. The first is model improvement. A model fine-tuned on the organization’s documentation and incident history improves as that corpus grows, while a cloud API stays tied to general training data. The gap between them widens each year, but no defensible dollar figure captures it. The second is breach exposure. IBM’s 2025 data puts the average cost of breaches spanning multiple environments at $5.05 million, against $4.01 million for data held on premises [20]. The added breach probability attributable to cloud AI specifically cannot be estimated with rigor, so the paper treats it as a qualitative benefit. A chief financial officer (CFO) who wants it in the model should build that probability estimate explicitly.
The case fails only if cloud prices stay flat. No vendor’s pricing in 2025–2026 moved that way.
A Phased Build
The paper recommends buying the capability in three increments over five-plus years, never as one Phase-3-scale purchase on day one. A single large build buys capacity the organization has not yet learned to operate. It also commits capital to a generation of silicon hardware that will be outdated before the workload grows into it. Each phase’s hardware fits into the next phase’s chassis, rack-power envelope, and software stack, so each purchase adds to the build rather than replacing it. The specific silicon hardware generation depreciates. What is built on it keeps gaining value: workflows, datasets, fine-tuned models, and the staff expertise that runs the stack.
| Phase | Timeline | Principal hardware additions | Directional capital expenditure | Concurrent users and agents |
|---|---|---|---|---|
| Phase 1 | Months 0–12 | One or two NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs (96 GB) in a Dell PowerEdge XE7745 or Supermicro SYS-422GL-NR; storage with a separate backup tier; 100 gigabit Ethernet (GbE) switching | $60K–$150K; $150K–$300K for the H200 NVL alternative | 5–15 |
| Phase 2 | Months 12–24+ | 8-GPU training server (H200 by default, B300 where a identified workload requires it); dedicated agentic compute server; embeddings GPU; machine learning operations (MLOps) platform | $500K–$950K (H200 path); $1.3M–$1.8M (B300 path) | 25–50 |
| Phase 3 | Months 24+ to 60+ | Second inference server; InfiniBand or RoCE fabric and Non-Volatile Memory Express over Fabrics (NVMe-oF) storage where multi-node training requires them | $215K–$950K | 50–100+ |
| Five-plus years | — | — | $0.8M–$1.9M (H200 arc); $1.8M–$3.0M (B300 arc) | — |
All figures are directional mid-2026 estimates from the Phased Hardware Deployment Roadmap section and need a current vendor quotation before commitment.
Capacity will not be the limiting factor in Phase 1. In independent testing, a single RTX PRO 6000 sustained 8,425 tokens per second on a 30-billion-parameter workload [34]. Fifteen concurrent users reading at 30 tokens per second need 450 tokens per second, so the card has roughly 18 times the required throughput. Neither Phase 1 configuration requires a cooling retrofit.
Organization size determines how far the build goes. An organization below 100 employees often stops at Phase 1 or Phase 2. At that size, fine-tuning demand is usually intermittent, so a $400K–$800K training server would sit idle between runs. The full three-phase build is sized for a medium organization of 100 to 499 employees. It extends to 999 employees by adding inference nodes, with no redesign.
Staffing grows with the phases. One to a few practitioners already familiar with Linux, Python, and containers can build and operate Phase 1. Specialist contractors fill the three points where the expertise gap is widest: the first inference server configuration, the first fine-tuning run, and an independent security review before any agent with tool access reaches production. Concentrating that expertise in so few people creates a single point of failure on a multi-year commitment in the six to seven figures. The case for additional practitioners therefore belongs in the Phase 1 plan.
NVIDIA is the recommended Phase 1 vendor. Not necessarily because their hardware performs better, but because it introduces less operational risk. Every candidate system from NVIDIA or AMD exceeds the throughput requirement many times over. For a team of one to a few practitioners, the difference lies in the maturity of the serving software and the amount of published troubleshooting help. Dell has announced support for AMD’s Instinct MI350P, a 144 GB PCIe card, in the same XE7745 chassis, with availability from July 2026 [35]. That makes the twelve-month AMD re-evaluation a decision to add a card rather than build a new server. The re-evaluation still depends on independent benchmark data, which did not exist as of mid-2026.
Isolated and Critical Infrastructure Environments
For air-gapped facilities and critical infrastructure, on-premises AI is the only architecture that qualifies, because a true air gap excludes every cloud API by definition. That exclusion no longer means giving up modern capability. The Phase 1 hardware and its recommended open-weight models run with no external connectivity. The infrastructure strategy that serves a corporate deployment therefore also serves a fully isolated one. The air gap shifts cost into operating process. Every model update requires a physical transfer: preparation, hash verification, approved-media transit, staging, and validation. Under strict media-handling rules, that cycle may takes days to weeks. It sets the pace of both model updates and incident remediation.
Demand for AI already exists in these environments. In December 2025, the Cybersecurity and Infrastructure Security Agency (CISA), Australia’s Cyber Security Centre, and partner agencies issued joint guidance. That guidance is written for critical infrastructure owners and operators who are already integrating machine learning, large language models, and AI agents into operational technology [36]. An operator’s practical choice is between a sanctioned model inside the air gap and unsanctioned cloud tools used outside it.
Procurement Timing and Supply Risk
Ownership turns a recurring cost set by outside parties into a one-time capital cost. A cloud customer absorbs tariff increases and export-control changes as a repriced bill every quarter. A hardware owner absorbs them once, at purchase, on a depreciable asset. Timing is what matters most. Many AI components carry lead times of 36 to 52 weeks [37]. Through mid-2026, HGX Blackwell supply stayed allocation-constrained, while H200 and RTX PRO 6000 hardware remained readily available. Ordering Phase 1 on available hardware in the current budget cycle leaves enough lead time to bring in Phase 2 hardware on schedule.
Ownership changes the form of supply-chain exposure but does not remove it. Driver updates, firmware patches, CUDA releases, and replacement hardware still flow through the same constrained channels. Longer-term resilience depends on diversifying at two layers: an AMD option for the hardware and an open-weight model strategy that limits dependence on any single software vendor.
The Phase 1 order commits the organization to nothing it cannot expand, which leaves timing as the one open variable. How long the order can wait depends on how fast the underlying technology moves. The gap between major capability shifts has shrunk from five years to roughly eighteen months; described in the The Rapid Evolution of AI and Why Its Acceleration Is the Strategy section under The Acceleration Argument subsection. That is shorter than the annual planning cycle most organizations use to make decisions of this kind.
References
-
I. Yotzov et al., “Firm Data on AI,” NBER Working Paper No. 34836, National Bureau of Economic Research, Feb. 2026. [Online]. Available: https://www.nber.org/papers/w34836. Revised Mar. 2026. doi:10.3386/w34836. [Accessed: 16-Jun-2026]
-
E. Brynjolfsson, D. Li, and L. R. Raymond, “Generative AI at Work,” The Quarterly Journal of Economics, May 2025. [Online]. Available: https://academic.oup.com/qje/article/140/2/889/7990658. Vol. 140, no. 2, pp. 889–942. doi:10.1093/qje/qjae044. [Accessed: 24-Jul-2026]
-
Faros AI, “The AI Engineering Report 2025: The AI Productivity Paradox,” Faros AI, July 23, 2025. [Online]. Available: https://www.faros.ai/ai-productivity-paradox. Vendor first-party telemetry; full report gated. [Accessed: 16-Jun-2026]
-
Society for Human Resource Management, “Survey: Onboarding Programs Are Too Short,” SHRM, Nov. 30, 2018. [Online]. Available: https://www.shrm.org/topics-tools/news/talent-acquisition/survey-onboarding-programs-short. [Accessed: 25-Jun-2026]
-
T. Slok, “Waiting for the AI J-Curve,” Apollo Academy, Apollo Global Management, Feb. 14, 2026. [Online]. Available: https://www.apolloacademy.com/waiting-for-the-ai-j-curve/. [Accessed: 16-Jun-2026]
-
E. Brynjolfsson, “The AI Productivity Take-off Is Finally Visible,” Financial Times, Feb. 14, 2026. [Online]. Available: https://www.ft.com/content/4b51d0b4-bbfe-4f05-b50a-1d485d419dc5. Subscription required. Figures corroborated by Fortune/Yahoo Finance (Feb. 15, 2026) and American Enterprise Institute commentary (Feb. 23, 2026). Yahoo Finance report (Feb. 15, 2026): https://finance.yahoo.com/news/one-stanford-original-ai-gurus-205316027.html. [Accessed: 16-Jun-2026]
-
J. Becker, N. Rush, E. Barnes, and D. Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” arXiv, July 12, 2025. [Online]. Available: https://arxiv.org/abs/2507.09089. arXiv:2507.09089v2, revised Jul. 25, 2025. [Accessed: 16-Jun-2026]
-
Microsoft Corporation, “Microsoft 365 Copilot Tuning Admin Guide (Early Access Preview),” Microsoft Learn, Mar. 10, 2026. [Online]. Available: https://learn.microsoft.com/en-us/microsoft-365/copilot/copilot-tuning-admin-guide. [Accessed: 18-May-2026]
-
OpenAI, “Deprecations,” OpenAI API, 2026. [Online]. Available: https://developers.openai.com/api/docs/deprecations. [Accessed: 28-May-2026]
-
Amazon Web Services, “Fine-Tuning for Anthropic's Claude 3 Haiku Model in Amazon Bedrock Is Now Generally Available,” AWS Blog, Nov. 8, 2024. [Online]. Available: https://aws.amazon.com/blogs/aws/fine-tuning-for-anthropics-claude-3-haiku-model-in-amazon-bedrock-is-now-generally-available/. [Accessed: 18-May-2026]
-
Google, “About Supervised Fine-Tuning for Gemini Models,” Google Cloud Documentation, Apr. 29, 2026. [Online]. Available: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini-supervised-tuning. [Accessed: 28-May-2026]
-
Microsoft Corporation, “Microsoft Copilot Studio Documentation,” Microsoft Learn, 2026. [Online]. Available: https://learn.microsoft.com/en-us/microsoft-copilot-studio. [Accessed: 18-May-2026]
-
Artificial Analysis, “Kimi K3 Achieves #3 in the Artificial Analysis Intelligence Index, Comparable to Opus 4.8 and GPT-5.5,” Artificial Analysis, July 16, 2026. [Online]. Available: https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5. [Accessed: 18-Jul-2026]
-
J. Edwards and L. Emberson, “Open Models Lag State-of-the-Art Closed Models by 4 Months,” Epoch AI Data Insights, May 29, 2026. [Online]. Available: https://epoch.ai/data-insights/open-closed-eci-gap. [Accessed: 21-Jun-2026]
-
Anthropic, “Is My Data Used for Model Training?” Anthropic Privacy Center, Mar. 16, 2026. [Online]. Available: https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training. [Accessed: 18-May-2026]
-
OpenAI, “Sharing Feedback, Evaluation and Fine-Tuning Data, and API Inputs and Outputs with OpenAI,” OpenAI Help Center, 2026. [Online]. Available: https://help.openai.com/en/articles/10306912-sharing-feedback-evaluation-and-fine-tuning-data-and-api-inputs-and-outputs-with-openai. [Accessed: 18-May-2026]
-
Amazon Web Services, “Security, Privacy, and Responsible AI — Amazon Bedrock,” AWS, 2026. [Online]. Available: https://aws.amazon.com/bedrock/security-privacy-responsible-ai/. [Accessed: 18-May-2026]
-
Anthropic, “Updates to Consumer Terms and Privacy Policy,” Anthropic, Aug. 28, 2025. [Online]. Available: https://www.anthropic.com/news/updates-to-our-consumer-terms. [Accessed: 28-May-2026]
-
IBM Security and Ponemon Institute, “Cost of a Data Breach Report 2025,” IBM Corporation, July 2025. [Online]. Available: https://www.ibm.com/reports/data-breach. [Accessed: 25-Jun-2026]
-
IBM, “Cost of a Data Breach,” IBM Think Insights, 2026. [Online]. Available: https://www.ibm.com/think/insights/data-matters/cost-of-a-data-breach. [Accessed: 25-Jun-2026]
-
Netskope Threat Labs, “Cloud and Threat Report: 2026,” Netskope, Jan. 6, 2026. [Online]. Available: https://www.netskope.com/resources/cloud-and-threat-reports/cloud-and-threat-report-2026. Vendor telemetry, Oct. 2024 to Oct. 2025. [Accessed: 29-Jul-2026]
-
Modulos AI, “Does the EU AI Act Apply to US Companies? Extraterritorial Scope Explained,” Apr. 2026. [Online]. Available: https://www.modulos.ai/blog/eu-ai-act-us-companies/. [Accessed: 18-May-2026]
-
Morgan Lewis, “The EU AI Act Is Here — With Extraterritorial Reach,” July 26, 2024. [Online]. Available: https://www.morganlewis.com/pubs/2024/07/the-eu-artificial-intelligence-act-is-here-with-extraterritorial-reach. [Accessed: 18-May-2026]
-
Anthropic, “Redeploying Fable 5,” Anthropic News, June 30, 2026. [Online]. Available: https://www.anthropic.com/news/redeploying-fable-5. Updated Jul. 1, 2026. [Accessed: 15-Jul-2026]
-
GitHub, Inc., “About Billing for GitHub Copilot in Organizations and Enterprises,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/copilot/concepts/billing/organizations-and-enterprises. [Accessed: 25-Jun-2026]
-
Microsoft Corporation, “Microsoft 365 Copilot Plans and Pricing — Enterprise,” Microsoft, 2026. [Online]. Available: https://www.microsoft.com/en-us/microsoft-365-copilot/pricing/enterprise. [Accessed: 25-Jun-2026]
-
Microsoft Corporation, “Microsoft 365 Plans and Pricing — Enterprise,” Microsoft, 2026. [Online]. Available: https://www.microsoft.com/en-us/microsoft-365/enterprise/microsoft-365-plans-and-pricing. [Accessed: 25-Jun-2026]
-
OpenAI, “ChatGPT Pricing,” OpenAI. [Online]. Available: https://openai.com/business/chatgpt-pricing/. [Accessed: 25-Jun-2026]
-
Anthropic, “Plans & Pricing,” Anthropic PBC, 2026. [Online]. Available: https://claude.com/pricing. [Accessed: 25-Jun-2026]
-
Anthropic, “Pricing,” Claude Platform Docs, 2026. [Online]. Available: https://platform.claude.com/docs/en/about-claude/pricing. [Accessed: 25-Jun-2026]
-
The Register, “Anthropic Ejects Bundled Tokens From Enterprise Seat Deal,” The Register, Apr. 16, 2026. [Online]. Available: https://www.theregister.com/2026/04/16/anthropic_ejects_bundled_tokens_enterprise/. [Accessed: 25-Jun-2026]
-
GitHub, “GitHub Copilot Is Moving to Usage-Based Billing,” The GitHub Blog, 2026. [Online]. Available: https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/. [Accessed: 25-Jun-2026]
-
S Morris, R McMorrow, H Murphy, R Rosner-Uddin, and M Acton, “Google outpaces rivals as Big Tech's AI spending plans rise to $725bn,” Financial Times, Apr. 29, 2026. [Online]. Available: https://www.ft.com/content/2138e81c-4d86-46f4-8ca0-287f8b737cdf. Subscription required. Tom's Hardware report: https://www.tomshardware.com/tech-industry/big-tech/big-techs-ai-spending-plans-reach-725-billion. [Accessed: 18-May-2026]
-
D. Trifonov, “RTX 4090 vs 5090 vs PRO 6000: LLM Inference Benchmark,” CloudRift AI, Oct. 9, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx-gpus-for-llm-inference. [Accessed: 04-Jun-2026]
-
Dell Technologies, “Dell and AMD Are Expanding What's Possible for On-Premises AI,” Dell Technologies Blog, May 7, 2026. [Online]. Available: https://www.dell.com/en-us/blog/dell-and-amd-are-expanding-what-s-possible-for-on-premises-ai/. [Accessed: 14-Jun-2026]
-
Cybersecurity and Infrastructure Security Agency and Australian Signals Directorate's Australian Cyber Security Centre, “Principles for the Secure Integration of Artificial Intelligence in Operational Technology,” Dec. 3, 2025. [Online]. Available: https://www.cisa.gov/resources-tools/resources/principles-secure-integration-artificial-intelligence-operational-technology. Joint cybersecurity guidance. [Accessed: 24-Sep-2026]
-
Instrumental Inc., “AI Data Centers Paid $6B+ in Tariffs in 2025 — A Cost to U.S. AI Competitiveness?” Mar. 17, 2026. [Online]. Available: https://instrumental.com/resources/build-better-news/ai-data-centers-paid-6b-in-tariffs-in-2025-a-cost-to-u-s-ai-competitiveness/. [Accessed: 18-May-2026]