Section19

The ROI Case: Time and Cost Savings

The financial case for on-premises AI does not depend on the productivity gains that dominate most AI ROI pitches. It rests on a certain avoided-cost floor and a strategic asset that compounds; surviving even the assumption that AI delivers zero measured productivity, because the avoided-cost layer is arithmetic from current vendor pricing rather than a forecast.

For an organization already paying for a cloud AI stack (GitHub Copilot, Microsoft 365 Copilot, ChatGPT Business, Claude Team, production-volume Anthropic application programming interface (API) usage), the figures below replace existing budget lines and reconcile against current invoices. For an organization that has not yet adopted cloud AI, the same figures show what the cloud path costs at the moment of adoption; returns run through opportunity cost and the competitive trajectory the rebuttal to vendor-bundled AI develops, not through a budget line already on the books. Most mid-sized organizations sit in the first category whether or not they have inventoried the spend, so the avoided-cost calculation leads.

The Primary Argument: Avoided Cloud Costs

A representative 75-person organization deploying cloud AI at the seat layer faces a five-year spend in the high six figures before any agentic API workload is added. The figures below come from published vendor pricing current as of mid-2026. OpenAI renamed ChatGPT Team to ChatGPT Business on August 29, 2025, then cut the annual rate from $25 to $20 per seat per month on April 2, 2026 [1], [2]. Anthropic restructured Enterprise pricing between November 2025 and February 2026, dropping the seat fee to roughly $20 per user per month while moving all token consumption to full API rates with no built-in discount [3]. GitHub moved all Copilot plans from request-based to usage-based billing on June 1, 2026, preserving the $19 Business seat as a floor while billing agentic workflows against a monthly credit allotment with overflow at published API rates [4], [5].

Working from the verified mid-2026 prices and the 75-person reference organization, the base-seat layer produces these five-year figures. GitHub Copilot Business for 75 developers at $19 per seat per month annual is $85,500 over 60 months, the accurate floor under the new usage-based system [5]. The Microsoft 365 Copilot enterprise add-on for 75 users at $30 per seat per month annual is $135,000 over 60 months [6]. ChatGPT Business for 75 users at the current $20 per seat per month annual rate is $90,000 over 60 months. Claude Team Standard for 75 users at $20 per seat per month annual is $90,000 over 60 months [7]. These are five-year base-seat figures, averaged to roughly $100K each, with capped / limited usage.

That total understates real spending for the work this paper exists to support. Agentic workloads bill at per-token API rates rather than per-seat caps with fast compounding rates. Agentic work is input-heavy, since large contexts, retrieved documents, and tool results dominate the token count, so cost concentrates on the input side at a lower unit price but enormous volume. Take ten production agentic workflows running against Claude Sonnet 4.6 at $3 per million input and $15 per million output tokens [8]. A single workflow that processes four million tokens a day, roughly 80 percent of them input, costs about $22 a day (4M × 0.8 * $3 + 4M × 0.2 × $15); ten of them run near $6,600 a month ($22 × 10 × 30), or about $396,000 over five years, and heavier multi-agent fleets push past that. Anthropic’s Enterprise unbundling produces the same outcome by a different route: the lower base fee invites adoption, and the unmetered token billing captures the agentic upside [3]. GitHub’s June 2026 transition does it for Copilot, with GitHub’s own product chief conceding that a power user orchestrating agentic workflows against frontier models can cost an order of magnitude more than a quiet user [4]. The direction across vendors is uniform with a direct consequence for the avoided-cost calculation: the correct figure for an agentic-workload organization is not just the base-seat cost, it is the base-seat cost plus the additional API token cost consumed by agentic workflows and users surpassing their alloted token limits.

A 75-person organization on Microsoft E3 carries a cost the $135,000 add-on figure hides. Copilot is an add-on, not a standalone product, so the underlying E3 license at $39 per user per month (up from $36/user/month) is mandatory, which puts the all-in enterprise cost at $66 per user per month, or $310,500 over five years for 75 users [6], [9]. Organizations already on E3 for non-AI reasons pay that base license regardless, so for them the marginal cost of adding Copilot is the $135,000 figure; organizations not yet on E3 must budget the full stack. A Microsoft-partner deployment analysis reports that 30 to 40 percent of enterprise Copilot licenses sit unused in the first 90 days, which at 60 percent utilization would put the per-active-user cost at roughly 1.67 times the headline seat rate [10]. Treat that as a consultancy estimate rather than independently audited data. The avoided-cost figure is what the organization pays the vendor, not what it extracts from the seats.

These figures are certain in a precise sense: arithmetic from published pricing, not predictions about utilization, productivity, or strategic value. They move as vendor pricing moves, and given the GitHub and Anthropic restructurings, the direction for organizations doing serious agentic work is up. The averaged $100K figure is the floor of the base-seat argument; the realistic figure for the workloads this architecture serves easily runs $500K to $1M+ over five years.

Layered Upside: Productivity Gains on Favorable Task Types

The productivity layer is real and task-conditional. The productivity-paradox analysis in this paper sets out the evidence: AI delivers measurable gains on generation-heavy, quickly validated tasks with redesignable downstream workflows, and flat or negative results on judgment-intensive work in mature systems. Brynjolfsson, Li, and Raymond’s Quarterly Journal of Economics study of 5,172 customer-support agents found a 15 percent average gain, 30 percent among novice and lower-skilled workers, and minimal effect on the most experienced [11]. The METR randomized trial on experienced developers in mature codebases found a 19 percent slowdown in July 2025 [12]; METR’s February 2026 follow-up could not reproduce that result, because wider AI adoption broke the experiment’s design, and the group judged its central estimate unreliable, with raw point estimates leaning slightly positive but straddling zero [13]. Both findings hold for their own task profiles.

The Phase 1 use cases this architecture serves (retrieval-augmented question-answering on internal documentation, code generation against well-defined patterns, document drafting from structured templates, log analysis) sit on the favorable side of that line. A 10 percent productivity improvement on the subset of tasks where AI is deployed, below the 15 percent Brynjolfsson average and excluding the judgment-intensive work where METR found no gain, is the conservative anchor. Applied to 50 knowledge workers at a fully loaded cost of $90 per hour with AI active on 40 percent of working hours: 2,080 working hours per full-time equivalent (FTE) times a 0.40 task share times a 0.10 gain yields 83.2 hours saved per FTE annually, worth $7,488 per FTE, or $374,400 a year across 50 staff and $1.87 million over five years.

That figure is the illustrative upside, not a baseline. Three assumptions carry it. The first is that AI is deployed on tasks where it actually helps. The second is that the workflow around the AI step is redesigned to absorb rather than bottleneck on the higher throughput, where Faros AI’s reported 91 percent jump in pull-request review time, drawn from its own engineering-analytics telemetry, is the canonical failure mode [14]. The third is that the recovered time is captured into output rather than dissipated as the digital leisure that Blank, Schubert, and Zhang documented at the household level [15]. Hit all three and the productivity layer lands; miss any and it shrinks. The case absorbs that risk because the avoided-cost layer beneath it depends on none of the three.

Compounding Model Value

A model fine-tuned on internal data improves as more proprietary data enters it. By Year 3 and Year 5, an internal model trained on the organization’s accumulated engineering documentation, incident records, and operational logs outperforms any general-purpose cloud API on that specific domain; widening the gap while the cloud alternative stays tied to general training corpora. That advantage resists the precision the avoided-cost figures carry, which is why it stays out of the conservative scenario. The strategic claim is that it is the highest-value long-term return in the investment, the asset the on-premises infrastructure exists to produce. A CFO should read this layer as the difference between leasing a capability that holds flat and building one that compounds; weighing that distinction against the organization’s competitive position rather than against a dollar figure the underlying data cannot support.

Data Sovereignty Risk: Qualitative, Not a Line Item

IBM’s 2025 Cost of a Data Breach Report puts the global average breach cost at $4.44 million, down 9 percent from $4.88 million the prior year and the first decline in five years, which IBM attributes to faster AI-assisted detection and containment [16]. Two findings in the same report bear more directly on this section than the headline. Breaches involving data spread across multiple environments averaged $5.05 million, against $4.01 million for data held on premises; and high levels of shadow AI, where employees use unsanctioned AI tools, added about $670,000 to the average breach [17]. AI-processed data such as internal strategy documents, employee records, and proprietary engineering output is a material exposure when it routes through cloud APIs. The figures are real and citable. They stay out of the quantitative ROI model because the incremental breach probability attributable specifically to cloud AI exposure, as distinct from every other vector the organization already manages, cannot be specified with rigor. A defensible bound runs from $0 to several million dollars, depending on probability assumptions the organization would have to build itself.

A CFO who wants this in the financial case should construct an explicit probability estimate rather than inherit one buried in someone else’s model. This paper carries data-sovereignty value as qualitative upside supporting the architectural argument, not as a dollar figure pretending to a precision the probability does not have.

Opportunity Cost for Organizations Not Yet Building

An organization that defers internal AI capability through the 2024–2026 window faces three compounding disadvantages by 2028–2030. The operational one is a learning curve that cannot be bought back: a competitor that began Phase 1 deployment in 2025–2026 will hold two to four years of workflow design, fine-tuning experience, and incident-recovery knowledge, none of it hireable retroactively from a thin labor market. The financial one is the direction of travel, since cloud AI pricing rises as hyperscalers recover the roughly $725 billion in 2026 capital expenditure the rebuttal section documents [18], and the early signals (Anthropic’s Enterprise unbundling, GitHub’s usage-based transition, and the share of organizations spending over $100,000 a month on cloud AI roughly doubling between 2024 and 2025 in CloudZero’s survey) all point the same way [19]. The strategic one is the widest: proprietary fine-tuned models a competitor builds on its own data over 2025–2028 are a moat, because the training data is itself the advantage, and a late mover acquires neither the data history nor the iterative refinement that earlier deployment produces.

None of the three quantifies to the precision of the avoided-cost calculation. All three are real. A CFO weighing an on-premises investment against indefinite cloud deployment should price the optionality the on-premises path keeps open, namely the ability to fine-tune on proprietary data, to host frontier-class capability without per-token billing, and to run inference under direct audit control, against the optionality the all-cloud path closes off.

Five-Year Total Cost of Ownership

The on-premises capital expenditure (CapEx) the phased hardware roadmap documents runs from roughly $1 million in the conservative configuration, built on an H200 Phase 2 training tier, to roughly $2.5 million in the higher-capability configuration built on a B300 training tier. Five-year operating expense (OpEx) for power and enterprise support runs $300,000 to $450,000, dominated by Dell ProSupport Plus-class maintenance rather than electricity; Phase 1 power alone runs about $2,800 to $3,680 a year at $0.12 per kilowatt-hour and a power usage effectiveness (PUE) of 1.4. Total five-year cost of ownership (TCO) therefore lands near $1.3 to $1.45 million on the H200 path and near $2.8 to $2.95 million on the B300 path.

The conservative scenario credits only the certain avoided cloud cost and holds productivity at zero. On the recommended H200 path, the $1.3 to $1.45 million TCO net of the $100K base-seat avoided cost leaves a five-year premium of roughly $1.2 million over an all-cloud baseline; crediting the realistic $500K to $1M in avoided seats and agentic API narrows that premium toward $500K. On the B300 path the premium is larger, roughly $1.5 to $2.0 million net, because the B300 buys aggregate memory and model-class headroom the H200 cannot reach. Either way, the conservative premium is the explicit price of data sovereignty, the customization ceiling, and capability ownership that no cloud architecture sells today.

The moderate scenario applies the 10 percent productivity assumption to the favorable-task use cases, layering $1.87 million in five-year productivity value on the avoided-cost offset. On the H200 path that turns the investment net positive: the deployment recovers more than its five-year cost before any strategic value is counted. On the B300 path the same productivity value covers most of the premium, leaving a residual of a few hundred thousand dollars that the model-class capability and the strategic asset justify rather than the productivity line alone.

The honest version of the ROI case is the layered one. Avoided cloud cost is certain and bounds the base case. Productivity gains on favorable task profiles are well-supported by task-conditional evidence but not guaranteed to be captured. Compounding model value is strategic, durable, and not quantifiable. Data-sovereignty value is real and qualitative. On the recommended H200 configuration, the investment is net positive at a conservative 10 percent productivity gain, and it carries a known, sovereignty-buying premium of about $500K even at zero productivity. The only assumption under which the case fails is that cloud costs stay flat, and no current vendor pricing trajectory supports it.

What This Says About the Investment Decision

The financial argument for on-premises AI is not that cloud spending is reckless or vendor pricing predatory. The vendors are making rational use of their capital, and their products serve specific workloads well, as the rebuttal to vendor-bundled AI sets out. The argument is narrower and harder to dodge: an organization doing serious agentic work over a five-year horizon will easily spend $500K to $1M+ on cloud AI seats and API consumption regardless, and its real choice is between paying that for capability that depreciates against the next pricing change and paying a premium for capability that compounds against its own accumulated data. On the recommended configuration, the on-premises path costs roughly twice the all-cloud path over five years and produces an owned, appreciating asset; the all-cloud path costs half as much and produces a recurring expense the vendor reprices on its own schedule. The decision a CFO is actually making is not whether AI pays for itself. It is whether the organization wants to own the capability or rent it.

References

  1. OpenAI, “ChatGPT Pricing,” OpenAI. [Online]. Available: https://openai.com/business/chatgpt-pricing/. [Accessed: 25-Jun-2026]

    ROI-1 Primary source Back to text

  2. OpenAI, “What Is ChatGPT Business?” OpenAI Help Center, 2026. [Online]. Available: https://help.openai.com/en/articles/8792828-what-is-chatgpt-business. [Accessed: 25-Jun-2026]

    ROI-2 Primary source Back to text

  3. The Register, “Anthropic Ejects Bundled Tokens From Enterprise Seat Deal,” The Register, Apr. 16, 2026. [Online]. Available: https://www.theregister.com/2026/04/16/anthropic_ejects_bundled_tokens_enterprise/. [Accessed: 25-Jun-2026]

    ROI-3 Contextual source Back to text

  4. GitHub, “GitHub Copilot Is Moving to Usage-Based Billing,” The GitHub Blog, 2026. [Online]. Available: https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/. [Accessed: 25-Jun-2026]

    ROI-4 Secondary source Back to text

  5. GitHub, Inc., “About Billing for GitHub Copilot in Organizations and Enterprises,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/copilot/concepts/billing/organizations-and-enterprises. [Accessed: 25-Jun-2026]

    ROI-5 Primary source Back to text

  6. Microsoft Corporation, “Microsoft 365 Copilot Plans and Pricing — Enterprise,” Microsoft, 2026. [Online]. Available: https://www.microsoft.com/en-us/microsoft-365-copilot/pricing/enterprise. [Accessed: 25-Jun-2026]

    ROI-6 Primary source Back to text

  7. Anthropic, “Plans & Pricing,” Anthropic PBC, 2026. [Online]. Available: https://claude.com/pricing. [Accessed: 25-Jun-2026]

    ROI-7 Primary source Back to text

  8. Anthropic, “Pricing,” Claude Platform Docs, 2026. [Online]. Available: https://platform.claude.com/docs/en/about-claude/pricing. [Accessed: 25-Jun-2026]

    ROI-8 Primary source Back to text

  9. Microsoft Corporation, “Microsoft 365 Plans and Pricing — Enterprise,” Microsoft, 2026. [Online]. Available: https://www.microsoft.com/en-us/microsoft-365/enterprise/microsoft-365-plans-and-pricing. [Accessed: 25-Jun-2026]

    ROI-9 Primary source Back to text

  10. E. O'Connor, “Microsoft 365 Copilot Pricing & Licensing: Enterprise Guide 2026,” EPC Group, Apr. 10, 2026. [Online]. Available: https://www.epcgroup.net/microsoft-365-copilot-pricing-licensing-enterprise-guide-2026. Last updated Sep. 27, 2026. [Accessed: 25-Jun-2026]

    ROI-10 Contextual source Back to text

  11. E. Brynjolfsson, D. Li, and L. R. Raymond, “Generative AI at Work,” The Quarterly Journal of Economics, May 2025. [Online]. Available: https://academic.oup.com/qje/article/140/2/889/7990658. Vol. 140, no. 2, pp. 889–942. doi:10.1093/qje/qjae044. [Accessed: 24-Jul-2026]

    ROI-11 Primary source Back to text

  12. J. Becker, N. Rush, E. Barnes, and D. Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” arXiv, July 25, 2025. [Online]. Available: https://arxiv.org/abs/2507.09089. arXiv:2507.09089v2. [Accessed: 25-Jun-2026]

    ROI-12 Secondary source Back to text

  13. J. Becker, N. Rush, T. Cunningham, D. Rein, and K. Mahamud, “We Are Changing Our Developer Productivity Experiment Design,” METR, Feb. 24, 2026. [Online]. Available: https://metr.org/blog/2026-02-24-uplift-update/. [Accessed: 25-Jun-2026]

    ROI-13 Secondary source Back to text

  14. Faros AI, “The AI Engineering Report 2025: The AI Productivity Paradox,” Faros AI, July 23, 2025. [Online]. Available: https://www.faros.ai/ai-productivity-paradox. Vendor first-party telemetry; full report gated. [Accessed: 25-Jun-2026]

    ROI-14 Secondary source Back to text

  15. M. Blank, G. Schubert, and M. B. Zhang, “The Household Impact of Generative AI: Evidence from Internet Browsing Behavior,” arXiv, Feb. 27, 2026. [Online]. Available: https://arxiv.org/abs/2603.03144. arXiv:2603.03144. Stanford Institute for Economic Policy Research Working Paper. [Accessed: 25-Jun-2026]

    ROI-15 Secondary source Back to text

  16. IBM Security and Ponemon Institute, “Cost of a Data Breach Report 2025,” IBM Corporation, July 2025. [Online]. Available: https://www.ibm.com/reports/data-breach. [Accessed: 25-Jun-2026]

    ROI-16 Secondary source Back to text

  17. IBM, “Cost of a Data Breach,” IBM Think Insights, 2026. [Online]. Available: https://www.ibm.com/think/insights/data-matters/cost-of-a-data-breach. [Accessed: 25-Jun-2026]

    ROI-17 Secondary source Back to text

  18. S Morris, R McMorrow, H Murphy, R Rosner-Uddin, and M Acton, “Google outpaces rivals as Big Tech's AI spending plans rise to $725bn,” Financial Times, Apr. 29, 2026. [Online]. Available: https://www.ft.com/content/2138e81c-4d86-46f4-8ca0-287f8b737cdf. Subscription required. Tom's Hardware report: https://www.tomshardware.com/tech-industry/big-tech/big-techs-ai-spending-plans-reach-725-billion. [Accessed: 18-May-2026]

    ROI-18 Contextual source Back to text

  19. CloudZero, “The State of AI Costs in 2025,” Aug. 4, 2025. [Online]. Available: https://www.cloudzero.com/state-of-ai-costs/. [Accessed: 25-Jun-2026]

    ROI-19 Secondary source Back to text

Contents