Section5

Rebuttal: “We Already Have [Vendor X]’s AI”

Vendor-bundled AI is the right tool for a large share of organizational work and the wrong place to stop. Microsoft Copilot, GitHub Copilot, Claude, ChatGPT, Google Workspace AI, and Amazon Bedrock are capable products, and an organization should keep paying for the ones it already uses. The objection this section answers is the one that treats those subscriptions as the finish line: we already have Vendor X’s AI, so the AI question is settled. It is not settled. These tools handle commodity work well and, by design, cannot do the work where AI becomes a durable advantage, and the gap between the two is widening, not closing. This argument assumes the organization has already decided that AI belongs in its operations; whether to adopt AI at all is a separate question this paper takes up elsewhere.

The Workload Framework Decides Where Each Tool Belongs

The correct architecture is not cloud AI versus on-premises AI. It is a risk-stratified assignment that runs through every AI-touching task in the organization.

Cloud AI is the right answer where data sensitivity is low, customization needs are minimal, and the work is genuinely commodity: drafting a meeting invite from a public agenda, summarizing a press release, generating boilerplate against an open-source library, brainstorming a marketing tagline, answering general questions that touch no internal context. The hyperscalers’ commercial-tier contracts exclude this data from model training [1], [2], [3], the per-token cost at these volumes is real but bounded, and the speed benefit is immediate. For this category, the per-seat sticker price is what it appears to be.

On-premises AI is the right answer for the opposite category: proprietary data whose corpus encodes years of organizational investment, custom training that produces a model carrying internal expertise, agentic workflows with persistent memory over organizational data, any task where audit control over each inference matters, and high-token workloads running over long durations. The risks the Cloud AI Providers section of this paper catalogues (data sovereignty, cost compounding, lock-in, the customization ceiling, outage exposure, and regulatory trajectory) bite in this category and not the first. That asymmetry, not a blanket verdict on cloud AI, is the point. A sensible reader should not conclude “abandon cloud AI.” The conclusion is “stop conflating the two categories.”

This resolves the question financial decision makers raise most often: if cloud AI carries those risks, why does this paper still recommend paying for Claude or Copilot at all? Because the risks attach to a specific subset of workloads, and that subset is exactly where on-premises infrastructure earns its return. The cloud AI line item does not disappear. It gets right-sized.

What Vendor-Bundled AI Does Well

Microsoft 365 Copilot in Word, Outlook, and Excel earns its place for drafting, summarization, and inline help against documents the user already has open. The product is not theatre; it is built where the work happens, and that placement is most of its value. Adoption tracks the usefulness: Microsoft reported 15 million paid Microsoft 365 Copilot seats by Q1 2026, up about 160% year over year [4]. GitHub Copilot and Claude Code measurably speed routine coding, especially code that follows established patterns, and the productivity research, read even at its most skeptical, finds real effects in well-scoped tasks; the Where AI Delivers Real Value section of this paper weighs that evidence in detail. Google Workspace AI is the natural fit for organizations already standardized on Google’s productivity suite, and ChatGPT Team and Enterprise cover broad knowledge work with the model family that holds the most consumer mindshare.

None of these products are bad. All of them are built for horizontal deployment across millions of customers, so they optimize for general capability, which is what the workload framework wants from them. The question is what happens at the ceiling.

The Customization Ceiling in 2026

The ceiling moved between 2024 and 2026, and the move changes the argument. Microsoft, Google, and AWS now offer some form of fine-tuning. The claim is no longer “fine-tuning is unavailable”; it is “fine-tuning is available under constraints that on-premises infrastructure removes.” Those constraints carry the argument, and they differ by vendor.

Microsoft 365 Copilot Tuning is the most relevant case for the organizations this paper addresses, because it excludes them by design. It entered public preview in 2025 with worldwide rollout planned for June 2026 [5], [6]. It fine-tunes agents on SharePoint content for a fixed set of task templates (document writing, summarization, expert answers, validation, and style editing), combining supervised fine-tuning, reinforcement learning, and reasoning fine-tuning. The binding constraint is eligibility: only tenants with at least 5,000 Microsoft 365 Copilot licenses can use it during preview [6]. For a sub-100-employee organization, that threshold puts Copilot Tuning out of reach entirely. Even above it, the capability is locked to Microsoft’s task templates, locked to SharePoint as the data source, and produces a model whose weights cannot leave Microsoft’s infrastructure. The organization leases a tuned capability; it never owns an asset.

OpenAI is closing its self-serve fine-tuning platform on a published schedule. Its deprecations notice, posted in early May 2026, blocked organizations that had not previously run fine-tuning jobs from creating new ones, and sets January 6, 2027 as the date after which no one can start a new fine-tuning job; inference on already-tuned models then continues only until their base models retire [7]. OpenAI’s stated reason is that GPT-5.5-class models handle, through prompt engineering and retrieval-augmented generation (RAG), most of what fine-tuning used to do — a defensible claim for format and tone work, and a weak one for a model that must carry a proprietary corpus. The direction keeps both the workflow and the artifact inside OpenAI’s stack.

Azure OpenAI fine-tunes GPT-4o, GPT-4o mini, the GPT-4.1 family, and o4-mini, and as of early 2026 offers reinforcement fine-tuning of GPT-5 under gated, invitation-only access [8]. Custom deployments bill hourly whether or not they serve traffic, which makes fine-tuning pay off only at high utilization. Every tuned model stays on Microsoft’s infrastructure.

Anthropic offers no fine-tuning through its Enterprise subscription or public API. The lone public pathway is Claude 3 Haiku on Amazon Bedrock: one older model, text-only, 32K-context, on AWS hardware [9], [10]. No standard commercial path fine-tunes Claude Sonnet, Opus, or the Claude 4.x generation, and Claude’s weights are not available for export or self-hosting.

Google Gemini Enterprise Agent Platform offers the widest cloud fine-tuning surface: supervised fine-tuning of Gemini on text, image, audio, video, and document data [11]. That is broader than Anthropic’s near-zero availability and broader than Azure’s OpenAI ceiling, and it leaves the architecture intact. Tuned Gemini models run on Google’s infrastructure, the training data transits Google’s systems, and the result is reachable only through Google APIs.

The pattern holds across every vendor. As of May 2026, none of them lets an organization fine-tune a model on arbitrary internal sources (engineering specifications, operational logs, incident records, proprietary documentation), then export the resulting weights and run them on its own hardware with full audit control over each inference. Self-hosted fine-tuning of open-weight models such as Gemma, Mistral, gpt-oss, Llama, and Qwen does all of that. That gap is the precise, defensible form of the ceiling argument, and it is where this paper’s strategy lives.

What Vendor AI Cannot Do

Four limits define what no vendor sells today:

Data sovereignty comes first. Every cloud AI deployment moves data through vendor infrastructure regardless of contract tier. The training-pipeline risk is largely settled: commercial-tier contracts now exclude that data, as the Cloud AI Providers section of this paper documents. Sovereignty is not. The data still leaves the building, vendor policy can change (Anthropic’s 2025 shift in its consumer terms, announced in August and enforced that October, pulled consumer-tier conversations into model training unless users opted out [12]), and employees using personal accounts create exposure no Enterprise contract reaches.

Agentic infrastructure on proprietary data at production scale comes second. Microsoft Copilot Studio runs multi-agent orchestration and autonomous tasks, but execution happens on Microsoft’s infrastructure, agent memory and state cannot be inspected or controlled independently of the platform, and the capability needs an Azure subscription plus Copilot Studio capacity packs (a $200 per-tenant-per-month baseline for 25,000 messages) on top of per-seat Copilot licensing [13]. Anthropic’s Claude Code and Cowork allow third-party inference gateway routing as of April 2026. This enables the possibility to point the Claude Code and Cowork clients at a self-hosted vLLM endpoint which routes tokens through a non-Claude model. [14].

Domain models that carry organizational expertise come third. GitHub Copilot Enterprise tailors code suggestions to a private codebase through repository indexing and retrieval, which handles code-pattern relevance well, but it does not reach engineering documentation, runbooks, specifications, or reasoning across non-code corpora, and whatever it produces stays tied to GitHub’s infrastructure and deprecation schedule.

Audit control over each inference comes fourth. In a cloud deployment, the organization cannot see what the model received, what it returned, or what was logged at the depth a self-hosted system allows, except through the vendor. For regulated work, that is not a peripheral concern.

The Investment Trajectory Argument

The four largest hyperscalers (Microsoft, Amazon, Alphabet, and Meta) set roughly $725 billion in combined 2026 capital expenditure after Q1 2026 earnings, up 77% from $410 billion the year before [15]. Counting Oracle, the five-company group’s 2026 spending roughly doubles the $443 billion it deployed in 2025 [16], [17]. Microsoft’s CFO attributed $25 billion of its $190 billion figure to memory and component price inflation alone [15]. Goldman Sachs projects total hyperscaler capital expenditure from 2025 through 2027 at $1.15 trillion, more than double the $477 billion spent from 2022 through 2024 [17]. About 75% of the 2026 aggregate goes to AI-specific infrastructure: GPUs, servers, data center construction, power, and cooling [17].

This spending is not philanthropy. It builds the subscription layer through which enterprise AI capability will be delivered for the next decade, and no firm commits a trillion dollars to give the resulting capability away. Gartner fills in the commercial picture: 40% of enterprise applications will embed task-specific AI agents by the end of 2026, up from under 5% in 2025 [18], and generative AI and agents will drive a $58 billion disruption of mainstream productivity tools over the next three years [19]. Lock-in is no longer confined to dedicated AI products. It is being built into enterprise applications the organization already cannot work without.

The organizational consequence is direct. An organization that builds no internal AI capability becomes one of the primary customers of all this infrastructure: dependent on it for capability, exposed to its pricing, and unable to differentiate from competitors buying the same capability from the same vendors. The cost compounding the Cloud AI Providers section of this paper documents (CloudZero’s 36% year-over-year rise in monthly AI spend, GitHub Copilot’s June 2026 move to usage-based billing, Anthropic’s Enterprise pricing restructuring) is the early signal of how that dependence prices itself once the build-out has to pay for itself. Building internal capability now buys optionality that the trajectory makes more expensive every quarter.

Where the Gap Becomes Competitive Advantage

The case is not against the vendors. Microsoft, Anthropic, OpenAI, Google, and AWS are making rational use of extraordinary capital, and their products are the right answer for the large share of organizational AI work that is genuinely commodity. The case is against mistaking that share for the whole. The work the framework assigns to on-premises infrastructure — proprietary corpora, owned model weights, agents with persistent memory over internal data, audited inference — is exactly the work no vendor sells today and none is on a credible path to selling tomorrow. Organizations that build that capability now converts a rising rental cost into an owned asset that compounds; organizations that do not will keep paying the same vendors for the same capabilities its competitors buy. The differentiation is in the work the cloud cannot do.

References

  1. Anthropic, “Is My Data Used for Model Training?” Anthropic Privacy Center, Mar. 16, 2026. [Online]. Available: https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training. [Accessed: 18-May-2026]

    REBT-1 Primary source Back to text

  2. OpenAI, “Sharing Feedback, Evaluation and Fine-Tuning Data, and API Inputs and Outputs with OpenAI,” OpenAI Help Center, 2026. [Online]. Available: https://help.openai.com/en/articles/10306912-sharing-feedback-evaluation-and-fine-tuning-data-and-api-inputs-and-outputs-with-openai. [Accessed: 18-May-2026]

    REBT-2 Primary source Back to text

  3. Amazon Web Services, “Security, Privacy, and Responsible AI — Amazon Bedrock,” AWS, 2026. [Online]. Available: https://aws.amazon.com/bedrock/security-privacy-responsible-ai/. [Accessed: 18-May-2026]

    REBT-3 Primary source Back to text

  4. IntuitionLabs, “Microsoft Copilot Pricing & Licensing Guide for Business,” Jan. 2026. [Online]. Available: https://intuitionlabs.ai/articles/microsoft-copilot-pricing-licensing. [Accessed: 18-May-2026]

    REBT-4 Contextual source Back to text

  5. Microsoft Corporation, “Microsoft 365 Copilot Tuning Overview (Early Access Preview),” Microsoft Learn, Mar. 10, 2026. [Online]. Available: https://learn.microsoft.com/en-us/microsoft-365/copilot/copilot-tuning-overview. [Accessed: 18-May-2026]

    REBT-5 Primary source Back to text

  6. Microsoft Corporation, “Microsoft 365 Copilot Tuning Admin Guide (Early Access Preview),” Microsoft Learn, Mar. 10, 2026. [Online]. Available: https://learn.microsoft.com/en-us/microsoft-365/copilot/copilot-tuning-admin-guide. [Accessed: 18-May-2026]

    REBT-6 Primary source Back to text

  7. OpenAI, “Deprecations,” OpenAI API, 2026. [Online]. Available: https://developers.openai.com/api/docs/deprecations. [Accessed: 28-May-2026]

    REBT-7 Primary source Back to text

  8. Microsoft Corporation, “Customize a Model with Fine-Tuning — Microsoft Foundry,” Microsoft Learn, Feb. 27, 2026. [Online]. Available: https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/fine-tuning. [Accessed: 18-May-2026]

    REBT-8 Primary source Back to text

  9. Amazon Web Services, “Fine-Tuning for Anthropic's Claude 3 Haiku Model in Amazon Bedrock Is Now Generally Available,” AWS Blog, Nov. 8, 2024. [Online]. Available: https://aws.amazon.com/blogs/aws/fine-tuning-for-anthropics-claude-3-haiku-model-in-amazon-bedrock-is-now-generally-available/. [Accessed: 18-May-2026]

    REBT-9 Secondary source Back to text

  10. ClaudeGuide, “Can You Fine-Tune Claude? What's Available Instead,” Apr. 2026. [Online]. Available: https://claudeguide.io/can-you-fine-tune-claude. [Accessed: 18-May-2026]

    REBT-10 Contextual source Back to text

  11. Google, “About Supervised Fine-Tuning for Gemini Models,” Google Cloud Documentation, Apr. 29, 2026. [Online]. Available: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini-supervised-tuning. [Accessed: 28-May-2026]

    REBT-11 Primary source Back to text

  12. Anthropic, “Updates to Consumer Terms and Privacy Policy,” Anthropic, Aug. 28, 2025. [Online]. Available: https://www.anthropic.com/news/updates-to-our-consumer-terms. [Accessed: 28-May-2026]

    REBT-12 Secondary source Back to text

  13. Microsoft Corporation, “Microsoft Copilot Studio Documentation,” Microsoft Learn, 2026. [Online]. Available: https://learn.microsoft.com/en-us/microsoft-copilot-studio. [Accessed: 18-May-2026]

    REBT-13 Primary source Back to text

  14. Anthropic, “LLM gateway configuration,” Claude Code Documentation. [Online]. Available: https://code.claude.com/docs/en/llm-gateway. [Accessed: 29-May-2026]

    REBT-14 Primary source Back to text

  15. S Morris, R McMorrow, H Murphy, R Rosner-Uddin, and M Acton, “Google outpaces rivals as Big Tech's AI spending plans rise to $725bn,” Financial Times, Apr. 29, 2026. [Online]. Available: https://www.ft.com/content/2138e81c-4d86-46f4-8ca0-287f8b737cdf. Subscription required. Tom's Hardware report: https://www.tomshardware.com/tech-industry/big-tech/big-techs-ai-spending-plans-reach-725-billion. [Accessed: 18-May-2026]

    REBT-15 Contextual source Back to text

  16. Futurum Group, “AI Capex 2026: The $690B Infrastructure Sprint,” Futurum Research, Feb. 12, 2026. [Online]. Available: https://futurumgroup.com/insights/ai-capex-2026-the-690b-infrastructure-sprint/. [Accessed: 18-May-2026]

    REBT-16 Secondary source Back to text

  17. Introl Research, “Hyperscaler Capex Hits $600B in 2026,” Introl, Jan. 2026. [Online]. Available: https://introl.com/blog/hyperscaler-capex-600b-2026-ai-infrastructure-debt-january-2026. [Accessed: 18-May-2026]

    REBT-17 Contextual source Back to text

  18. Gartner, Inc., “Gartner Predicts 40% of Enterprise Applications Will Feature Task-Specific AI Agents by 2026, Up from Less Than 5% in 2025,” press release, Aug. 26, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-08-26-gartner-predicts-40-percent-of-enterprise-apps-will-feature-task-specific-ai-agents-by-2026-up-from-less-than-5-percent-in-2025. [Accessed: 18-May-2026]

    REBT-18 Secondary source Back to text

  19. Gartner, Inc., “Gartner Unveils Top Predictions for IT Organizations and Users in 2026 and Beyond,” press release, Oct. 21, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-10-21-gartner-unveils-top-predictions-for-it-organizations-and-users-in-2026-and-beyond. [Accessed: 18-May-2026]

    REBT-19 Secondary source Back to text

Contents