Section4
Cloud AI Providers: Capabilities, Plans, and Strategic Risks
Cloud AI is the right tool for commodity work and the wrong foundation for an organization’s core capability. The hyperscalers offer the most capable foundation models on the planet, billed per seat and per token, accessible from any browser within minutes of contract signature. For drafting a meeting invite, summarizing a public PDF, or pair-programming against an open-source codebase, this is an excellent answer. For proprietary data processing, custom model training on confidential corpora, and persistent agentic workflows, the same architecture creates dependencies that compound across cost, control, and customization in ways the per-seat sticker price never reveals. The argument that follows is not that any of these vendors are doing something wrong. Their products are improving rapidly and their data handling commitments at the commercial tier are stronger than most critics acknowledge. The argument is that an organization whose competitive advantage runs through its data cannot rent that advantage from a third party.
Anthropic Claude
Anthropic’s Claude line sets the current frontier for reasoning, long-context comprehension, and the kind of safety-constrained behavior that matters in engineering, legal, medical, and financial settings. The commercial plan structure as of May 2026 consists of Individual Pro at $17 per seat per month annual ($20 monthly), Individual Max 5x and 20x at $100 and $200 per month, Team Standard at $25 per seat per month billed monthly ($20/mo if billed annually) with a five-seat minimum, Team Premium at $125 per seat per month ($100/mo if billed annually) with Claude Code and Cowork included, and Enterprise at custom pricing starting at roughly $50,000 per year with SSO, SCIM, audit logging, HIPAA BAA, custom retention, and a 500K-token context window [1], [2]. Commercial-tier data is excluded from model training as a contractual guarantee under the Commercial Terms — not an opt-in toggle [3]. By default, API logs are retained for seven days and Zero Data Retention is available to qualifying Enterprise customers [4].
Two findings deserve direct attention from financial decision makers. First, Anthropic restructured Enterprise pricing between November 2025 and February 2026 to eliminate bundled token discounts. The base seat fee dropped, but all token consumption now bills at standard API rates. For a 100-seat mid-market deployment at realistic usage, the unbundling adds $15,000–$40,000 to annual total cost of ownership compared to the legacy structure [5]. This is reflected in a later section of this paper discussing the ROI model. Second, and structurally more important: Anthropic does not offer Claude fine-tuning through the public API as of May 2026, and no timeline has been announced [6],[7]. The only public fine-tuning pathway for any Claude model is Claude 3 Haiku on Amazon Bedrock — one older model, text-only, capped at 32K context, available solely through AWS infrastructure [8]. Fine-tuning Claude Sonnet, Opus, or any of the Claude 4.x generation is not available through any standard commercial pathway. Enterprise customers approach domain adaptation through retrieval-augmented generation (RAG), large cached system prompts, and Anthropic’s open Agent Skills standard — techniques that Anthropic itself argues approximate most fine-tuning use cases [9]. For text classification or tone calibration on commodity tasks, that argument holds. For training a model that internalizes decades of proprietary engineering documentation, it does not.
Claude’s lead concentrates in two kinds of work: building and maintaining software, and reasoning over large document sets. Claude Opus 4.8, released May 28, 2026, builds on Opus 4.7 at the same price and pairs a further coding gain with a measurable improvement in honesty: Anthropic reports the model is roughly four times less likely than Opus 4.7 to let a flaw in its own code pass unremarked [10], [11]. The engineering evidence is direct. On SWE-bench Verified, the human-validated 500-problem subset of the SWE-bench software-engineering benchmark introduced by Jimenez and colleagues at Princeton in October 2023 [12], which hands a model a real GitHub issue inside a full repository and grades whether its patch makes the project’s own tests pass, Opus 4.8 resolves 88.6% of problems, up from 87.6% for Opus 4.7 [13], [11]. The trajectory is worth stating plainly: in the original 2023 evaluation, the best model the authors tested resolved under 2% of the benchmark. On SWE-bench Pro, the harder variant spanning four programming languages, Opus 4.8 reaches 69.2%, a gain of nearly five points over Opus 4.7 (64.3%) in six weeks. One caveat to note: Anthropic’s own memorization screen flags a fraction of these problems as possibly present in training data, and the company reports its margin over the prior model holds once those problems are excluded [11]. The knowledge-work results follow the same line. Opus 4.8 posts leading scores, by Anthropic’s measurement, on Finance Agent, a test of multi-step financial analysis over source documents, and reaches a state-of-the-art 1890 Elo on the independent Artificial Analysis reproduction of GDPval, the benchmark from Patwardhan and colleagues that grades model deliverables against the real work product of professionals averaging fourteen years of experience across 44 occupations [14], [10]. The one category where it does not lead is terminal-centric coding, where GPT-5.5 retains an edge on Terminal-Bench 2.1 [10].
Anthropic has built these capabilities into two agent products that represent what Claude looks like in daily practice. Claude Code is a command-line agent for software developers: it navigates repositories, writes and patches code across multiple files, runs tests, and sustains autonomous work across multi-session coding projects without requiring the developer to direct each step [15]. With Opus 4.8, Claude Code added dynamic workflows, a research-preview capability in which the model plans a large task, runs hundreds of parallel subagents, and verifies their output before reporting back, aimed at repository-scale migrations [10]. The product is the full engineering session, not just the code suggestion. Claude Cowork, generally available for enterprise customers since April 2026, extends the same agentic architecture to non-engineering knowledge workers; it runs on the desktop, connects to local files and external services such as Google Drive and DocuSign, executes scheduled recurring tasks, and produces finished work rather than conversational suggestions [16]. Enterprise deployments add role-based access controls, spend limits, and OpenTelemetry streaming to a security information and event management (SIEM) system [16]. On Claude Code specifically, Menlo Ventures attributed Anthropic’s 42% enterprise coding share, more than double OpenAI’s 21%, to its adoption trajectory [17]. One constraint Anthropic states explicitly: Cowork is not cleared for HIPAA, FedRAMP, or financial services regulated workloads as of May 2026 [16]. Both products can use Anthropic’s cloud API inference infrastructure or a local offline language model inference server that supports the Anthropic API protocol.
That capability arrives where a mid-market buyer can reach it. Anthropic ships the model through its API and the three clouds these organizations already pay for: Amazon Bedrock, Google Vertex AI, and Microsoft Foundry [18]. The market has followed the capability. Menlo Ventures estimates Anthropic took 40% of enterprise large language model spend in 2025, up from 12% in 2023, and a 42% share of code generation specifically, more than double the next vendor; Menlo holds a stake in Anthropic, which is worth weighing against its own numbers [17].
A development in June 2026 made the strategic stakes concrete. On June 9, Anthropic released Claude Fable 5 and Claude Mythos 5, the first models in a new Mythos-class tier that sits above the Opus class in capability [19], [20]. The two share the same underlying model and the same specifications, including a 1M-token context window: Fable 5 is generally available with safety classifiers that route a small fraction of sensitive sessions to Opus 4.8, while Mythos 5 ships without those classifiers and is restricted to vetted organizations through Anthropic’s Project Glasswing [20]. Both are priced at $10 per million input tokens and $50 per million output tokens, double the Opus 4.8 rate, and both are designated Covered Models carrying a mandatory 30-day data-retention window with no Zero Data Retention option, a stricter posture than the rest of the Claude line [20]. Three days later, on June 12, the picture changed entirely. Citing national security authorities, the US government issued an export-control directive ordering Anthropic to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States, including Anthropic’s own foreign-national employees [21]. Unable to filter access at that granularity in real time, Anthropic disabled both models for every customer worldwide to ensure compliance; access to its other models, including Opus 4.8, was unaffected [21], [22].
The suspension ran from June 12 to July 1. The directive followed a report in which Amazon researchers bypassed Fable 5’s safeguards, prompting the model to identify software vulnerabilities and, in one case, produce code demonstrating how a vulnerability could be exploited; Anthropic’s own testing found that every model it checked, including Claude Opus 4.8, GPT-5.5, and Kimi K2.7, could produce the same demonstration [23]. Anthropic retrained the relevant safety classifier to block the reported technique in over 99% of cases, accepted a higher false-positive rate on benign coding requests as the cost, and submitted both the prior and new safeguards to testing by researchers from the Commerce Department’s Center for AI Standards and Innovation (CAISI) [23]. The government approved a limited redeployment of Mythos 5 to US critical-infrastructure operators on June 26, lifted the export controls on June 30, and Fable 5 returned to worldwide availability on July 1 on Anthropic’s own platforms, with subscription plans including it only through July 19 before it converts to metered usage credits, and cloud-marketplace access following afterward [23]. Anthropic paired the redeployment with a proposed industry framework, developed with Amazon, Microsoft, Google, and other Glasswing partners, for scoring jailbreak severity, and with commitments to expanded pre-release government access to frontier models and their safeguards [23]. The operative facts for a buyer are two. A frontier model deployed to a global customer base became unavailable to all of them within seventy-two hours of launch, by an action no customer could anticipate, contest, or route around [21]. And when it returned nineteen days later, it returned on terms negotiated between the vendor and the government, with stricter classifiers that refuse more benign requests and subscription access converted to metered credits, none of which any customer had a part in setting [23].
This paper cites Claude throughout as representing the current capability frontier. The strategic concern is not the model’s quality. The concern is that the most capable reasoning system available cannot be self-hosted, cannot be fine-tuned on proprietary data through its native delivery channel, and cannot be operated without data transiting Anthropic infrastructure. The position this white paper argues for is open-weight models that approximate this frontier with evidence documenting how close the open-weight gap has closed.
OpenAI and ChatGPT Enterprise
OpenAI remains the broadest-capability generalist, particularly in coding and multimodal work, and its enterprise traction is substantial. Its flagship line advanced twice in 2026. GPT-5.5, released April 23, brought particular gains in agentic coding, computer use, knowledge work, and extended-context reasoning [24], [25]. GPT-5.6 succeeded it on July 9, shipping as a three-variant family: Sol as the most capable version, Terra as the balanced everyday option, and Luna built for speed and cost, with OpenAI claiming Sol is 54% more token-efficient on agentic coding tasks than its predecessor [26], [27]. Two details of the GPT-5.6 launch bear on this section’s argument. OpenAI describes GPT-5.6 as its strongest cybersecurity model to date, and the US government asked OpenAI to stagger the model’s rollout before broad public release; the same pre-release lever applied to Anthropic’s Fable 5 weeks earlier, now exercised on a second vendor [26], [27]. Alongside the models, OpenAI launched ChatGPT Work, an agent that gathers context across connected apps and files to produce finished documents, spreadsheets, and presentations, its direct answer to Claude Cowork [26], [27].
ChatGPT’s subscription structure spans seven tiers: Free ($0), Go ($8/month), Plus ($20/month), two Pro variants at $100 and $200 per month, Business ($20/user/month billed annually), and Enterprise (custom) [28], [29]. Enterprise requires a 150-seat minimum and mandatory annual commitment, putting the contract floor at approximately $108,000 per year at the reported ~$60/user/month rate, though negotiated contracts range considerably below that. Enterprise customers receive virtually unlimited GPT-5.5 Instant messages, with context windows of 128K tokens for GPT-5.5 Instant and 196K for GPT-5.5 Thinking. API access to GPT-5.5 carries a 1M token context window, priced at $5 per million input tokens and $30 per million output tokens — double the GPT-5.4 rate. For high-volume workloads, prompts exceeding 272K input tokens are billed at 2x input and 1.5x output for the full session. Enterprise and API customers remain outside OpenAI’s training data pipeline by default [30]; Pro, Plus, and lower tiers require manual opt-out.
The more consequential development for enterprise buyers is OpenAI’s exit from self-serve model customization. On May 7, 2026, OpenAI notified developers that organizations which had not previously run fine-tuning jobs could no longer create new ones. Existing active users retain the ability to create training jobs for the coming months, and all fine-tuned models remain available for inference until their underlying base models are deprecated. The full shutdown timeline runs through January 6, 2027, at which point no new fine-tuning jobs can be initiated by anyone [31], [32]. OpenAI’s rationale, stated directly, is that new-generation base models like GPT-5.5 cover the vast majority of scenarios previously handled by fine-tuning through prompt engineering and retrieval — at lower cost and faster iteration cycles.
That rationale has merit in many cases. It is less convincing for organizations that need behavioral guarantees beyond what prompt engineering can enforce — specific output formats, domain terminology, proprietary reasoning patterns, or suppression of behaviors the base model exhibits by default. For teams willing to work with open-source models, parameter-efficient techniques like LoRA and QLoRA remain fully available and unaffected by OpenAI’s platform decision, which is itself an argument for the on-premises deployment path this paper examines. The trajectory is clear: OpenAI is steering workflow customization inside its own managed stack [33]. For an organization that had planned to treat fine-tuned model weights as an internal asset, that option closes on a defined schedule.
OpenAI’s Codex is the primary competitor to Anthropic’s Claude Code [34]. Codex is used by more than 4 million developers each week and has been recognized as a Leader in the Gartner Magic Quadrant for Enterprise AI Coding Agents [35], with OpenAI positioning it as a platform for broader enterprise workflows beyond software development alone. Codex is available as a command-line tool, a desktop application for macOS and Windows, and through the ChatGPT web interface, with agents running in sandboxed environments that limit file and network access by default.
Microsoft: M365 Copilot, GitHub Copilot, Azure OpenAI, Copilot Studio
Microsoft sells the most thoroughly bundled AI stack in the market. M365 Copilot Enterprise is an add-on at $30 per user per month annual on top of a qualifying Microsoft 365 enterprise base license: E3 @ +$36/user/mo, E5 @ +$57/user/mo and increasing to +$60 effective July 1, 2026, E7 @ +$99/user/mo [36], [37], [38]. The base 365 license is mandatory; Copilot does not activate without one. Copilot Business for SMB tenants under 300 seats lists at $21 per user per month annual standard. Copilot Studio, the platform for building custom agents and workflows, runs $0.01 per tenant per month per message with pay-as-you-go usage ($250 for 25k messages) or prepaid capacity packs at $200 per tenant per month for 25,000 messages, and requires an Azure subscription for agent capabilities — neither cost is included in M365 Copilot licensing [39]. Microsoft consolidated the previously separate Copilot for Sales, Service, and Finance add-ons into the base $30 license in October 2025.
GitHub Copilot prices at $19 per user per month for Business and $39 per user per month for Enterprise, with Enterprise requiring GitHub Enterprise Cloud at an additional $21 per user per month — total real cost $60 per user per month for new Enterprise customers [40]. On June 1, 2026, GitHub moved all Copilot plans from premium-request billing to usage-based token billing [41], [42]. The new system replaces the prior premium request unit with GitHub AI Credits, where one credit equals one US cent, consumed against the input, output, and cached tokens each interaction uses at published per-model API rates; code completions and Next Edit Suggestions remain unmetered, while Copilot code review, now agentic, additionally draws down GitHub Actions minutes [42]. Base seat prices held and each plan includes a monthly AI Credit allotment matched to its price, but the included allowance effectively shrank, the prior fallback to cheaper models when quota is exhausted was removed, and existing Business and Enterprise customers receive a larger allotment only through a promotional window running June 1 to September 1, 2026 [41], [42]. GitHub’s CPO articulated the rationale plainly: a quiet user costs almost nothing to serve, while a power user orchestrating agentic workflows with frontier models can cost an order of magnitude more. The change drew immediate pushback from developers, some of whom projected heavy agentic sessions costing several multiples of the flat plan they replaced [43]. The $19 per seat figure that anchors most Copilot ROI models — including this paper’s avoided-cost calculation discussed in the ROI Case section — now represents only the floor. For engineering teams using Copilot in agentic mode against frontier models, actual costs will run materially higher and will fluctuate month to month.
Customization through the Microsoft stack runs along two paths, each with specific limits as of May 2026. Azure OpenAI / AI Foundry supports fine-tuning of GPT-4o and GPT-4o mini, the GPT-4.1 family, o4-mini, gpt-5, and several open-weight models [44]. Currently, GPT-5 support for reinforcement fine-tuning is generally available, but access is gated and available by invitation only. Training cost is dependent on the chosen model, quantity of tokens, and training duration; starting at ~$0.10 to ~$2.00 per million input tokens, ~$0.30 to ~$8.00 per million output tokens, and ~$0.75 to ~$27 per million training tokens with newer models billing $100+ per training hour [45]. Custom model deployments incur hourly hosting charges regardless of usage with automatic deletion of deployments inactive longer than fifteen days. Azure AI Foundry also supports fine-tuning of 1,600+ third-party models (Llama, Mistral, Phi, Cohere) through managed compute, which requires the customer to provide GPU VM quota. Copilot Studio, despite its name and positioning, is not a fine-tuning platform — it customizes through knowledge bases, topics, actions, and system prompts but does not modify model weights [39]. For an organization that wants Microsoft Copilot to behave like a domain expert on its proprietary corpus, the available tools are RAG and prompt engineering, not weight modification.
One regulatory development matters for executive readers. The Australian Competition and Consumer Commission sued Microsoft in October 2025, alleging that bundling Copilot into rising Microsoft 365 subscription prices misled 2.7 million Australian customers by not disclosing a cheaper “Classic” plan option [46]. Microsoft began contacting affected customers with refund offers in November 2025; the case remains in early stages. As documented evidence of how bundled AI pricing can force customers into higher tiers without clear opt-out, it is the cleanest example currently on record.
Google Gemini, Workspace AI, and the Gemini Enterprise Agent Platform
Google integrated their Gemini AI models into most of their consumer and business products. The same Gemini models sit inside the productivity tools a mid-market company already runs, behind a developer API, and inside a managed platform for building agents. As of early 2026, Gemini ships inside every paid Google Workspace plan rather than as a separate add-on. Google folded the old standalone Gemini Business and Enterprise SKUs into the base subscriptions in 2025 and raised seat prices to absorb the cost, so Gemini in Gmail, Docs, Sheets, Slides, Chat, and Meet, along with the NotebookLM research tool, now comes bundled. Workspace Business Standard lists at $14 per seat per month on an annual commitment with Gemini across the apps included; Business Starter at $7 carries Gemini in Gmail only [47].
Above Workspace sits Gemini Enterprise, a separate agentic platform Google positions as the front door to AI at work. It combines permission-aware search across company data, prebuilt and custom agents, a no-code agent builder, and central governance, with prebuilt connectors to systems like SharePoint, Confluence, Jira, and ServiceNow [48]. The Business edition serves individuals and teams up to 300 seats with no IT setup starting at $21 per seat per month. The Standard and Plus editions starting at $30 per seat per month add unlimited seats, third-party agent access, Gemini Code Assist, and the controls a regulated buyer needs: VPC Service Controls, customer-managed encryption keys, Access Transparency, data residency, and support for HIPAA and FedRAMP High workloads [48]. Google publishes a self-serve price only for the Business edition and routes Standard and Plus through sales, so a buyer pricing those tiers should get a written quote rather than rely on a list figure [49].
For teams that build rather than buy, the Gemini Enterprise Agent Platform, the service formerly called Vertex AI, is where the model work happens. It provides a Model Garden spanning Gemini, Anthropic’s Claude, Google’s open Gemma models, and third-party options, plus agent tooling, model evaluation, and tuning [50]. Published API rates run from Gemini 2.5 Flash-Lite at $0.10 input and $0.40 output per million tokens, through Gemini 2.5 Pro at $1.25 and $10.00, to the Gemini 3.1 Pro preview at $2.00 and $12.00 and the latest Gemini 3.5 Flash at $1.50 and $9.00, all quoted at the standard paid tier [51].
Google offers the broadest cloud fine-tuning surface of the major providers. Supervised fine-tuning adjusts model weights against a labeled dataset and accepts text, image, audio, video, and document training data, wider than any competing cloud product [52]. One limit a buyer should confirm before committing: supervised tuning currently targets the Gemini 2.5 model family, and as of May 2026 the Gemini 3.1 Flash-Lite model is the only offering within the Gemini 3.x family, so a tuned production endpoint may have to migrate as older base models retire [52]. None of this changes the architectural fact the rest of this section keeps returning to. A fine-tuned Gemini model runs on Google infrastructure, the proprietary data used to train it passes through Google’s systems, and the resulting model is a Google asset reachable only through Google APIs. The customization is real. The independence is not.
Amazon Bedrock and Amazon Q
Amazon Bedrock is the most explicit of the cloud platforms on data isolation. AWS states that when you tune a foundation model on Bedrock the job runs against a private copy of that model, so your data is not shared with the model provider and is not used to improve the base model [53], [54]. Bedrock is in scope for ISO, SOC, and CSA STAR Level 2, is HIPAA eligible, supports GDPR compliance, and is FedRAMP High authorized in the AWS GovCloud (US-West) region; AWS PrivateLink keeps traffic between a customer VPC and Bedrock off the public internet [54]. The whole product is a single API in front of a large model catalog. As of May 2026 that catalog spans Amazon’s own Nova family, Anthropic’s Claude, Meta’s Llama, Mistral, Cohere, AI21, DeepSeek, Google’s Gemma, OpenAI’s open-weight models, and others, all reached through the same request shape, IAM roles, KMS keys, and CloudTrail logging as the rest of an AWS account.
Pricing is region-dependent, tier-dependent, and consumption-based per million tokens with no seat fee. For example, Amazon Nova Micro US-East Standard Tier runs about $0.035 input and $0.14 output; Claude Opus 4.x on Bedrock runs about $5.00 input and $25.00 output; batch inference takes 50% off on-demand rates, and Provisioned Throughput reserves capacity by the hour [55]. The costs that surprise teams sit next to the tokens. Bedrock Knowledge Bases, the managed retrieval-augmented generation (RAG) pipeline, defaults to Amazon OpenSearch Serverless (Classic collection type) as its vector store via the Quick Create flow [56], which enforces a two-OCU compute minimum—comprising 1 OCU for primary-and-standby indexing and 1 OCU for high-availability search replicas—at $0.24 per OCU-hour [57], thereby establishing an idle floor of approximately $350 per month ($0.24 × 2 OCUs × 730 hrs) regardless of query volume [57], [58]. Since December 2025, the generally available Amazon S3 Vectors offers the same retrieval at up to 90% lower vector storage and query cost and now serves as the cheaper default for new Knowledge Bases [59]. At its June 2026 New York Summit, AWS introduced a managed Knowledge Base for Bedrock that adds native data connectors, automatic multi-format parsing, and an agentic retriever for multi-step queries, integrated with the AgentCore Gateway [60]. Guardrails, the content-filtering and policy layer, bills separately per text unit [55].
Bedrock’s customization surface is wider than the Claude-only story most buyers know. Bedrock provides managed fine-tuning, continued pre-training, and model distillation without the customer running training infrastructure [61]. For Amazon’s Nova family it supports supervised fine-tuning, reinforcement fine-tuning, and distillation directly, with continued pre-training handled in SageMaker for teams teaching a model domain knowledge from raw text [62]. Model Distillation transfers behavior from a large teacher model into a smaller, cheaper student, for example Nova Premier into Nova Micro, and supports Claude and Llama teachers [63]. Customers can also import models trained elsewhere on SageMaker. For Anthropic’s models the ceiling is low and unchanged: Claude 3 Haiku is the only Claude model Bedrock will fine-tune, text-only, capped at 32K context, with training data confined to the customer’s AWS environment [8]. Every customized model, whatever its origin, runs inside Bedrock and is reachable only through Bedrock.
For agent workloads Amazon ships Bedrock AgentCore, generally available since October 2025, a framework- and model-agnostic runtime for production agents. It gives each session an isolated microVM with execution windows up to eight hours, managed short- and long-term memory, a gateway that turns existing APIs and Model Context Protocol servers into agent tools, identity-aware authorization, a sandboxed code interpreter and browser tool, and OpenTelemetry-compatible observability, all behind VPC and PrivateLink [64]. At the same June 2026 summit, AWS added a managed web-search tool to AgentCore that grounds agent responses in cited web results with no data egress from the customer’s AWS environment [60].
Amazon Q is the packaged assistant layer on top. Amazon Q Developer, an AWS-integrated coding assistant offering agentic task execution, code transformation, and security scanning, provides a perpetual Free Tier and a Pro Tier at $19 per user per month that enables teams to customize suggestions to their own codebase via enterprise access controls [65]. Q Business, a workplace assistant that answers questions across connected enterprise systems, runs $3 per user per month for Lite and $20 for Pro, with index-capacity charges on top and data handling that inherits Bedrock’s isolation guarantees [65], [54]. For an organization already standardized on AWS, Bedrock is the most flexible cloud AI consumption layer on the market. For an organization deciding where to base a multi-year AI strategy, Bedrock’s honesty about data isolation does not remove the underlying dependency: the models, the tuned weights, and the agents all live on AWS infrastructure and bill at AWS rates.
Cohere
Cohere is the one major commercial provider whose architecture is built around the deployment model this paper argues for. The Toronto company, founded in 2019 and focused on regulated industries and the public sector, sells its enterprise models to run inside the customer’s own environment rather than only behind a multi-tenant API [66]. Its generation lineup is the Command family, served through a single Chat endpoint with native retrieval-augmented generation (RAG) and span-level source citations: Command A+, Command A, Command R+, Command R, and the budget Command R7B [67]. On self-serve production keys, the flagship Command R+ and Command A price at $2.50 per million input tokens and $10 per million output tokens, below the Claude Opus rate cited earlier in this section and comparable to OpenAI’s GPT-5.4 on input, while Command R7B runs $0.0375 and $0.15 per million for high-volume retrieval work [68]. The newest model has a dual nature worth noting: Command A+, released May 20, 2026, is a 218-billion-parameter sparse mixture-of-experts model with 25 billion active parameters, published under a full Apache 2.0 license and able to run on as few as two H100 GPUs, which makes it at once an open-weight model an organization can self-host and a hosted product Cohere serves through its API and private-deployment tiers [69]. It appears among the production-ready open-weight models catalogued elsewhere in this paper for that reason.
Around the Command models Cohere builds a retrieval stack aimed at enterprise search. Embed 4 is a multimodal embedding model that maps text, images, and interleaved document content such as PDF screenshots, slides, and tables into one vector space, and its binary and integer quantization combined with Matryoshka dimension truncation can cut vector-storage footprint by roughly 96 percent against full-precision embeddings, a direct reduction in vector-database cost [70], [71]. Rerank 4, released in December 2025, is a reranking foundation model that raises retrieval precision across more than 100 languages and mixed formats including documents, email, tables, JSON, and code, and its 32K-token context window lets it weigh roughly fifty pages of a contract or filing in a single pass [72]. These components feed North, Cohere’s agentic enterprise platform, generally available since August 2025, which unites secure agents, the Compass enterprise-search layer, and generation from the Command, Embed, and Rerank models [73]. North runs in production at Royal Bank of Canada, which co-developed an on-premises North for Banking that keeps data in its own facilities, at Dell, which ships North inside the Dell AI Factory stack, and at Bell Canada, which both runs it internally and resells it as a sovereign AI service; the platform installs on-premises, in a virtual private cloud, or fully air-gapped on as few as two GPUs [73], [74].
The difference is how Cohere serves the commercial models themselves. Through Model Vault, launched in September 2025, Cohere deploys Command, Embed, and Rerank as managed inference inside an isolated virtual private cloud or fully on-premises, bringing the model to the data so that prompts and proprietary corpora never leave the organization’s network [72], [66]. The same models also reach customers through the public marketplaces they already use, including AWS Bedrock and SageMaker, Microsoft Azure AI Foundry, Google Cloud and Gemini Enterprise Agent Platform, and Oracle Cloud Infrastructure [66]. This is the capability that OpenAI, Anthropic, and Google do not offer: where their frontier models stay API-only and their weights never leave vendor infrastructure, Cohere will run its commercial models, weights included, inside an environment the customer controls. One line of products sits outside this picture. Cohere Labs, the company’s research arm, publishes the multilingual Aya family, including Aya Expanse, Aya Vision, and the February 2026 Tiny Aya 3.35-billion-parameter edge models that span more than 70 languages, but these carry non-commercial research licenses (CC BY-NC for the Aya Expanse and Vision releases) and so belong to neither the production open-weight tier nor the commercial-offering tier this paper evaluates [67].
The corporate trajectory reinforces the sovereignty framing. Cohere reported roughly $240 million in annual recurring revenue for 2025, reached a $7 billion valuation in a September 2025 round backed by Nvidia, AMD, Oracle, and Salesforce, and carried more than 500 employees into 2026 with a public listing widely anticipated [72], [75], [74]. In April 2026 it agreed to acquire the German lab Aleph Alpha in a transaction reported to value the combined company near $20 billion, anchored by a roughly $600 million Series E investment led by Germany’s Schwarz Group, which adds a European headquarters in Berlin and a foothold in EU public-sector and sovereign-AI markets [75], [76]. The acquisitions continued in July 2026, when Cohere announced the purchase of Reliant AI to extend its sovereign-deployment offering into biopharma and healthcare, two of the regulated sectors where the private-deployment argument bites hardest [77]. For the argument of this paper, Cohere is the instructive case among the commercial vendors. A provider whose models approach the enterprise frontier and can run on the customer’s own hardware narrows the distance between cloud convenience and on-premises control further than any API-only competitor. The caveats are real, since a vendor relationship, per-model licensing on the hosted tier, and a capability gap against the largest closed models all persist, but Cohere shows that private and on-premises deployment of capable commercial models is a shipping product rather than a theoretical option.
Risk Analysis: Where Cloud AI Becomes Untenable
Six risks structure the case against treating cloud AI as the foundation for an organization’s core AI capability. They are not equally weighted, and the most important is not the one most often cited.
Data sovereignty supersedes data training as the operative risk. The standard critique of cloud AI focuses on prompt content being used to train vendor models. As of May 2026, every major provider — Anthropic, OpenAI, Microsoft, Google, AWS — contractually excludes commercial-tier data from training [3], [30], [53]. The training-pipeline argument was strongest two years ago. Today the sharper edge of the same argument is sovereignty: data still transits vendor infrastructure regardless of training policy, vendor policies can change (Anthropic’s October 2025 consumer-tier opt-in mechanism is a real precedent), prompt logs exist even when not used for training, and accidental use of consumer-tier credentials by employees creates exposure that no Enterprise contract prevents. The security architecture section of this paper addresses the Shadow AI policy implications. The point here is that the on-premises argument on this dimension reads cleanest when framed as data never leaving the building, not as a training-pipeline argument that the major providers have substantially addressed.
Cost unpredictability is structural, not incidental. Per-token pricing appears modest for individual use and compounds rapidly across users, long-context tasks, and agentic workflows that generate dozens of model calls per user action. CloudZero’s 2025 survey of 500 engineering professionals reported average monthly AI spending rising from $62,964 to $85,521 — a 36% jump in a single year — with the share of organizations spending over $100,000 per month doubling from 20% to 45%, and only 51% confident they could evaluate AI ROI at all [78]. GitHub Copilot’s switch to usage-based billing, live since June 1, 2026, is the cleanest current illustration of the dynamic: GitHub’s own CPO acknowledged that a power user orchestrating agentic workflows with frontier models can cost an order of magnitude more than a quiet user, which is precisely why GitHub moved to convert the predictable per-seat cost into a variable inference bill [41]. Anthropic’s Enterprise unbundling does the same thing through a different mechanism, and the July 2026 redeployment of Fable 5 extended it further by moving the most capable Claude model out of subscription inclusion and onto metered usage credits after a one-week window [23]. OpenAI is running the same conversion inside ChatGPT Enterprise: workspace agents moved from a free period to credit-based, token-metered pricing on July 6, 2026, and its spreadsheet and presentation add-ins follow the same token-billed model as their free previews expire [79]. The direction across vendors is consistent: as agentic and long-context workloads become the dominant usage pattern, the pricing models that made cloud AI look cheap at the seat level are being replaced by models that scale with inference volume.
Vendor lock-in operates at three depths. Lock-in at the API surface — provider-specific function calling schemas, prompt formats, and response structures — is the visible layer. Below it sits ecosystem lock-in to provider-specific tooling like OpenAI’s Responses API and Assistants migration, Anthropic’s MCP ecosystem, and Azure AI Foundry’s pipelines [80]. The deepest layer is the one most often missed: fine-tuned models built on vendor infrastructure cannot be extracted and self-hosted. A fine-tuned Gemini model is a Google asset. A fine-tuned GPT-4.1 model is an Azure asset. The fine-tuned weights themselves reside on vendor infrastructure and are accessible only through vendor APIs. Menlo Ventures data from late 2025, cited by Waehner, shows Anthropic at roughly 40% of enterprise LLM API spend and OpenAI at 27% (down from approximately 50% in 2023) — evidence that the API switching cost is real even when a competitive alternative becomes available.
The customization ceiling is the single least negotiable risk. Across the entire cloud AI market as of May 2026, no provider enables self-hosting of fine-tuned models. The matrix breaks down as follows: Anthropic direct API offers no fine-tuning at all; Anthropic via Amazon Bedrock offers Claude 3 Haiku fine-tuning, text-only, 32K context; OpenAI’s self-serve fine-tuning platform is winding down with new-user access already closed and existing-user access ending January 6, 2027; Azure OpenAI fine-tunes GPT-4.1 and earlier with gated access to GPT-5 fine-tuning; Google Gemini Enterprise Agent Platform fine-tunes Gemini across multiple modalities; Amazon Bedrock fine-tunes multiple non-flagship models. In every case the fine-tuned artifact lives on vendor infrastructure. Domain adaptation through RAG and prompt engineering can approximate fine-tuning for many use cases — Anthropic and Microsoft both argue this position credibly — but for an organization wanting to train a model that has internalized its proprietary corpus and run that model in isolation on its own hardware, no cloud product on the market today provides the answer. This is not a temporary capability gap. It is the architecture.
Outage and availability exposure is documented and trending worse. A major AWS outage in 2025 disrupted services across industries and became a focal point for enterprise resilience planning; Azure also suffered high-profile 2025 outages [81]. Forrester’s November 2025 Predictions 2026 report forecasts at least two major multi-day cloud outages in 2026, attributing the risk to hyperscalers diverting investment from legacy infrastructure to GPU-centric data centers while aging systems strain under growing complexity [82]. Outages are only one way a cloud-hosted model can disappear. In June 2026, Anthropic disabled its two most capable models, Fable 5 and Mythos 5, for every customer worldwide within three days of their launch, complying with a US government export-control directive it had no power to refuse [21], [22]. The cause was neither an outage nor a contract dispute, and no service-level agreement would have covered it; when the model is the vendor’s to grant, it is also the vendor’s, or a regulator’s, to withdraw. Access returned on July 1, 2026 after a nineteen-day suspension, and the return makes the point a second time: the model came back with a retrained classifier that refuses more benign requests and with subscription access converted to metered usage credits, terms set between the vendor and the government while every customer waited [23]. Organizations without on-premises fallback are exposed to service degradation events outside their control during exactly the periods — high-demand periods — when those events are most likely to occur.
Regulatory trajectory points one direction. White & Case wrote in their September 2025 AI Watch Insight “there are currently no comprehensive federal laws or regulations in the US that have been enacted specifically to regulate AI” [83]. The trajectory in the EU, UK, and US is toward stricter controls on AI data handling. The June 2026 Fable 5 and Mythos 5 suspension shows the lever is not limited to data handling: export-control authority can now disable a cloud-delivered model for an entire customer base overnight, an exposure that does not arise for open weights already running inside an organization’s own perimeter [84]. That lever hardened into standing practice within weeks. A June 2, 2026 Executive Order, Promoting Advanced Artificial Intelligence Innovation and Security, formalized federal involvement in frontier-model security, including an interagency cybersecurity vulnerability clearinghouse in which Anthropic has committed to participate [85], [23]; the Fable 5 directive followed ten days after the order, and in July the US government asked OpenAI to stagger the public rollout of GPT-5.6 [26]. Pre-release government review of cloud-delivered frontier models is now demonstrated practice across at least two vendors, and it applies at the delivery channel, which is exactly the layer a cloud customer does not control. Organizations that have already externalized AI workloads to cloud providers face future remediation costs that organizations building on-premises capability today will not. Deloitte’s 2026 State of AI in the Enterprise survey of 3,235 senior leaders across six industries and 24 countries found that only one in five companies has a mature governance model for autonomous AI agents — a gap that regulators are likely to close before enterprises do [86].
Where This Leads
Cloud AI is the right answer for some workloads and the wrong answer for others. Summarizing a public press release, drafting a meeting invite, generating boilerplate code against an open-source library, brainstorming a marketing tagline — these are tasks where data sensitivity is low, customization needs are minimal, and the marginal cost of a token is genuinely cheap. For proprietary data processing, training on confidential corpora, building agentic workflows with persistent memory over organizational data, any task where the model itself needs to encode internal expertise, and workloads consuming large quantities of tokens over long durations, the same cloud-based architecture, with the varying cost considerations, is structurally inadequate. The risks catalogued here do not make cloud AI categorically dangerous. They make it categorically insufficient as a sole foundation for the AI workloads mentioned. This paper defines exactly which workloads belong on which side of that line and constructs the layered architecture in which cloud AI and on-premises AI each serve the work they are suited to.
References
-
Anthropic, “Plans & Pricing,” Anthropic PBC, May 2026. [Online]. Available: https://claude.com/pricing. [Accessed: 18-May-2026]
-
Vantage Point, “Anthropic's Enterprise AI Tiers Explained: Pro, Max, Team, and Enterprise,” May 2026. [Online]. Available: https://vantagepoint.io/blog/sf/anthropic/enterprise-ai-tiers-explained. [Accessed: 18-May-2026]
-
Anthropic, “Is My Data Used for Model Training?” Anthropic Privacy Center, Mar. 16, 2026. [Online]. Available: https://privacy.claude.com/en/articles/7996868-is-my-data-used-for-model-training. [Accessed: 18-May-2026]
-
Anthropic, “API and data retention,” Claude API Docs. [Online]. Available: https://platform.claude.com/docs/en/manage-claude/api-and-data-retention. [Accessed: 28-May-2026]
-
The Register, “Anthropic Ejects Bundled Tokens from Enterprise Seat Deal,” The Register, Apr. 16, 2026. [Online]. Available: https://www.theregister.com/2026/04/16/anthropic_ejects_bundled_tokens_enterprise/. [Accessed: 18-May-2026]
-
Anthropic, “Glossary: Fine-tuning,” Claude API Docs, 2025. [Online]. Available: https://platform.claude.com/docs/en/about-claude/glossary#fine-tuning. [Accessed: 28-May-2026]
-
ClaudeGuide, “Can You Fine-Tune Claude? What's Available Instead,” Apr. 2026. [Online]. Available: https://claudeguide.io/can-you-fine-tune-claude. [Accessed: 18-May-2026]
-
Amazon Web Services, “Fine-Tuning for Anthropic's Claude 3 Haiku Model in Amazon Bedrock Is Now Generally Available,” AWS Blog, Nov. 8, 2024. [Online]. Available: https://aws.amazon.com/blogs/aws/fine-tuning-for-anthropics-claude-3-haiku-model-in-amazon-bedrock-is-now-generally-available/. [Accessed: 18-May-2026]
-
IntuitionLabs, “Claude Enterprise Guide 2026: Deployment & Training Specs,” 2026. [Online]. Available: https://intuitionlabs.ai/articles/claude-enterprise-deployment-training-guide-2026. [Accessed: 18-May-2026]
-
Anthropic, “Introducing Claude Opus 4.8,” Anthropic PBC, May 28, 2026. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-8. [Accessed: 18-Jun-2026]
-
Anthropic, “Claude Opus 4.8 System Card,” Anthropic PBC, May 28, 2026. [Online]. Available: https://www-cdn.anthropic.com/0b4915911bb0d19eca5b5ee635c80fef830a37ea.pdf. [Accessed: 18-Jun-2026]
-
C. E. Jimenez et al., “SWE-bench: Can Language Models Resolve Real-World GitHub Issues?” Proc. 12th Int. Conf. Learn. Represent. (ICLR), May 2024. [Online]. Available: https://arxiv.org/abs/2310.06770. Vienna, Austria. [Accessed: 28-May-2026]
-
Anthropic, “Claude Opus 4.7 System Card,” Anthropic PBC, Apr. 16, 2026. [Online]. Available: https://anthropic.com/claude-opus-4-7-system-card. [Accessed: 28-May-2026]
-
T. Patwardhan, R. Dias, E. Proehl, G. Kim, M. Wang, and O. Watkins et al., “GDPval: Evaluating AI Model Performance on Real-World Economically Valuable Tasks,” arXiv, OpenAI, Oct. 2025. [Online]. Available: https://arxiv.org/abs/2510.04374. arXiv:2510.04374. [Accessed: 28-May-2026]
-
Anthropic, “Claude Code,” Anthropic PBC, 2025. [Online]. Available: https://claude.com/product/claude-code. [Accessed: 28-May-2026]
-
Anthropic, “Claude Cowork,” Anthropic PBC, Apr. 2026. [Online]. Available: https://claude.com/product/cowork. [Accessed: 28-May-2026]
-
Menlo Ventures, “2025: The State of Generative AI in the Enterprise,” Jan. 2026. [Online]. Available: https://menlovc.com/perspective/2025-the-state-of-generative-ai-in-the-enterprise/. [Accessed: 28-May-2026]
-
Anthropic, “Introducing Claude Opus 4.7,” Anthropic PBC, Apr. 16, 2026. [Online]. Available: https://www.anthropic.com/news/claude-opus-4-7. [Accessed: 28-May-2026]
-
Anthropic, “Claude Fable 5 and Claude Mythos 5,” Anthropic PBC, June 9, 2026. [Online]. Available: https://www.anthropic.com/news/claude-fable-5-mythos-5. [Accessed: 18-Jun-2026]
-
Anthropic, “Introducing Claude Fable 5 and Claude Mythos 5,” Claude API Docs, June 9, 2026. [Online]. Available: https://platform.claude.com/docs/en/about-claude/models/introducing-claude-fable-5-and-claude-mythos-5. [Accessed: 18-Jun-2026]
-
Anthropic, “Statement on the US Government Directive to Suspend Access to Fable 5 and Mythos 5,” Anthropic PBC, press release, June 12, 2026. [Online]. Available: https://www.anthropic.com/news/fable-mythos-access. [Accessed: 18-Jun-2026]
-
CNBC, “Anthropic Disables Access to Fable 5 and Mythos 5 to Comply with Government Directive,” CNBC, June 12, 2026. [Online]. Available: https://www.cnbc.com/2026/06/12/anthropic-disables-access-to-fable-5-and-mythos-5-to-comply-with-government-directive.html. [Accessed: 18-Jun-2026]
-
Anthropic, “Redeploying Fable 5,” Anthropic News, June 30, 2026. [Online]. Available: https://www.anthropic.com/news/redeploying-fable-5. Updated Jul. 1, 2026. [Accessed: 15-Jul-2026]
-
OpenAI, “Introducing GPT-5.5,” OpenAI, Apr. 23, 2026. [Online]. Available: https://openai.com/index/introducing-gpt-5-5/. [Accessed: 28-May-2026]
-
OpenAI, “GPT-5.5 Model,” OpenAI API Reference, 2026. [Online]. Available: https://developers.openai.com/api/docs/models/gpt-5.5. [Accessed: 28-May-2026]
-
I. Fried and M. Mills, “OpenAI Releases GPT-5.6 and ChatGPT Work Tool,” Axios, July 9, 2026. [Online]. Available: https://www.axios.com/2026/07/09/ai-openai-gpt-release. [Accessed: 15-Jul-2026]
-
L. Ropek, “OpenAI Launches Its New Family of Models with GPT-5.6,” TechCrunch, July 9, 2026. [Online]. Available: https://techcrunch.com/2026/07/09/openai-launches-its-new-family-of-models-with-gpt-5-6/. [Accessed: 15-Jul-2026]
-
OpenAI, “ChatGPT Plans | Free, Go, Plus, Pro, Business, and Enterprise,” chatgpt.com. [Online]. Available: https://chatgpt.com/pricing/. [Accessed: 18-May-2026]
-
Finout, “OpenAI Pricing in 2026 for Individuals, Orgs & Developers,” May 2026. [Online]. Available: https://www.finout.io/blog/openai-pricing-in-2026. [Accessed: 18-May-2026]
-
OpenAI, “Sharing Feedback, Evaluation and Fine-Tuning Data, and API Inputs and Outputs with OpenAI,” OpenAI Help Center, 2026. [Online]. Available: https://help.openai.com/en/articles/10306912-sharing-feedback-evaluation-and-fine-tuning-data-and-api-inputs-and-outputs-with-openai. [Accessed: 18-May-2026]
-
OpenAI, “Deprecations,” OpenAI API, 2026. [Online]. Available: https://developers.openai.com/api/docs/deprecations. [Accessed: 18-May-2026]
-
OpenAI, “Supervised Fine-Tuning,” OpenAI API, 2026. [Online]. Available: https://developers.openai.com/api/docs/guides/supervised-fine-tuning. [Accessed: 18-May-2026]
-
Startup Fortune, “OpenAI Is Winding Down Fine-Tuning and That Changes the Startup Playbook,” May 2026. [Online]. Available: https://startupfortune.com/openai-is-winding-down-fine-tuning-and-that-changes-the-startup-playbook/. [Accessed: 18-May-2026]
-
OpenAI, “Introducing the Codex App,” OpenAI, 2026. [Online]. Available: https://openai.com/index/introducing-the-codex-app/. [Accessed: 28-May-2026]
-
OpenAI, “OpenAI Named a Leader in Enterprise Coding Agents by Gartner,” OpenAI, May 2026. [Online]. Available: https://openai.com/index/gartner-2026-agentic-coding-leader/. [Accessed: 28-May-2026]
-
Microsoft Corporation, “Microsoft 365 Copilot Plans and Pricing — Enterprise,” Microsoft, May 2026. [Online]. Available: https://www.microsoft.com/en-us/microsoft-365-copilot/pricing/enterprise. [Accessed: 18-May-2026]
-
Microsoft Corporation, “Microsoft 365 Plans and Pricing — Enterprise,” Microsoft, May 2026. [Online]. Available: https://www.microsoft.com/en-us/microsoft-365/enterprise/microsoft-365-plans-and-pricing. [Accessed: 18-May-2026]
-
E. O'Connor, “Microsoft 365 Copilot Pricing & Licensing: Enterprise Guide 2026,” EPC Group, Apr. 10, 2026. [Online]. Available: https://www.epcgroup.net/microsoft-365-copilot-pricing-licensing-enterprise-guide-2026. Last updated Sep. 27, 2026. [Accessed: 18-May-2026]
-
Microsoft Corporation, “Microsoft Copilot Studio Documentation,” Microsoft Learn, 2026. [Online]. Available: https://learn.microsoft.com/en-us/microsoft-copilot-studio. [Accessed: 18-May-2026]
-
GitHub, Inc., “About Billing for GitHub Copilot in Organizations and Enterprises,” GitHub Docs, May 2026. [Online]. Available: https://docs.github.com/en/copilot/concepts/billing/organizations-and-enterprises. [Accessed: 18-May-2026]
-
M. Rodriguez, “GitHub Copilot Is Moving to Usage-Based Billing,” The GitHub Blog, May 2026. [Online]. Available: https://github.blog/news-insights/company-news/github-copilot-is-moving-to-usage-based-billing/. [Accessed: 18-May-2026]
-
GitHub, Inc., “What Changed with Copilot Billing,” GitHub Docs, June 2026. [Online]. Available: https://docs.github.com/en/copilot/reference/copilot-billing/request-based-billing-legacy/what-changed-with-billing. [Accessed: 18-Jun-2026]
-
L. Ropek, “'What a Joke': GitHub Copilot's New Token-Based Billing Spurs Consternation Among Devs,” TechCrunch, May 30, 2026. [Online]. Available: https://techcrunch.com/2026/05/30/what-a-joke-github-copilots-new-token-based-billing-spurs-consternation-among-devs/. [Accessed: 18-Jun-2026]
-
Microsoft Corporation, “Customize a Model with Fine-Tuning — Microsoft Foundry,” Microsoft Learn, Feb. 27, 2026. [Online]. Available: https://learn.microsoft.com/en-us/azure/foundry/openai/how-to/fine-tuning. [Accessed: 18-May-2026]
-
Microsoft, “Azure OpenAI Service - Pricing,” Microsoft Azure. [Online]. Available: https://azure.microsoft.com/en-us/pricing/details/azure-openai/. [Accessed: 28-May-2026]
-
IntuitionLabs, “Microsoft Copilot Pricing & Licensing Guide for Business,” Jan. 2026. [Online]. Available: https://intuitionlabs.ai/articles/microsoft-copilot-pricing-licensing. [Accessed: 18-May-2026]
-
Google, “Compare Flexible Pricing Plan Options — Google Workspace,” Google LLC, 2026. [Online]. Available: https://workspace.google.com/pricing. [Accessed: 28-May-2026]
-
Google, “Gemini Enterprise App: Best of Google AI for Business,” Google Cloud, 2026. [Online]. Available: https://cloud.google.com/gemini-enterprise. [Accessed: 28-May-2026]
-
Google, “Gemini Enterprise App FAQs,” Google Cloud, 2026. [Online]. Available: https://cloud.google.com/gemini-enterprise/faq. [Accessed: 28-May-2026]
-
Google, “Gemini Enterprise Agent Platform (formerly Vertex AI),” Google Cloud, 2026. [Online]. Available: https://cloud.google.com/products/gemini-enterprise-agent-platform. [Accessed: 28-May-2026]
-
Google, “Gemini Developer API Pricing,” Google AI for Developers, May 28, 2026. [Online]. Available: https://ai.google.dev/gemini-api/docs/pricing. [Accessed: 28-May-2026]
-
Google, “About Supervised Fine-Tuning for Gemini Models,” Google Cloud Documentation, Apr. 29, 2026. [Online]. Available: https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/gemini-supervised-tuning. [Accessed: 28-May-2026]
-
Amazon Web Services, “Security, Privacy, and Responsible AI — Amazon Bedrock,” AWS, 2026. [Online]. Available: https://aws.amazon.com/bedrock/security-privacy-responsible-ai/. [Accessed: 18-May-2026]
-
Amazon Web Services, “Amazon Bedrock Security and Privacy,” Amazon Web Services, May 13, 2026. [Online]. Available: https://aws.amazon.com/bedrock/security-compliance/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon Bedrock Pricing,” AWS, 2026. [Online]. Available: https://aws.amazon.com/bedrock/pricing/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Create a knowledge base by connecting to a data source in Amazon Bedrock Knowledge Bases,” AWS Documentation, 2025. [Online]. Available: https://docs.aws.amazon.com/bedrock/latest/userguide/knowledge-base-create.html. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon OpenSearch Service Pricing,” AWS, 2026. [Online]. Available: https://aws.amazon.com/opensearch-service/pricing/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon OpenSearch Serverless slashes entry cost in half for all collection types,” AWS What's New, June 5, 2024. [Online]. Available: https://aws.amazon.com/about-aws/whats-new/2024/06/amazon-opensearch-serverless-entry-cost-half-collection-types/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon S3 Vectors Is Now Generally Available with 40 Times the Scale of Preview,” AWS What's New, Dec. 2025. [Online]. Available: https://aws.amazon.com/about-aws/whats-new/2025/12/amazon-s3-vectors-generally-available/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Top Announcements of the AWS Summit in New York, 2026,” AWS News Blog, June 18, 2026. [Online]. Available: https://aws.amazon.com/blogs/aws/top-announcements-of-the-aws-summit-in-new-york-2026/. [Accessed: 18-Jun-2026]
-
Amazon Web Services, “Customize a Model with Fine-Tuning or Continued Pre-Training in Amazon Bedrock,” Amazon Bedrock User Guide, 2026. [Online]. Available: https://docs.aws.amazon.com/bedrock/latest/userguide/custom-model-fine-tuning.html. [Accessed: 28-May-2026]
-
Amazon Web Services, “Customizing Amazon Nova Models,” Amazon Nova User Guide, 2026. [Online]. Available: https://docs.aws.amazon.com/nova/latest/userguide/customization.html. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon Bedrock Model Distillation,” AWS, 2026. [Online]. Available: https://aws.amazon.com/bedrock/model-distillation/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon Bedrock AgentCore Is Now Generally Available,” AWS What's New, Oct. 13, 2025. [Online]. Available: https://aws.amazon.com/about-aws/whats-new/2025/10/amazon-bedrock-agentcore-available/. [Accessed: 28-May-2026]
-
Amazon Web Services, “Amazon Q Pricing,” AWS, 2026. [Online]. Available: https://aws.amazon.com/q/pricing/. [Accessed: 28-May-2026]
-
Cohere, “AI Security and Data Protection,” 2026. [Online]. Available: https://cohere.com/security. [Accessed: 21-Jun-2026]
-
Cohere, “An Overview of Cohere's Models,” Cohere Documentation, 2026. [Online]. Available: https://docs.cohere.com/docs/models. [Accessed: 21-Jun-2026]
-
Cohere, “Pricing: Secure and Scalable Enterprise AI,” 2026. [Online]. Available: https://cohere.com/pricing. [Accessed: 21-Jun-2026]
-
Cohere, “Introducing Command A+: Making sovereign agentic capabilities available to all,” May 20, 2026. [Online]. Available: https://cohere.com/blog/command-a-plus. [Accessed: 21-Jun-2026]
-
Cohere, “Announcing Embed Multimodal v4,” Cohere Documentation, Apr. 15, 2025. [Online]. Available: https://docs.cohere.com/changelog/embed-multimodal-v4. [Accessed: 21-Jun-2026]
-
Cohere, “Introduction to Embeddings at Cohere,” Cohere Documentation, 2026. [Online]. Available: https://docs.cohere.com/v2/docs/embeddings. [Accessed: 21-Jun-2026]
-
N. Patience, “Cohere's Multilingual and Sovereign AI Moat Ahead of a 2026 IPO,” Futurum Group, Feb. 20, 2026. [Online]. Available: https://futurumgroup.com/insights/coheres-multilingual-sovereign-ai-moat-ahead-of-a-2026-ipo/. [Accessed: 21-Jun-2026]
-
Constellation Research, “Cohere North Generally Available,” Aug. 6, 2025. [Online]. Available: https://www.constellationr.com/blog-news/insights/cohere-north-generally-available. [Accessed: 21-Jun-2026]
-
The Globe and Mail, “Canadian AI Firm Cohere, Germany's Aleph Alpha Announce Merger,” Apr. 24, 2026. [Online]. Available: https://www.theglobeandmail.com/business/article-canadian-ai-firm-cohere-germanys-aleph-alpha-announce-merger/. [Accessed: 21-Jun-2026]
-
CNBC, “Cohere to Acquire German AI Company Aleph Alpha as It Looks to Expand in Europe,” Apr. 24, 2026. [Online]. Available: https://www.cnbc.com/2026/04/24/cohere-aleph-alpha-germany-ai-europe-expansion.html. [Accessed: 21-Jun-2026]
-
TechCrunch, “Why Cohere Is Merging with Aleph Alpha,” Apr. 25, 2026. [Online]. Available: https://techcrunch.com/2026/04/25/why-cohere-is-merging-with-aleph-alpha/. [Accessed: 21-Jun-2026]
-
Cohere, “Cohere Acquires Reliant AI to Expand Sovereign Enterprise AI for the Global Biopharma and Healthcare Sectors,” Cohere Newsroom, press release, July 7, 2026. [Online]. Available: https://cohere.com/blog/cohere-acquires-reliant-ai-expand-sovereign-enterprise-ai. [Accessed: 15-Jul-2026]
-
CloudZero, “The State of AI Costs in 2025,” Aug. 4, 2025. [Online]. Available: https://www.cloudzero.com/state-of-ai-costs/. [Accessed: 18-May-2026]
-
OpenAI, “ChatGPT Enterprise & Edu — Release Notes,” OpenAI Help Center, July 2026. [Online]. Available: https://help.openai.com/en/articles/10128477-chatgpt-enterprise-edu-release-notes. [Accessed: 15-Jul-2026]
-
K. Waehner, “Enterprise Agentic AI Landscape 2026: Trust, Flexibility, and Vendor Lock-In,” Apr. 6, 2026. [Online]. Available: https://www.kai-waehner.de/blog/2026/04/06/enterprise-agentic-ai-landscape-2026-trust-flexibility-and-vendor-lock-in/. [Accessed: 18-May-2026]
-
Data Center Knowledge, “2025 Cloud Highlights: AI, Outages, and Future Infrastructure,” Data Center Knowledge, Jan. 6, 2026. [Online]. Available: https://www.datacenterknowledge.com/cloud/2025-cloud-highlights-ai-outages-and-the-future-of-infrastructure. [Accessed: 18-May-2026]
-
Forrester Research, “Predictions 2026: Cloud Outages, Private AI on Private Clouds, and the Rise of the Neoclouds,” Forrester, Nov. 20, 2025. [Online]. Available: https://www.forrester.com/blogs/predictions-2026-cloud-outages-private-ai-on-private-clouds-and-the-rise-of-the-neoclouds/. [Accessed: 18-May-2026]
-
White & Case LLP, “AI Watch: Global regulatory tracker – United States,” Sept. 24, 2025. [Online]. Available: https://www.whitecase.com/insight-our-thinking/ai-watch-global-regulatory-tracker-united-states. [Accessed: 28-May-2026]
-
K. M. Bombach, K. Dwyer, J. E. Roberson, S. Dohale, and M. R. Carnes, “AI Company Anthropic Suspends Access to Claude Fable 5, Claude Mythos 5 Following US Export Control Directive,” The National Law Review, Greenberg Traurig LLP, June 13, 2026. [Online]. Available: https://natlawreview.com/article/ai-company-anthropic-suspends-access-claude-fable-5-claude-mythos-5-following-us. [Accessed: 18-Jun-2026]
-
Executive Office of the President, “Promoting Advanced Artificial Intelligence Innovation and Security,” The White House, June 2, 2026. [Online]. Available: https://www.whitehouse.gov/presidential-actions/2026/06/promoting-advanced-artificial-intelligence-innovation-and-security/. Executive Order. [Accessed: 15-Jul-2026]
-
Deloitte AI Institute, “The State of AI in the Enterprise,” Deloitte, 2026. [Online]. Available: https://www.deloitte.com/us/en/what-we-do/capabilities/applied-artificial-intelligence/content/state-of-ai-in-the-enterprise.html. [Accessed: 18-May-2026]