Section8

Open-Weight Models Ready for Production

Open-weight models in 2026 are production-ready for on-premises deployment on the hardware this paper recommends for small to medium organizations, which this paper defines as fewer than 500 employees: small below 100, medium from 100 to 499. That claim was speculative in 2023 and contested through most of 2024; it is now demonstrable. Google’s Gemma 4 31B holds the #3 position on the Arena AI open-model leaderboard [1]. Mistral Large 3, a 675-billion-parameter mixture-of-experts model released under Apache 2.0, debuted at #2 in the open-source non-reasoning category on LMArena [2]. OpenAI’s gpt-oss-120b reaches near-parity with o4-mini on core reasoning benchmarks while running on a single 80 GB GPU [3], and independent measurement places several open-weight models within measurable range of frontier closed systems [4]. The gap narrowed again in June 2026: Z.ai released GLM-5.2 to open weights on June 16, and at 51 on the Artificial Analysis Intelligence Index (v4.1) it became the leading open-weight model on that benchmark [5]. July compressed the remaining distance further. Thinking Machines Lab released Inkling on July 15 under Apache 2.0, the highest-scoring open-weight model yet published by a US lab [6]. Moonshot AI announced Kimi K3 on July 16 at 2.8 trillion parameters, with weights available on July 27, 2026, and Artificial Analysis measured it at 57 on the Intelligence Index, fourth of 189 models tracked and above Claude Opus 4.8 [7], [8]. The leading closed models as of mid-July 2026 (Claude Fable 5 at 60, GPT-5.6 Sol at 59, Gemini 3.1 Pro) still top the most demanding composite benchmarks [9]. The gap is real. It is no longer wide enough to justify a closed-source-only strategy for the workloads an organization below 500 employees actually runs.

Measured on the Artificial Analysis Intelligence Index, the leading open-weight models now sit a small number of points below the closed frontier. Kimi K3 scores 57 against 60 for Anthropic’s Claude Fable 5 and 59 for OpenAI’s GPT-5.6 Sol, a spread of three points [8], [9], with the next leading open-weight model, GLM-5.2, at 51 [5], [7]. Epoch AI puts a number on the lag using its Epoch Capabilities Index (ECI), a composite over many benchmarks: across January 1 to May 28, 2026, the best available open-weight model trailed the closed state of the art by an average of four months and 8 ECI points, with a 90% confidence interval of 7 to 11 points [10]. Under a stricter test that requires the open model’s point estimate to exceed the closed model outright, that lag grows to six months, and it has widened from the three-month figure Epoch measured over January 2023 through October 2025 [10]. Epoch names two reasons its own estimate probably understates the true gap: open-weight models tend to score worse on private benchmarks than public ones, which is consistent with harder hillclimbing on public suites, and leading closed labs do not always release their strongest models for measurement [10]. The planning figure for a team standardizing on open weights is therefore four to six months of published lag, with real-workload lag likely somewhat longer than that. That lag binds a lab racing the frontier. It does not bind a small to medium organization, where the schedule is set by integration and evaluation work measured in quarters rather than by a model’s position on an index. One methodological caution applies to every Intelligence Index figure in this section: Artificial Analysis revises the index periodically, and scores are comparable only within a single index version. Figures below are v4.1 unless the text notes otherwise.

The Models That Matter

The deployable models divide along two axes: capability tier and provenance. Capability tier sets the hardware a model needs and the workloads it can carry. Provenance sets whether procurement must run a supply-chain security review before the model reaches production.

Google Gemma 4

Gemma 4, released April 2, 2026, ships under Apache 2.0, a break from prior Gemma versions, which carried Google’s custom Gemma Terms of Use [1], [11]. The family spans five tiers. The E2B and E4B variants (roughly 2.3B and 4.5B effective parameters via Per-Layer Embeddings) target edge and mobile hardware. The 12B variant is a unified, encoder-free multimodal model with multi-step reasoning and agentic workflow performance nearing the 26B model running locally with just 16GB of VRAM or unified memory. [12]. The 26B A4B variant is a mixture-of-experts (MoE) design (26B total, about 4B active per token) that runs on a single-GPU server inference and fits within a 24 GB card at INT4 quantization. The 31B dense variant is the maximum-quality server model, fitting within 24–32 GB at INT4 and roughly 62 GB at FP16. Modality and context length split by tier rather than running uniformly across the family. Every variant accepts image input; E2B, E4B, and the 12B Unified model also accept audio, and the 31B dense model does not. Context length runs to 128K tokens on E2B and E4B and to 256K on the 12B, 26B A4B, and 31B, carried by a hybrid attention scheme that interleaves local sliding-window layers with full global attention and always closes on a global layer [11].

Gemma 4 31B scores about 85% on MMLU-Pro and 89% on AIME 2026, and posts a Codeforces ELO of 2150, up from Gemma 3’s 110 [11]. For a Phase 1 deployment optimizing hardware efficiency at production quality on knowledge work and coding, Gemma 4 31B at FP8 on a single RTX PRO 6000 Blackwell is the closest match between capability and silicon currently available.

Mistral

Mistral AI, a French AI company founded in 2023 and headquartered in Paris, offers a catalog of mostly Apache 2.0 licensed AI models. That licensing posture is the operational headline. Mistral Large 3 (December 2025) is a 675B-total, 41B-active MoE under Apache 2.0, trained from scratch on 3,000 NVIDIA H200 GPUs, with a 256K context window, native image input via a 2.5B vision encoder, 73.11% on MMLU-Pro, and 93.60% on MATH-500 [2]. Mistral Small 4 (March 2026, Apache 2.0) consolidates instruct, reasoning, multimodal, and agentic coding into a single MoE model with 119B total and 6.5B active parameters and a configurable reasoning-effort parameter [13]. The Ministral 3 family (3B, 8B, 14B dense; Apache 2.0) covers edge and local deployment [2].

Two exceptions sit outside Apache 2.0. Mistral Medium 3.5 (April 2026) is a 128B dense multimodal flagship that merges instruction-following, reasoning, and agentic coding into one set of weights, with a 256K context window and per-request reasoning-effort control; it scores 77.6% on SWE-Bench Verified and ships under a modified MIT license that gates high-revenue commercial use through a paid channel [14]. Voxtral TTS, Mistral’s open-weights text-to-speech model, ships under CC BY-NC 4.0 and needs a separate agreement for commercial use. For an organization with EU exposure that wants to avoid the Llama 4 multimodal restriction discussed below, the Apache-licensed core of Mistral is the strongest substitute at the same capability tier, with no monthly-active-user caps, no geographic carve-outs, and no acceptable-use restrictions beyond standard Apache terms.

NVIDIA Nemotron 3

Nemotron 3 is a hybrid Mamba-Transformer MoE line in which Mamba-2 layers provide linear-time sequence processing, which is what makes a 1-million-token context window practical rather than theoretical [15], [16]. Three tiers shipped across the first half of 2026.

Nemotron 3 Nano (31.6B total, 3.2B active) is the cost-efficient inference tier; NVIDIA reports 3.3× the throughput of Qwen3-30B-A3B (a popular Chinese open-weight model variant) and 2.2× that of gpt-oss-20b on a single H200 GPU at 8K input / 16K output [16]. Nemotron 3 Super (120B total, 12B active), released March 11, 2026, targets multi-agent systems and high-throughput production. Independent evaluation by Artificial Analysis placed Super at 36 on its Intelligence Index at release, ahead of gpt-oss-120b at 33 but behind Qwen3.5 122B at 42; NVIDIA reports roughly 2.2× the inference throughput of gpt-oss-120b at that intelligence tier [17], [18]. Nemotron 3 Ultra (550B total, 55B active) was open-sourced June 4, 2026 with full post-training (supervised fine-tuning, reinforcement learning, and multi-teacher distillation); Artificial Analysis scored it 48 at release, the highest-rated US-developed open-weight model on that index at the time [19]. Both positions have since moved. Under the index revision current in July 2026 Artificial Analysis lists Ultra at 38 and gpt-oss-120b at 24 with Thinking Machines Inkling passing Ultra on July 15 [6].

The license is the NVIDIA Nemotron Open Model License: not Apache 2.0 and not approved by the Open Source Initiative (OSI), but unusually transparent for a custom license, since NVIDIA releases open weights, training datasets, and recipes alongside each model [20]. Nemotron’s advantage is tight coupling to NVIDIA hardware and the NVIDIA Inference Microservices (NIM) path with built-in TensorRT-LLM optimization. Where NIM is already in the production stack, Nemotron is the natural model choice.

OpenAI gpt-oss

OpenAI’s August 2025 release of gpt-oss-120b and gpt-oss-20b is the company’s first open-weight release since GPT-2 in 2019, and its first under a fully permissive license [3]. Both models ship under Apache 2.0. The 120B variant carries 116.8B total parameters and activates 5.1B per token; its MXFP4 quantization fits the entire model in roughly 61 GB, so it runs on a single 80 GB GPU (an H100 or AMD MI300X) and comfortably on the 96 GB RTX PRO 6000. It supports a 128K context window and exposes configurable reasoning effort with full chain-of-thought output. Near-parity with o4-mini on core reasoning is the independently corroborated result [3], [4]; OpenAI’s own evaluation also reports wins over o4-mini on HealthBench and AIME and over o3-mini on Codeforces and tool-calling, which should be read as vendor-reported until reproduced.

For an organization with provenance concerns about Chinese-origin models (addressed below) and EU constraints on Llama 4 multimodal, gpt-oss-120b is a US-origin, Apache-2.0 alternative at the 100B-plus class, and its single-GPU footprint is a direct Phase 1 sizing advantage.

Meta Llama 4

Meta released Llama 4 on April 5, 2025 as a natively multimodal MoE family [21]. Scout (17B active, 16 experts, 109B total) fits a single H100 at INT4 and ships with a 10-million-token context window. Maverick (17B active, 128 experts, 400B total) targets a single H100 DGX host. Behemoth (2T total, 288B active) remains unreleased; Meta used it as a teacher model for the released variants through codistillation and has given no release window. On NVIDIA Blackwell B200, NVIDIA’s optimized stack reports Scout above 40,000 output tokens per second, and above 12,000 on H200 [21], [22].

The license is the consequential detail. Llama 4 ships under the Llama 4 Community License Agreement, a custom Meta license that is neither Apache 2.0 nor OSI-approved, which permits royalty-free commercial use below 700 million monthly active users, requires “Built with Llama” attribution on distributed derivatives, and permits training on Llama 4 outputs with attribution [23]. The material restriction is geographic: the multimodal capabilities cannot be licensed to individuals domiciled in, or companies with principal place of business in, the European Union; text-only paths remain unrestricted in the EU [24]. Any organization with EU subsidiaries or EU-based staff using the multimodal capabilities needs legal counsel to clear this restriction before standardizing on Llama 4 multimodal variants.

Cohere

Cohere, a Toronto-based enterprise AI company founded in 2019 and merged with Germany’s Aleph Alpha in 2026, changed its open-model posture in mid-2026 by releasing two models under Apache 2.0. That is a departure from the CC-BY-NC research licensing that still governs its earlier Command research weights and its Aya and Tiny Aya multilingual families, all of which remain non-commercial and therefore out of scope for production. Both Apache-licensed releases target the sovereign, on-premises deployment this paper argues for, and both carry clean, allied-Western provenance that simplify the supply-chain review.

Command A+ (command-a-plus-05-2026), released May 20, 2026, is a 218B-total, 25B-active sparse MoE with vision input, native citation grounding, agentic tool use, and a 128K context window, published on Hugging Face in BF16, FP8, and W4A4 quantizations across 48 languages. It is the first frontier-scale model a major enterprise vendor has released under a fully permissive license. At W4A4 it runs on a single B200 or two H100s/200s, which places it above the single-RTX PRO 6000 Phase 1 target but inside the enterprise tier; Cohere applies NVFP4 to the MoE experts while holding the attention path at full precision and reports the result as near lossless against BF16 [25], [26]. For an organization that wants a general-purpose, multimodal, agentic open model without the licensing or provenance overhead of the Chinese flagships, Command A+ is a strong clean-provenance option at its scale.

North Mini Code (June 9, 2026), the first model in Cohere’s North code-agent family, is a 30B-total, 3B-active MoE built for agentic software engineering, with a 256K context window and a 64K maximum generation length, also under Apache 2.0 [27]. Its small active footprint runs comfortably on a single RTX PRO 6000 or H100 at FP8, and independent evaluation by Artificial Analysis placed it at 33.4 on its Coding Index, ahead of several larger open models though slightly behind Qwen3.6-35B-A3B [28]. The documented tradeoff is verbosity: in independent testing it generated roughly three times the output tokens of comparable models, a cost that compounds in high-volume serving [29]. For a Phase 1 coding-assistant workload that values permissive licensing and single-GPU deployment, it is a Western-origin counterpart to Qwen3.6-35B-A3B and DeepSeek V4-Flash (non-reasoning).

Thinking Machines Inkling

Inkling, released July 15, 2026, is the first production model from Thinking Machines Lab and the highest-scoring open-weight model published by a US lab [6], [30]. It is a 975B-total, 41B-active MoE under Apache 2.0, pretrained on 45 trillion tokens, with a 1-million-token context window on the downloadable checkpoint, native text, image, and audio input, text-only output, and a thinking-effort control the caller sets per request across a 0.2 to 0.99 range [30]. Weights shipped to Hugging Face on release day rather than on a promised future date, and vLLM, SGLang, llama.cpp, and Transformers support arrived with the launch.

Artificial Analysis scored Inkling at 41 on the Intelligence Index, ahead of Nemotron 3 Ultra at 38, Gemma 4 31B at 29, and gpt-oss-120b at 24, and behind the Chinese open-weight leaders [6]. Two results matter more than the headline rank. Inkling is markedly token-efficient, averaging 25K output tokens per index task against 43K for GLM-5.2, 38K for Kimi K2.6, and 37K for DeepSeek V4-Pro, which changes serving economics on a fixed GPU budget [6]. Against that, it scores +2 on AA-Omniscience with 40% accuracy and a 63% hallucination rate, so it guesses rather than abstains more often than the open-weight leaders, and it should not carry accuracy-critical retrieval work without external grounding and citation enforcement [6].

The deployment footprint places Inkling above Phase 1. The BF16 checkpoint requires more than 2 TB of aggregate VRAM; the NVFP4 checkpoint reduces that to roughly 600 GB, which is an 8-GPU H200 or B200 node [31]. Thinking Machines has previewed Inkling-Small at 276B total and 12B active on a similar training recipe, with weights pending, and that variant is the one to watch for mid-tier hardware. Inkling’s strategic value is the combination of Apache 2.0 terms, US provenance, and an explicit design intent as a customization base rather than a leaderboard entry. For an organization that needs a permissively licensed, clean-provenance foundation for domain fine-tuning at frontier-adjacent scale, it is the strongest candidate currently available.

The Chinese-Origin Models

Five Chinese open-weight families now compete at the frontier of capability: Qwen (Alibaba Cloud), DeepSeek (High-Flyer), GLM (Z.ai), Kimi (Moonshot AI), and MiniMax (Shanghai). Most ship under permissive Apache or MIT licenses and post benchmarks within striking distance of the leading closed models on coding and agentic tasks; MiniMax is the exception, releasing its M3 flagship under a more restrictive custom license. The procurement question is whether that capability justifies the supply-chain review the origin demands.

Qwen’s open ceiling is Qwen3.5-397B-A17B (February 2026, Apache 2.0), a 397B sparse MoE that pairs Gated DeltaNet linear attention with standard gated attention and supports 262K native context [32]. The smaller Qwen3.6-35B-A3B (April 2026, Apache 2.0) carries 35B total and 3B active, fits a single 80 GB GPU, and is the better default for most deployments; Artificial Analysis rates it among the leading models for its size while noting it’s more expensive per token and verbose [33], [34]. The Qwen3.7-Max flagship, announced May 2026, is API-only, with no released weights [35]. DeepSeek shipped V4 on April 24, 2026 in two MIT-licensed tiers: V4-Flash at 284B total with 13B active and a 1-million-token context window, and V4-Pro at 1.6T total with 49B active; V4 is approximately tied with the strongest closed models on SWE-Bench Verified near 80% and leads on LiveCodeBench [36]. The prior generation remains widely deployed, including V3.1-Terminus and the January 2025 R1 reasoning model that first matched OpenAI o1 among open weights [37]. GLM-5 (Z.ai, February 2026) is a 744B MoE with 40B active, a 200K context window, MIT licensing, and end-to-end training on Huawei Ascend neural processing units (NPUs), the first frontier-scale model trained with no NVIDIA hardware in the pipeline [38]; Z.ai has since shipped GLM-5.1 (754B, MIT) and, on June 16, 2026, the MIT-licensed GLM-5.2, which keeps the GLM-5.1 footprint (roughly 753B total, 40B active), extends the context window to 1 million tokens, and reached 51 on the Artificial Analysis Intelligence Index to become the leading open-weight model on that benchmark; GLM-5.2 is text-only, with Z.ai’s vision capability shipping separately as the non-open GLM-5V line [39], [5]. Kimi K2.6 (Moonshot AI, April 2026) is a 1-trillion-parameter MoE activating 32B per token, natively multimodal, under a modified MIT license that adds one attribution requirement above 100M monthly active users or $20M monthly revenue and behaves as standard MIT below those thresholds [40]; it edges out the strongest closed coding models on SWE-Bench Pro (about 58.6%) at a fraction of their per-token cost. Moonshot refreshed the line on June 12, 2026 with Kimi K2.7-Code, a coding-specialized build on the same trillion-parameter, 256K-context architecture that cuts reasoning-token usage by roughly 30% versus K2.6 under the same modified MIT terms; its launch benchmarks are vendor-reported, with independent results still pending [41]. MiniMax (Shanghai) joined the frontier tier on June 1, 2026 with M3, a natively multimodal MoE carrying about 428B total and 23B active parameters, with open weights reaching Hugging Face on June 7 and the architecture’s technical report on June 11 [42], [43]. M3’s MiniMax Sparse Attention (MSA), a block-sparse mechanism layered on grouped-query attention, is what makes its 1-million-token context window economical, cutting per-token compute at full context to about one-twentieth of the prior M2 generation with roughly 9× faster prefill and 15× faster decode [42], [43]. M3 posts a vendor-reported 59.0% on SWE-Bench Pro and an independently measured 44 on the Artificial Analysis Intelligence Index, tied with DeepSeek V4-Pro and behind GLM-5.2’s 51 [5]. Its licensing is the procurement caveat: unlike the Apache- and MIT-licensed Chinese frontier models, M3 ships under MiniMax’s custom “minimax-community” license, which conditions commercial use and drew criticism at launch, so it warrants the same license review as Llama 4 or Voxtral before any standardization [42].

Kimi K3 reset the open ceiling in July 2026. Moonshot announced it on July 16 with the API live at launch and the full weights available in July 27 alongside a technical report. Its architecture has 2.8 trillion total parameters in a highly sparse MoE routing 16 of 896 experts per token, Kimi Delta Attention carrying a 1-million-token context window, text and image input, and reasoning effort fixed at maximum with lower settings promised later [7], [44]. Artificial Analysis scored it 57 on the Intelligence Index, fourth of 189 models tracked, level with Claude Opus 4.8 at 56 and above every open-weight peer, with GLM-5.2 at 51 and DeepSeek V4-Pro at 44 [8], [9]. Cost per completed task, rather than price per token, is the number that governs a total-cost comparison: at $3.00 per million input tokens and $15.00 per million output tokens with cached input discounted 90% to $0.30, Artificial Analysis measured $0.94 per Intelligence Index task against $1.04 for GPT-5.6 Sol and $1.80 for Opus 4.8, so K3 undercuts the closed frontier at comparable intelligence while costing roughly three times GLM-5.2 and more than twenty times DeepSeek V4-Pro on the same basis, and it serves at 62 output tokens per second against a 72 t/s median for its price tier [8], [45], [46]. At 2.8 trillion parameters, Kimi K3 is quite a large model. Quantized to 4 bits, it requires roughly 1.4 TB of high-bandwidth memory before factoring in key-value cache and runtime overhead, which puts self-hosting into multi-node territory and outside any single-node build this paper recommends.

The capability is real. The provenance review is mandatory. By February 2025, the U.S. House of Representatives, the Pentagon, NASA, the U.S. Navy, and multiple state governments had restricted DeepSeek’s cloud product on government devices [47]. Cisco’s Robust Intelligence team, with University of Pennsylvania researchers and the standard HarmBench jailbreak suite, recorded a 100% attack success rate against DeepSeek R1, since every one of 50 test prompts elicited a harmful response, against 26% for OpenAI’s o1-preview on the same battery [48]. Wiz Research separately found a publicly accessible DeepSeek database exposing chat history, API tokens, and system logs, and DeepSeek’s terms state that user data is stored on Chinese servers under Chinese law [49]. Representatives Gottheimer and LaHood introduced the No DeepSeek on Government Devices Act (HR 1121) on February 7, 2025 [50], and an April 2025 House Select Committee report documented ties between DeepSeek and a state-backed research institute with military-research adjacency [51].

These concerns attach to the cloud product, not automatically to the self-hosted weights. A DeepSeek V4, GLM-5.2, or MiniMax M3 deployment running on-premises, behind the firewall, on hardware the organization controls, transmits nothing to the originating vendor. MIT-licensed weights do not become a backdoor by virtue of origin; as of mid-2026 no peer-reviewed publication documents a cryptographic backdoor in any of these weights. The risk that travels with the weights is behavioral: observable content censorship aligned with Chinese government policy, training-data-encoded bias, and the institutional affiliations of the originating lab. The Cisco jailbreak finding applies whether the model runs on the vendor’s servers or yours, and it matters for any deployment where untrusted users can craft prompts.

A supply-chain review for a Chinese-origin model documents model origin and corporate structure; any government or military research affiliations; published behavior research on the specific model; an observable behavior audit covering content restrictions and geopolitical framing; and legal counsel’s read on applicable procurement regulations. Until that review is signed off, initial production use belongs in non-sensitive, air-gap-isolated contexts. For organizations whose posture does not require this scrutiny, Qwen3.6-35B-A3B, DeepSeek V4, GLM-5.2, Kimi K2.7-Code, and MiniMax M3 cover different points on the size-versus-capability curve. For organizations with U.S. government contracts, defense adjacency, or sensitive data classifications, the U.S.-origin pair of Gemma 4 31B and gpt-oss-120b covers the same capability range without the procurement overhead, with Inkling available above them for organizations that can host it.

Allen Institute OLMo

Allen Institute for AI (Ai2), a Seattle based non-profit AI research institute founded in 2014 by the late Paul Allen, released Olmo 2 and Olmo 3 under Apache 2.0. These models are the most fully open frontier line from a US institution, releasing weights, the Dolma training dataset, training code, the evaluation framework, and intermediate checkpoints. Olmo 2 32B (March 2025) was the first fully open model to outperform GPT-3.5-Turbo and GPT-4o-mini on a multi-skill suite, trained to 6T tokens at one-third the compute of Qwen 2.5 32B [52], [53]; Olmo 3 (November 2025) adds OlmoTrace, which links model outputs to specific training-data decisions and supports behavior provenance at a level no other model in this roster documents [54]. For regulated industries and contractors that must trace behavior to training data for compliance, Olmo is the right choice. Its value is auditability, not raw capability; on general reasoning, Gemma 4, Mistral Large 3, Nemotron 3 Super, gpt-oss, and Llama 4 outperform it at comparable scale.

The Small-Model Tier

Most of the tokens a production deployment generates will not come from its largest model. Agentic systems decompose work into bounded steps. Belcak et al. argue that most of those steps are format-constrained rather than reasoning-constrained [55]. For an organization below 500 employees, where Phase 1 is one-to-few GPU accelerators, that distribution decides what the hardware covers: every call a 4B specialist model absorbs is concurrency the 31B model keeps. Organizations should design the sub-10B model tier into Phase 1 rather than treat small models as a concession to hardware constraints.

Defining the Class

Belcak et al. at NVIDIA Research define a small language model (SLM) by deployability rather than by parameter count: a model that fits on a common consumer device and processes one user’s requests with latency low enough to be practical [55]. They anchor that to a rule of thumb of roughly 10 billion parameters as of 2025 and note the threshold moves with hardware [55]. That definition is useful for procurement because it tracks what a model costs to serve rather than what its parameter count signals in marketing. Vendor naming does not follow it. Nemotron 3 Nano carries 31.6 billion total parameters and activates 3.2 billion [16], so it sits well outside the class by total footprint even though its active-parameter count looks small.

The Belcak paper is a position piece and should be read as one. Its authors argue that SLMs are already powerful enough for the language work agents actually perform, that they suit agentic systems better operationally, and that they cost less by virtue of size [55]. The supporting evidence is assembled from prior published results rather than generated by a controlled study. The appendix estimates that 40% to 70% of language-model calls in the MetaGPT, Open Operator, and Cradle agents could move to specialized SLMs; those figures are the authors’ own assessments rather than measurements [55]. The authors also grant that the counter-argument from centralized inference economies of scale is valid and unsettled [55]. What the paper contributes is a well-specified hypothesis with named mechanisms. The measurements below are what test it.

Capability Compression

Google’s Gemma 4 model card supplies the cleanest test available, because it holds the vendor, the evaluation harness, and the benchmark suite constant across a single generation step. Gemma 4 E4B carries 4.5 billion effective parameters. Gemma 3 27B carries roughly six times that. On MMLU-Pro, Gemma 4 E4B scores 69.4% against Gemma 3 27B’s 67.6%. The margin widens on the benchmarks that predict agentic behavior: 42.2% against 16.2% on Tau2 tool use, and 52.0% against 29.1% on LiveCodeBench v6 [11]. Even E2B at 2.3 billion effective parameters beats the prior-generation 27B on Tau2, 24.5% against 16.2%, while losing on MMLU-Pro at 60.0% against 67.6% [11].

One caveat worth noting with those numbers is that the model card scores Gemma 3 27B in its non-thinking configuration while the Gemma 4 entries use the family’s configurable thinking mode [11], so part of the margin is test-time compute rather than pretrained capability. The direction of the result survives that caveat while its magnitude does not. A team sizing hardware on these numbers should re-measure at its own inference settings.

IBM’s Granite 4.0 line shows the same compression expressed as an architecture decision. The family runs from a 350M dense model through a 1.5B hybrid, a 3B hybrid, a 7B mixture of experts activating 1B parameters, and a 32B mixture of experts activating 9B [56]. The hybrid variants interleave Mamba-2 state-space layers with transformer attention with IBM reporting more than 70% lower memory use and twice the inference speed against comparable models, concentrated in multi-session and long-context serving [56]. Read that as a vendor measurement against an unnamed comparison set. The mechanism is sound, since state-space layers avoid both the quadratic attention cost and the growing key-value cache that dominate memory at long context, but the specific multiple is IBM’s and has no independent replication in the published record. Granite 4.1 then returned to dense architectures at 3B, 8B, and 30B, in base and instruction-tuned variants with optional FP8 checkpoints [57]. That reversal is the more instructive fact: hybrid state-space designs win on memory while the tooling maturity of dense transformers still wins often enough that IBM ships both and tells users to pick the dense variant wherever Mamba-2 support is immature [56].

Specialization and the Fine-Tuning Return

The small tier’s strongest case rests on fine-tuning rather than on out-of-the-box capability. Google’s FunctionGemma is a 270-million-parameter model built on the Gemma 3 270M backbone, trained only for function calling, with a 32K context window and an explicit statement in its model card that it is not intended for use as a dialogue model [58]. Out of the box it is weak where deployment would need it: 36.2% on Berkeley Function Calling Leaderboard (BFCL) Live Simple and 20.8% on Live Parallel Multiple [58]. Fine-tuned on Google’s published Mobile Actions dataset, accuracy on that task moves from 58% to 85% [58]. At dynamic INT8 the artifact occupies 288 MB, peaks at 551 MB resident, and decodes 125.9 tokens per second on the CPU of a Samsung S25 Ultra with a 0.3-second time to first token [58]. Nothing in a Phase 1 rack strains against those numbers. The transferable finding is the 27-point gap between base model and fine-tune, which is the return on a task-specific dataset rather than on parameters.

The economics of collecting that return follow from size. Belcak et al. report that parameter-efficient methods including Low-Rank Adaptation (LoRA) and its quantized variant QLoRA, along with full-parameter fine-tuning at small scale, need only a few GPU-hours, which moves specialization from a quarterly project to an overnight one, and that 10,000 to 100,000 curated examples suffice for a small model [55]. They also report that serving a 7B model costs 10 to 30 times less than a 70B to 175B model in latency, energy, and floating-point operations, a figure the paper carries from secondary industry sources rather than measures [55]. Treat the order of magnitude as sound and the specific multiple as unverified until an organization measures it on its own traffic.

Where the Small Tier Fails

Two important limits worth mentioning that affect production deployments.

The first is long-horizon execution, which is not a small-model problem alone. A published synthesis of agent failure studies reports METR’s time-horizon finding that frontier models succeed on close to 100% of tasks a skilled human completes in under roughly four minutes and on under 10% of tasks taking a human more than roughly four hours, with the task length completable at 50% reliability doubling about every seven months [59]. The HORIZON diagnostic benchmark, which collected more than 3,100 trajectories from current GPT-5 and Claude agents across four domains, attributes the collapse to agents losing the original objective and repeating failed actions rather than to retrieval failure [60]. Both are preprints and neither is SLM-specific, which is precisely why they bind here: if the frontier degrades this way, the orchestrating role belongs to the largest model a deployment can afford, and the small tier takes the bounded steps that orchestrator dispatches.

The second limit is breadth; stated plainly in the model cards. Microsoft trained Phi-4-Reasoning-Vision-15B primarily on English and says it is not intended to support multilingual use [61]. IBM trained and tested Granite Guardian 4.1 only on English, and warns that the model must be used strictly in its prescribed scoring mode, that deviation may produce unsafe output, and that the model is susceptible to adversarial attack [62]. Google’s FunctionGemma safety evaluations covered English prompts only [58]. An organization with multilingual users cannot staff a small tier from these models without measuring the degradation on its own languages first.

What the Small Tier Changes in Phase 1

Nothing in this tier displaces the 31B-to-120B model on the RTX PRO 6000. It changes what else runs on that card and the concurrency arithmetic. Weight footprints at these sizes are arithmetic rather than measurement: a 4B model at FP8 holds roughly 4 GB of weights, a 270M task model under half a gigabyte, and a 2B speech model about 4 GB at BF16, all before key-value cache.

Co-location is not free. Every gigabyte a specialist holds is a gigabyte the primary model’s key-value cache does not get, which sets maximum concurrency. A 96 GB card running a 70B model at FP8 with roughly 26 GB of cache headroom has no room for a specialist tier at all. The same card running Gemma 4 26B A4B at INT4 has tens of gigabytes left over. That trade, fewer concurrent sessions on the largest model against a specialist tier that absorbs the calls which never needed the largest model, is the real Phase 1 decision; settled with measured traffic rather than assumed ratios.

Measured traffic has a deadline. Belcak et al. describe a migration path that begins with logging every non-interface model call, filtering personally identifiable information out of the captured data, clustering the logged prompts into recurring task types, and fine-tuning one specialist per cluster [55]. The logging step is the one that cannot be deferred. An organization that does not capture its agentic traffic from the first day of Phase 1 has no dataset to specialize on in Phase 2; no amount of future hardware buys that data back.

Adjacent Model Classes

A production deployment is a set of models, each with their own specialization, where typically only one of them is the chat model. The classes below carry work a general-purpose language model does badly, expensively, or not at all. Most of them run under the same vLLM serving stack the chat tier already requires.

Embedding models convert text into vectors and constitute the retrieval half of any retrieval-augmented generation system. Google’s EmbeddingGemma carries 300 million parameters, accepts a 2,048-token input, and emits 768-dimension vectors that Matryoshka Representation Learning lets the caller truncate to 512, 256, or 128 dimensions and renormalize, trading index size against accuracy along a published curve: 61.15 mean task score on multilingual MTEB at 768 dimensions, falling to 58.23 at 128 [63], [64]. Its quantization behavior matters for the same reason FP8 matters upstream. The 4-bit quantization-aware-training checkpoint scores 60.62 against 61.15 at full precision on that benchmark, a loss small enough to disappear into retrieval noise [63]. Model choice here is more consequential than teams expect, because a retrieval miss is invisible to the generation model, which will answer fluently from whatever it was handed.

Guardrail and judge models score text against criteria instead of generating it. IBM’s Granite Guardian 4.1 8B ships under Apache 2.0 and returns a yes/no judgement against pre-baked criteria covering harmful content, jailbreak attempts, and profanity; against retrieval criteria covering context relevance, groundedness, and answer relevance; and against a function-calling hallucination criterion that checks whether an agent’s tool call is syntactically and semantically consistent with both the tool definition and the user query [62]. Callers can also supply arbitrary criteria in natural language. The model runs in a thinking mode that emits reasoning traces and a no-thinking mode that returns the score alone. The latter is the one to use when the guardrail sits in the request path and its latency becomes the user’s latency [62]. Where throughput constraints rule out an 8B judge, IBM points to Granite-Guardian-HAP-38M, a 38-million-parameter classifier for hate, abuse, and profanity [62]. A guardrail model reduces risk. It does not discharge it.

Speech models transcribe and translate audio. IBM’s Granite Speech 4.1 family runs at 2 billion parameters across three Apache 2.0 variants: one balanced for recognition and translation, one adding speaker attribution, timestamps, and keyword-prompted recognition for names and technical jargon, and one non-autoregressive variant built for low latency. All three train on 174,000 hours of audio and cover English, French, German, Spanish, Portuguese, and Japanese [65]. IBM reports the 2B model ranking first for accuracy on the public Open ASR Leaderboard and the non-autoregressive variant third [65]; that is a vendor’s account of an independent leaderboard, and a procurement team should confirm the standing at evaluation time. A 2-billion-parameter specialist at the top of an open leaderboard for its task is the small-model argument restated in another domain. Granite Speech runs under vLLM, so it shares the Phase 1 serving stack rather than demanding a second one [65].

Vision-language models read images, documents, and screens. Microsoft’s Phi-4-Reasoning-Vision-15B, released March 4, 2026 under MIT, pairs the Phi-4-Reasoning backbone with a SigLIP-2 encoder in a mid-fusion design, spends up to 3,600 visual tokens on a single image, and decides per request whether to run a chain of thought or answer directly [61]. It reaches 88.2% on ScreenSpot-V2 interface grounding and 83.3% on ChartQA. Microsoft trained it on 240 B200 GPUs in four days, a modest budget at this capability [61]. Two limits sit in the same model card and deserve equal weight. Its context window is 16,384 tokens, an order of magnitude below the language models in the roster above, which constrains multi-page document work. Qwen3-VL-8B-Instruct, a smaller model, outscores it on OCRBench at 89.2% against 76.0% and on MMMU at 60.7% against 54.3% [61]. Parameter count does not settle vision-language selection. The specific task does.

Diffusion language models replace token-by-token autoregression with iterative denoising over a block of tokens. DiffusionGemma, released under Apache 2.0 on the Gemma 4 26B A4B mixture-of-experts architecture, denoises a 256-token canvas in parallel, emits 15 to 20 tokens per forward pass, and exceeds 1,100 tokens per second for a single user on an H100 at FP8 [66]. Google’s own comparison table prices that speed honestly. Against the autoregressive Gemma 4 26B A4B it gives up 5 points on MMLU-Pro (77.6% against 82.6%), 19.2 points on AIME 2026 (69.1% against 88.3%), 12 points on Tau2 (56.2% against 68.2%), and 19.5 points on MMMU Pro (54.3% against 73.8%) [66]. The deployment caveat is sharper than the accuracy cost. Google states the model is engineered for small-batch inference [66], which is the opposite of the continuous-batching, high-concurrency regime a Phase 1 node operates in. Diffusion decoding buys single-stream latency, and a shared multi-user server is not where that purchase pays.

Two further classes ship under the same governance and belong on the inventory even though this section does not evaluate them: IBM’s Granite Docling for document conversion, and Granite Time Series for forecasting, the latter a reminder that not every model in an on-premises deployment is a language model at all [56].

One governance property cuts across these classes and belongs in the procurement checklist regardless of which models an organization selects. IBM cryptographically signs every Granite checkpoint, publishes a verification procedure, and holds ISO/IEC 42001 certification for the Granite 4.0 family, the management-system standard for artificial intelligence [56]. Signature verification answers the one question a supply-chain review must ask that a Hugging Face download cannot answer on its own: whether the file on disk is the file the publisher released. The control exists today and costs nothing to apply.

Hardware Tier Performance Matrix

The matrix below is the evidence base for Phase 1 hardware selection. An accurate headcount of the expected first-wave users is the starting point determining hardware sizing before network traffic becomes available. This paper defines small as fewer than 100 employees and medium between 100 and 500 with the recommendations here applicable through 1,000 employees.

Headcount alone does not produce a concurrency number. Three demand classes load a GPU differently.

Interactive chat is a duty-cycle problem that scales with the number of users. A person composing a prompt, reading the answer, and thinking about it occupies the card for the small fraction of time their session is open. Plan on 30 to 60 percent of the staff having an account (a seat) within the first year of a deployment and 5 to 10 percent of those having a request in flight during peak usage. That puts a 60-person organization at roughly 2 to 4 concurrent interactive requests, a 400-person firm at 6 to 24, and a 900-person firm at 14 to 54. Those percentages are planning assumptions, not measurements, and they are the first two numbers a pilot should replace with observed values.

Power users break the duty-cycle assumption. A software developer using a coding assistant, an analyst iterating over documents, or a support specialist working through a queue. These types of users have a request in flight on nearly every action. Count each one as a full concurrent stream rather than a fraction of a seat. Ten engineers are ten streams. They also run the longest contexts, which is where the cost compounds: KV cache grows with context length multiplied by concurrency (batch size) and is different for every model since its calculation depends on the model’s architecture and optimizations. Refer to the KV Cache subsection within the Generative AI Fundamentals section for calculation examples: at FP8 with 32K context and batch size of one, Qwen3-8B has a KV cache size of 2.25 GB, Google’s Gemma 4 31B at 2.8 GB, and OpenAI’s gpt-oss-120B at 0.5 GB. For ten concurrent requests, these KV cache memory requirements jump to 22.5 GB, 28 GB, and 5 GB on top of the memory required to load the entire model. At 64K context size, these values double.

Agent fan-out decouples from headcount entirely. One orchestrated task that dispatches sub-calls in parallel may create 5 to 20 simultaneous requests from a single human or scheuled pipline action. A 40-person firm running three agentic workflows can exceed the peak concurrency of a 450-person firm doing interactive chat only. Agents also do not read, so the reading-speed floor below does not govern them; budget agent load by tokens per completed task against the deadline the workflow has to meet.

For interactive human use, 20 to 30 output tokens per second per stream keeps generation ahead of average human reading speed. Anything below it turns a working pilot into an assistant nobody uses. Applying it to a medium size organization of 400 employees: 200 seats at a 10 percent peak is 20 interactive streams and adding 15 engineers on coding assistants results in 35 streams at a minimum of 25 tok/s which sets an aggregate design target near 875 tok/s. A small organization under the same assumptions lands between 100 and 350 tok/s. Don’t read the 8,425 tok/s benchmark figure below as the margin against those targets. That number comes from a 3B-active mixture-of-experts model at 8K context under saturating load. A 31B dense model at long context on the same card will generate tokens at a fraction of that speed.

Two methodologies appear, each internally consistent and each backed by a public reproduction repository. The consumer and mid-range rows use a single-GPU workload: Qwen3-Coder-30B-A3B AWQ, 8K context, vLLM with FP8 key-value (KV) cache, 400 concurrent requests [67]. The datacenter rows use an 8-GPU workload at long context: GLM-4.6 at FP8, 8-way tensor parallelism, 16K context [68]. The two groups run different model classes, so their absolute numbers are not directly comparable; the datacenter rows are therefore expressed as multipliers over an 8× RTX PRO 6000 baseline on the same heavy workload.

TierHardwareVRAM70B at FP8Verified throughputPhase 1 verdict
Consumer1× RTX 509032 GBNo4,570 tok/s on the 30B-MoE workload [67]Fits ≤30B MoE at INT4/AWQ; no room for 70B at FP8
Consumer pair2× RTX 5090 (PCIe)64 GBNo~9,000 tok/s, near-2× replica scaling, same workload [67]More replicas of a ≤30B model, not a larger one
Mid-range1× RTX PRO 6000 Blackwell96 GBYes8,425 tok/s, same workload [67]; 1.63× an H100 NVL at FP4 vs FP8 [69]Single-GPU 70B at FP8, the Phase 1 target
Enterprise8× H100 SXM640 GBYes1.7× the 8× PRO 6000 on GLM-4.6-FP8, 8-way tensor parallelism, 16K context [68]Overprovisioned for Phase 1
Enterprise8× H200 SXM1,128 GBYes3.4× the 8× PRO 6000, same workload; holds up best under long context [68]Overprovisioned for Phase 1
Frontier8× B200 SXM1,440 GBYes4.9× the 8× PRO 6000, same workload; also the lowest cost per token [68]Overprovisioned for Phase 1

Three findings drive Phase 1 hardware selection.

The RTX 5090 covers ≤30B-class models at FP4/INT4/AWQ but cannot host a 70B model at FP8. A 70B model at FP8 needs roughly 70 GB of memory for the weights plus about 26 GB of KV-cache headroom at modest context, well above the 5090’s 32 GB. Hosting a 70B at FP4/INT4 quantization is possible but unsafe for production reasoning unless the algorithm and calibration are tightly controlled, as the next subsection shows. Stacking 5090s adds replica instances of the same model (independent copies behind a load balancer), not the ability to host a larger one. Stacking 5090s is appropriate for a small organization whose load is many short interactive sessions. It does not work for a medium organization that needs 70B-class capability on a single endpoint. Tensor parallelism over PCIe Gen 5 works on moderate models but is communication-bound, and PCIe-only configurations cannot match NVLink-class scaling once a model genuinely needs 8-way parallelism [68].

The RTX PRO 6000 Blackwell at 96 GB is the inflection point for single-GPU 70B deployment. On the 30B-MoE workload it reaches 8,425 tok/s on a single card at 28% lower cost per token than the H100 SXM on single-GPU work [67], [70]. Akamai’s controlled comparison measured the PRO 6000 at FP4 delivering 1.63× the throughput of an H100 NVL at FP8, sustained past 100 concurrent requests [69]. The 96 GB ceiling fits a 70B model at FP8 (about 70 GB of weights) with roughly 26 GB left for KV cache, the minimum viable single-GPU configuration for any Phase 1 deployment that needs 70B-class capability at production quality. One card is capable of supporting a small organization outright. A medium organization scales it by adding replicas behind the model-selection layer, which should be included in the deployment from the first day. The card has no NVLink, so multi-card builds communicate over PCIe Gen 5; that is the constraint that draws the line between Phase 1 and Phase 2.

The 8× datacenter tier reclaims throughput leadership at the cost of capital expense, and the H200/B200 choice turns on workload, not prestige. On the heaviest workload (GLM-4.6-FP8 at 16K context with 8-way tensor parallelism), an 8× H100 node delivers 1.7× the 8× PRO 6000, an 8× H200 node 3.4×, and an 8× B200 node 4.9×, with the B200 also posting the lowest cost per token [68]. The H200 (141 GB HBM3e, 4.8 TB/s) is the long-context inference workhorse: as context grows from 2K to 16K, its throughput falls 47% against the H100’s 64%, a difference that follows from its larger, higher-bandwidth memory under KV-cache-heavy loads, and its Multi-Instance GPU (MIG) partitioning suits multi-tenant serving [68]. The B200 (180 GB HBM3e per GPU in a DGX cluster, near 8 TB/s, with FP4 tensor cores) is the throughput and training leader: its 180 GB fits 70–180B models on a single GPU without tensor parallelism, and FP4 roughly doubles inference throughput over FP8. Choose the B200 when maximum node throughput or on-box training dominates; choose the H200 when long-context inference economics and proven availability dominate. For interactive human use on 30–120B models, both tiers stay overprovisioned across this paper’s entire size range, including the 500 to 999 organization size band. Two conditions change that verdict: on-premises training or continuous fine-tuning in Phase 2, and agent fleets whose fan-out increase concurrency and throughput demand.

Quantization and the Quality Floor

Quantization choice is the other half of the deployment equation. Kurtic et al. evaluated FP8, INT8, and INT4 across the entire Llama-3.1 family through more than 500,000 evaluations, and the peer-reviewed result remains the most rigorous reference available [71]. Three findings translate directly into policy.

FP8 is effectively lossless. W8A8-FP recovers between 99.3% and 100.1% of BF16 accuracy across 8B, 70B, and 405B Llama-3.1, often within evaluation noise [71]. On any hardware that supports FP8 (NVIDIA Hopper, Blackwell, the RTX PRO 6000, the RTX 5090), FP8 is the default, and its throughput gain over FP16 is free.

Well-tuned INT8 sits within 1–3% of BF16. W8A8-INT recovers 97.3% to 101.5% of BF16, low enough to be invisible on most workloads [71]. INT8 is the right choice where FP8 hardware support is missing, particularly on Ampere-class GPUs (A100, A6000), and it carries throughput in asynchronous continuous batching, the standard multi-user serving mode.

INT4 is more capable than its reputation, but only when the algorithm and calibration are right. W4A16-INT (4-bit weights via GPTQ with MSE-optimal clipping and properly tuned calibration data) recovers 96.1% to 99.98% of BF16 across Llama-3.1 sizes and rivals INT8 on real-world tasks, including coding [71]. GPTQ-INT4 done correctly is production-grade. GGUF Q4_K_M and other consumer-format INT4 variants use different group sizes and calibration regimes tuned for llama.cpp on CPU or single-user GPU inference; they show wider variance and degrade more on reasoning-heavy work.

FP4 quantization is a Blackwell-architecture-native format (RTX 5090, RTX PRO 6000 Blackwell, B200) absent on Hopper-era hardware (H100, H200, A100), and it delivers throughput gains over FP8 that Akamai’s benchmark captures at 1.63× on the PRO 6000 [69] and NVIDIA’s FP4 tensor-core design targets at 2× on the same chip. Three production deployments establish the quality baseline. OpenAI post-trained gpt-oss-120b’s MoE weights in MXFP4 (the Open Compute Project Microscaling FP4 format, which adds block-level scaling to preserve dynamic range at 4-bit floating-point precision), ran every benchmark at that precision, and reached near-parity with o4-mini on core reasoning, which establishes that MoE models trained natively in MXFP4 sustain frontier-adjacent quality [3]. NVIDIA pretrained Nemotron 3 Super in NVFP4 on the same principle [18]. Command A+ supplies the third data point, and it takes a different route. Rather than train natively in the 4-bit format, Cohere quantizes a trained model selectively: it applies NVFP4 W4A4 to the MoE experts only, holds the attention path (the Q/K/V/O projections, the KV cache, and attention compute) at full precision, and uses distillation to close the residual gap, shipping the result as near lossless against BF16 [25], [26]. That selective, distillation-backed recipe is the practical bridge between native-format training and the naive post-hoc compression the Kurtic study warns against. The critical difference from INT4 is where the quality evidence sits. Kurtic et al. characterized GPTQ-INT4 applied post-hoc to BF16 models and found it production-grade when calibration is tight [71]; researchers have not yet produced an equivalent systematic study of post-hoc FP4 compression across model families and sizes. FP4’s strongest quality record still comes from native post-training, though Command A+ shows a co-designed selective conversion can reach the same bar. The operational implication is narrow: run MXFP4 or NVFP4 on Blackwell hardware when the model ships in that format, and treat FP4 as a precision available for new deployments and fine-tuning workflows, not a compression retrofit for existing BF16 weights.

The same discipline extends to the small tier and the adjacent classes. Google publishes quantization-aware-training checkpoints for EmbeddingGemma and measures the cost directly: 60.62 against 61.15 on multilingual MTEB at 4 bits [63]. FunctionGemma’s on-device numbers are quoted at dynamic INT8 [58]. When a vendor ships a quantized checkpoint with published evaluations at that precision, use it rather than compressing the full-precision weights independently.

The operational rule follows from the evidence: for production multi-user serving, run Activation-aware Weight Quantization (AWQ) or FP4/FP8 weights under vLLM, and reserve GGUF Q4_K_M for developer workstations and single-user inference where llama.cpp is the target. The two formats solve different problems and treating them as interchangeable is how quality regressions reach production.

What This Means for Phase 1

Phase 1 converges on one-to-few GPU accelerators and a short model list. For organizations whose posture requires a documented supply-chain review before any Chinese-origin model reaches production (the default for regulated data, government adjacency, or sensitive workloads) the strongest options on Phase 1 hardware are Gemma 4 31B, Nemotron 3 Super, and gpt-oss-120b on a single RTX PRO 6000 Blackwell 96 GB. Gemma 4 31B at FP8 covers general reasoning, competitive-class math, and coding at single-GPU scale. gpt-oss-120b at MXFP4 and Nemotron 3 Super at NVFP4 fit the same hardware for workloads that need the 120B tier with configurable chain-of-thought. Gemma 4 and gpt-oss-120b ship under Apache 2.0; Nemotron 3 Super ships under the NVIDIA Nemotron Open Model License, permissive in substance but not OSI-approved, so it requires a license review before standardization [20]. All three deploy under vLLM. Cohere’s Apache-2.0 releases extend the same clean-provenance set: North Mini Code, a 30B/3B coding specialist, fits the single RTX PRO 6000, and Command A+ (218B/25B) carries the agentic, multimodal tier where two H100/200s or one B200 are available, neither requiring the strict provenance review the Chinese-origin flagships require.

For organizations with no provenance constraints, adding Qwen3.6-35B-A3B, DeepSeek V4-Flash, GLM-5.2, or Kimi K2.7-Code widens the available capability range, contingent on the supply-chain review being documented and signed off; MiniMax M3 belongs in the same catalog but pairs that review with a non-permissive custom license that conditions commercial use. Skipping that review trades a signed-off risk assessment for regulatory, contractual, and reputational exposure that no benchmark score offsets. Organizations of different sizes will evaluate that trade differently and eventually come to the same conclusion. The review consumes roughly the same counsel hours at 60 employees as at 4,000, and a firm below 500 employees rarely has in-house counsel to conduct a thorough review. This is why the clean-provenance set is preferred at this scale and worth the benchmark tradeoff between it and the Chinese flagships.

The small and specialized tiers belong in the same build. Gemma 4 E4B or Granite 4.0-H-Tiny covers narrow general work; FunctionGemma or Granite 4.0-H-Micro handles tool-call formatting; EmbeddingGemma carries retrieval; Granite Speech 4.1 2B covers audio where the workload touches it; Granite Guardian 4.1 8B in no-thinking mode sits on the request path. All carry US or allied provenance, so none adds significant supply-chain review to the Phase 1 schedule, though EmbeddingGemma and FunctionGemma ship under the Gemma Terms of Use rather than Apache 2.0. What this tier does add is a logging requirement, which has a deadline rather than a budget.

The July 2026 frontier releases change the roadmap rather than the Phase 1 build. Inkling at NVFP4 needs roughly 600 GB, an eight-card node, and Kimi K3 at 4-bit is roughly 1.4 TB in size, which requires a multi-node system. For an organization below 500 employees, an eight-card node is a Phase 2 to Phase 3 deployment, and a multi-node system sits outside the hardware range this paper recommends at any phase. The practical consequence for Phase 1 is architectural rather than financial: build the serving stack on vLLM or SGLang with tensor and expert parallelism configured from the start, route requests through a model-selection layer rather than a single endpoint, and keep model-specific assumptions out of the application layer. An organization that does this can adopt Inkling-Small, add specialists as its logged traffic identifies them, or move to a K3-class model at Phase 3, without rebuilding the deployment.

References

  1. Google DeepMind, “Gemma 4: Byte for Byte, the Most Capable Open Models,” Google, Apr. 2, 2026. [Online]. Available: https://blog.google/innovation-and-ai/technology/developers-tools/gemma-4/. [Accessed: 04-Jun-2026]

    OWML-1 Secondary source Back to text

  2. Mistral AI, “Introducing Mistral 3,” Dec. 2025. [Online]. Available: https://mistral.ai/news/mistral-3. [Accessed: 04-Jun-2026]

    OWML-2 Secondary source Back to text

  3. OpenAI, “Introducing gpt-oss,” Aug. 5, 2025. [Online]. Available: https://openai.com/index/introducing-gpt-oss/. [Accessed: 04-Jun-2026]

    OWML-3 Secondary source Back to text

  4. Artificial Analysis, “Comparison of AI Models Across Intelligence, Performance, and Price,” 2026. [Online]. Available: https://artificialanalysis.ai/models. [Accessed: 04-Jun-2026]

    OWML-4 Secondary source Back to text

  5. Artificial Analysis, “GLM-5.2 Is the New Leading Open Weights Model on the Artificial Analysis Intelligence Index,” June 16, 2026. [Online]. Available: https://artificialanalysis.ai/articles/glm-5-2-is-the-new-leading-open-weights-model-on-the-artificial-analysis-intelligence-index. [Accessed: 18-Jun-2026]

    OWML-5 Secondary source Back to text

  6. Artificial Analysis, “Thinking Machines Has Released Inkling, the New Leading U.S. Open Weights Model,” July 15, 2026. [Online]. Available: https://artificialanalysis.ai/articles/thinking-machines-has-released-inkling-the-new-leading-u-s-open-weights-model. [Accessed: 18-Jul-2026]

    OWML-6 Secondary source Back to text

  7. Moonshot AI, “Kimi K3: Open Frontier Intelligence,” Kimi Blog, July 16, 2026. [Online]. Available: https://www.kimi.com/blog/kimi-k3. [Accessed: 18-Jul-2026]

    OWML-7 Secondary source Back to text

  8. Artificial Analysis, “Kimi K3 Achieves #3 in the Artificial Analysis Intelligence Index, Comparable to Opus 4.8 and GPT-5.5,” Artificial Analysis, July 16, 2026. [Online]. Available: https://artificialanalysis.ai/articles/kimi-k3-achieves-3-in-the-artificial-analysis-intelligence-index-comparable-to-opus-4-8-and-gpt-5-5. [Accessed: 18-Jul-2026]

    OWML-8 Secondary source Back to text

  9. Artificial Analysis, “Artificial Analysis Intelligence Index,” July 2026. [Online]. Available: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index. [Accessed: 18-Jul-2026]

    OWML-9 Secondary source Back to text

  10. J. Edwards and L. Emberson, “Open Models Lag State-of-the-Art Closed Models by 4 Months,” Epoch AI Data Insights, May 29, 2026. [Online]. Available: https://epoch.ai/data-insights/open-closed-eci-gap. [Accessed: 21-Jun-2026]

    OWML-10 Secondary source Back to text

  11. Google DeepMind, “Gemma 4 model card,” Google AI for Developers, Apr. 17, 2026. [Online]. Available: https://ai.google.dev/gemma/docs/core/model_card_4. [Accessed: 04-Jun-2026]

    OWML-11 Primary source Back to text

  12. Google DeepMind, “Introducing Gemma 4 12B: a unified, encoder-free multimodal model,” Google Blog, June 3, 2026. [Online]. Available: https://blog.google/innovation-and-ai/technology/developers-tools/introducing-gemma-4-12b/. [Accessed: 04-Jun-2026]

    OWML-12 Secondary source Back to text

  13. Mistral AI, “Models — From Cloud to Edge,” 2026. [Online]. Available: https://mistral.ai/models. [Accessed: 04-Jun-2026]

    OWML-13 Primary source Back to text

  14. Mistral AI, “Mistral Medium 3.5 (mistral-medium-3-5-26-04) — Model Card,” Mistral AI Documentation, Apr. 2026. [Online]. Available: https://docs.mistral.ai/models/model-cards/mistral-medium-3-5-26-04. [Accessed: 04-Jun-2026]

    OWML-14 Primary source Back to text

  15. NVIDIA Corporation, “NVIDIA Debuts Nemotron 3 Family of Open Models,” NVIDIA Newsroom, press release, Dec. 2025. [Online]. Available: https://nvidianews.nvidia.com/news/nvidia-debuts-nemotron-3-family-of-open-models. [Accessed: 04-Jun-2026]

    OWML-15 Secondary source Back to text

  16. NVIDIA Corporation, “Nemotron 3 Family of Models,” NVIDIA Research, 2026. [Online]. Available: https://research.nvidia.com/labs/nemotron/Nemotron-3/. [Accessed: 04-Jun-2026]

    OWML-16 Primary source Back to text

  17. Artificial Analysis, “NVIDIA Nemotron 3 Super: The New Leader in Open, Efficient Intelligence,” Mar. 11, 2026. [Online]. Available: https://artificialanalysis.ai/articles/nvidia-nemotron-3-super-the-new-leader-in-open-efficient-intelligence. [Accessed: 04-Jun-2026]

    OWML-17 Secondary source Back to text

  18. NVIDIA Corporation, “NVIDIA Nemotron 3 Super,” NVIDIA Research, 2026. [Online]. Available: https://research.nvidia.com/labs/nemotron/Nemotron-3-Super/. [Accessed: 04-Jun-2026]

    OWML-18 Primary source Back to text

  19. NVIDIA Corporation, “NVIDIA Nemotron 3 Ultra,” NVIDIA Research, June 4, 2026. [Online]. Available: https://research.nvidia.com/labs/nemotron/Nemotron-3-Ultra/. [Accessed: 04-Jun-2026]

    OWML-19 Primary source Back to text

  20. NVIDIA Corporation, “NVIDIA Nemotron Open Model License,” 2026. [Online]. Available: https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-nemotron-open-model-license/. [Accessed: 04-Jun-2026]

    OWML-20 Primary source Back to text

  21. Meta AI, “The Llama 4 Herd: The Beginning of a New Era of Natively Multimodal AI Innovation,” Meta Platforms, Inc., Apr. 5, 2025. [Online]. Available: https://ai.meta.com/blog/llama-4-multimodal-intelligence/. [Accessed: 04-Jun-2026]

    OWML-21 Secondary source Back to text

  22. A. Srivastava, “NVIDIA Accelerates Inference on Meta Llama 4 Scout and Maverick,” NVIDIA Technical Blog, Apr. 5, 2025. [Online]. Available: https://developer.nvidia.com/blog/nvidia-accelerates-inference-on-meta-llama-4-scout-and-maverick/. [Accessed: 04-Jun-2026]

    OWML-22 Secondary source Back to text

  23. Meta Platforms, Inc., “Llama 4 Community License Agreement,” Apr. 5, 2025. [Online]. Available: https://www.llama.com/llama4/license/. [Accessed: 04-Jun-2026]

    OWML-23 Primary source Back to text

  24. Meta Platforms, Inc., “Llama 4 Acceptable Use Policy,” Apr. 5, 2025. [Online]. Available: https://www.llama.com/llama4/use-policy/. [Accessed: 04-Jun-2026]

    OWML-24 Primary source Back to text

  25. Cohere, “Introducing Command A+: Making sovereign agentic capabilities available to all,” May 20, 2026. [Online]. Available: https://cohere.com/blog/command-a-plus. [Accessed: 18-Jun-2026]

    OWML-25 Secondary source Back to text

  26. Cohere Labs, “Command A+ (command-a-plus-05-2026) Model Card,” Hugging Face, May 20, 2026. [Online]. Available: https://huggingface.co/CohereLabs/command-a-plus-05-2026-bf16. [Accessed: 18-Jun-2026]

    OWML-26 Primary source Back to text

  27. Cohere, “Introducing North Mini Code: Cohere's First Model for Developers,” June 9, 2026. [Online]. Available: https://cohere.com/blog/north-mini-code. [Accessed: 18-Jun-2026]

    OWML-27 Secondary source Back to text

  28. Artificial Analysis, “North Mini Code — Intelligence, Performance & Price Analysis,” June 9, 2026. [Online]. Available: https://artificialanalysis.ai/models/north-mini-code. [Accessed: 18-Jun-2026]

    OWML-28 Secondary source Back to text

  29. M. Nuñez, “Cohere Open-Sources a Coding Agent That Runs on a Single H100,” VentureBeat, June 9, 2026. [Online]. Available: https://venturebeat.com/technology/cohere-open-sources-a-coding-agent-that-runs-on-a-single-h100. [Accessed: 18-Jun-2026]

    OWML-29 Contextual source Back to text

  30. VentureBeat, “Thinking Machines Open Sources First Multimodal Language Model, Inkling, Focused on Low Cost and 'Resistance to Censorship',” July 16, 2026. [Online]. Available: https://venturebeat.com/technology/thinking-machines-open-sources-first-multimodal-language-model-inkling-focused-on-low-cost-and-resistance-to-censorship. [Accessed: 18-Jul-2026]

    OWML-30 Contextual source Back to text

  31. Implicator AI, “Thinking Machines Inkling Takes U.S. Open Model Lead With 41 Score,” Implicator, July 16, 2026. [Online]. Available: https://www.implicator.ai/thinking-machines-inkling-takes-u-s-open-model-lead-with-41-score/. [Accessed: 18-Jul-2026]

    OWML-31 Contextual source Back to text

  32. Qwen Team, Alibaba Cloud, “Qwen3.5-397B-A17B Model Card,” Hugging Face, Feb. 16, 2026. [Online]. Available: https://huggingface.co/Qwen/Qwen3.5-397B-A17B. [Accessed: 04-Jun-2026]

    OWML-32 Primary source Back to text

  33. Qwen Team, Alibaba Cloud, “Qwen3.6-35B-A3B Model Card,” Hugging Face, Apr. 2026. [Online]. Available: https://huggingface.co/Qwen/Qwen3.6-35B-A3B. [Accessed: 04-Jun-2026]

    OWML-33 Primary source Back to text

  34. Artificial Analysis, “Qwen3.6 35B A3B (Reasoning) — Intelligence, Performance & Price Analysis,” Apr. 2026. [Online]. Available: https://artificialanalysis.ai/models/qwen3-6-35b-a3b. [Accessed: 04-Jun-2026]

    OWML-34 Secondary source Back to text

  35. Qwen Team, Alibaba Cloud, “Qwen3.7: The Agent Frontier,” Qwen Blog, May 20, 2026. [Online]. Available: https://qwen.ai/blog?id=qwen3.7. [Accessed: 04-Jun-2026]

    OWML-35 Secondary source Back to text

  36. DeepSeek-AI, “DeepSeek-V4-Pro Model Card,” Hugging Face, Apr. 24, 2026. [Online]. Available: https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro. [Accessed: 04-Jun-2026]

    OWML-36 Primary source Back to text

  37. DeepSeek-AI, “DeepSeek-V3 Technical Report,” arXiv, Dec. 26, 2024. [Online]. Available: https://arxiv.org/abs/2412.19437. arXiv:2412.19437. [Accessed: 04-Jun-2026]

    OWML-37 Primary source Back to text

  38. Z.ai, “GLM-5: From Vibe Coding to Agentic Engineering,” Z.ai Blog, Feb. 12, 2026. [Online]. Available: https://z.ai/blog/glm-5. [Accessed: 04-Jun-2026]

    OWML-38 Secondary source Back to text

  39. Z.ai (Zhipu AI), “GLM-5.2,” Hugging Face, June 16, 2026. [Online]. Available: https://huggingface.co/zai-org/GLM-5.2. [Accessed: 18-Jun-2026]

    OWML-39 Primary source Back to text

  40. Moonshot AI, “Kimi-K2.6,” Hugging Face, Apr. 20, 2026. [Online]. Available: https://huggingface.co/moonshotai/Kimi-K2.6. [Accessed: 04-Jun-2026]

    OWML-40 Primary source Back to text

  41. Moonshot AI, “Kimi-K2.7-Code Model Card,” Hugging Face, June 12, 2026. [Online]. Available: https://huggingface.co/moonshotai/Kimi-K2.7-Code. [Accessed: 18-Jun-2026]

    OWML-41 Primary source Back to text

  42. MiniMax AI, “MiniMax-M3 Model Card,” Hugging Face, June 7, 2026. [Online]. Available: https://huggingface.co/MiniMaxAI/MiniMax-M3. [Accessed: 18-Jun-2026]

    OWML-42 Primary source Back to text

  43. X. Lai, W. Xu, and Y. Yang et al., “MiniMax Sparse Attention,” arXiv, June 11, 2026. [Online]. Available: https://arxiv.org/abs/2606.13392. arXiv:2606.13392 [cs.AI]. [Accessed: 18-Jun-2026]

    OWML-43 Secondary source Back to text

  44. Moonshot AI, “Kimi K3 Quickstart,” Kimi API Platform Documentation, July 2026. [Online]. Available: https://platform.kimi.ai/docs/guide/kimi-k3-quickstart. [Accessed: 18-Jul-2026]

    OWML-44 Primary source Back to text

  45. Artificial Analysis, “Kimi K3 — Intelligence, Performance & Price Analysis,” July 2026. [Online]. Available: https://artificialanalysis.ai/models/kimi-k3. [Accessed: 18-Jul-2026]

    OWML-45 Secondary source Back to text

  46. M. Bastian, “Kimi's Open Model K3 Nears GPT-5.6 Sol and Fable 5 While Signaling the End of Super Cheap Chinese AI,” The Decoder, July 17, 2026. [Online]. Available: https://the-decoder.com/kimis-open-model-k3-nears-gpt-5-6-sol-and-fable-5-while-signaling-the-end-of-super-cheap-chinese-ai/. [Accessed: 18-Jul-2026]

    OWML-46 Contextual source Back to text

  47. N. Lee, R. Huffman, R. Burnette, A. Gweon, A. Shah, and B. Adetula, “U.S. Federal and State Governments Moving Quickly to Restrict Use of DeepSeek,” Inside Government Contracts, Covington & Burling LLP, Feb. 17, 2025. [Online]. Available: https://www.insidegovernmentcontracts.com/2025/02/u-s-federal-and-states-governments-moving-quickly-to-restrict-use-of-deepseek/. [Accessed: 04-Jun-2026]

    OWML-47 Secondary source Back to text

  48. P. Kassianik and A. Karbasi, “Evaluating Security Risk in DeepSeek and Other Frontier Reasoning Models,” Cisco Blogs, Jan. 31, 2025. [Online]. Available: https://blogs.cisco.com/security/evaluating-security-risk-in-deepseek-and-other-frontier-reasoning-models. [Accessed: 04-Jun-2026]

    OWML-48 Secondary source Back to text

  49. G. Nagli, “Wiz Research Uncovers Exposed DeepSeek Database Leaking Sensitive Information, Including Chat History,” Wiz Blog, Jan. 29, 2025. [Online]. Available: https://www.wiz.io/blog/wiz-research-uncovers-exposed-deepseek-database-leak. [Accessed: 04-Jun-2026]

    OWML-49 Secondary source Back to text

  50. U.S. Congress, “H.R. 1121 — No DeepSeek on Government Devices Act,” 119th Congress, Feb. 7, 2025. [Online]. Available: https://www.congress.gov/bill/119th-congress/house-bill/1121. [Accessed: 04-Jun-2026]

    OWML-50 Primary source Back to text

  51. U.S. House Select Committee on the Chinese Communist Party, “DeepSeek Unmasked: Exposing the CCP's Latest Tool for Spying, Stealing, and Subverting U.S. Export Control Restrictions,” U.S. House of Representatives, Apr. 16, 2025. [Online]. Available: https://chinaselectcommittee.house.gov/media/reports/deepseek-unmasked-exposing-the-ccp-s-latest-tool-for-spying-stealing-and-subverting-us-export-control-restrictions. [Accessed: 04-Jun-2026]

    OWML-51 Primary source Back to text

  52. Allen Institute for AI, “OLMo 2: The Best Fully Open Language Model to Date,” Ai2, Nov. 26, 2024. [Online]. Available: https://allenai.org/blog/olmo2. [Accessed: 04-Jun-2026]

    OWML-52 Secondary source Back to text

  53. Allen Institute for AI, “OLMo 2 32B: First Fully Open Model to Outperform GPT-3.5 and GPT-4o Mini,” Ai2, Mar. 13, 2025. [Online]. Available: https://allenai.org/blog/olmo2-32B. [Accessed: 04-Jun-2026]

    OWML-53 Secondary source Back to text

  54. Team Olmo and A. Ettinger et al., “OLMo 3,” arXiv, Dec. 2025. [Online]. Available: https://arxiv.org/abs/2512.13961. arXiv:2512.13961 [cs.CL]. [Accessed: 04-Jun-2026]

    OWML-54 Primary source Back to text

  55. P. Belcak et al., “Small Language Models are the Future of Agentic AI,” arXiv, NVIDIA Research, June 2, 2025. [Online]. Available: https://arxiv.org/abs/2506.02153. arXiv:2506.02153v2 [cs.AI], revised Sep. 15, 2025. [Accessed: 22-Jul-2026]

    OWML-55 Secondary source Back to text

  56. IBM Corporation, “Granite 4.0,” IBM Granite Documentation, 2026. [Online]. Available: https://www.ibm.com/granite/docs/models/granite. [Accessed: 22-Jul-2026]

    OWML-56 Primary source Back to text

  57. IBM Corporation, “Granite 4.1,” IBM Granite Documentation, 2026. [Online]. Available: https://www.ibm.com/granite/docs/models/granite4-1. [Accessed: 22-Jul-2026]

    OWML-57 Primary source Back to text

  58. Google DeepMind, “FunctionGemma model card,” Google AI for Developers, Apr. 16, 2026. [Online]. Available: https://ai.google.dev/gemma/docs/functiongemma/model_card. [Accessed: 22-Jul-2026]

    OWML-58 Primary source Back to text

  59. W. Albayaydh, R. Zhao, and I. Flechais, “Beyond the Leaderboard: A Synthesis of Tool-Use, Planning, and Reasoning Failures in Large Language Model Agents,” arXiv, July 2026. [Online]. Available: https://arxiv.org/abs/2607.05775. arXiv:2607.05775 [cs.AI]. [Accessed: 22-Jul-2026]

    OWML-59 Secondary source Back to text

  60. H. Bai et al., “The Long-Horizon Task Mirage? Diagnosing Where and Why Agentic Systems Break,” arXiv, Apr. 13, 2026. [Online]. Available: https://arxiv.org/abs/2604.11978. arXiv:2604.11978 [cs.AI]. [Accessed: 22-Jul-2026]

    OWML-60 Secondary source Back to text

  61. Microsoft Corporation, “Phi-4-reasoning-vision-15B,” Hugging Face, Mar. 4, 2026. [Online]. Available: https://huggingface.co/microsoft/Phi-4-reasoning-vision-15B. [Accessed: 04-Jun-2026]

    OWML-61 Primary source Back to text

  62. IBM Corporation, “Granite Guardian 4.1,” IBM Granite Documentation, 2026. [Online]. Available: https://www.ibm.com/granite/docs/models/guardian. [Accessed: 22-Jul-2026]

    OWML-62 Primary source Back to text

  63. Google DeepMind, “EmbeddingGemma model card,” Google AI for Developers, Sept. 25, 2025. [Online]. Available: https://ai.google.dev/gemma/docs/embeddinggemma/model_card. [Accessed: 22-Jul-2026]

    OWML-63 Primary source Back to text

  64. H. Schechter Vera, S. Dua, and EmbeddingGemma Team, “EmbeddingGemma: Powerful and Lightweight Text Representations,” arXiv, Google DeepMind, Sept. 2025. [Online]. Available: https://arxiv.org/abs/2509.20354. arXiv:2509.20354. [Accessed: 22-Jul-2026]

    OWML-64 Primary source Back to text

  65. IBM Corporation, “Granite Speech,” IBM Granite Documentation, 2026. [Online]. Available: https://www.ibm.com/granite/docs/models/speech. [Accessed: 22-Jul-2026]

    OWML-65 Primary source Back to text

  66. Google DeepMind, “DiffusionGemma model card,” Google AI for Developers, June 10, 2026. [Online]. Available: https://ai.google.dev/gemma/docs/diffusiongemma/model_card. [Accessed: 22-Jul-2026]

    OWML-66 Primary source Back to text

  67. D. Trifonov, “RTX 4090 vs 5090 vs PRO 6000: LLM Inference Benchmark,” CloudRift AI, Oct. 9, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx-gpus-for-llm-inference. [Accessed: 04-Jun-2026]

    OWML-67 Secondary source Back to text

  68. N. Trifonova, “Blackwell Dominates. Benchmarking LLM Inference on NVIDIA B200, H200, H100, and RTX PRO 6000,” CloudRift AI, Jan. 21, 2026. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-b200. [Accessed: 04-Jun-2026]

    OWML-68 Secondary source Back to text

  69. M. Tabares and C. Lutzer, “Benchmarking NVIDIA RTX PRO 6000 Blackwell on Akamai Cloud,” Akamai Blog, Oct. 30, 2025. [Online]. Available: https://www.akamai.com/blog/cloud/benchmarking-nvidia-rtx-pro-6000-blackwell-akamai-cloud. [Accessed: 04-Jun-2026]

    OWML-69 Secondary source Back to text

  70. D. Trifonov, “RTX PRO 6000 vs H100, H200, and L40S: LLM Inference,” CloudRift AI, Nov. 27, 2025. [Online]. Available: https://www.cloudrift.ai/blog/benchmarking-rtx6000-vs-datacenter-gpus. [Accessed: 04-Jun-2026]

    OWML-70 Secondary source Back to text

  71. E. Kurtic, A. N. Marques, S. Pandit, M. Kurtz, and D. Alistarh, “'Give Me BF16 or Give Me Death'? Accuracy-Performance Trade-Offs in LLM Quantization,” Proc. 63rd Annu. Meeting of the Association for Computational Linguistics (ACL 2025), Vol. 1: Long Papers, July 2025. [Online]. Available: https://aclanthology.org/2025.acl-long.1304/. Vienna, Austria, pp. 26872–26886. [Accessed: 04-Jun-2026]

    OWML-71 Primary source Back to text

Contents