Section17

Expert Voices: External Validation

The experts who disagree most sharply about artificial intelligence agree on a narrow proposition that carries the whole weight of this paper’s case: current models produce measurable value on tasks whose output can be validated quickly. The constraint on capturing that value is deployment design rather than model capability. Nothing in the investment case requires resolving whether autoregressive large language models (LLMs) reach human-level reasoning, or whether economy-wide productivity gains land in 2026 or 2036. Executives who read the productivity-paradox literature as proof that waiting was prudent have read the past three years correctly and the design space ahead incorrectly.

The witnesses fall into four groups. Two frontier researchers hold opposite views on where the technology is going. Yann LeCun, a Turing Award laureate who invented the convolutional network the field ran on for two decades, has put a billion dollars of investor capital behind the position that today’s transformer-based family is a dead end on the path to human-level intelligence [1], [2]. Dario Amodei, chief executive of Anthropic, argues from inside a frontier lab that the public systematically underestimates how fast capability is arriving [3]. Richard Sutton, who took the 2024 Turing Award for founding reinforcement learning, reaches LeCun’s conclusion by a different route: these models cannot learn from experience once training ends, and something else will replace them [4]. Two economists read the same macroeconomic record and reach opposite conclusions: Daron Acemoglu, the 2024 Nobel laureate in economics, prices the decade’s productivity gain at roughly half a percent of total factor productivity [5], while Erik Brynjolfsson, director of Stanford’s Digital Economy Lab, argues the take-off is already visible in the 2025 data [6]. Satya Nadella, whose company sells more of this technology than any other named here, set the macroeconomic bar himself and says it has not been cleared [7], [8]. Between them sit the practitioners who build, deploy, operate, and educate on this software: Andrej Karpathy, Erik Meijer, Gene Kim, Birgitta Böckeler, Martin Fowler, Kelsey Hightower, Linus Torvalds, and Simon Willison. Dr. Eric Topol provides the outside perspective on advances in AI and medicine[9].

Their agreement is narrow and sufficient. The value current models produce on well-scoped tasks survives whatever architecture comes next, because graphics processing units (GPUs), networking, storage, and operational capability are not specific to transformers.

Karpathy: Verifiability as the Selection Rule

Andrej Karpathy was a founding member of OpenAI, ran Tesla’s Autopilot vision program, built and taught Stanford’s first deep learning course, and coined the term “vibe coding” in early 2025 [10]. In May 2026 he joined Anthropic’s pre-training team under team lead Nick Joseph, tasked with building a group that applies Claude to pre-training research [11].

What makes him worth citing is the direction of his errors. Through 2025 he publicly deflated expectations at a moment when enthusiasm was the profitable position, calling the industry’s presentation of agentic capability “slop” and arguing that the correct framing was a decade of agents rather than a year of them [12]. A researcher who talks down the technology he works on is a better witness than one who talks it up.

His 2017 essay “Software 2.0” remains the most cited framing of neural networks as a way to build software: Software 1.0 is explicit logic a human writes, Software 2.0 is a program encoded in network weights and optimized against an objective [13]. In his June 2025 keynote he extended this to Software 3.0, where the programmer instructs a model through prompts, context, tools, and examples with the context window becoming the control surface [14].

The important contribution is the verifiability thesis, which he set out in a November 2025 essay and restated at Sequoia Ascent in April 2026. Software 1.0 automates what a programmer can specify; in his formulation, “Software 2.0 easily automates what you can verify” [15]. The mechanism is concrete and testable. Reinforcement learning can practice a task at scale only when the environment is resettable, efficient enough to permit many attempts, and rewardable by an automated signal [15]. Mathematics, coding against tests, and retrieval-augmented generation (RAG) against ground-truth documents satisfy all three. Creative, strategic, and common-sense tasks satisfy none; improving slowly and unevenly as a result.

This is the same task-conditional model the productivity analysis in this paper derives from the empirical literature, reached from the machine-learning side rather than the economics side. Brynjolfsson, Li, and Raymond’s customer-support field experiment anchors the affirmative case with generation-heavy work that the customer validates immediately [16]. That is precisely the profile the verifiability thesis predicts will pay out.

Between 2025 and 2026, Karpathy calibrated his agentic expectations as the technology matured. In October 2025 he argued that agents capable of acting as trusted employees on open-ended work were roughly a decade away; naming autocomplete as his own sweet spot [12]. Then, at Sequoia in April 2026, he described December 2025 as “a stark transition” after which he began delegating larger units of work to coding agents, such as implementing a feature or refactoring a subsystem [10]. Both statements can hold, because they describe different objects. His October 2025 decade-scale claim concerns agents as autonomous employees; the December claim concerns agents as supervised coding tools. The revision is the useful part: a public skeptic moved his practical estimate on evidence and said so. For procurement, the consequence is that the agentic compute capacity specified in the phased hardware roadmap enters a regime just past its inflection point. Useful enough to plan around, but not reliable to staff against.

LeCun: The Strongest Architectural Skeptic

In 2018, Yann LeCun, Geoffrey Hinton and Yoshua Bengio were awarded the Association for Computing Machinery (ACM) Turing Award, often referred to as the Nobel Prize of Computing, for their conceptual and engineering breakthroughs that established deep neural networks as a critical pillar of computing. His contributions to computing date back to 1989 with his introduction of the Convolutional Neural Network (CNN) architecture that carried computer vision for over twenty years. He founded Meta’s Fundamental AI Research lab in 2013 and directed it for twelve years while holding a New York University professorship. He is the reason to engage architectural skepticism seriously since he was right about neural networks when the consensus held they were a dead end. That is the structure of the claim he now makes about the transformer architecture behind modern LLMs.

His argument appears in its most complete form in his 2022 position paper, which the peer-reviewed literature cites heavily [1]. Next-token prediction over text is confined to fast, reactive processing. Human-level reasoning requires a configurable predictive world model, hierarchical planning, and persistent memory, none of which token prediction produces, because most of what humans know is not encoded in language. His alternative is the Joint Embedding Predictive Architecture (JEPA), which predicts abstract representations of future states rather than pixels or tokens. Meta instantiated it in V-JEPA 2, a 1.2-billion-parameter world model trained on more than a million hours of video and released in June 2025 with three new physical-reasoning benchmarks [17].

His independence is now beyond question, because his commercial position runs against his own conclusion. LeCun departed from Meta in November 2025 and went on to co-found AMI Labs with Alexandre LeBrun, who serves as chief executive, and holds the chairman role while retaining his NYU position. In March 2026 the company announced a $1.03 billion seed round at a $3.5 billion pre-money valuation, with NVIDIA among the backers [2]. His commercial incentive now points toward large language models not being able to reach general intelligence, which inverts the conflict the independence test usually screens for. Holding that incentive, LeBrun still frames the company’s horizon in years and concedes that a viable JEPA-based alternative will take time to arrive [2].

LeCun answers a different question than this paper does. His claim concerns whether the current architectural family can reach human-level reasoning. This paper’s claim concerns whether current models deliver value on code generation, retrieval-augmented question answering over internal documentation, document drafting, log analysis, troubleshooting support, and other generation-heavy, easily verifiable tasks. Both propositions can be true at once. The part of his view that bears on procurement strengthens the long-horizon case: if JEPA-style world models displace autoregressive transformers, the on-premises stack specified in the hardware roadmap runs the successor family without modification, because GPUs, high-bandwidth memory, networking, and operational capability are architecture-agnostic. The investment thesis never depended on the transformer being terminal.

However, there is a version of LeCun’s position that would damage the case. If transformer capability plateaus before the models become reliable on the specific enterprise tasks planned here, and if the successor architecture demands hardware this build cannot host, then the Phase 1 asset depreciates faster than the return period assumes. Nothing in the current evidence points that way. The re-evaluation criteria in the vendor comparison exist to catch it if it starts to.

Sutton: The Limit the Deployment Has to Design Around

Richard Sutton took the 2024 ACM Turing Award for the body of work that established reinforcement learning as a field along with his his 2019 essay “The Bitter Lesson” as the argument the scaling era cited to justify itself. In September 2025 he turned that argument against the models built on it. His position is that large language models learn to predict what a person would say instead of what the world does, that they cannot learn on the job once training ends, and that a continual-learning architecture will therefore supersede them regardless of how much compute the current family absorbs [4]. The interviewer put the opposing view to him directly, that language models could serve as the prior on which experiential learning is built, with which Sutton disagreed [4].

The reason to treat this as more than a strong opinion is that the mechanism has peer-reviewed support from Sutton’s own group. In Nature, Dohare and colleagues showed that standard deep-learning methods progressively lose plasticity when training continues across a sequence of tasks, degrading until the network learns no better than a shallow one; on ImageNet repurposed for continual learning, binary classification accuracy fell from roughly 88% on an early task to roughly 77% by the two-thousandth task, and the authors’ continual backpropagation algorithm maintained plasticity where standard methods did not [18]. That is a measured failure mode with a published method, not a forecast.

The deployment consequence is specific and it cuts both ways. Institutional knowledge cannot be expected to accumulate in the weights of a deployed model, which is why the retrieval corpora and context engineering described in this paper’s model-selection guidance carry the organization’s facts and fine-tuning is reserved for format and behavior. The offsetting benefit is rarely stated: a model that does not learn in place also does not drift in place. Performance measured at acceptance testing remains valid until the operator deliberately changes the model, the prompt, or the corpus, which makes the system auditable in a way a continually learning one would not be. Sutton is describing a ceiling on where the technology can go. For a fixed deployment on a fixed set of tasks, the same property is a stability guarantee.

Amodei: The Capability Trajectory

Dario Amodei led research at OpenAI before co-founding Anthropic in 2021, where he serves as chief executive. He holds a Princeton doctorate in biophysics and was among the first to document and track the scaling laws that relate model capability to compute and training data [19]. His forecasts carry weight for one specific reason: he belongs to the small group that watches capability emerge from a training run before anyone outside the lab sees it.

His October 2024 essay “Machines of Loving Grace” is the strongest documented case from a sitting lab leader that the public underestimates the trajectory. He defines powerful AI as a system smarter than a Nobel laureate across most fields, able to work autonomously on tasks lasting days or weeks, and runnable in millions of parallel instances. His central prediction is that such systems would compress 50 to 100 years of biological and medical progress into 5 to 10 years, a scenario he calls the compressed twenty-first century [20]. His January 2026 essay “The Adolescence of Technology” is the risk-side companion, reaffirming the timeline with the claim that powerful AI could be “as little as 1–2 years away,” while allowing that it could also be considerably further out [3]. In the same essay he predicts that AI could displace half of all entry-level white-collar jobs within one to five years [3].

Three qualifications govern how this paper uses him. First, he states his own uncertainty in the text, writing that nothing in the essay is meant to communicate certainty or even likelihood, and that AI may simply not advance as fast as he imagines [3]. Second, these are forward projections about future capability, not descriptions of today’s models, and every deployment recommendation in this paper is calibrated to current capability. The urgency argument borrows his rate-of-change framing because that rate is what the organization must plan against.

Third, the disclosure. This paper presents Anthropic as one of several cloud-model vendors for the augmentation strategy, noting Dario Amodei as head of the company, since they have constantly produced some of the most capable models to date while also being a leader in AI safety efforts. With Karpathy being an Anthropic employee since May 2026 [11], that makes two of the voices in this section carry a direct financial relationship with a vendor the paper presents. Part of the argument’s force survives that disclosure, because Amodei’s position, whose company revenue depends on selling AI capability, is arguing that the upside is underestimated, and because Karpathy’s cited positions predate his hire and cut against his employer’s commercial interest. Readers should nonetheless discount both accordingly since this paper’s position does not rest on either.

Meijer and Kim: Verification Is the Scarce Skill

Erik Meijer is a programming-language researcher whose work shaped Visual Basic, C#, .NET, LINQ, Hack, and others. He spent years at Microsoft and Meta, holds a part-time professorship in cloud programming at Delft University of Technology, and now runs a venture building program-synthesis tooling [21]. That last fact is a disclosed commercial interest: he sells a framework premised on models writing the code.

His formal position appears in a peer-reviewed venue, which is why it belongs here. His ACM Queue paper describes Universalis, a language for instructing models, and Automind, a runtime that treats the LLM as the virtual machine executing those programs. The design embeds pre-condition and post-condition contracts directly in the language semantics, so a developer can reason about whether a generated program does what it should and raise an error when a condition is violated [21]. That is a verification framework, not a productivity claim, and it is the supporting citation in this subsection.

Gene Kim supplies the organizational half. He wrote The Phoenix Project and The DevOps Handbook, has spent two decades studying why technology change succeeds or fails inside organizations, and co-authored a book on AI-assisted development with Steve Yegge. In March 2026 he published an account of a private conversation with Meijer in which Meijer argued that coding should be treated as a solved problem for planning purposes. Meijer’s reasoning, as Kim recounts it, is that across the universe of computable functions “we only use a tiny, tiny subset,” concentrated in the low-complexity classes that survive production review at firms like Meta and Microsoft [22]. Two caveats apply and both are the reader’s to weigh. First, the account is secondhand, unrecorded, and appeared in a post promoting a conference at which Meijer was speaking. Second, the claim is directionally consistent with Meijer’s published work; however, it is not independently corroborated.

The implication this paper draws concerns workforce strategy. As code production gets cheaper, the human capability worth funding is design, architecture, agentic orchestration, and the verification discipline Meijer’s contracts formalize. The productivity analysis in this paper puts an empirical floor under that conclusion: in vendor-collected telemetry from roughly 10,000 developers, teams that adopted AI heavily merged 98% more pull requests while delivery metrics stayed flat, because the bottleneck migrated from writing code to reviewing it. That measurement comes from a company selling engineering analytics and should be read with the discount that implies, but the direction matches everything else in the record. Generation got cheaper. Verification did not.

Böckeler and Fowler: Field Reports on Non-Deterministic Tools

Martin Fowler wrote Refactoring, co-authored the Agile Manifesto, and holds the title of Chief Scientist at Thoughtworks, where he draws a salary and holds equity in a firm that sells AI-assisted delivery consulting [23]. His June 2025 article states the framing this paper adopts: LLMs change software development on the scale of the move from assembler to the first high-level languages, with one genuinely new wrinkle. They raise the abstraction level and are simultaneously “forcing us to consider what it means to program with non-deterministic tools” [24]. A Fortran function compiled a hundred times reproduces the same bug. A prompt stored in version control guarantees nothing about the output generated by the next run. Fowler is candid that he had done little more than dabble with the tools himself at the time of writing [24], which is why the field evidence below carries the weight.

That evidence comes from Birgitta Böckeler, a Distinguished Engineer at Thoughtworks with more than twenty years as a developer and architect, who coordinates the firm’s work on AI-assisted delivery and writes the “Exploring Generative AI” memo series that Fowler edits and publishes [25]. The series is the reason to read this group: it documents what failed on real engagements instead of only publishing conclusions.

Two July 2026 memos speak directly to the question this paper exists to answer. Böckeler ran open-weight coding models locally for four weeks and reported the results without softening them. Response speed had improved substantially against a year earlier. Tool calling, the capability agentic workflows depend on, still failed often, though models usually recovered from their own malformed calls. Output quality was “hit and miss” and well below frontier hosted models. Her conclusion is that local models are not yet a plug-and-play experience for developers unwilling to invest setup time, and her working choice after the exercise was Qwen3.6 35B mixture-of-experts at 4-bit quantization, at roughly 22 GB [26], [27].

The constraint she hit is the one this paper’s hardware specification removes. Böckeler ran on Apple M3 Max and M5 Pro laptops with 48 GB and 64 GB of unified memory at roughly 300 GB/s of bandwidth, and she reports memory capacity as the binding limit: models near 30 GB crowded out the context window, and a 48 GB model crashed outright [26]. The Phase 1 specification in the hardware roadmap is a server-class card with 96 GB of dedicated memory and an order of magnitude more bandwidth, hosting the same Gemma 4 31B class she tested plus 70B-class models at FP8. Her findings are a fair warning about laptop-class local inference and a direct argument for the server-class build. Read as a caution against on-premises deployment generally, they would be misapplied.

The engineering implication Fowler and Böckeler share is that generated code is a hypothesis requiring validation rather than a deliverable requiring acceptance. Architectural judgment and refactoring discipline gain value in an AI-assisted environment, because unreviewed generated code accumulates technical debt faster than hand-written code does. That lands where Meijer’s contracts and Karpathy’s verifiability framing both point.

Torvalds: A Skeptic Who Changed Position

Linus Torvalds created the Linux kernel in 1991 and the Git version control system in 2005. The kernel is the operating system behind most of the world’s internet computing servers, smartphones, and other high-tech devices along with the running the inference runtime and container platforms this paper specifies, so his judgment about generated code applies to the same environment most technical organizations operate within. He also sells nothing and holds no position within any AI vendor organization.

In October 2024 he told an interviewer at Open Source Summit Vienna that the industry was “90% marketing and 10% reality,” said he would ignore it, and gave the technology five years before real workloads would show what it was good for [28]. On July 14, 2026, replying on the linux-media list to a maintainer who wanted the output of an agentic patch-review system triaged before it reached patch authors, he wrote that “Linux is not one of those anti-AI projects” and told anyone who disagreed to fork the kernel or walk away [29]. In the same message he described AI as a tool that is now clearly useful, said the open question is the economics of it rather than the utility, and observed that the review system under dispute keeps surfacing bugs the maintainers had missed. He conceded that these tools raise maintainer workload, and argued that the answer is to make the tools serve maintainers instead of making things more difficult [29].

The interval between those two statements is roughly twenty-one months. That is the same pattern this section credits in Karpathy, from a person with less to gain and a larger installed base to affect. Note what did not change: Torvalds never claimed the tools were reliable. He claimed they were useful, which is a lower and more defensible bar; setting the accountability rules that make usefulness safe to act on.

Those rules are the transferable artifact an organization can adopt for free. The kernel now carries a process document stating that generated code must comply with the GPL-2.0-only license and carry appropriate SPDX identifiers, that “AI agents MUST NOT add Signed-off-by tags” because only a human can certify the Developer Certificate of Origin (DCO), and that contributions should carry an Assisted-by trailer naming the agent and the model version, with specialized analysis tools listed and ordinary tools such as git and gcc omitted [30]. The named human reviews the generated code, confirms licensing compliance, and takes full responsibility for the contribution [30]. A companion document sets the disclosure norm: contributors describe which tools produced which portions, include the prompts when a change came from a short prompt set or a summary of them for longer sessions, and state how the change was tested. It also tells contributors to expect scrutiny in proportion to how much of the submission was generated, and preserves maintainer discretion to accept, reject, deprioritize, demand extra testing, or ask the submitter to explain the code before review [31]. Anyone who cannot defend what they submitted is told not to submit it.

That structure answers the accountability problem without banning the tools or pretending disclosure is optional. It is also auditable with a simple git log query over a fixed trailer, which is a lower governance cost than any policy requiring a review board. One important distinction is that the kernel has more active reviewers than most small and medium organizations have employees, so the scale does not necessarily transfer over. The premise does, however: the kernel wrote these rules because reviewer bandwidth is its scarce resource, which is the same conclusion the verification argument reaches from the opposite direction.

Hightower: The Location Decision Is an Economics Decision

Kelsey Hightower is a renowned software engineer and former Google Cloud Distinguished Engineer who co-authored Kubernetes: Up and Running and created the foundational teaching materials used by most system operators. Widely celebrated for his human-first approach to technology and exceptional live demos, he helped shape the cloud-native ecosystem before retiring from Google in 2023. Two facts make him the right witness on location economics. He helped build the abstractions that closed the operational gap between a rented data center and an owned one, and he no longer draws income from either side of that choice.

At a 2024 industry roundtable he stated the conclusion bluntly: “There is no hybrid anything. There are just data centers” [32]. His point is that modern tooling has closed most of the management gap between on-premises and cloud environments, which makes location an economics and governance question rather than a technical one. He paired that with the observation that the major clouds have converged on near-identical core offerings and that many organizations pay considerably more than they need to [32]. That is a practitioner statement of the avoided-cost argument the return-on-investment analysis in this paper quantifies. One caveat belongs on the record: the roundtable was hosted by a data-center networking vendor and the write-up was published by that vendor before republication, so the venue had a commercial interest in the conclusion. Hightower’s independence is not in question; the venue’s is.

His more recent position sets a condition on the deployment case. In a June 2026 interview he argued that agents should not be turned loose on raw infrastructure without guardrails and supplied context, putting it as a comparison to what people already do with unrestricted console access: “I’ve seen what humans do when you just give them the AWS console. Watch what Claude’s going to do!” [33]. In the same conversation he described challenging founders to explain what their company does without saying “AI,” on the grounds that the exercise often reveals a cheaper, simpler path to the same outcome [33]. Applied here, that is the governance position this paper takes: AI deployed without operational discipline generates containment costs capable of exceeding the productivity gains. Not every workload that could run through a model should.

Willison: The Failure Mode That Bounds Agent Design

Simon Willison co-created the Django web framework and coined the term prompt injection in September 2022 to describe what happens when trusted and untrusted content share a context window [34]. He named it after SQL injection because the underlying defect is the same. In June 2025 he gave the operational form of the problem a name that the security field adopted quickly: the “lethal trifecta” [34].

The three capabilities are access to private data, exposure to content an attacker can control, and any channel that can carry data outward. An agent holding all three can be induced to exfiltrate, because models cannot reliably rank the authority of an instruction by where it came from; operator instructions and retrieved documents arrive as the same token stream [34]. Willison catalogues the pattern in production systems across many vendors and distrusts guardrail products that advertise catching a high share of attacks, on the ground that a 95% catch rate is a failing grade in application security [34]. The research literature agrees on the shape of the answer. Beurer-Kellner and colleagues propose six design patterns that resist injection by constraining agents so they cannot execute arbitrary tasks. They conclude that general-purpose agents are unlikely to offer reliable safety guarantees with the current class of language models [35].

For an on-premises program this sets a boundary the hardware cannot move. Running the model on owned infrastructure removes the vendor from the data path, which is a real benefit the security architecture in this paper claims. It does nothing about the injection path, because the untrusted content is the organization’s own inbound email, support tickets, supplier documents, and retrieved web pages. The agentic workloads specified in the roadmap should be scoped one workflow at a time after first closing the exfiltration channel since that is a vulnerability an operator can actually resolve without destroying the agent’s usefulness.

Nadella: The Test a Business Leader Set for Himself

Satya Nadella runs the company with one of the largest commercial exposure to AI infrastructure of anyone quoted in this section. Microsoft sells access to the compute, the models, and the applications. His framing is worth reading precisely because it fails the independence test in the direction that favors this paper’s skeptics.

In February 2025 he rejected self-declared artificial general intelligence (AGI) milestones as benchmark hacking and named a macroeconomic bar instead: “The real benchmark is: the world growing at 10%” [7]. He put the developed-world figure at about 7% and inflation-adjusted growth at about 5%, tying the shortfall to workflow rather than capability, and used the arrival of the spreadsheet and email as the analogy: corporate forecasting changed when the work artifact and the workflow changed, not when the computing got faster [7]. He has held that position through 2026. When asked in July 2026 whether AI is a bubble, he did not deny it, but instead stated that productivity has to show up as broad, economy-wide growth; otherwise, “we’re not going to have this movie end well” [8].

Nadella’s own diagnosis of the gap is workflow redesign, which is this paper’s core claim for effective AI adoption. The economic growth indicator he mentions is restated below by two leading economists. Their economics disagreement provides helpful insights to test whether the organization’s investment into AI is worth its cost.

Acemoglu and Brynjolfsson: An Economics Disagreement

Daron Acemoglu and Erik Brynjolfsson provide two economic perspectives worth considering. Acemoglu is an Institute Professor at MIT and took the 2024 Nobel Memorial Prize in Economic Sciences for work on institutions and prosperity. His method is transparent task-level accounting by showing what his model includes and excludes. Brynjolfsson directs Stanford’s Digital Economy Lab and originated the productivity J-curve framework both sides of this dispute now argue within. He co-founded Workhelix, which sells plans for measuring and capturing generative-AI value; a commercial interest worth keeping in mind while reading his conclusions [36].

Acemoglu’s task-based model projects total factor productivity (TFP) gains no greater than 0.66% over a decade, derived from the modest fraction of tasks AI affects multiplied by the average task-level cost saving [37]. In June 2026 he restated the estimate at roughly 0.55% of TFP over the decade, with about 5% of tasks profitably automated in the near term and a 1% to 1.5% lift to gross domestic product (GDP) [5]. His own gloss provides the calibration: “I don’t think we should belittle 0.5 percent in 10 years” [38]. The number refutes the trillion-dollar investment narratives, not AI.

His sharper objection sets current AI capabilities against the requirements for effectively working within the messy realities of the world. Acemoglu holds that most research showing AI productivity gains is overblown precisely because it concentrates on easy, well-defined tasks with clear context that are not representative of the economy as a whole [5]. He further argues that the large gains investors are pricing in would require something close to artificial general intelligence, which he does not believe is near [5]. Every affirmative study this paper leans on sits inside the well-defined task category he is dismissing.

The objection is correct from a national economics point, but it does not affect this paper’s position. Acemoglu is answering a question about the aggregate: what does AI add to national output when averaged across all tasks in the economy, most of which are messy and ill-defined. This paper is answering a narrower one: does deploying AI against a deliberately selected set of well-defined, quickly validated tasks inside one organization return more than the infrastructure costs. His finding that such tasks are a small share of the economy is entirely compatible with an organization choosing to work only inside that share. What his objection does rule out is any claim that AI deployments will lift organization-wide productivity by a figure resembling the vendor projections, which this paper does not claim.

Brynjolfsson reads the recent record the other way. He points to the Bureau of Labor Statistics (BLS) revision of 2025 job creation down to 181,000 from an initial 584,000, against fourth-quarter GDP growth of 3.7%, which his analysis translates to roughly 2.7% productivity growth for the year, nearly double the prior decade’s 1.4% average [6]. He flags that several more periods are needed to confirm a trend and that monetary or geopolitical shocks could offset it [6]. The Census Bureau working paper he co-authored supplies microfoundations from manufacturing data, finding causal J-curve-shaped returns in which short-run losses in productivity and profitability precede longer-run gains, with the losses concentrated among older establishments and mitigated by growth-oriented strategies and within-firm spillovers [39]. Kristina McElheran, the paper’s lead author, states the operative finding plainly: “AI isn’t plug-and-play” [40].

One tension inside Brynjolfsson’s own work deserves pointing out, because it complicates the affirmative case. His customer-support field experiment found the largest gains among novice and lower-skilled agents [16]. His 2025 payroll study found a 16% relative employment decline, as of September 2025, for workers aged 22 to 25 in the most AI-exposed occupations, concentrated in roles where AI automates instead of augmenting [41]. Both can be true: a tool that raises a novice’s output also reduces how many novices a firm needs. An organization deploying AI to compress the gap between junior and senior staff should expect that second effect and plan for it.

The disagreement between the two economists does not change the decision organizations are presented with, because both readings describe an average across all firms, including those that deploy AI badly. That average is a floor for the population, not a forecast for a specific deployment. The evidence that firms can land above it is specific: Brynjolfsson reports finding “a small cohort of power users” who automate end-to-end workstreams with agents and complete work in hours that previously took weeks [6], and the Census manufacturing data shows the J-curve trough shrinking for firms with growth-oriented strategies and internal spillovers [39]. The investment case needs only that the expected gain for a competent implementer exceeds the cost of the infrastructure, and the 15% average productivity gain with a roughly 30% novice gain in the customer-support field experiment supplies a defensible anchor for a task profile on-premises deployments can actually target [16].

Topol and the Clinical Trial Record

Eric Topol is a cardiologist, professor and executive vice president at Scripps Research, and among the most-cited researchers in medicine. He is one of the clearest voices outside of software arguing that the near-term return from AI is administrative: relieving clinicians of documentation, ordering, and authorization work, which he describes as “The gift of time from AI” [9]. The profile matches the verifiability thesis exactly. The AI model drafts a visit note from the conversation between the patient and clinician, which is then reviewed and validated within minutes by the same clinician immediately after the visit.

A 2025 study tested that claim and arrived at the same conclusion. Lukac and colleagues ran a three-group pragmatic randomized trial at a large California academic health system, assigning 238 outpatient physicians across 14 specialties 1:1:1 to Microsoft Dragon Ambient eXperience (DAX) Copilot, Nabla, or usual care, from November 4, 2024 to January 3, 2025 [42]. Nabla users cut time-in-note by 9.5% versus the control group, with a 95% confidence interval of −17.2% to −1.8%. DAX users showed no significant change against control, at −1.7% with a 95% confidence interval of −9.4% to +5.9% [42]. Burnout, task load, and work exhaustion measures improved when using the AI tools, but the authors state those secondary findings need confirmation in larger multicenter trials. Physicians reported occasional clinically significant inaccuracies on both platforms [42]. It is important to note the date when this study was conducted and the rate of improvement and utility in AI tools since then.

Three findings from that trial affect procurement. Product choice decided the primary outcome: two tools in the same category, on the same task, at the same institution, produced opposite results, which is the strongest available argument for evaluating candidate models and vendors on the organization’s own workload before standardizing. Adoption, not capability, set the ceiling on realized value: DAX was used in 33.5% of 24,696 eligible visits and Nabla in 29.5% of 23,653 [42]. Roughly seven in ten eligible patient encounters did not use any ambient AI scribing tool. The realized gain is the per-use gain multiplied by the adoption rate, which is a deployment-design variable the organization controls. The persistence of occasional inaccuracies makes human review a continuous operating cost, which is expected to decrease as these tools improve.

The magnitude frames the affirmative case. A 9.5% improvement in the affected activity, using the best resulting tool at the time, in a task with a clear verification loop, sits below the 15% average from the customer-support field experiment [16]. Taken together, the two studies bracket what a well-chosen task may return: a real single-digit to mid-teens improvement in the activity itself, which is consistent with Acemoglu’s claim.

What the Convergence Establishes

These observers disagree about artificial general intelligence, about pace, about architecture, and about the size of the macroeconomic effect, but agree that AI produces measurable value on tasks whose output validates quickly, current architectures stay useful for those tasks across any reasonable planning horizon, and the hardware outlives the architecture regardless. The constraint on capturing that value is deployment design.

Karpathy and Meijer reach that position through verifiability and program synthesis, Böckeler and Fowler through field practice, Torvalds through the governance a project adopts when review capacity is its scarce resource, Willison through the failure mode that bounds what an agent may safely touch, Hightower through infrastructure economics, Brynjolfsson through measurement, Topol and the clinical trial record through adoption and verification, and Acemoglu through macro economic modeling. Nadella, who profits the most from the opposite conclusion, names workflow redesign as the missing input. LeCun, the strongest architectural skeptic, holds that language models are useful today and inadequate as a path to human-level intelligence, which is a position this paper agrees with; Sutton reaches the same verdict from continual learning, with peer-reviewed evidence for the mechanism.

What would negatively impact AI’s ability to positively perform within the organization is specific enough to watch for. If the tasks given to AI lack the fast and cheap validation the verifiability thesis requires, then those tasks should not be worked on by AI. If review and testing capacity cannot absorb the higher generation volume, then throughput gains disappear into a growing backlog, which is a workflow failure. If the tools are made available, but not used at anything close to the rates the clinical trial measured, the modeled return will be off. The remedy is workflow redesign and staff training. All three are failures the organization can detect early and fix. Which tail of the distribution each deployment land in is set by task selection, workflow redesign, and the capability of the people utilizing the infrastructure. These are all variables inside the organization’s control.

References

  1. Y. LeCun, “A Path Towards Autonomous Machine Intelligence,” OpenReview, June 27, 2022. [Online]. Available: https://openreview.net/pdf?id=BZ5a1r-kVsf. Version 0.9.2, working paper. [Accessed: 24-Jul-2026]

    EXVX-1 Secondary source Back to text

  2. A. Heim, “Yann LeCun's AMI Labs raises $1.03B to build world models,” TechCrunch, Mar. 9, 2026. [Online]. Available: https://techcrunch.com/2026/03/09/yann-lecuns-ami-labs-raises-1-03-billion-to-build-world-models/. [Accessed: 24-Jul-2026]

    EXVX-2 Contextual source Back to text

  3. D. Amodei, “The Adolescence of Technology: Confronting and Overcoming the Risks of Powerful AI,” darioamodei.com, Jan. 2026. [Online]. Available: https://www.darioamodei.com/essay/the-adolescence-of-technology. [Accessed: 24-Jul-2026]

    EXVX-3 Contextual source Back to text

  4. D. Patel and R. S. Sutton, “Richard Sutton — Father of RL Thinks LLMs Are a Dead End,” Dwarkesh Podcast, Sept. 26, 2025. [Online]. Available: https://www.dwarkesh.com/p/richard-sutton. Host: D. Patel. Recorded at the Alberta Machine Intelligence Institute; transcript published with the episode. [Accessed: 04-Aug-2026]

    EXVX-4 Contextual source Back to text

  5. N. Lichtenberg, “Nobel Laureate Daron Acemoglu on the 'Brainless' AI Discourse, the Myth of Capitalism and the Gen Z Revolution Risk,” Fortune, June 21, 2026. [Online]. Available: https://fortune.com/2026/06/21/nobel-laureate-daron-acemoglu-ai-productivity-capitalism-democracy/. [Accessed: 24-Jul-2026]

    EXVX-5 Contextual source Back to text

  6. E. Brynjolfsson, “The AI Productivity Take-off Is Finally Visible,” Financial Times, Feb. 14, 2026. [Online]. Available: https://www.ft.com/content/4b51d0b4-bbfe-4f05-b50a-1d485d419dc5. Subscription required. Figures corroborated by Fortune/Yahoo Finance (Feb. 15, 2026) and American Enterprise Institute commentary (Feb. 23, 2026). Fortune report (Feb. 15, 2026): https://fortune.com/2026/02/15/ai-productivity-liftoff-doubling-2025-jobs-report-transition-harvest-phase-j-curve/. [Accessed: 16-Jun-2026]

    EXVX-6 Contextual source Back to text

  7. D. Patel and S. Nadella, “Satya Nadella — Microsoft's AGI Plan and Quantum Breakthrough,” Dwarkesh Podcast, Feb. 19, 2025. [Online]. Available: https://www.dwarkesh.com/p/satya-nadella. Host: D. Patel. Transcript published with the episode. [Accessed: 04-Aug-2026]

    EXVX-7 Contextual source Back to text

  8. A.-M. Stanciuc, “Nadella won't call it a bubble. He'll only say how it ends badly,” TNW, July 27, 2026. [Online]. Available: https://thenextweb.com/news/nadella-microsoft-compute-crunch-azure-customers-copilot-priority. Reports S. Nadella interviewed by F. Zakaria, GPS, CNN, Jul. 26, 2026. [Accessed: 04-Aug-2026]

    EXVX-8 Contextual source Back to text

  9. S. Markey, “Dr. Eric Topol on How AI Is Transforming Health and Medicine,” Clinical Center News, National Institutes of Health, Dec. 4, 2024. [Online]. Available: https://www.cc.nih.gov/news/2024/nov-dec/eric-topol-ai. Reports a Grand Rounds lecture delivered Sep. 2024. [Accessed: 04-Aug-2026]

    EXVX-9 Contextual source Back to text

  10. A. Karpathy, “Sequoia Ascent 2026 summary,” karpathy.bearblog.dev, Apr. 30, 2026. [Online]. Available: https://karpathy.bearblog.dev/sequoia-ascent-2026/. Author-reviewed machine-generated summary followed by an edited transcript of the fireside chat with S. Zhan; quotations are from the transcript. [Accessed: 24-Jul-2026]

    EXVX-10 Contextual source Back to text

  11. J. Novet, “Anthropic hires OpenAI co-founder Andrej Karpathy, former Tesla AI leader,” CNBC, May 19, 2026. [Online]. Available: https://www.cnbc.com/2026/05/19/anthropic-hires-openai-cofounder-andrej-karpathy-former-tesla-ai-lead.html. [Accessed: 24-Jul-2026]

    EXVX-11 Contextual source Back to text

  12. D. Patel and A. Karpathy, “Andrej Karpathy — AGI Is Still a Decade Away,” Dwarkesh Podcast, Oct. 17, 2025. [Online]. Available: https://www.dwarkesh.com/p/andrej-karpathy. Host: D. Patel. [Accessed: 24-Jul-2026]

    EXVX-12 Contextual source Back to text

  13. A. Karpathy, “Software 2.0,” Medium, Nov. 11, 2017. [Online]. Available: https://karpathy.medium.com/software-2-0-a64152b37c35. [Accessed: 24-Jul-2026]

    EXVX-13 Contextual source Back to text

  14. A. Karpathy, “Software Is Changing (Again),” Y Combinator AI Startup School, June 17, 2025. [Online]. Available: https://www.ycombinator.com/library/MW-andrej-karpathy-software-is-changing-again. Keynote, San Francisco, CA. [Accessed: 24-Jul-2026]

    EXVX-14 Contextual source Back to text

  15. A. Karpathy, “Verifiability,” karpathy.bearblog.dev, Nov. 17, 2025. [Online]. Available: https://karpathy.bearblog.dev/verifiability/. [Accessed: 24-Jul-2026]

    EXVX-15 Contextual source Back to text

  16. E. Brynjolfsson, D. Li, and L. R. Raymond, “Generative AI at Work,” The Quarterly Journal of Economics, May 2025. [Online]. Available: https://academic.oup.com/qje/article/140/2/889/7990658. Vol. 140, no. 2, pp. 889–942. doi:10.1093/qje/qjae044. [Accessed: 24-Jul-2026]

    EXVX-16 Primary source Back to text

  17. Meta AI, “Introducing the V-JEPA 2 world model and new benchmarks for physical reasoning,” Meta AI Blog, June 11, 2025. [Online]. Available: https://ai.meta.com/blog/v-jepa-2-world-model-benchmarks/. [Accessed: 24-Jul-2026]

    EXVX-17 Secondary source Back to text

  18. S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, and R. S. Sutton, “Loss of plasticity in deep continual learning,” Nature, Aug. 2024. [Online]. Available: https://www.nature.com/articles/s41586-024-07711-7. Vol. 632, pp. 768–774. doi:10.1038/s41586-024-07711-7. [Accessed: 04-Aug-2026]

    EXVX-18 Primary source Back to text

  19. J. Kaplan et al., “Scaling Laws for Neural Language Models,” arXiv, Jan. 2020. [Online]. Available: https://arxiv.org/abs/2001.08361. arXiv:2001.08361 [cs.LG]. [Accessed: 10-May-2026]

    EXVX-19 Secondary source Back to text

  20. D. Amodei, “Machines of Loving Grace: How AI Could Transform the World for the Better,” darioamodei.com, Oct. 2024. [Online]. Available: https://www.darioamodei.com/essay/machines-of-loving-grace. [Accessed: 24-Jul-2026]

    EXVX-20 Contextual source Back to text

  21. E. Meijer, “Unleashing the Power of End-User Programmable AI: Creating an AI-First Program-Synthesis Framework,” ACM Queue, 2025. [Online]. Available: https://queue.acm.org/detail.cfm?id=3746223. Vol. 23, no. 3. doi:10.1145/3746223. [Accessed: 24-Jul-2026]

    EXVX-21 Primary source Back to text

  22. G. Kim, “Dr. Erik Meijer on Coding as a Solved Problem,” LinkedIn, Mar. 12, 2026. [Online]. Available: https://www.linkedin.com/posts/realgenekim_enterprise-ai-summit-april-9-10-2026-activity-7437762134593196033-PjGL. Secondhand account of an unrecorded private conversation, published in a post promoting a conference at which the subject was a speaker. [Accessed: 24-Jul-2026]

    EXVX-22 Contextual source Back to text

  23. M. Fowler, “About Martin Fowler,” martinfowler.com, 2026. [Online]. Available: https://martinfowler.com/aboutMe.html. Role and financial disclosures. [Accessed: 24-Jul-2026]

    EXVX-23 Contextual source Back to text

  24. M. Fowler, “LLMs Bring a New Nature of Abstraction,” martinfowler.com, June 24, 2025. [Online]. Available: https://martinfowler.com/articles/2025-nature-abstraction.html. [Accessed: 24-Jul-2026]

    EXVX-24 Contextual source Back to text

  25. B. Böckeler, “Exploring Generative AI,” martinfowler.com. [Online]. Available: https://martinfowler.com/articles/exploring-gen-ai.html. Series index, 2023–2026; ed. M. Fowler. [Accessed: 24-Jul-2026]

    EXVX-25 Contextual source Back to text

  26. B. Böckeler, “Viability of Local Models for Coding,” Exploring Generative AI, martinfowler.com, July 7, 2026. [Online]. Available: https://martinfowler.com/articles/exploring-gen-ai/local-models-for-coding-factors.html. [Accessed: 24-Jul-2026]

    EXVX-26 Contextual source Back to text

  27. B. Böckeler, “Experiences with Local Models for Coding,” Exploring Generative AI, martinfowler.com, July 8, 2026. [Online]. Available: https://martinfowler.com/articles/exploring-gen-ai/local-models-for-coding-experiences.html. [Accessed: 24-Jul-2026]

    EXVX-27 Contextual source Back to text

  28. L. Proven, “Linus Torvalds: 90% of AI marketing is hype so 'I ignore it',” The Register, Oct. 29, 2024. [Online]. Available: https://www.theregister.com/2024/10/29/linus_torvalds_ai_hype/. Reports an interview recorded by TFiR at Open Source Summit Europe, Vienna, Oct. 2024. [Accessed: 04-Aug-2026]

    EXVX-28 Contextual source Back to text

  29. L. Torvalds, “Re: Linking Patchwork with Sashiko?” linux-media mailing list, July 14, 2026. [Online]. Available: https://lore.kernel.org/linux-media/CAHk-=wi4zC+Ze8e+p3tMv8TtG_80KzsZ1syL9anBtmEh5Z40vg@mail.gmail.com/. [Accessed: 04-Aug-2026]

    EXVX-29 Contextual source Back to text

  30. The kernel development community, “AI Coding Assistants,” The Linux Kernel Documentation. [Online]. Available: https://docs.kernel.org/process/coding-assistants.html. Documentation/process/coding-assistants.rst, merged Dec. 23, 2025; retrieved from the 7.2.0-rc6 documentation build. [Accessed: 04-Aug-2026]

    EXVX-30 Primary source Back to text

  31. The kernel development community, “Kernel Guidelines for Tool-Generated Content,” The Linux Kernel Documentation. [Online]. Available: https://docs.kernel.org/process/generated-content.html. Documentation/process/generated-content.rst. [Accessed: 04-Aug-2026]

    EXVX-31 Primary source Back to text

  32. Cloud Native Computing Foundation, “The 2024 Trends on Cloud Computing by Kelsey Hightower and Alex Saroyan,” CNCF, Mar. 28, 2024. [Online]. Available: https://www.cncf.io/blog/2024/03/28/the-2024-trends-on-cloud-computing-by-kelsey-hightower-and-alex-saroyan/. Member post originally published by Netris. [Accessed: 24-Jul-2026]

    EXVX-32 Contextual source Back to text

  33. G. Orosz and K. Hightower, “Kubernetes and Retiring at the Top with Kelsey Hightower,” The Pragmatic Engineer Podcast, June 3, 2026. [Online]. Available: https://newsletter.pragmaticengineer.com/p/kubernetes-and-retiring-at-the-top. Host: G. Orosz. [Accessed: 24-Jul-2026]

    EXVX-33 Contextual source Back to text

  34. S. Willison, “The lethal trifecta for AI agents: private data, untrusted content, and external communication,” simonwillison.net, June 16, 2025. [Online]. Available: https://simonwillison.net/2025/Jun/16/the-lethal-trifecta/. [Accessed: 04-Aug-2026]

    EXVX-34 Contextual source Back to text

  35. L. Beurer-Kellner et al., “Design Patterns for Securing LLM Agents against Prompt Injections,” arXiv, June 2025. [Online]. Available: https://arxiv.org/abs/2506.08837. arXiv:2506.08837 [cs.CR]. [Accessed: 04-Aug-2026]

    EXVX-35 Secondary source Back to text

  36. Stanford Digital Economy Lab, “Erik Brynjolfsson,” Stanford University, 2026. [Online]. Available: https://digitaleconomy.stanford.edu/person/erik-brynjolfsson/. Faculty biography and affiliations, including co-founder of Workhelix, Inc. [Accessed: 24-Jul-2026]

    EXVX-36 Primary source Back to text

  37. D. Acemoglu, “The Simple Macroeconomics of AI,” Economic Policy, Jan. 2025. [Online]. Available: https://www.nber.org/papers/w32487. Vol. 40, no. 121, pp. 13–58. Preprint: NBER Working Paper No. 32487, May 2024, doi:10.3386/w32487. [Accessed: 24-Jul-2026]

    EXVX-37 Primary source Back to text

  38. MIT News Office, “Daron Acemoglu: What Do We Know About the Economics of AI?” Massachusetts Institute of Technology, Dec. 6, 2024. [Online]. Available: https://economics.mit.edu/news/daron-acemoglu-what-do-we-know-about-economics-ai. [Accessed: 24-Jul-2026]

    EXVX-38 Contextual source Back to text

  39. K. McElheran, M.-J. Yang, Z. Kroff, and E. Brynjolfsson, “The Rise of Industrial AI in America: Microfoundations of the Productivity J-Curve(s),” Center for Economic Studies, U.S. Census Bureau, Working Paper No. CES-25-27, Apr. 2025. [Online]. Available: https://www.census.gov/library/working-papers/2025/adrm/CES-WP-25-27.html. [Accessed: 24-Jul-2026]

    EXVX-39 Secondary source Back to text

  40. MIT Sloan School of Management, “The 'Productivity Paradox' of AI Adoption in Manufacturing Firms,” Ideas Made to Matter, Jan. 29, 2026. [Online]. Available: https://mitsloan.mit.edu/ideas-made-to-matter/productivity-paradox-ai-adoption-manufacturing-firms. [Accessed: 24-Jul-2026]

    EXVX-40 Contextual source Back to text

  41. E. Brynjolfsson, B. Chandar, and R. Chen, “Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence,” Stanford Institute for Economic Policy Research / Stanford Digital Economy Lab Working Paper, Aug. 2025. [Online]. Available: https://digitaleconomy.stanford.edu/publications/canaries-in-the-coal-mine/. [Accessed: 24-Jul-2026]

    EXVX-41 Secondary source Back to text

  42. P. J. Lukac et al., “Ambient AI Scribes in Clinical Practice: A Randomized Trial,” NEJM AI, Dec. 2025. [Online]. Available: https://ai.nejm.org/doi/abs/10.1056/AIoa2501000. Vol. 2, no. 12. doi:10.1056/AIoa2501000. ClinicalTrials.gov NCT06792890. [Accessed: 04-Aug-2026]

    EXVX-42 Primary source Back to text

Contents