Section6
The AI Productivity Paradox: Why Most Deployments Underdeliver
The productivity paradox is a deployment problem, not a capability problem. Hundreds of billions of dollars have gone into AI infrastructure, nearly every knowledge worker can reach a frontier model, and the macroeconomic data shows almost none of it. That gap between investment and measured output is the honest problem any case for on-premises deployment has to engage before it recommends a single GPU. The evidence is solid; the interpretation is contested. What the interpretation turns on is specific and within an organization’s control. Firms that bolted AI onto unchanged workflows got flat or negative results. Firms that redesigned workflows around tasks AI handles well got measurable gains. The use cases planned here, generation-heavy work with fast validation and redesignable downstream steps, sit in the second category. Executives who read the paradox literature as proof that waiting was prudent have read it correctly about the past three years and incorrectly about the design space ahead.
The Macro Evidence
The most credible primary source on the paradox is a National Bureau of Economic Research working paper from February 2026, revised that March, led by Ivan Yotzov of the Bank of England with twelve co-authors, among them Nicholas Bloom of Stanford and Steven Davis of the Hoover Institution [1]. Its authority rests on its data, not its bylines. Roughly 6,000 senior business executives across the United States, United Kingdom, Germany, and Australia answered identical questions between November 2025 and January 2026 through panels their national central banks administer. That design yields the first representative international firm-level dataset of its kind, and its central finding is not a fringe claim. It is the baseline any optimistic argument has to beat.
The picture is wide adoption with shallow use and little measured effect. About 69% of firms actively use AI, and more than two-thirds of executives (mostly CEOs, CFOs, and senior finance managers) use it in a typical week, but average executive usage runs 1.5 hours a week [1]. On impact the paper is blunt: more than 90% of executives report no effect of AI on their own firm’s employment over the past three years, and 89% report no effect on labor productivity [1]. The same executives expect more ahead, forecasting that AI will raise their firms’ productivity 1.4% and output 0.8% over the next three years while cutting employment 0.7%, with US executives expecting a larger 2.25% productivity gain [1]. Asked the same question, employees expect employment to rise 0.5%, a gap between what bosses and workers anticipate that is itself worth watching [1].
Torsten Slok, chief economist at Apollo Global Management, gave the pessimistic read its sharpest form in a February 2026 note: “AI is everywhere except in the incoming macroeconomic data” [2]. The line echoes Robert Solow’s 1987 quip that the computer age was visible everywhere except in the productivity statistics. Slok points out that AI is absent from employment, productivity, and inflation data, and that for the S&P 493, the S&P 500 stripped of its seven largest technology firms, it is absent from profit margins and earnings expectations [2]. He stays agnostic on the J-curve. The flat data may be the investment trough before a surge, or it may not.
Erik Brynjolfsson reads the same period the opposite way. In a February 2026 Financial Times op-ed, the director of Stanford’s Digital Economy Lab argued that the take-off is now visible. The Bureau of Labor Statistics (BLS) revised 2025 job gains down to 181,000 from an initial 584,000 while fourth-quarter gross domestic product (GDP) held at 3.7%, a decoupling of output from labor input that his own analysis translates to roughly 2.7% US productivity growth for 2025, nearly double the 1.4% average of the prior decade [3]. He calls this the move from AI’s investment phase to its harvest phase, and cautions that several more quarters are needed to confirm a trend [3]. The revisions are preliminary and Brynjolfsson himself flags the need for confirmation, so the harvest reading is a live hypothesis, not a settled result.
Whether Slok or Brynjolfsson is right depends on data not yet in. The disagreement is the honest position to occupy: two of the field’s most credentialed observers read the same numbers in opposite directions, and the sound response is to build a deployment strategy that holds up under either.
Daron Acemoglu, the 2024 economics Nobel laureate, sits between them. His 2024 Economic Policy paper estimates that AI will add no more than 0.66% to US total factor productivity (TFP) over ten years, about 0.064% a year, and as little as 0.53% once harder-to-automate tasks are counted; the matching GDP effect is roughly 0.93% to 1.16% over the decade [4]. He builds the number from three measured inputs: about 20% of US labor tasks are exposed to AI, only about 23% of computer vision tasks can be automated profitably within ten years, and the average task-level cost saving is about 27% [5]. His own gloss is “I don’t think we should belittle 0.5 percent in 10 years” [5]. The figure is not a refutation of AI. It is a refutation of the trillion-dollar productivity narratives.
The Micro Evidence: The METR Study and Its Update
The most rigorous controlled experiment in this area is METR’s randomized controlled trial (RCT), published in July 2025 [6]. Sixteen experienced developers, each with about five years in the specific mature open-source codebases they worked in, completed 246 tasks randomly assigned to permit or forbid early-2025 AI tools, mainly Cursor Pro with Claude 3.5 and 3.7 Sonnet. Before starting, they predicted AI would make them 24% faster. After finishing, they estimated it had made them 20% faster. It had made them 19% slower, with a confidence interval running from 2% to 39% longer [6]. The roughly 39-point gap between the developers’ estimate and the measured result is the more durable finding, because it holds regardless of where the productivity number itself lands. Experienced developers cannot reliably judge AI’s effect on their own work.
The 19% figure travels widely and needs two qualifications. The first is scope. The result is about experienced developers, complex tasks, mature codebases, and early-2025 tools. METR did not claim it generalizes to junior developers, well-scoped tasks, greenfield projects, or current models, and the contribution is the demonstration that the effect is task-conditional and self-assessment unreliable, not that AI slows everyone down.
The second is what METR found when it tried to repeat the study. A follow-up published in February 2026 reports that the design broke down [7]. Starting in August 2025, METR ran a second experiment with 57 developers across 143 repositories and more than 800 tasks, paying $50 an hour rather than the original $150. Wider AI adoption wrecked the comparison. Developers increasingly refused to work without AI at all, and when surveyed, 30% to 50% said they were declining to submit certain tasks because they did not want to do them unassisted. Both effects strip out exactly the developers and tasks where AI helps most, so METR judges its central estimate an unreliable signal and only very weak evidence [7]. The raw point estimates do lean positive: an 18% speedup among the returning developers (interval −38% to +9%) and a 4% speedup among the newly recruited ones (interval −15% to +9%), both straddling zero [7]. METR’s stated read is qualitative, that developers are probably faster with AI in early 2026 than in early 2025 but that the data cannot size the gain, and the group is redesigning the experiment [7]. A separate METR survey in May 2026 found a median self-reported 1.4-to-2x gain in the value of work among 349 technical workers, a number METR itself treats cautiously given a 2% response rate and the same self-assessment problem the RCT exposed [8].
Telemetry from Faros AI puts a mechanism under the macro numbers, with a caveat about its source. Faros sells engineering-analytics software, and its 2025 report sits behind a lead-capture form, so this is vendor first-party measurement, not independent research, and its commercial interest runs toward dramatizing a problem it sells the fix for. With that discount applied, the dataset is large and the direction is consistent with the rest of the evidence. Across about 10,000 developers on 1,255 enterprise teams, teams with heavy AI adoption completed 21% more tasks and merged 98% more pull requests (PRs) than baseline, yet review time rose 91% and DevOps Research and Assessment (DORA) delivery metrics, deployment frequency and lead time for changes, stayed flat [9]. More code entered the pipeline, none of it reached production faster, and the report describes added quality pressure from the larger volume. The bottleneck moved from writing code to reviewing it.
This is Amdahl’s Law operating on a delivery pipeline [10]. Speeding up one stage lifts the whole system only in proportion to that stage’s share of total effort, and only if the stages downstream can absorb the new volume. If the AI-accelerated step is, say, a third of total developer effort, even doubling its speed raises end-to-end throughput by roughly 20%, and that ceiling holds only when review and deployment keep pace. The Faros data shows they do not. Generation capacity outran review capacity, the system sped up where AI touched it and slowed where humans still gated the work, and the second effect swallowed the first.
Why Deployments Underdeliver
The failure modes follow a pattern where every single one is a deployment choice rather than a statement about what AI can do.
The largest is the absence of workflow redesign. The Faros pattern is the clean case: teams inserted AI into the coding step and changed nothing downstream, so review demand rose 91% against constant review capacity and total throughput barely moved [9]. Amdahl predicts this without any reference to AI [10]. Generation that outruns review does not speed delivery; it relocates the queue. The remedy is structural, redesigning review, testing, and deployment to absorb the higher volume, and it is a design decision rather than a property of the model.
Close behind is the wrong tool aimed at the wrong task. AI is strong at pattern-based generation, drafting, summarizing, completing well-understood code, answering questions against supplied documents, and weak at novel reasoning, precise recall outside its training distribution, and reliable multi-step tool use at the edge of current agent capability. Point it at the second category, measure aggregate output, and the result is the flat-to-negative reading the 6,000-firm survey captures [1].
Two human factors finish the list, working in opposite directions. Most workers get the tool without training in how to prompt it, supply context, or check its output, so they treat it as a search engine and get search-engine results: uneven, sometimes confidently wrong, not worth the disruption. At the other extreme, over-reliant users ship unchecked output that someone has to repair downstream, converting an apparent speed gain into negative net work. Skeptics and over-trusters both produce no gain, by opposite routes.
The Affirmative Case
The strongest evidence for real gains comes from the mirror image of the failure analysis: generation-heavy tasks, fast output validation, and workflows built around the tool.
The canonical positive result is Brynjolfsson, Li, and Raymond’s study of 5,172 customer-support agents at a Fortune 500 software firm, published in the Quarterly Journal of Economics in 2025 [11]. A generative-AI assistant offering real-time chat suggestions raised issues resolved per hour by 15% on average, with about a 30% gain for novice and lower-skilled agents and only small speed gains, alongside small quality declines, for the most experienced [11]. Customer sentiment improved, requests for managerial escalation fell, and retention rose. The structure matters more than the headline. Support work is generation-heavy and structured, the customer validates the answer immediately, and the gain concentrates among novices because the model hands them the patterns experienced agents already carry. That is the profile where AI works.
The same logic shows up far from the enterprise. A Stanford-affiliated study by Michael Blank, Gregor Schubert, and Miao Ben Zhang, using internet-browsing data from more than 200,000 US households between 2021 and 2024, finds that ChatGPT adoption makes households markedly more efficient at productive digital chores such as job hunting, travel planning, and online shopping [12]. The headline 76% to 176% efficiency range is model-implied rather than directly measured: households spend the same time on productive tasks and more on leisure after adopting, and the efficiency gain is what a standard time-allocation model infers from that shift [12]. The scope is consumer, not enterprise. Users bank the saved time as leisure, not as additional output or skill-building [12]. At home, AI generates free time more than measured productivity. Whether that carries into an organization depends entirely on whether the organization captures the recovered hours into work.
These results do not conflict. The customer-support study, the household study, and the leaning-positive METR follow-up all describe AI delivering on tasks that fit its strengths. The original METR slowdown, the 6,000-firm no-impact finding, and the Faros review bottleneck all describe AI failing where tasks do not fit or where downstream stages cannot absorb the new volume. Same finding from different angles: the effect is task-conditional, and the task and the surrounding workflow decide the sign.
Implications for AI Deployments Within Technical Teams
The use cases technical teams would have for the AI infrastructure proposed in this paper sit in the favorable half. Code generation against well-defined internal patterns is structured generation. Retrieval-augmented question answering over internal documentation is the close enterprise analog to the customer-support deployment, generation-heavy, fast to validate, and with the same upward compression of less-experienced staff. Drafting documents from structured templates is pattern-based generation. Log analysis is information retrieval over structured text, the household study’s digital-chore profile moved into operations.
These tasks share four important properties. The critical step is generation rather than judgment, which is where the model is strong. Output validates fast and clearly, so errors surface before they propagate rather than after. The workflow around the AI step can be rebuilt to absorb higher volume, which is what keeps the Amdahl bottleneck from minimizing the gain. And the user population includes novices, who gain the most because the model distributes expertise they have not yet built. None of these is the METR profile of experienced engineers doing complex work in mature codebases, and nothing here claims AI will accelerate that work uniformly.
The paradox and the optimism are both real. What separates them is not the AI model. It is the set of deployment decisions made before the model goes into production, and those decisions are the organization’s to make.
References
-
I. Yotzov et al., “Firm Data on AI,” NBER Working Paper No. 34836, National Bureau of Economic Research, Feb. 2026. [Online]. Available: https://www.nber.org/papers/w34836. Revised Mar. 2026. doi:10.3386/w34836. [Accessed: 16-Jun-2026]
-
T. Slok, “Waiting for the AI J-Curve,” Apollo Academy, Apollo Global Management, Feb. 14, 2026. [Online]. Available: https://www.apolloacademy.com/waiting-for-the-ai-j-curve/. [Accessed: 16-Jun-2026]
-
E. Brynjolfsson, “The AI Productivity Take-off Is Finally Visible,” Financial Times, Feb. 14, 2026. [Online]. Available: https://www.ft.com/content/4b51d0b4-bbfe-4f05-b50a-1d485d419dc5. Subscription required. Figures corroborated by Fortune/Yahoo Finance (Feb. 15, 2026) and American Enterprise Institute commentary (Feb. 23, 2026). Yahoo Finance report (Feb. 15, 2026): https://finance.yahoo.com/news/one-stanford-original-ai-gurus-205316027.html. [Accessed: 16-Jun-2026]
-
D. Acemoglu, “The Simple Macroeconomics of AI,” Economic Policy, Jan. 2025. [Online]. Available: https://www.nber.org/papers/w32487. Vol. 40, no. 121, pp. 13–58. Preprint: NBER Working Paper No. 32487, May 2024, doi:10.3386/w32487. [Accessed: 16-Jun-2026]
-
MIT News Office, “Daron Acemoglu: What Do We Know About the Economics of AI?” Massachusetts Institute of Technology, Dec. 6, 2024. [Online]. Available: https://economics.mit.edu/news/daron-acemoglu-what-do-we-know-about-economics-ai. [Accessed: 16-Jun-2026]
-
J. Becker, N. Rush, E. Barnes, and D. Rein, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity,” arXiv, July 12, 2025. [Online]. Available: https://arxiv.org/abs/2507.09089. arXiv:2507.09089v2, revised Jul. 25, 2025. [Accessed: 16-Jun-2026]
-
J. Becker, N. Rush, T. Cunningham, D. Rein, and K. Mahamud, “We are Changing our Developer Productivity Experiment Design,” METR, Feb. 24, 2026. [Online]. Available: https://metr.org/blog/2026-02-24-uplift-update/. [Accessed: 16-Jun-2026]
-
J. Becker, “Measuring the Self-Reported Impact of Early-2026 AI on Technical Worker Productivity,” METR, May 11, 2026. [Online]. Available: https://metr.org/blog/2026-05-11-ai-usage-survey/. [Accessed: 16-Jun-2026]
-
Faros AI, “The AI Engineering Report 2025: The AI Productivity Paradox,” Faros AI, July 23, 2025. [Online]. Available: https://www.faros.ai/ai-productivity-paradox. Vendor first-party telemetry; full report gated. [Accessed: 16-Jun-2026]
-
G. M. Amdahl, “Validity of the Single Processor Approach to Achieving Large Scale Computing Capabilities,” Proc. AFIPS Spring Joint Computer Conf., 1967. [Online]. Available: https://dl.acm.org/doi/10.1145/1465482.1465560. Vol. 30, pp. 483–485. doi:10.1145/1465482.1465560. [Accessed: 16-Jun-2026]
-
E. Brynjolfsson, D. Li, and L. R. Raymond, “Generative AI at Work,” The Quarterly Journal of Economics, May 2025. [Online]. Available: https://academic.oup.com/qje/article/140/2/889/7990658. Vol. 140, no. 2, pp. 889–942. doi:10.1093/qje/qjae044. [Accessed: 24-Jul-2026]
-
M. Blank, G. Schubert, and M. B. Zhang, “The Household Impact of Generative AI: Evidence from Internet Browsing Behavior,” arXiv, Feb. 27, 2026. [Online]. Available: https://arxiv.org/abs/2603.03144. arXiv:2603.03144. Stanford Institute for Economic Policy Research Working Paper. [Accessed: 16-Jun-2026]