Section20

Effective AI Use: Practical Guidance for Teams

The documented productivity gains are real, conditional, and learnable. They do not arrive with the hardware. Teams capture them by doing four things deliberately: choosing the right interface for each task, prompting with structure instead of hope, validating outputs before any consequential action, and building habits of sharing what works. None of this requires unusual talent. All of it requires practice, discipline, and management attention, which most organizations do not give enough time to. McKinsey finds that fundamentally redesigning workflows correlates more strongly with real business impact than almost any other practice it tests, and that the organizations capturing the most value from AI are roughly three times as likely as their peers to have done it [1]. Yet in McKinsey’s prior survey, only 21% of organizations using generative AI had redesigned even some workflows, even though workflow redesign showed the largest effect on earnings before interest and taxes (EBIT) of the 25 organizational attributes tested [2]. The distance between those two findings is the opportunity this section addresses, and the practices below are how a team closes it.

Prompting: The Skill That Sets the Ceiling

A prompt is an instruction set written for a system that has no shared context with the writer. Most disappointing AI outputs disappoint because the prompt underspecified the task, not because the model could not do the work. Anthropic’s and OpenAI’s developer documentation converge on the same hierarchy of techniques, ordered by how broadly each one helps: be clear and direct, provide examples, give the model room to reason, structure the input with delimiters, and assign a role only when role-specific behavior is needed [3], [4].

Clarity is the largest lever and the most underused. Anthropic’s guidance offers a test that needs no technical background: before submitting a prompt, ask whether a new colleague with zero context on your work could follow it [3]. Most failed prompts fail that test. “Summarize this document” does not say for whom, at what length, with what emphasis, or in what format. “Summarize this engineering postmortem for an executive audience in 200 words, leading with customer impact and ending with the three highest-priority remediation actions” specifies all four and produces a far better first draft.

Few-shot prompting, providing two to five worked examples of the input-output pattern you want, is the second-largest lever. Anthropic’s documentation states that examples can dramatically improve accuracy and consistency [3]; OpenAI frames the same move as providing reference text [4]. The mechanism is identical: examples disambiguate intent in ways prose instructions cannot. Asking for “a professional email tone” leaves the model to guess what professional means in your organization. Providing two emails you sent last week as examples and asking for the same voice clarifies the intent.

Structure matters more than most users expect. Both Anthropic and OpenAI recommend XML-style tags as delimiters to separate instructions from content [3], [4], wrapping the document in <document> tags, prior examples in <examples> tags, and the actual instruction outside both. This removes a common failure where the model treats part of the source as instruction, or part of the instruction as source. For any prompt longer than a paragraph, that structure earns its keep.

Iteration is a planned step, not a fallback. Anthropic’s course materials treat prompt refinement as a development phase: a first draft establishes task structure, and later passes tune context, examples, and format against real outputs [5]. The practical consequence is that the right comparison is not “the AI’s output versus my expectations” but “the output after one prompt versus the output after three.” Users who quit after the first attempt are measuring the wrong thing.

The common mistakes are predictable. Underspecified context tops both vendors’ lists. Reaching for elaborate role-play or formatting before establishing clarity and examples is Anthropic’s named anti-pattern. OpenAI flags the failure to pin model versions in production, which lets prompts drift as the model updates beneath them.

One distinction bears on the next subsection. Anthropic engineers, in a September 2025 post, separate prompt engineering from context engineering [6]. Prompt engineering writes effective instructions for a single turn of interaction. Context engineering curates everything the model sees across an agentic loop: system prompt, retrieved documents, tool outputs, and conversation history. For chat-interface work, prompt engineering is the right frame and the skill above is the right skill. For agentic and pipeline work, the operative question shifts from how the instruction is worded to what configuration of context produces the behavior you want.

Choosing the Interface

Three interaction modes exist, and picking the wrong one for the task is among the most common adoption errors. The chat interface suits exploratory work, iterative drafting, and quick question-and-answer where the user reviews each output immediately. The agent suits multi-step tasks where intermediate results inform later steps and the work needs tool access such as search, code execution, file reading, or application programming interface (API) calls. The pipeline suits high-volume, structured, repetitive work where the process is well understood and outputs can be validated by rule rather than by reading.

One question resolves most cases: how will you know the output is right? If you will read and judge each result before acting, chat is fine. If the work involves several dependent steps you would rather not supervise one at a time, an agent applies, but only once you have defined what it may do without asking. If you will run the same operation hundreds or thousands of times and check outputs programmatically, build a pipeline.

This is a live decision being made unevenly inside most organizations, not a theoretical one. McKinsey’s November 2025 survey found 23% of organizations scaling an agentic AI system somewhere in the enterprise and another 39% experimenting with agents, with most of those who are scaling confined to one or two functions [1]. Clear team-level guidance on when each mode applies prevents the failure where workers default to whichever interface they encountered first and force every task into it.

The Limitations, Stated Honestly

AI tools have specific, documented failure modes, and adoption succeeds in proportion to how clearly the workforce understands them. Four matter most day to day: hallucination, context degradation, agentic tool unreliability, and latency.

Hallucination, the generation of plausible but false content, carries the highest cost per incident and the lowest visibility. Models fabricate citations to papers that do not exist, invent API signatures that look correct, and assert thin or contradicted claims with full confidence. The Open Worldwide Application Security Project (OWASP) Top 10 for Large Language Model Applications ranks prompt injection as the top risk for 2025 and treats hallucination under its Misinformation category, with the explicit position that no method fully prevents these failures; mitigation, not elimination, is the realistic goal [7]. Hallucination rates vary widely by model and task, and even the strongest models produce factual errors at non-trivial rates on recall-style questions. The implication is firm: any factual claim, citation, or specification an AI produces must be checked against a primary source before it enters consequential use.

Context degradation is the second failure mode. Long-context research that Anthropic’s engineering guidance draws on documents what its authors call context rot: as the number of tokens in the window grows, the model’s ability to accurately recall any specific item in it declines, even within technically supported window sizes [8], [6]. A model with a one-million-token window will not reliably surface a fact buried at position 800,000 of an 850,000-token document. Pasting an entire 200-page document and asking about page 47 is therefore a worse strategy than extracting the relevant 1,500 words and submitting those. The operating principle Anthropic states is to find the smallest set of high-signal tokens that produces the result [6]. Retrieval-augmented generation (RAG) addresses this at the system level; at the user level, selective extraction beats wholesale paste.

Agentic tool reliability is the third, and it drives the validation guidance below. OWASP added Excessive Agency as a 2025 category to cover agents that hold tool access and can take irreversible actions, breaking the risk into three causes: excessive functionality, where the agent reaches tools beyond its task; excessive permissions, where those tools run with broader privileges than needed; and excessive autonomy, where high-impact actions proceed without a human in the loop [7]. Agents fail on complex tool chains at rates that require human review before any consequential action. That is the current state of the technology, not a passing limitation, and the architecture has to account for it.

Latency is the fourth and the least discussed. On-premises inference of large models is not instantaneous, and the figure depends heavily on the deployment. A 70-billion-parameter model running an agentic loop with several tool calls and intermediate reasoning can take tens of seconds to return a final result. Workflows with hard time budgets, such as interactive customer-facing applications or real-time monitoring, must plan for this. Most internal knowledge work has no hard time budget, but where the constraint applies, it is real.

Validation Workflows

Treat AI output as a first draft from a fast but fallible collaborator. The framing sets the expectation correctly: the output is worth reviewing, often worth keeping, and never worth shipping unverified. Defining when model outputs need human validation is one of the practices that most separates organizations capturing real value from AI from those that are not [1]. The discipline below is not bureaucratic overhead; it is what the successful adopters actually do.

The validation discipline varies by output type. AI-generated code passes the same automated tests as human code, gated before review; any test that would catch a human error catches the equivalent machine error, and the team that disables tests because “the AI wrote it” is choosing a worse outcome. For documents, every factual claim, citation, and specific number gets verified against the cited source, and AI-generated citations in particular fail often enough that the only safe posture treats them as suspect by default. For agentic actions with real-world consequences, such as sending email, executing database changes, deploying code, deleting files, or making external API calls with persistent side effects, an explicit human approval gate must precede execution. OWASP’s Excessive Agency guidance names exactly this class of action as one requiring architectural human-in-the-loop controls [7].

The National Institute of Standards and Technology (NIST) AI Risk Management Framework supplies the governance language for this. Its four functions, Govern, Map, Measure, and Manage, translate directly to team practice [9]. Govern sets the policies behind AI use: which tools are approved for which purposes, who is accountable for those decisions, and how risk management aligns with organizational priorities. NIST treats Govern as a cross-cutting function that informs the other three rather than a one-time step. Map establishes the context for each team’s AI use cases and identifies the associated risks. Measure uses quantitative and qualitative methods to assess and monitor those risks against trustworthy-AI characteristics such as safety, security, fairness, and privacy. Manage allocates resources to the mapped and measured risks, prioritizes them, and treats each one by mitigating, transferring, avoiding, or accepting it, with response and recovery plans for incidents. The framework is voluntary, but it is among the most widely referenced frameworks for AI risk governance, and it can be implemented at any team size.

One technique pays for itself faster than most: ask the model to evaluate its own output against stated criteria before you accept it. Anthropic’s materials document self-evaluation prompting, adding an instruction along the lines of “before your final answer, list the factual claims in your response, rate your confidence in each, and flag any that should be independently verified” [5]. The model is often better at catching its own uncertain claims when asked directly than when generating freely. It does not replace human verification on consequential work, but it is a useful intermediate gate.

Building Adoption That Works

McKinsey’s adoption research identifies twelve practices correlated with measurable AI value, six of which work on workforce behavior rather than structure or tooling: regular internal communication about the value created by AI solutions, senior leaders role-modeling its use, role-based capability training, building employee trust, a compelling change story, and incentives that reinforce adoption [2]. Organizations that treat AI as a procurement event, buying the tools, distributing them, and declaring success, capture less than those that treat it as a behavior change. The reason is plain. Workers who never see leadership using AI, never hear specific value stories, and never get help with tasks they actually do will not build the skill on their own.

Start with each team’s highest-value, lowest-risk use case and show concrete impact before expanding. The pattern produces evidence in the form that wins skeptics: a specific task that took four hours and now takes one, with the workflow documented and the output validated. McKinsey’s data shows that high performing organizations capturing measurable AI value with the adoption initiative championed by their leaders are three times more likely than their peers to strongly agree that senior leaders demonstrate ownership of and commitment to AI initiatives along with saying their organization intends to use AI to bring about transformative change to their businesses [1]. That ownership shows up as leaders using the tools visibly on their own work and talking specifically about what worked, not as a slogan.

Build internal case studies in one consistent format: before, the task took X hours and looked like this; after, it takes Y hours and looks like this; here is the prompt and the workflow. Share them widely. This is the lowest-cost, highest-signal communication available, because it shows the gain in a form workers can copy rather than describing it in the abstract.

Address replacement fears directly, because the current evidence shows AI augmenting work rather than eliminating it. McKinsey’s 2026 survey of more than 10,000 senior leaders finds that they expect AI to act mainly as a support tool in the near term [10], and most organizations using AI report little or no change in headcount over the past year [1]. The productivity paradox section’s evidence, the METR developer study and the Faros engineering telemetry, shows AI changing the nature of work and in some cases slowing specific early-2025 tasks, not replacing the workers doing them. Saying this honestly to a skeptical workforce beats vague reassurance, because the honest version is what workers will discover as they engage with the tools anyway.

Establish an internal community of practice. A chat channel, a regular meeting, a shared library of prompts and workflows; the form matters less than the existence. This is among the most cost-effective adoption levers available, because it converts every successful use into shared knowledge and every failure into shared learning, and it implements three of the twelve high-performer practices at once: internal communication about the value AI creates, leader role-modeling, and role-based capability building [2].

The community of practice also closes a security gap. ISACA’s 2025 reporting documents that when organizations fail to provide approved, well-supported AI tools, employees adopt unsanctioned ones, pasting proprietary data into consumer chat interfaces, using personal accounts for work, and building shadow workflows that IT and security never see [11]. Gartner projects that more than 40% of enterprises will experience a security or compliance incident linked to unauthorized shadow AI by 2030 [12], and IBM’s 2025 breach research found that one in five organizations had already suffered a breach traceable to unsanctioned AI use, at roughly $670,000 in added cost per incident [13]. A sanctioned channel where workers can ask questions, see which tools are approved, and surface their workarounds is itself a control. This paper’s governance section covers the policy side of shadow AI; this section covers the adoption side, and the two work together or not at all.

Winning Skeptics

The most reliable conversion is a concrete demonstration on a task the skeptic personally finds painful. Abstract arguments about AI capability rarely change minds. Watching a tedious task finish in thirty seconds usually does. The goal is a personal experience of the gain, not a won debate.

The specific targeting heuristic is to find the highest-friction, lowest-judgment task in the skeptic’s actual work and build the demonstration there. High friction means repetitive, boring, or time-consuming relative to its value. Low judgment means the skeptic can validate the result on sight. The combination is what works, because the friction makes them want the help and the fast validation lets them trust the result without an act of faith. This is the same task-conditional model the productivity paradox section laid out: generation-heavy work with fast output validation is where AI delivers.

Skeptics who decline official tools are often not avoiding AI at all; many are using unsanctioned alternatives, or their colleagues are, or both. Reco’s 2025 telemetry, drawn from its own customer base, found that 27% of employees at companies with 11 to 50 workers use unsanctioned AI tools [14]. That figure reflects what one vendor observes across its customers rather than a representative sample, so treat it as directional, but the direction is unambiguous: small organizations show the densest shadow-AI use and the fewest resources to govern it. The choice for management is not whether AI enters the organization. It is whether AI enters through supported, governed channels or unsupported ones, which makes skeptic engagement part of channel control as much as productivity capture.

Teams that succeed at this make it the manager’s job rather than a volunteer effort. Identify the skeptic, identify the painful task, build the prompt with them, watch them run it once, and ask what they would change. The exercise takes thirty minutes and converts more skeptics per hour than any abstract training program.

The paradox and the optimism are both real, and the productivity paradox section showed why. What separates them is not the AI model. It is the set of deployment decisions a team makes before the model reaches production, and those decisions belong to the organization. The hardware in this paper’s roadmap buys the capability; the practices in this section are what convert it into measured output. An organization that funds the first without the second has bought a faster way to produce work it still has to throw away.

References

  1. A. Singla, A. Sukharevsky, B. Hall, L. Yee, M. Chui, and T. Balakrishnan, “The State of AI in 2025: Agents, Innovation, and Transformation,” McKinsey & Company (QuantumBlack), Nov. 5, 2025. [Online]. Available: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai. Survey Jun. 25 to Jul. 29, 2025, n = 1,993. [Accessed: 25-Jun-2026]

    EAIU-1 Secondary source Back to text

  2. A. Singla, A. Sukharevsky, L. Yee, and M. Chui et al., “The State of AI: How Organizations Are Rewiring to Capture Value,” McKinsey & Company (QuantumBlack), Mar. 12, 2025. [Online]. Available: https://www.mckinsey.com/capabilities/quantumblack/our-insights/the-state-of-ai-how-organizations-are-rewiring-to-capture-value. 2024 survey, n = 1,491. [Accessed: 25-Jun-2026]

    EAIU-2 Secondary source Back to text

  3. Anthropic, “Prompt Engineering Overview,” Anthropic Documentation, 2025. [Online]. Available: https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/overview. [Accessed: 25-Jun-2026]

    EAIU-3 Primary source Back to text

  4. OpenAI, “Prompt Engineering Guide,” OpenAI Documentation, 2025. [Online]. Available: https://platform.openai.com/docs/guides/prompt-engineering. [Accessed: 19-May-2026]

    EAIU-4 Primary source Back to text

  5. R. Dakan, J. Feller, and Anthropic, “Prompt Engineering Techniques,” Anthropic, 2025. [Online]. Available: https://www-cdn.anthropic.com/62df988c101af71291b06843b63d39bbd600bed8.pdf. Course materials. [Accessed: 19-May-2026]

    EAIU-5 Primary source Back to text

  6. Anthropic Applied AI Team, “Effective Context Engineering for AI Agents,” Anthropic Engineering Blog, Sept. 29, 2025. [Online]. Available: https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents. [Accessed: 25-Jun-2026]

    EAIU-6 Secondary source Back to text

  7. OWASP Foundation, “OWASP Top 10 for Large Language Model Applications 2025,” OWASP Foundation, 2024. [Online]. Available: https://owasp.org/www-project-top-10-for-large-language-model-applications/. [Accessed: 25-Jun-2026]

    EAIU-7 Primary source Back to text

  8. K. Hong, A. Troynikov, and J. Huber, “Context Rot: How Increasing Input Tokens Impacts LLM Performance,” Chroma Research, July 2025. [Online]. Available: https://research.trychroma.com/context-rot. [Accessed: 25-Jun-2026]

    EAIU-8 Secondary source Back to text

  9. National Institute of Standards and Technology, “Artificial Intelligence Risk Management Framework (AI RMF 1.0),” NIST AI Resource Center, U.S. Department of Commerce, Jan. 2023. [Online]. Available: https://airc.nist.gov/airmf-resources/airmf/. The page states that AI RMF 1.0 is being updated and a revised version is in progress. [Accessed: 19-May-2026]

    EAIU-9 Primary source Back to text

  10. A. Krivkovich, D. Klingler, D. Maor, and P. Guggenberger, “The State of Organizations 2026,” McKinsey & Company, Feb. 19, 2026. [Online]. Available: https://www.mckinsey.com/capabilities/people-and-organizational-performance/our-insights/the-state-of-organizations. Survey June to Sep. 2025, n = 10,018. [Accessed: 25-Jun-2026]

    EAIU-10 Secondary source Back to text

  11. ISACA, “The Rise of Shadow AI: Auditing Unauthorized AI Tools in the Enterprise,” ISACA Industry News, Sept. 26, 2025. [Online]. Available: https://www.isaca.org/resources/news-and-trends/industry-news/2025/the-rise-of-shadow-ai-auditing-unauthorized-ai-tools-in-the-enterprise. [Accessed: 19-May-2026]

    EAIU-11 Contextual source Back to text

  12. Gartner, “Gartner Identifies Critical GenAI Blind Spots That CIOs Must Urgently Address,” Gartner Newsroom, press release, Nov. 19, 2025. [Online]. Available: https://www.gartner.com/en/newsroom/press-releases/2025-11-19-gartner-identifies-critical-genai-blind-spots-that-cios-must-urgently-address0. [Accessed: 25-Jun-2026]

    EAIU-12 Secondary source Back to text

  13. IBM Security and Ponemon Institute, “Cost of a Data Breach Report 2025,” IBM Corporation, July 30, 2025. [Online]. Available: https://www.ibm.com/reports/data-breach. IBM press release (Jul. 30, 2025): https://newsroom.ibm.com/2025-07-30-ibm-report-13-of-organizations-reported-breaches-of-ai-models-or-applications,-97-of-which-reported-lacking-proper-ai-access-controls. [Accessed: 25-Jun-2026]

    EAIU-13 Secondary source Back to text

  14. Reco AI, “The State of Shadow AI Report 2025,” Aug. 5, 2025. [Online]. Available: https://www.reco.ai/state-of-shadow-ai-report. Vendor first-party telemetry across Reco's customer base; not a representative sample. [Accessed: 25-Jun-2026]

    EAIU-14 Secondary source Back to text

Contents