Insights
The AI Adoption Gap
August 17, 2026
You have seen the number: 95% of enterprise AI projects fail. It counts something far narrower than that. The evidence underneath supports a more useful claim about why deployments stall, and about what your people have already done without asking.
52%
of US employees use AI in their role
15%
use it daily
Your people adopted AI without waiting for a program. Whether it becomes how the work gets done is the question your organization still answers.
Gallup, Organizational AI Adoption Jumps Six Points, fielded May 6 to 20, 2026. Probability-based panel of 22,573 employed US adults, weighted to Current Population Survey targets, margin of error 0.9 points. Self-reported frequency rather than telemetry, and “use AI in your role” counts any tool the respondent reaches for, including personal ones outside sanctioned software.
What the 95% actually counts
Project NANDA at the MIT Media Lab published The GenAI Divide in July 2025. One exhibit tracks task-specific and custom generative AI tools through a funnel: 60% of organizations investigated one, 20% piloted, 5% reached production. That 5% became “95% of AI fails.”
Four things about the figure belong in the same breath as the figure.
It covers custom and task-specific tools. The same exhibit puts general-purpose tools at 40% implemented, so the headline says nothing about the ChatGPT and Copilot deployments most companies actually run.
“Successfully implemented” means users or executives remarked on a marked and sustained productivity or profit impact. No financial statement was examined. A project returning a small real gain codes the same as one returning nothing, because the test is whether someone volunteered a comment about it.
The report labels its own contents “Preliminary Findings,” carries the version number v0.1, and lists one reviewer who is also one of its four authors. Its base is 52 organizations interviewed and 153 senior leaders surveyed at four conferences. Inside a single opening paragraph, 95% attaches to organizations, then to integrated pilots, and a few pages later to enterprise AI solutions.
The original MIT-hosted URL for the PDF now redirects to a Project NANDA overview page that does not mention the report.
No retraction exists. No institutional endorsement exists either. Wharton’s Kevin Werbach asked NANDA to release its supporting data or withdraw the report, arguing that the figure has no traceable derivation in the document (as quoted by Futuriom). The Fortune article that carried the number into every board deck reported 150 interviews and a survey of 350 employees, against the report’s 52 organizations and 153 senior leaders. No correction appears on that page.
Your people adopted AI. They did not wait for you.
Gallup’s May 2026 panel puts AI use at 52% of US employees, with 15% using it daily. The Real-Time Population Survey read about 41% of workers using generative AI for their job in November 2025, with about 12% using it daily. Two instruments framed the question differently and converged on daily use.
Most of that use runs outside the sanctioned path. Two-thirds of office professionals at large firms say they have used AI at work while believing it broke company policy, and 88% have put work information into public AI tools: correspondence, meeting notes, customer data, financial documents (PagerDuty and Wakefield Research, June 2026, a vendor-sponsored panel measuring a perception of policy rather than a verified violation). When Software AG asked shadow-AI users why, a third answered that IT does not offer the tools they need.
The NANDA report carries the same finding about its own sample. Workers at more than 90% of the companies it surveyed use personal AI tools for work, and the report states that shadow AI often delivers better returns than formal initiatives. The document that produced “95% of AI fails” also reports value being captured outside procurement.
Six defensible answers to one adoption question
Ask how widely AI has been adopted and the answer moves with the unit you count.
About 20% of US firms
All US firms, from the only nationally representative probability sample of businesses
37% of US firms with 250 or more employees
The same survey, filtered to firms of the size that run procurement
78% of the labor force
Share of workers employed at an adopting firm
41% of workers
Individuals reporting any work-related generative AI use
52% of employees
Individuals reporting AI use in their role at any frequency
88% of survey respondents
Individuals reporting their employer uses AI in at least one function
Federal Reserve staff reviewing three of these surveys attribute the spread to differences in sampling distributions and units of analysis, then add question framing, the materiality of reported usage, and social desirability bias.
The 20% and the 88% do not belong on the same axis. The Census Bureau counts all US firms, most of them very small. McKinsey counts individual respondents, most of them at large organizations. The comparison that survives scrutiny sits inside the Census data: 37% at firms with 250 or more employees against about 20% nationally. A number this sensitive to its own denominator cannot carry an argument, which is what happened to the 95%.
What has been measured
Three studies measured output rather than opinion, and they agree with each other.
Brynjolfsson, Li and Raymond tracked 5,172 customer support agents through a staggered rollout and found 15% more issues resolved per hour, published in the Quarterly Journal of Economics in May 2025. The outcome came from system logs. Gains concentrated among less experienced agents, and the most experienced saw small speed gains alongside small quality declines.
Three randomized trials at Microsoft, Accenture and a Fortune 100 firm pooled 4,867 developers and measured 26.08% more completed tasks. Completed tasks means weekly pull requests. A standard error of 10.3 points puts the true effect somewhere between roughly 6% and 46%.
A randomized trial of 758 BCG consultants found 12.2% more tasks completed, 25.1% faster, and more than 40% higher quality on work inside the model’s capability range. On one task built so the model would get it wrong, consultants using GPT-4 scored 60% and 70% correct against a control group at about 84.5%.
The study most often quoted for the other side belongs here too. METR found 16 experienced open-source developers took 19% longer with AI assistance while forecasting a 24% speedup. METR’s own follow-up with 57 developers and more than 800 tasks pointed the other way, and METR describes that result as only very weak evidence. Both confidence intervals cross zero. Anyone quoting either number as settled is misrepresenting what is known.
Uplift is real, bounded, and shaped like a task rather than like a business.
The value gap is largely a measurement gap
McKinsey’s November 2025 survey found 39% of respondents attributing any EBIT impact to AI, and most of those put it under 5%. Carry McKinsey’s own caution with the number: non-attribution can mean not measured. Deloitte found 25% of its respondents had moved 40% or more of their pilots into production, with 54% expecting to reach that inside six months. IBM’s 2025 CEO study reported 25% of initiatives delivering expected returns and 16% scaled enterprise-wide.
Read one more survey and the picture inverts. Wharton’s Human-AI Research group, fielding inside McKinsey’s own window, found 74% of large-enterprise leaders reporting positive generative AI returns. Both can be true. McKinsey asked what share of enterprise EBIT is attributable to AI. Wharton asked whether returns are positive. No agreed way to ask the question exists.
Wharton’s internal split explains which answer you tend to hear. 81% of respondents at vice-president level and above called returns positive, against 69% of mid-managers, and 45% against 27% on significantly positive. BCG’s 2026 survey found 62% of CEOs confident against 48% of non-C-suite executives. The surveys reporting the best returns are surveys of the people who approved the spend.
Brynjolfsson, Rock and Syverson published a mechanism for this exact pattern. Adjusting for intangible investment in the computing era put total factor productivity 15.9% above official measures by the end of 2017. The model was validated on computers rather than on AI, and it cuts against optimists too, because the same curve implies a later stretch of overestimation.
Direction, redesign, and room to learn
BCG surveyed 11,749 employees in 2026 and split them four ways. With limited strategic clarity and limited tools, 55% reported measurable business improvement. Give them tools and no direction: 60%. Give them direction and no tools: 80%. Direction moved the number 25 points. Access moved it 5.
The caveat belongs in this paragraph rather than a footnote. Both halves came from the same employee on the same questionnaire, so the people who feel good about their company’s strategy are also the people likelier to report it working. “Measurable impact” means the employee believes metrics improved. BCG’s definition of strategic clarity bundles a written strategy together with guidance on what to do with time saved, so it never was a clean contest between strategy and tools.
McKinsey’s self-identified high performers, about 6% of respondents, are close to three times likelier to have fundamentally redesigned workflows and three times likelier to strongly agree that senior leaders demonstrate ownership. Of 25 organizational attributes McKinsey tested in an earlier edition, workflow redesign had the largest effect on the ability to see EBIT impact. 21% of adopters had redesigned any workflow at all. That same analysis names head count reduction among the attributes with the largest bottom-line effect, which deserves saying out loud rather than burying.
Room to learn has not moved in a year. 72% of BCG’s respondents say AI changed the skills expected of them, and 36% feel they received adequate training. BCG notes both figures are unchanged from the prior year. Employees with more than five hours of training showed up as regular users at 79% against 67%, an association with adoption rather than with return. Among regular AI users, 66% receive limited or no guidance on what the freed time is for, and more than half say they are not reinvesting it in more strategic work. Only 28% see a connection between what leaders say about AI and what the organization does.
Gallup asked the people who decline, inside organizations that do make AI available. 46% prefer their existing way of working, 43% cite data privacy and compliance, 43% object on ethical grounds, and 39% do not believe AI can assist with the work they do. Gallup declines to rank these and concludes that usefulness and ethics remain the most significant barriers to initial adoption. The reason your rollout stalled may be that two in five non-users looked at the tool and judged it irrelevant to their job.
One finding sits awkwardly against the rest and belongs at the end of this section anyway. BCG found 46% of employees at companies redesigning workflows expect their job to disappear within ten years, against 34% at companies rolling out tools. Managers and leaders carry more of that fear than frontline employees, 43% against 36%. The approach the evidence ties to impact is the approach that frightens people most, and no rollout plan gets to skip that.
Governance and constraints come before the recommendation
Security and risk lead the barriers to scaling agentic AI, named by nearly two-thirds of respondents who hold AI governance responsibility, ahead of both regulatory uncertainty and technical limitations (McKinsey, March 2026).
Data comes next, and it is more prosaic than the model conversation suggests. 43% of data leaders name data readiness as the most significant barrier to aligning AI with business objectives, in a survey where 88% also claim their data is ready (Drexel LeBow and Precisely, commissioned by a data-integrity vendor). 18% of enterprises say their data is fully governed. 96% of IT professionals call company-specific content important for agents, and 36% have connected an agent to trusted internal content.
The compliance floor moves faster than deployments do. NIST’s AI Risk Management Framework stays voluntary, names 12 generative AI risk categories in its 2024 profile, and has no version 2.0. ISO/IEC 42001 remains at first edition and gets certified by external bodies, which is why it travels through procurement while NIST does not. The EU postponed high-risk obligations for standalone systems from August 2026 to December 2027, blaming, in its own recital, the delayed preparation of standards and the delayed establishment of national conformity assessment. Colorado repealed its AI Act before the law took effect and replaced it with a narrower disclosure regime starting January 2027. California’s automated decision-making obligations land the same month.
Assess governance before the recommendation for one honest reason. The evidence cannot show that governance produces returns, and every link between the two is correlational, usually with the organizations able to fund heavy governance already large and already mature. The evidence does show that the constraint set changes faster than the deployment does, and a recommendation written against last year’s constraints arrives obsolete.
The constraints stay visible after go-live
METR states that its published task horizons sit at a 50% success rate, and that reliability-critical work can require 98% or better. The gap between those two numbers is where most enterprise deployments live.
Reliability falls under repetition and under conversation. Requiring the same customer-service task to succeed eight times running dropped one agent from above 60% to below 25% (τ-bench, on two-generation-old models, cited here as the origin of the finding rather than as current performance). On a benchmark built from realistic CRM data, leading agents scored about 58% single-turn and about 35% when the task required multiple turns, while exceeding 83% on prescribed workflow execution. Researchers at Microsoft and Salesforce found every top model performed 39% worse on average when a task arrived across several turns instead of all at once, and decomposed the loss into a minor drop in aptitude and a large rise in unreliability.
That 83% is the design instruction. Bounded, specified, verified work holds. Open-ended, long-horizon, multi-turn work comes apart, and it comes apart through inconsistency rather than through incapacity, which is the harder failure to notice in a demo.
Watching costs less than the failure it catches, and few teams do it. Among teams building agents, 89% have some form of observability and 52% run offline evaluations. Roughly a quarter of teams with agents already in production evaluate nothing. That survey drew from a developer-tool audience, so the real cross-enterprise figure is likely worse.
What would prove this wrong
People adopt what helps them work, and yours already have. The open question is whether your organization gives that behavior somewhere to land.
The sequence the evidence supports: decide what work is supposed to change, check whether your data, your governance and the reliability of the models can carry that change, then choose the tool, then keep all three in view after go-live.
One outcome would falsify the argument. If the value gap is a measurement artifact and the productivity J-curve holds, a great many companies reporting nothing today will report a great deal in two years without changing anything about how they lead. The claim here is that the conditions matter, not that the tools do not.
A note on these sources
Almost everything above is self-reported by executives or employees about their own organizations, which makes it evidence about belief as much as about outcome. The measured exceptions are the customer support study, the developer trials, the consultant trial, and the benchmark results. The most-quoted number in the field comes from a preliminary v0.1 document whose original host no longer serves the file.
Where good sources disagree, Wharton’s 74% against McKinsey’s 39%, or METR against its own earlier finding, this piece reports the disagreement instead of resolving it. Nobody has resolved it.
