A meta-review of 10 studies showing why AI transformation is a governance challenge involving ownership, workflows, data, risk, and accountability.
AI can immediately improve an individual task. Transforming an enterprise, however, requires decisions about where AI belongs, who owns it, which data it can access, how work must change, how outputs are checked, and how value and risk are measured. This denotes a governance problem, and across our ten-study evidence base, each research identifies at least one organizational decision between AI capability and enterprise value.
Nicklpass conducted an evidence-weighted review of ten studies on enterprise AI outcomes to test whether "AI transformation" is best understood as a technology problem or a governance problem.
The ten studies in this analysis fall into two groups. Seven directly examine enterprise operating conditions (e.g., leadership, workflow design, data, adoption, governance maturity, and risk controls). The remaining three are workplace experiments that show why these conditions matter (i.e., how AI outcomes change depending on task fit, system design, and the way humans and AI work together).
Research limitation: None of the ten studies randomly assigns organizations to "strong governance" and "weak governance." Instead, this review checks whether the studies point to the same conclusion, gives more weight to stronger evidence such as experiments, and links each finding to a clear governance decision.
Governance points to a system of decisions that controls how AI enters and changes an organization.
Hence, each study was coded against these six questions, derived from the world's leading AI research and governance authorities, to ensure flow from AI ideation to enterprise value:
A study is counted as governance-concordant when at least one reported finding showed that an AI outcome depended on one or more of these organizational decisions.
The list below maps all ten studies: what each one measured, who it studied, and the governance decision it exposes.
Randomized field experiment; 758 BCG consultants. AI improved suitable-task performance but reduced correctness outside its capability frontier.
Signal: Task fit, human review
Staggered workplace deployment; 5,172 agents and more than 3 million chats. Productivity increased 15% in an organization-specific, human-controlled system.
Signal: Data, workflow, adoption, assurance
Randomized field experiment; 776 P&G professionals. AI changed the value of team size, expertise, and cross-functional collaboration.
Signal: Work and team design
65 practitioner interviews, including 50 industry participants. 84% of industry interviewees cited leadership-driven causes.
Signal: Fit, ownership, metrics, workflow
Global survey; 1,491 participants. Workflow redesign had the strongest association with reported EBIT impact.
Signal: Ownership, workflow, measurement
Survey of 721 companies plus 16 interviews at nine enterprises. Higher maturity appeared as a cumulative bundle of policies, data, metrics, and new ways of working.
Signal: Capability system
Survey of 2,773 leaders in 14 countries. Spending rose faster than scaling and governance readiness.
Signal: Governance, adoption, risk
Survey of 1,000 executives across 59 countries. 74% had yet to show tangible value; BCG attributed most implementation difficulty to people and processes.
Signal: Portfolio, workflow, adoption
Survey of 840 AI-adopting enterprises across G7 countries plus interviews. Adoption depended on data, skills, infrastructure, partnerships, information, and institutions.
Signal: Ecosystem governance
Survey of 1,500 senior executives plus 15 expert interviews. Responsible-AI maturity correlated with lower incident cost and severity.
Signal: Assurance, risk, accountability
A study of 758 BCG consultants reveals that on suitable tasks, AI users completed 12.2% more work and finished 25.1% faster. But on a task beyond GPT-4's capabilities, they were 19 percentage points less likely to find the correct answer.
With local and cloud hosting, loop engineering, and a plethora of new paid and free LLMs, workflows, agents, and harnesses further exacerbating the situation, an immediate governance requirement is quite evident.
Now more than ever, an enterprise must decide:
Source: Harvard Business School
In the peer-reviewed study of 5,172 customer-support agents and more than three million chats, access to an organization-specific AI assistant increased issues resolved per hour by 15% on average.
The AI system was built around the organization's own work. First, it learned from previous customer-agent conversations. Second, it gave greater weight to successful conversations from top performers. Third, it used approved internal documentation as a knowledge source. Fourth, it recommended an answer only when it found enough supporting evidence; otherwise, it stayed silent. Fifth, employees received structured training on when and how to use it. Finally, the employee — not the AI — made the decision to accept, edit, or reject each suggestion.
The largest gains accrued to less-skilled and less-experienced workers. Their issues resolved per hour increased by approximately 30%, and agents with two months of experience began performing at the level of untreated agents with more than six months of experience.
The study also found small quality declines among some highly skilled workers. Consequently, the most experienced worker may need a different interface, incentive, or review rule than a new employee.
Source: NBER
The randomized P&G field experiment involving 776 product-development professionals compared individuals and cross-functional teams, with and without AI.
Researchers spent a year fitting the experiment to real product-development routines, standardized AI training, recorded usage, used independent evaluators under a manager-validated protocol, and kept human judgment in idea selection.
Individuals using AI matched the quality of traditional two-person teams and finished faster, about 16% faster for individuals and 13% faster for teams. AI-assisted teams were also three times more likely to produce top 10% ideas and helped R&D and commercial teams work more effectively across silos.
Source: NBER
RAND interviewed 65 experienced AI practitioners to identify recurring causes of AI-project failure. The sample included 50 industry practitioners and 15 academics.
Among the 50 industry interviewees, 84% cited at least one leadership-driven issue as a primary cause of failure.
Those problems included:
Common problems included choosing the wrong business problem, tracking the wrong metrics, using AI where simpler tools would work, poor workflow fit, and changing priorities too early.
RAND also identified data quality, infrastructure, and technical limitations. Governance was not the only source of failure. But leadership-driven problems were the most frequently cited category.
Source: RAND
McKinsey's global survey of 1,491 participants examined how organizations were rewiring themselves to capture AI value.
Among 25 organizational attributes, fundamental workflow redesign had the largest association with self-reported EBIT impact from generative AI. CEO oversight of AI governance was also among the factors most correlated with stronger financial impact.
Yet only 21% of respondents reporting generative-AI use said their organizations had fundamentally redesigned at least some workflows. Fewer than one in five said their organizations tracked well-defined KPIs for generative-AI solutions.
Source: McKinsey
MIT CISR's enterprise AI maturity research combined a survey of 721 companies with 16 executive interviews at nine enterprises. It divided companies into four stages:
Together, 62% of enterprises remain stuck in the first two stages — and both perform below industry average on growth and profitability.
Most companies remained in the first two stages. Companies in stages three and four also reported financial performance above their industry averages (Stage 3: growth +11.3 pp; Stage 4: growth +11.9 pp and profit +10.4 pp above average), while those in stages one and two reported below-average performance (Stage 1: growth roughly −12 pp and profit −2.2 pp; Stage 2: profit around −3.2 pp below average).
The important finding is structural. Advanced AI adoption appears as a bundle of mutually reinforcing organizational capabilities. Missing data, unclear ownership, absent value metrics, or skipped capabilities can produce downstream failure. Companies in stages three and four reported above-industry-average performance, but the study shows association, not causation.
Source: MIT CISR
Deloitte's fourth State of Generative AI in the Enterprise survey covered 2,773 director-to-C-suite respondents across 14 countries.
The report found that the technology and investment clock runs fast. 78% of leaders expected their organizations to increase AI spending, chasing more models and more experiments. But the organizational change clock runs slow: 69% believed that fully implementing a governance strategy would take more than a year.
The report found that:
That gap creates what Deloitte calls an absorption bottleneck. More than two-thirds of respondents expected 30% or fewer of their experiments to scale fully within the next three to six months (i.e., the organization, not the technology, is the throughput constraint). Moreover, regulatory compliance had become the single most frequently reported barrier, followed by data deficiencies, workforce issues, and risk management.
The gap is still visible in Deloitte's 2026 State of AI update, which was not added to the ten-study denominator. It reports that only 34% of organizations are truly reimagining the business; only one in five has a mature governance model for autonomous AI agents.
Source: Deloitte
BCG's survey of 1,000 C-suite and senior executives across 59 countries and more than 20 sectors found that 74% of companies had yet to demonstrate tangible value of AI transformation.
The broader finding is that companies that scale value distinguish themselves through change management, product development, workflow optimization, talent, and governance, while data quality remains the major technical foundation.
Leaders frequently spend the most executive attention on the model because the model is visible, novel, and easy to purchase. The harder constraints sit in the company: fragmented ownership, weak data, unchanged work, and unclear incentives.
Source: BCG
The OECD, BCG, and INSEAD produced a 196-page study of AI adoption built around 840 AI-adopting enterprises across G7 countries.
It found that an enterprise may control its employees but still inherit risk and dependency from model vendors, cloud providers, data suppliers, consultants, research partners, and regulators. Accountability, data rights, and service continuity cross organizational boundaries.
An internal committee cannot govern systems, vendors, data flows, and consequences it cannot see. That is a governance problem.
Source: OECD
Infosys' Responsible Enterprise AI in the Agentic Era study surveyed 1,500 senior executives and interviewed 15 responsible-AI leaders.
The report found that incidents are nearly universal: 95% of surveyed executives reported at least one negative AI-related incident in the preceding two years, and of those experiencing damage, 77% reported direct financial loss. Yet fewer than 2% of organizations met the complete responsible-AI benchmark created for the study.
The study noted:
The small group that came closest fared measurably better. Organizations classified as responsible-AI leaders reported 39% lower incident costs and 18% lower incident severity compared with all other organizations.
Source: Infosys
Uncontrolled AI cannot scale; the failure rates, incident costs, and absorption bottlenecks catch up with it. But AI that is safe yet unused or unprofitable cannot transform the business either — governance that only says "no" produces the below-average performers stuck in the early maturity stages.
The organizations pulling ahead are the ones treating governance and value as the same discipline: every study in the set, from MIT CISR's accumulating capabilities to Infosys' validate-protect-escalate-monitor pipeline, describes control and scale advancing together, not trading off.
It is for this reason that Nicklpass Buyer Club is built — to tackle AI governance and implementation issues in organizations. It shows which AI and SaaS tools your company owns, uses, and pays for, and what could be the most optimized stack for your enterprise.
Nicklpass brings your AI tools, usage, spend, and licenses into one place, so you know what to prioritize first.