Stanford's Digital Economy Lab spent five months inside 51 enterprise AI deployments that actually reached production and actually paid for themselves. Across 41 organisations and seven countries, one finding keeps repeating: the technology was the easy part. This topic is about the other 77%.
AI does not fix a broken process. It automates it.
01 / Eleven findings
What fifty-one working deployments had in common.
77%of the hardest problems
01
The technology was the easy part.
Change management, data quality and process redesign — that is where the difficulty actually lived. Practitioners described the model itself as the least of it.
So whatBudget the process work, or what you have budgeted is a demo.
61%had already failed once
02
Every success has a failure buried under it.
Three in five of these projects had at least one earlier attempt that did not work. The cost of that attempt sits in somebody else's budget and never reaches the final ROI.
So whatThe return you are shown is the third attempt, priced as though it were the first.
weeks / yearssame use case, different company
03
The clock is organisational, not technical.
Comparable use cases reached production in weeks at one company and were still unfinished years later at another. The variables were executive sponsorship, the processes already in place, and whether end users wanted it at all.
So whatCount the approval layers between a business problem and production. That number is your real timeline.
71 / 30median gain — escalation vs approval
04
Let it run, and review the exceptions.
Escalation models — AI handles 80% or more, humans see only exceptions and samples — showed a 71% median productivity gain, against 30% where a person approved every output. The report is careful here, and so should you be: those two groups were not doing the same work.
So whatThis is not evidence that less oversight is better. It is evidence that the two belong to different tasks.
35%against 23% from end users
05
The veto is upstairs, not downstairs.
Legal, HR, Risk and Compliance were the most frequent source of resistance — ahead of the people who would actually use the thing. Users can decline to adopt. Staff functions can stop the launch.
So whatPut them in the room at design time. A gatekeeper invited early becomes a co-designer.
45%of deployments cut headcount
06
Fewer than half of them cut anybody.
Reduction was the single most common workforce outcome, but the alternatives together — hiring avoided, redeployment to higher-value work, a deliberate decision to hold the line — accounted for 55%.
So whatIf the only line in your business case is FTEs removed, you will miss every case where the same team simply did far more.
20%were agentic
07
Agents are still the minority, and the top of the table.
Only a fifth of the sample used agentic workflows, but those posted the highest median gains: 71%, against 40% for conventional high automation.
So whatMost of the value here came from a workflow, a model, an API and a human handling the exceptions.
$1 : $10visible cost to invisible
08
A dollar of technology can carry ten of everything else.
The Productivity J-Curve research the report builds on finds that one dollar of tangible technology investment can require up to ten dollars of intangibles — process redesign, retraining, reorganisation — before the return arrives.
So whatThe dip comes before the curve. Say that out loud, at the start, to the person funding it.
messyis not a blocker
09
Stop waiting for the data to be clean.
Models read PDFs, email threads and call transcripts far better than the software that came before them. Governance still matters. “Finish the warehouse first” no longer does.
So whatStore everything, connect it, let the model structure it. What you never captured is the only part you cannot fix later.
0projects killed by security
10
Security did not kill a single one.
Not one deployment in the sample died because of a security requirement. Isolation, access control, audit trails and PII redaction turned out to be the things that let AI near valuable data in the first place.
So whatSecurity is a tax paid at the front. What it buys is permission to work on something that matters.
42%say the model is swappable
11
Commodity for the routine. Decisive for the hard.
Across all implementations, 42% treated model choice as fully interchangeable. Split by task the picture sharpens: on routine work 71% said interchangeable and not one called it a critical differentiator; on advanced work only 18% said interchangeable, and 35% called it decisive.
So whatBuild the seam that lets you swap. Then stop arguing about which model, except where it genuinely decides the outcome.
02 / The same use case, twice
Weeks at one company. Years at another.
Two organisations, comparable problems, the same available technology. The difference between these two columns is not a model, a budget or an engineer. It is how many places a decision can stall.
Weeks5
A business owner names the problem
The process is already written down
Data access is granted
Legal is in the room from day one
Live with one team
Years11
The demo lands well
A POC is approved
The business likes it
The data turns out to be unreachable
Waiting on an IT interface
Security raises requirements
Legal begins its review
The sponsor changes job
Budget crosses a financial year
A second POC is commissioned
Nobody mentions the project any more
The technology worked on day one in both columns.
03 / Where the human sits
Not a question about the model. A question about the mistake.
How much oversight a system needs has almost nothing to do with how advanced it is. It follows from what a wrong answer costs, whether you can take it back, and how long it takes anyone to notice. Set the three and read the answer.
Cheap, reversible, quickly noticed — this is where the technology is genuinely good, and where the 71% came from.
04 / When the model actually matters
A commodity for the routine. Decisive for the hard.
Share of implementations that treated the choice of foundation model as fully interchangeable, split by the complexity of the task.
Routine tasks71%
Classification, document search, ordinary generation, first-pass screening. Not one case called the model a critical differentiator.
Advanced tasks18%
Complex coding, compliance analysis, clinical documentation, multi-step reasoning. 35% called the model decisive.
The durable advantage is the orchestration layer, not the foundation model. Workflow, data, integration, evaluation, governance — none of which changes when a better model ships.
Jack's extensions / What the report does not say
Six things we learned the harder way.
+01
Error cost is not a slider once there is hardware.
The oversight framing in the report assumes a wrong answer can be withdrawn. An arm that has already moved, a drone that has already flown, a line that has already run — none of those are recoverable by a rollback. On physical systems the escalation path has to be an interlock, not a dialog box.
+02
You cannot retro-fit a sensor to yesterday.
“Store everything” is good advice with a hardware ceiling. Software can start logging tomorrow; a product that shipped without the sensor recorded nothing, and never will. Decide at design review what the thing remembers, because that is the whole of what any future model can learn from.
+03
Legal is not one veto. It is several that disagree.
The report counts Legal, HR, Risk and Compliance as a single 35%. Ship into more than one market and it stops being one block: data residency, what counts as personal information, and who is allowed to sign off all differ — and they contradict each other. Design for the strictest one first. Retro-fitting a jurisdiction is a rewrite.
+04
A small team cannot build a model gateway. It can afford the seam.
The multi-model architecture in the report is an enterprise diagram. The part of it a five-person team can have this week is one function that every model call goes through. It costs an hour, and it buys the same thing the diagram buys: the ability to change your mind later.
+05
Adoption is measured by what stopped.
Usage dashboards are easy to make look healthy. The honest test is whether anything died — a spreadsheet nobody updates any more, a weekly meeting that got shorter, a queue that stopped being sorted by hand. If the old way is still running alongside, you have added work and called it adoption.
+06
Name the dip before you are standing in it.
The J-curve has a morale version and it arrives in month two: the new process is slower than the old one, the model is still wrong in ways nobody has catalogued, and the people who were sceptical are being proved right in public. Sponsors quit here. Predict the dip at kickoff, in writing, with a date on the far side of it.
Before the first line of code
Answer five questions honestly.
Which business number are we trying to move, and what is it today?
Can anyone in the room draw the current process end to end, including who owns each decision?
When the model is wrong, who carries it — and how long before they find out?
Which business owner is accountable for the result, rather than for the delivery?
What evidence will exist in ninety days that this was worth doing?
05 / Read the study honestly
What this research cannot tell you.
It only studied survivors.
Every case in the sample already worked. The report describes what success looked like and what it took. It says nothing about how often success happens.
The numbers are self-reported.
Interviews with executives and project leads, triangulated against internal documents where those were available. Nobody independently audited the productivity claims.
The sample leans digital.
Nine industries, but weighted towards manufacturing, financial services and technology — the sectors that were already instrumented. Do not paste these findings onto a small workshop or a heavily regulated public body.
Patterns are not causes.
Strong sponsorship shows up in successful projects. It is at least as likely that mature organisations produce both the sponsor and the success.
Source
Elisa Pereira, Alvin Wang Graylin and Erik Brynjolfsson, “The Enterprise AI Playbook: Lessons from 51 Successful Deployments.” Stanford Digital Economy Lab, April 2026.
Read the report (PDF) ↗Figures on this page are taken from the report itself. Stanford's publication page titles it “…51 Successful Developments”; the PDF cover reads “Deployments”.