Literature review
Written by a person. Last read by a person on 2026-09-07, 1 day ago. Its facts were checked by the eval suite on 2026-09-07.
You want to know whether the strategy rests on anything, or you are checking a cited figure.
Scenario-based sample. Halden Systems is invented. The research below is not, and it is
recorded in data/literature.yaml with the date each source was read.
For the Halden Systems capability program.
Bottom line. The evidence says the gap is not effort and cannot be closed by more training of the kind already delivered. 88 percent of organizations have adopted AI and 6 percent are getting real value from it.
The one-line version
88 percent of organizations have adopted AI. 6 percent are getting real value from it. The gap is not effort.
What the evidence says, in seven lines
- One program for everyone is worse than none. Support that helps a beginner makes an expert perform worse. Not a preference. A measured effect.
- Recordings teach the steps and skip the thinking. That is the half that transfers.
- Sitting near someone who knows beats being trained. It predicts capability better than any course does.
- Using the assistant does not teach judgment about it. Usage is up. Trust is down. More usage will not fix that.
- Fear makes people retreat to what they know. That is the exact behavior we are trying to change, and half the staff are afraid.
- Paying people to finish a course buys finished courses. We already have those. What we lack is understanding.
- The time does not exist. 3 in four have no learning hours that are not stolen from delivery.
The three numbers to remember
| 88 / 6 | Adopted AI / getting real value from it. McKinsey, 2025. |
| 84 / 33 | Developers using AI / trusting its accuracy. Stack Overflow, 49,000 people. |
| 278 / 27 | The lift in organizational performance from version control, for teams with above-average documentation against below-average. DORA. |
The third one is the argument for this program in a single ratio, and it is worth being exact about what it compares. DORA measured how much each technical capability lifts organizational performance, then split the teams by the quality of their internal documentation.
Version control returned 278 percent for the teams whose documentation was above average, and 27 percent for the teams below it. The same practice, adopted by both groups, paid back roughly ten times more where the writing was good.
Good means quality here rather than volume, and the distinction is the whole finding. DORA scores documentation on eight attributes, among them whether it is clear, whether people can find it, and whether it can be relied on. Nobody was asked how much of it they had.
The claim is therefore not that teams should write more, which is what a learning function is usually assumed to be proposing. It is that the quality of the writing already being produced decides how much the engineering investment beside it returns. Documentation is a multiplier on that spend, not a line item competing with it.
What we got wrong, and are keeping in
PwC surveyed 50,000 workers and found the opposite of what we assumed. Daily AI users feel more job-secure, not less, at 58 percent against 36 for occasional users.
Fear is therefore not a wall to clear before people start. Use appears to reduce it.
That flips the plan. We do not need to resolve the anxiety first. We need to get people using the thing safely, in small ways, quickly. The fear comes down as a result.
Both things are true at once. Fear makes people retreat, and using the tool reduces fear. It is a loop, and our job is to break into it rather than to argue with it.
Where the fear actually sits
Half the staff think this tool may cut their jobs. 2 in five think that getting good at it would prove they can be replaced.
1 in eight would say that to a manager. So every conversation leadership has had about this has been with the 1 in eight.
There is a hard finding here. Vague reassurance makes it worse, not better. We told people nobody is losing their job. Under a quarter believe we have been straight with them. Those are the same fact.
What replaces vague reassurance is a commitment specific enough that somebody could hold us to it. That means three things a reader can check, and the reassurance we already gave has none of them.
A date, so that the promise covers a stated period rather than an indefinite one. A scope, so that it names which roles and which sites it applies to instead of gesturing at everybody. Third, a notice period, so that if the position changes, people hear it on a known timetable rather than on the day it takes effect.
The promise we made, that nobody is losing their job, cannot be kept, because it was never specific enough to be broken, and staff read that accurately as costing leadership nothing to say. A commitment that could be breached is the only kind worth making, because it is the only kind that puts something at risk for the person making it.
What we will not cite
The claim that 70 percent of change programs fail. It is everywhere and it has no data behind it. It traces back through citations to an assertion.
We are naming it as unsupported on purpose. It is the most tempting number available to us, and saying so is what earns the right to make an evidence argument at all.
How to read the consultant reports
McKinsey, BCG, Deloitte, PwC and the analyst houses are all in here. They carry
type: consultancy research in data/literature.yaml, so a reader can see exactly which entries
this section is about instead of taking our word for the grouping.
Read them differently, because each is published by a firm that sells the remedy its own research calls for. The clearest case is the McKinsey survey we quote most often. It reports that adoption has run far ahead of value capture, and it is published under the firm's AI practice, which is visible in the URL recorded against it.
The finding names a gap and the publisher sells the work of closing that gap. PwC surveys the workforce and sells workforce transformation. The pattern holds across the tier, and it is the reason these sources sit in a group of their own.
None of that makes the findings wrong, and their samples are far bigger than anything we could field. What it means is that the choice of which problem to measure is not neutral, and the size of a number is no evidence that its subject is the thing most worth fixing.
We cite them for two things. Scale, because nobody else surveys 50,000 people, and what the sponsor has already read, which matters more: the COO has seen the McKinsey number before we walk in, and a director who cannot handle it accurately loses the room to a figure they never checked.
We do not cite them for whether something works. That is what the research tier is for.
Gartner and Forrester get a narrower job still. They belong in the vendor section as a market map. A procurement that cites a quadrant as its reason will be asked why, and will deserve it.
Verification status
Six sources opened and read on 2026-09-07: DORA, Stack Overflow, McKinsey, PwC, GitLab, and the paper debunking the 70 percent claim. Every figure above comes from those.
Eight more are named and unread. They carry no figure, and none of the argument rests on them.
The 13 research findings are stated at a level we can defend in a room, which is a lower bar than the papers behind them would support and the right one for a funding meeting.
What the evidence does not say
It does not say this will work. Every finding is about mechanism, and mechanism shapes a design without predicting a result here.
It says nothing useful about what multi-site coordination costs, which is the largest unpriced risk in the plan.
The work on trusting automated output also predates assistants of this kind. The effect is old and solid. The systems are new. We treat the carry-over as likely rather than proven, and the strategy document should say so in those words.