The leadership set, on one page
Written by a person. Last read by a person on 2026-09-07, 1 day ago. Its facts were checked by the eval suite on 2026-09-07.
You want to read the whole case in one scroll, you are sending it to somebody who will not click, or you want to print it.
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Every document in the set, in order, generated from the same sources as the individual pages. It cannot fall out of step with them, and a new document appears here without anybody remembering to add it. Each one also stands alone at its own address.
The contents list at the top of the page carries all 14 of them. A second one here would be the same list twice, which is what looking at the built page showed.
The situation at Halden Systems
1 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. We taught 3,000 people the steps and never taught them the system. 78 percent can follow our documented procedures; 31 percent can say what those procedures do. That 47-point gap between following and understanding is what produces the escalations, the fear and the stalled adoption, it costs about $6.1M a year, and no better tool will close it.
They can follow the steps. They do not understand the system.
The organization
Halden Systems builds and sells engineering software.
| Staff | 5,000 |
| Sites, across 4 continents | 9 |
| Hours between the first site's morning and the last site's evening | 17 |
| Technical staff | 2,000 |
| Non-technical staff | 3,000 |
The timezone spread is the figure that decides the delivery model, and the split between the last two rows is the one that decides who the program is for.
2 groups inside those 3,000 non-technical staff need naming, because the program treats them differently.
The 1,150 who work in engineering systems daily. Writers, designers, researchers, program managers and the senior half of support. Their work lives in repositories, issue trackers and build pipelines. They are the group whose capability gap costs money today.
The other 1,850, who were given an AI assistant anyway. Their work does not touch repositories, and the assistant went to every employee without regard to whether it suited the job. They are counted here because a tool nobody assessed them for is now in their hands.
The 4 earlier migrations at least moved people whose work had to move. This one reached 1,850 people with no case for reaching them, and it repeats the same error more widely: capability was never measured before the tool arrived.
5 of the 9 sites have nobody who can answer a question about repositories, pipelines or issue trackers during the local working day.
What is failing
Over 4 years Halden moved most of its non-engineering work into engineering systems. Documentation became docs-as-code. Design tokens went into a repository. Program management moved into the issue tracker. Support began authoring runbooks in markdown.
Every one of those moves was run as a tooling migration: a cutover, a demonstration, a recording, and a channel for questions. None was run as a capability problem, and our own survey shows what that produced.
78 percent of the 1,150 can follow the documented steps. 31 percent can say what those steps do. Our completion dashboard reports the seventy-eight and calls the program finished.
A person who follows steps without understanding them is fine until the screen stops matching the instructions. Then they escalate. That happens a median 2.3 times a week each, which across the 1,150 comes to roughly 760 engineer hours absorbed every week.
64 percent say they are afraid of breaking something that cannot be undone. Given that they cannot predict what any step will do, that fear is accurate rather than timid, and it is the reason the escalation rate is a floor rather than a peak.
The AI assistant is the fifth rollout, not a new problem
The assistants went to the same people, through the same workflows, on the same playbook.
The survey asked the same 1,150 people four questions about it.
| Share | |
|---|---|
| Have used the assistant | 71% |
| Would put their name on its output | 38% |
| Can tell which answers need checking | 22% |
| Have been given an answer that was wrong and sounded confident | 57% |
The third row is the one that matters, because an assistant is only safe for somebody who can judge when it is wrong. On that measure fewer than a quarter of them are equipped to use it, and more than half have already met the failure it produces.
Somebody who cannot yet judge the output learns the wrong lesson from a confident wrong answer, which is that they cannot tell good answers from bad ones. The reasonable response is to stop using the assistant, and to say nothing about having stopped, because the rollout was announced as a success.
That is why usage will keep looking healthy while the benefit does not arrive. We are counting who opened the tool, not who can use it safely.
What was already tried, and what it produced
1 site ran a pilot at its own cost. Dublin, 24 people from the inner population, 6 weeks, a written runbook and one facilitated change each, with a named person to ask.
Median days to a first merged change fell from forty-one to nine for that group. The pilot cost nothing beyond existing staff time, which matters for one reason: it answers the objection that nothing can begin before the funding decision. Something already did.
What the scale changes
At 1,200 people in one building, one determined person holds a program together by force of will. At 5,000 across 9 sites, 3 constraints bind that would not otherwise.
1. No hour of the working day is shared by every site. 17 hours separate the first site's morning from the last site's evening, and 18 percent say they would attend an optional live session. Any plan that rests on live attendance is refuted by our own survey before anyone costs it.
2. Proximity to help predicts capability better than job title does. The confidence gap between sites with a local expert and sites without one is 22 points, wider than the gap between job families. Median days to a first merged change is 41 company-wide and 58 at the 5 sites with nobody to ask.
3. 9 sites are each buying their own training, at about $410,000 a year of overlap. No site director was wrong to do that. Waiting for a central answer costs them more than buying locally, and the duplication is a symptom of having no shared owner rather than of poor judgment.
Where the money is
Three lines, and they do not come from the same population.
Revenue, which comes from customers rather than staff. Halden sells engineering software, and its customers need the same kind of teaching its own people do. Solution engineers currently spend about 6 days per new enterprise customer delivering that teaching by hand, across 240 new enterprise customers a year. 95 partner implementations a year depend on partner staff holding a certification we do not currently offer.
The connection to the internal problem is the machinery, not the audience. The content, the assessment and the certification that would close the internal gap are the same assets a customer program sells, which is why one function should own both rather than two functions building them twice.
Cost, which comes from the 1,150 and their colleagues. Roughly 760 engineer hours a week at $145 fully loaded, plus the $410,000 of duplicated regional spend, plus about 1,900 internal tickets a month whose answer already exists in writing.
Efficiency, which comes from time. 41 median days to a first merged change, 58 at the sites without an expert, and nine in the Dublin pilot.
Who decides
The Chief Operating Officer funds it. The problem crosses every function, and no single department can pay for it from its own budget.
The VP of Engineering is its loudest advocate, because these are their engineers absorbing 760 hours a week. They are also the wrong owner. A program owned by engineering gets built for people who think like engineers, and those people are not the ones failing.
The nine regional site directors decide whether it works. They hold the budgets that currently pay for the duplication, they employ the people the program serves, and they are the ones who have been coping without help. A plan that takes their budgets before it has earned their trust turns the constituency it needs into the opposition that stops it.
Literature review
2 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented. The research below is not, and it is
recorded in data/literature.yaml with the date each source was read.
For the Halden Systems capability program.
Bottom line. The evidence says the gap is not effort and cannot be closed by more training of the kind already delivered. 88 percent of organizations have adopted AI and 6 percent are getting real value from it.
The one-line version
88 percent of organizations have adopted AI. 6 percent are getting real value from it. The gap is not effort.
What the evidence says, in seven lines
- One program for everyone is worse than none. Support that helps a beginner makes an expert perform worse. Not a preference. A measured effect.
- Recordings teach the steps and skip the thinking. That is the half that transfers.
- Sitting near someone who knows beats being trained. It predicts capability better than any course does.
- Using the assistant does not teach judgment about it. Usage is up. Trust is down. More usage will not fix that.
- Fear makes people retreat to what they know. That is the exact behavior we are trying to change, and half the staff are afraid.
- Paying people to finish a course buys finished courses. We already have those. What we lack is understanding.
- The time does not exist. 3 in four have no learning hours that are not stolen from delivery.
The three numbers to remember
| 88 / 6 | Adopted AI / getting real value from it. McKinsey, 2025. |
| 84 / 33 | Developers using AI / trusting its accuracy. Stack Overflow, 49,000 people. |
| 278 / 27 | The lift in organizational performance from version control, for teams with above-average documentation against below-average. DORA. |
The third one is the argument for this program in a single ratio, and it is worth being exact about what it compares. DORA measured how much each technical capability lifts organizational performance, then split the teams by the quality of their internal documentation.
Version control returned 278 percent for the teams whose documentation was above average, and 27 percent for the teams below it. The same practice, adopted by both groups, paid back roughly ten times more where the writing was good.
Good means quality here rather than volume, and the distinction is the whole finding. DORA scores documentation on eight attributes, among them whether it is clear, whether people can find it, and whether it can be relied on. Nobody was asked how much of it they had.
The claim is therefore not that teams should write more, which is what a learning function is usually assumed to be proposing. It is that the quality of the writing already being produced decides how much the engineering investment beside it returns. Documentation is a multiplier on that spend, not a line item competing with it.
What we got wrong, and are keeping in
PwC surveyed 50,000 workers and found the opposite of what we assumed. Daily AI users feel more job-secure, not less, at 58 percent against 36 for occasional users.
Fear is therefore not a wall to clear before people start. Use appears to reduce it.
That flips the plan. We do not need to resolve the anxiety first. We need to get people using the thing safely, in small ways, quickly. The fear comes down as a result.
Both things are true at once. Fear makes people retreat, and using the tool reduces fear. It is a loop, and our job is to break into it rather than to argue with it.
Where the fear actually sits
Half the staff think this tool may cut their jobs. 2 in five think that getting good at it would prove they can be replaced.
1 in eight would say that to a manager. So every conversation leadership has had about this has been with the 1 in eight.
There is a hard finding here. Vague reassurance makes it worse, not better. We told people nobody is losing their job. Under a quarter believe we have been straight with them. Those are the same fact.
What replaces vague reassurance is a commitment specific enough that somebody could hold us to it. That means three things a reader can check, and the reassurance we already gave has none of them.
A date, so that the promise covers a stated period rather than an indefinite one. A scope, so that it names which roles and which sites it applies to instead of gesturing at everybody. Third, a notice period, so that if the position changes, people hear it on a known timetable rather than on the day it takes effect.
The promise we made, that nobody is losing their job, cannot be kept, because it was never specific enough to be broken, and staff read that accurately as costing leadership nothing to say. A commitment that could be breached is the only kind worth making, because it is the only kind that puts something at risk for the person making it.
What we will not cite
The claim that 70 percent of change programs fail. It is everywhere and it has no data behind it. It traces back through citations to an assertion.
We are naming it as unsupported on purpose. It is the most tempting number available to us, and saying so is what earns the right to make an evidence argument at all.
How to read the consultant reports
McKinsey, BCG, Deloitte, PwC and the analyst houses are all in here. They carry
type: consultancy research in data/literature.yaml, so a reader can see exactly which entries
this section is about instead of taking our word for the grouping.
Read them differently, because each is published by a firm that sells the remedy its own research calls for. The clearest case is the McKinsey survey we quote most often. It reports that adoption has run far ahead of value capture, and it is published under the firm's AI practice, which is visible in the URL recorded against it.
The finding names a gap and the publisher sells the work of closing that gap. PwC surveys the workforce and sells workforce transformation. The pattern holds across the tier, and it is the reason these sources sit in a group of their own.
None of that makes the findings wrong, and their samples are far bigger than anything we could field. What it means is that the choice of which problem to measure is not neutral, and the size of a number is no evidence that its subject is the thing most worth fixing.
We cite them for two things. Scale, because nobody else surveys 50,000 people, and what the sponsor has already read, which matters more: the COO has seen the McKinsey number before we walk in, and a director who cannot handle it accurately loses the room to a figure they never checked.
We do not cite them for whether something works. That is what the research tier is for.
Gartner and Forrester get a narrower job still. They belong in the vendor section as a market map. A procurement that cites a quadrant as its reason will be asked why, and will deserve it.
Verification status
Six sources opened and read on 2026-09-07: DORA, Stack Overflow, McKinsey, PwC, GitLab, and the paper debunking the 70 percent claim. Every figure above comes from those.
Eight more are named and unread. They carry no figure, and none of the argument rests on them.
The 13 research findings are stated at a level we can defend in a room, which is a lower bar than the papers behind them would support and the right one for a funding meeting.
What the evidence does not say
It does not say this will work. Every finding is about mechanism, and mechanism shapes a design without predicting a result here.
It says nothing useful about what multi-site coordination costs, which is the largest unpriced risk in the plan.
The work on trusting automated output also predates assistants of this kind. The effect is old and solid. The systems are new. We treat the carry-over as likely rather than proven, and the strategy document should say so in those words.
The capability survey
3 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. 1,237 Halden Systems staff answered, a 41% response rate. 78% can follow our documented steps and 31% can say what those steps do. The items least convenient for this program are in the results below, because a survey where every item supports the proposal is a survey nobody ran.
The instrument
| Fielded | 4 to 22 May 2026, 3 weeks |
| Population | 3,000, all non-technical staff |
| Responses | 1,237 |
| Response rate | 41% |
| Languages | 5, the working languages of the 9 sites |
| Design | 14 Likert items, 4 behavioral frequency items, 3 free-text |
Two design decisions are worth stating, because both change what the numbers mean.
We surveyed all 3,000, not the 1,150 who work in engineering systems daily. Sampling only the inner population would have measured the program's convenience rather than the company's problem, and the assistant went to everybody regardless of whether their work touches a repository.
Free text came last. A respondent who has already answered 18 structured items has been reminded what the subject is, and writes something specific instead of something general.
The headline
| Item | Result |
|---|---|
| Can complete common tasks by following documented steps | 78% |
| Understand what those steps actually do | 31% |
| Worry about breaking something that cannot be undone | 64% |
| Times a week they ask a technical colleague for help, median | 2.3 |
The 47-point distance between the first 2 rows is the finding the program exists to answer. Every dashboard we currently run reports the 78 and calls it success.
The third row is not timidity. Somebody who cannot predict what a step will do is correct to fear an irreversible one, and the fourth row is the cost of that fear landing on an engineer.
The assistant
| Item | Result |
|---|---|
| Have used it at least once | 71% |
| Use it weekly or more | 44% |
| Would put their name on its output | 38% |
| Have had an answer that was wrong and sounded confident | 57% |
| Know which of its answers need verifying | 22% |
The drop from 71% to 44% is where the rollout actually stands, and neither figure is the one that matters. Judging when the output is wrong is the skill, and 22% report having it.
What people will not say to a manager
| Item | Result |
|---|---|
| Believe the assistant will reduce headcount in their area | 46% |
| Would say that to their manager | 12% |
| Have avoided asking a question because of how it would look | 51% |
| Believe becoming skilled with it would prove they are replaceable | 39% |
| Believe leadership has been straight about it | 23% |
The 34-point gap between the first 2 rows is why leadership's read of this is wrong. Every conversation on the subject has been with the 12%.
The last row sits against a reassurance we already gave. We told people nobody is losing their job, and under a quarter believe we have been straight with them. Those are the same fact.
Time, and what gets in the way
| Item | Result |
|---|---|
| Report more simultaneous tool or process changes than they can absorb | 68% |
| Have no learning time that is not taken from delivery | 74% |
| Do this learning outside working hours | 41% |
The third row is an equity problem before it is a capability problem. It selects for people without caring responsibilities, and any plan that assumes it will continue is choosing that.
Incentives, and what people would actually do
| Item | Result |
|---|---|
| Believe effort to build these skills is noticed | 19% |
| Would answer colleagues' questions if that counted for something | 61% |
| Would be motivated by a reward for completing training | 16% |
| Would accept being publicly identified as learning this | 27% |
Rows 2 and 3 point in opposite directions and settle the incentive design. Paying for completion motivates 16%, and completion is already at 78%. Recognizing contribution reaches 61%, and contribution is the thing we have none of.
Row 4 constrains how. Any scheme that requires somebody to be visibly a learner reaches about a quarter of the population, and loses the reset expert entirely.
By site and by language
| Item | Result |
|---|---|
| Confidence gap, sites with a local expert against sites without | 22 points |
| Response rate at the smallest South American site | 19% |
| Confidence gap for respondents working in English as a second language | 14 points |
The 22-point gap is wider than the gap between job families, which is the finding that makes proximity to help a better predictor of capability than role.
The 19% is reported rather than hidden. It is the site we understand least, and it is one of the 5 with nobody to ask, so the true picture there is probably worse than these results suggest.
The results that are inconvenient for this program
Kept because removing them would make the rest untrustworthy.
| Item | Result |
|---|---|
| Rated previous training as useful | 29% |
| Believe more training would help | 34% |
| Would attend an optional live session | 18% |
29% found our previous training useful. That is this function's own report card, and it is the strongest argument against giving this function more money.
Only a third think more training is the answer. They are largely right, which is why the program is built around local help, contribution and time rather than courses.
18% would attend an optional session. Any plan resting on voluntary synchronous attendance is refuted by its own survey before anybody costs it.
Free text, coded
3 free-text items, 1,237 responses, 7 themes. Counts are responses coded to each theme, and a response can carry more than one.
| Theme | Coded | What it sounds like |
|---|---|---|
| I was good at my job and now I am not | 218 | competence reset |
| Elaborate workarounds nobody can see | 174 | coping machinery |
| I might break something and not be able to undo it | 161 | fear of permanence |
| It was confidently wrong and I stopped | 143 | burned once |
| If I get good at it, I have proved they do not need me | 189 | the replacement bind |
| This is the fourth thing this year | 152 | no time to absorb |
| The person who could tell me is asleep | 128 | no local help |
"I have been a writer for fifteen years. I was the person people asked. Now I copy commands from a document and hope."
The largest theme is not about tools. It is about having been competent and no longer being so, and a program that treats it as a skills gap will be answering a different question than the one people are asking.
The people this is for
4 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. 6 learner personas and 1 sponsor, drawn from Halden Systems staff and derived from named survey themes. The 6 account for 93% of the population, not 100%, and the missing 7% is left visible. Every persona names what it rules out, because a persona that refutes nothing is decoration.
The six, and what each one refutes
| Persona | Role | Share | Rules out |
|---|---|---|---|
| The reset expert | Senior technical writer, 15 years, HQ | 19% | Anything labeled introductory |
| The stranded newcomer | Researcher, 8 months, no local expert | 18% | Any plan that resolves to synchronous delivery |
| The workaround builder | Program manager, 6 years, has local help | 16% | Measuring need by who asks for help |
| The quiet complier | Marketing operations, 9 years, HQ | 16% | Completion and attendance as evidence |
| Burned by the assistant | Support lead, 4 years, second-language English | 14% | Adoption targets as a measure of success |
| Never knew different | Designer, 14 months, joined after the migration | 10% | Treating the population as uniformly in deficit |
The right-hand column is the working half. A design that survives all 6 of those constraints is narrower than the one anybody would have proposed, which is the point of building them first.
What each one needs
The reset expert, 19%. She was the person others asked. The documentation moved into the repository and her expertise did not move with it. She needs a model rather than a course, and she learns fast once something is explained instead of demonstrated. Any material that requires her to publicly identify as a beginner will not be opened. Themes: competence reset, fear of permanence.
The stranded newcomer, 18%. Her question is answered 18 hours later, by which time she has worked around it. She needs help that does not depend on somebody being awake, and written material good enough to answer the second question as well as the first. Themes: no local help, fear of permanence.
The workaround builder, 16%. He keeps a private file of exact commands and a standing favor with one engineer. He is productive and he is invisible to every measure we have. His private file is most of a good runbook already. Theme: coping machinery.
The quiet complier, 16%. She attends every session, completes every module, and uses none of it. She is the reason 78% can follow the steps and 31% understand them. She needs a specific answer about what the tool will and will not be used for, not another reassurance. Themes: the replacement bind, no time to absorb, competence reset.
Burned by the assistant, 14%. It gave him a process that did not exist, in confident prose, and he caught it barely. He is a rational non-user, and counting him as an adoption failure would push him back toward the behavior that nearly cost him. He needs the skill nobody taught: how to tell which answers need checking. Themes: burned once, competence reset.
Never knew different, 10%. She learned the repository in her first week alongside everything else and does not understand why her colleagues struggle. She should be recruited rather than trained. A tenth of the population is already the answer to the other nine tenths, and it is the cheapest capacity in the building. Themes: none, and that absence is the finding.
The seventh, who is not a learner
The regional sponsor. Site director, Asia-Pacific, 200 staff, 0% of the learner population and the person who decides whether any of this works.
He has bought training twice from his own budget, because waiting for a central answer cost more than paying for a local one. Both purchases overlap with what other sites bought. He is not the cause of the $410,000 of duplication; he is a rational response to having no shared owner.
What he needs is a reason to stop buying, which means something arriving faster than his own procurement does. Consolidation that reaches him a year from now will be ignored, correctly.
He rules out the plan whose first year is centralization, and that single constraint removed the most obvious saving from the strategy.
The 7% nobody speaks for
The 6 learner personas sum to 93%. The remaining 7% is not assigned, and the gap is left in rather than distributed.
Personas that sum to exactly 100% have been rounded until they did. The people in that 7% are mostly at the 2 smallest sites, where the response rate was lowest, and the honest statement is that we do not know enough about them to write one. That is a finding about our sampling, and it belongs in the numbers rather than in a footnote nobody reads.
How these were derived
Each persona names the survey themes it rests on, and every theme in the survey reaches at least one persona. Both directions matter. A persona with no theme behind it is a character somebody invented, and a theme with no persona in front of it is evidence the design has quietly dropped.
"Never knew different" carries no themes on purpose. She was built from the absence of complaint in a cohort that joined after the migration, which is a different kind of evidence and is labeled as such rather than dressed up as the same kind.
The first 90 days
5 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
The entry plan for a director newly hired into Halden Systems.
Bottom line. The first 90 days end with a funding decision made and procurement in flight, not with a signed contract. 10 days to find out what is true, 20 to gather evidence, thirty to build the case, thirty to ask for the money.
Why a newly hired director writes this
These documents are written by somebody hired into the situation, not by an observer describing it. That frame is worth adopting for three reasons, and the third is the one that matters.
It is the frame a director interview already uses. Some version of "what would you do in your first 90 days" is asked in nearly every one, and answering it with artifacts rather than assertions is a better answer than the question expects.
It sequences the set. Ten documents with no narrator arrive in an arbitrary order. With a start date they arrive in the order somebody would actually produce them, and the sequence is itself an argument about how this work should be done.
It also answers the question the set would otherwise raise. A reader looking at a strategy document with a survey of 1,237 people in front of it will ask where that came from and who authorized it. The answer is that it was the first month's work, and that a proposal written before it would have been an opinion.
Ten, thirty, sixty, ninety
The usual plan has three phases. This one has four, and the extra phase is at the front.
Days 1 to 10 produce nothing that proposes anything. The output is a written list of what is not yet known. The sponsor, the VP of Engineering, all 9 site directors, and a sample of the population. Read the prior rollout materials, the ticket queue, and every line of regional training spend, and ask each site director what they bought and why without implying it was wrong, because it was not.
10 days is deliberately short. A plan produced inside it is one that arrived in the new director's luggage, and the fastest way to lose the 5 site directors who decide whether this program reaches their people is to arrive holding the answer.
Days 11 to 30 are evidence, and it arrives on day fifteen.
A thirty-day survey is a hard thing to ask a sponsor for, and it was the wrong design. Here is the replacement.
| When | What lands |
|---|---|
| Days 1 to 5 | Data we already hold: tickets, calendars, training spend |
| Days 3 to 8 | Twenty interviews |
| Days 8 to 15 | A six-question pulse, open 7 days |
| Day 15 | First readout, with a pilot already running |
| Days 16 to 30 | The full instrument, 10 days in the field |
The pilot starts on day 11 at 1 site with a local expert and one without. It already exists and costs nothing to begin, and that comparison is what the scaling case will need.
The long instrument is a second pass. It deepens the picture rather than gating it.
What does not happen here is committing to a number in front of the Chief Operating Officer before the pulse lands. The cost of guessing early is that the evidence then has to agree with the guess.
Days 31 to 60 are the case. The survey closes on day 41; analysis and personas to day 46; the strategy document to day 53; the project plan to day 58. Early pilot data arrives across the same period and goes into both.
No hiring. The shape of the team follows from the strategy and the procurement decision, and a requisition opened in month two is a guess that gets defended for two years.
Days 61 to 90 are the ask, with procurement started and not finished. Deck to day 64, readout around day 65, statement of work to day 70. The request goes to vendors around day 70 with responses due near day 100, and the evaluation weights are fixed and written down before any submission is opened. The scaling proposal completes by day 76.
The two things this plan will not claim
It does not close a vendor contract. A response window is 40 days and reference calls take longer than that. A ninety-day plan that ends with a signature tells anybody who has run a procurement that the writer has never run one, and it is the most common way these plans give themselves away.
It does not centralize regional budgets in year one. Roughly $410,000 a year is being spent twice across 9 sites, and it is the most obvious saving in the case. It is also the one that would cost the program the regional site directors, who are buying locally because waiting for a central answer costs them more. They stop when something arrives faster than their own procurement does, and not before.
Both refusals are in the plan on purpose. A ninety-day plan is read for what it promises, and the two most informative lines in one are usually the promises it declines to make.
Capability strategy
6 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. The 1,150 non-technical staff who work in our engineering systems can follow the steps and cannot read the system, which costs about $6.1M a year in absorbed engineer time and duplicated spend. We are asking for $2.1M over three years to close it, decided this quarter.
Two kinds of number appear below. Figures about Halden Systems and its staff come from
data/scenario.yamlanddata/survey.yaml, and are invented for a fictional company. Figures credited to DORA, Stack Overflow, McKinsey or PwC are real, were read on 2026-09-07, and are recorded indata/literature.yaml. They are never mixed unlabeled.
For the Chief Operating Officer. Prepared in the first 90 days, and the decision it asks for is a funding commitment, not an approval to plan.
The gap, and what a year of it costs
Over 4 years we moved most of our non-engineering work into engineering systems, and we ran every one of those moves as a tooling rollout rather than as a capability problem. The tooling arrived. The capability did not.
78 percent of the people who work in those systems can follow the documented steps. 31 percent can say what those steps do. That 47-point gap is the diagnosis, and our completion dashboard reports it as a success because completion is the only thing it measures.
A person who can follow steps without a model of what the steps do is a person who escalates whenever the steps stop matching what is on the screen. That happens 2.3 times a week per person, across roughly 1,150 people, and each escalation lands on an engineer.
The arithmetic is deliberately plain, because every term in it is already recorded somewhere we can check.
| Engineer hours absorbed each week | 760 |
| Fully loaded cost per hour | $145 |
| Annual cost of the interruptions alone | about $5.7M |
| Duplicate regional training spend | $410,000 a year |
| Annual cost of the gap | about $6.1M |
| Internal tickets a month whose answer already exists | about 1,900 |
Nothing in that table requires a new measurement. The interruptions are in calendars, the duplicate spend is in nine regional budgets, and the tickets are in a queue with their answers already written down somewhere else.
That is the cost of the year we just had. It is also the cost of the year ahead, because none of the four moves that produced it is being reversed and a fifth has already started.
What we are not solving
A strategy that addresses everything is a wish list, and an executive reading one correctly concludes that nobody has chosen. This program does not attempt the following, and each omission is a decision rather than an oversight.
We are not retraining the engineering organization. Their capability is not what is failing, and a program that arrives asking for their time on top of the interruptions will be read as a second tax.
We are not building a learning platform. Section 5 reaches a buy recommendation. We have no advantage in that market and no reason to acquire one.
We are not centralizing regional training budgets in year one. Nobody who bought duplicate training was wrong to: waiting for a central answer costs a site director more than buying locally. Taking that authority away before we have earned it converts our most useful constituency into an opposition, and the regional-sponsor persona refutes the plan that tries.
We are not promising a revenue number we can defend. See the ordering below.
We are not fixing the fifteen years of undocumented internal knowledge that sits behind those 1,900 tickets. That is a real content project and it is not this one.
How we will know, before what we will do
The measurement plan comes first on purpose. A reader who sees the solution before the measure has no way to tell what would count as failure, and by then the answer is designed to fit.
We report against the four levels our learning function already uses, and we say plainly which of them we can reach today.
Level 1, reaction. Whether people can use what we give them. Reachable now, from the same route counters the documentation already carries.
Level 2, learning. Whether somebody can now do the thing. This is the level the whole program turns on, because it is the level where the 47-point gap lives, and it is the one most programs skip because completion is cheaper to collect. Each unit publishes criteria you can observe, and a unit is admitted only when somebody who worked through it alone can meet them.
Level 3, behavior. Whether the work changes. Measured as escalations per person per week and as median days to a first merged change, both of which we are already able to count.
Level 4, results. What it costs the business. Engineer hours returned, tickets deflected, duplicate spend consolidated. Reachable for cost and efficiency, and honestly not reachable for revenue, which is why revenue is ordered where it is.
The failure conditions are stated with the same numbers. If escalations per person do not fall within two quarters of a site going live, the design is wrong rather than the rollout. If the completion rate rises while the understanding rate does not, we have bought more of the metric we already had too much of.
The case, in the order it will be heard
Revenue leads because it is the largest number. Cost and efficiency come second and carry the argument, because they are the ones we can prove.
We say that out loud rather than hoping nobody notices the ordering, and the reason is not modesty. Revenue attribution for an internal enablement program is the weakest evidence any learning function offers, and every function that has overclaimed it has been caught doing so by somebody in a finance review. A reader who spots the ordering before we name it stops trusting the rest of the document.
Revenue, and the weakest evidence. Solution engineers spend about 6 days per new enterprise customer teaching material a certification program would carry, across 240 new enterprise customers a year. 95 partner implementations a year depend on partner staff being certified, and each is delivery capacity we do not have to hire for.
This is the part of learning that earns rather than spends. It is also the part where attribution is an argument rather than a measurement, and we will be asked which it is.
Cost, and the easiest to verify. The $5.7M of absorbed engineer time is the largest single line and the one a skeptic can check without our help. The $410,000 of duplicate regional spend is the fastest saving available, because consolidating it needs an owner rather than a behavior change.
Efficiency, and the one that compounds. Median days to a first merged change is 41 company-wide and 58 at the 5 sites with nobody to ask. The pilot runs at nine. We are not targeting zero: a first change that takes an hour would mean the task was not real.
It compounds because the person who reaches a first merged change stops being a source of escalations and starts being somebody a colleague can ask. That is the same mechanism as the 22-point site gap, run forwards instead of backwards, and it is why the number to watch is the 5 sites without a local expert rather than the company median.
What the evidence says about the design
Four findings from the literature review shape the program, and each one rules something out.
One program for everyone is worse than none. Support that helps a beginner measurably degrades an expert's performance. The program is therefore tiered by what somebody can already do, and a single mandatory curriculum is refuted before it is proposed.
Sitting near somebody who knows beats being trained. Proximity to help predicts capability better than role does, and our own 22-point confidence gap between sites with and without a local expert is larger than the gap between job families. The unit of the program is therefore a person at a site, not a course in a catalog.
Recordings teach the steps and skip the thinking, which is the half that transfers. That is precisely how we produced a population at 78 percent completion and 31 percent understanding.
Paying people to finish a course buys finished courses. Completion is already at 78 percent. Buying more of a metric that is nearly full is the definition of spending without effect, so the incentive is built on contribution rather than completion. Somebody who answers a colleague's question, fixes a runbook, or lands a review is doing the thing we want more of, and all three are countable without a new system.
There is external evidence for the shape of this, and one figure carries it. DORA measured how much each technical capability lifts organizational performance, then split teams by the quality of their internal documentation. Version control returned 278 percent above average and 27 percent below it. The same practice, adopted by both groups, paid back roughly ten times more where the writing was good.
The commitment we owe people, and who owns it
Half our staff think this tool may cut their jobs, and 2 in five think that getting good at it would prove they can be replaced. 1 in eight would say either thing to a manager, so every conversation leadership has had about this has been with the 1 in eight.
We have already tried reassurance. We told people nobody is losing their job, and under a quarter believe we have been straight with them. Those are the same fact: a promise with no date, no scope and no notice period cannot be kept, because it was never specific enough to be broken.
The commitment is therefore a program component with a named owner, rather than a courtesy offered in a town hall. The Chief People Officer owns it, it states a period, the roles and sites it covers, and the notice we will give if the position changes, and it is published where anybody can hold us to it.
This is not a soft item. PwC surveyed 50,000 workers and found that daily AI users feel more job-secure than occasional ones, at 58 percent against 36. Fear makes people retreat to what they already know, and use appears to reduce the fear, so the two findings are a loop rather than a contradiction. Our job is to break into the loop safely and early, not to argue with it.
AI, which is a chapter and not a program
The assistants went to everybody, into the same workflows, with the same rollout playbook that produced the gap we are already paying for.
71 percent have used the assistant. 38 percent would sign their name to its output. 22 percent say they can tell which answers need checking, and that last figure is the actual skill. 57 percent have had an answer that was wrong and sounded certain, which is the best predictor we have of somebody having quietly stopped.
The market is not ahead of us here. McKinsey found 88 percent of organizations have adopted AI and 6 percent are capturing real value, and the gap did not close between their 2025 and 2026 surveys. Two years of a flat gap is a structural finding rather than an early-days one.
Stack Overflow found developer use rising from 76 to 84 percent while distrust of accuracy rose from 31 to 46 in the same population of 49,000 people. More adoption is not producing better judgment.
None of that argues for a separate AI program. It argues that the assistant is the fourth thing we have handed this population without a model of what it does, and it will not be the last. A strategy whose subject is a tool expires with the tool. This one is about the capability underneath, so the next arrival is a chapter rather than a rewrite.
Who owns it
The Chief Operating Officer funds it, because the problem crosses every function and no single department can pay for it out of its own budget.
It is not owned by engineering, and that needs saying because the VP of Engineering is its loudest advocate. It is their engineers absorbing the 760 hours, so their enthusiasm is rational. A program owned by engineering also gets built for people who think like engineers, which is the population that is not failing.
The nine regional site directors are the constituency that decides whether this works, because they are the ones currently paying for training out of their own budgets and coping without a shared owner. They are consulted on the design before it is funded, not briefed on it afterwards.
What we are asking for
$2.1M over three years, against an annual gap cost of $6.1M, decided this quarter. The model is in the budget document and the first year is recovered inside the first year, which is a lower claim than these programs usually make and one we can hold.
Procurement is already in flight and no vendor contract is signed inside the 90 days. A plan that signs inside the window tells anybody who has run a procurement that we have not.
The scaling proposal in section 6 carries the kill criteria. If the program cannot be run by somebody who did not design it, it is a consulting engagement with an internal invoice, and we should stop it rather than fund a second year.
The objections, and what answers them
7 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Two kinds of number appear below. Figures about Halden Systems and its staff come from
data/survey.yamlanddata/scenario.yaml, and are invented for a fictional company. Figures credited to DORA, Stack Overflow, McKinsey or PwC are real, were read on 2026-09-07, and are recorded indata/literature.yaml. They are never mixed without being labeled.
Bottom line. 3 objections arrive every time, all reasonable, and two of them are partly right. None of them changes the recommendation, and the answers are below.
Every strategy of this kind meets the same 3 objections, all of them reasonable, and two of them partly right. Each is answered below with the evidence rather than with reassurance.
"We do not need a study. We need to start."
This is the strongest objection and it deserves the strongest answer.
We are not proposing to wait. We are proposing to start on day one and hold the money until day sixty.
The pilot runs from day one. It already exists and costs nothing to begin. Eight other things start in the first week, and every one of them is listed below. Nothing is on hold except the spending commitment, which is the only decision that gets harder to reverse.
The external evidence
88 percent of organizations have adopted AI. 6 percent report real value. That is the action-first path, measured across the whole market by McKinsey in 2025. Two thirds have not begun to scale it. The difference between 88 and 6 is not who moved fastest.
Developer AI use went from 76 to 84 percent in one year. Distrust of its accuracy went from 31 to 46 percent. More adoption, worse outcome, same 49,000 people. Doing more of the same thing is not the missing ingredient.
The same investment returns 278 percent or 27 percent depending on documentation quality. That is DORA on version control, and the pattern repeats across every practice they studied. Acting before checking the precondition wastes the action rather than accelerating it.
The internal evidence, which is harder to argue with
Scenario figures from here to the end of this section.
We have already run this play three times.
Three migrations, each with a demonstration, a recording and a channel. 29 percent rate the training useful. Around $410,000 a year goes to regional training bought twice because waiting for a central answer cost more than paying locally.
The bias for action has already been exercised here. That is the invoice.
One more thing worth saying out loud
You get one credible relaunch. A second attempt at the same population meets more resistance than the first, because people remember. Choosing wrong in week one means defending it for two years.
"30 days for a survey is too long."
Correct, and the plan changed rather than being defended. There is real evidence on day fifteen, and the table below shows how it arrives that early.
| When | What lands |
|---|---|
| Days 1 to 5 | Data we already have: tickets, calendars, training spend, completion reports |
| Days 3 to 8 | Twenty interviews. Enough to know what to ask |
| Days 8 to 15 | A six-question pulse to a sample, open for 7 days |
| Day 15 | First readout. Numbers, quotes, and a pilot already running |
| Days 16 to 30 | The full instrument, 10 days in the field |
| Day 30 | It closes, and deepens the picture rather than replacing it |
The pulse is 6 questions open for 7 days, which is short enough that people answer it. The long instrument is a second pass that deepens the picture, and by the time it lands we have already been acting for a month.
"This is a training problem. Buy training."
Partly right, and the part that is wrong is the expensive half, because it buys more of the thing the organization already has too much of.
Completion is already at 78 percent. Understanding is at 31. We can buy more completion and it will change nothing, because completion was never the thing in short supply.
16 percent say a reward for finishing a course would motivate them. The population has worked this out already.
What starts in week one
Nine things start immediately, none of which waits for the survey to close or needs a budget line approved first.
- The runbook pilot, at 1 site with a local expert and one without. That comparison is what the scaling case will need.
- Stop reporting completion. While the dashboard says 78 percent, nobody believes there is a problem.
- Cancel one competing initiative. Two thirds of staff report more changes than they can absorb. Removing one is the only credible way to say this matters.
- Publish a checkable commitment on roles. A date, a scope, and notice. Not a reassurance.
- Name local experts at the five stranded sites, with time in the role rather than a favor asked.
- Open a wrongness log for cases where the assistant was confident and wrong. Turns private burns into shared knowledge.
- Senior people ask basic questions in public, weekly, in writing.
- Change the question managers ask from "did you finish it" to "what did you get stuck on".
- Pull the existing data. Tickets, interruptions and regional spend are sitting there already.
Five of those nine, items 2, 3, 4, 7 and 8, are decisions rather than projects, which means they can be made in an afternoon by somebody who already has the authority to make them.
What we deliberately do not do fast
We do not take the regional training budgets. That $410,000 is the most obvious saving in the case and it is the one that would cost us the 5 site directors who decide whether this reaches their people.
They buy locally because waiting for a central answer costs them more. They stop when something arrives faster than their own purchasing does, and not before.
We do not sign a vendor inside 90 days. A response window is 40 days and reference calls run longer. A plan that ends in a signature tells anyone who has bought software that we have not.
What to change that is not training
8 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Two kinds of number appear below. Figures about Halden Systems and its staff come from
data/survey.yamlanddata/scenario.yaml, and are invented for a fictional company. Figures credited to DORA, Stack Overflow, McKinsey or PwC are real, were read on 2026-09-07, and are recorded indata/literature.yaml. They are never mixed without being labeled.
Bottom line. Ten changes worth making that are not training. Most cost nothing, and the cheapest are the ones that decide whether any training lands at all.
10 recommendations. None of them is a course. Most cost nothing, and the cheap ones are the ones that move the numbers.
The reasoning is the same each time. People are behaving sensibly given what the company rewards and what it has told them. Change those and the behavior follows. Leave them and no amount of good material will help.
The four that matter most
1. Make a promise about roles that can be checked
We told people nobody is losing their job. Under a quarter believe we have been straight with them.
The research on job worry is blunt about why: a reassurance you cannot check reads as a reason to worry. It raises suspicion rather than settling it.
Replace it with something falsifiable. A scope, a date, and a commitment to give notice if it changes. "No role will be cut as a direct result of adopting the assistant before June, and if that changes you hear it from us 90 days ahead."
Half the staff think this tool may cut their jobs. 2 in five think that getting good at it proves they can be replaced. Until that is answered, we are asking people to build the case against themselves.
2. Cancel something
Two thirds report more simultaneous changes than they can absorb. Three quarters have no learning time that is not taken from delivery.
Asking for hours that do not exist gets ignored by people acting sensibly, and then gets reported as a motivation problem.
Remove one competing initiative. Nothing says this matters like taking something away to make room for it. It costs a decision and no money.
3. Stop reporting completion
78 percent can follow the steps. 31 percent can say what the steps do.
While the dashboard shows 78, the organization believes the problem is solved. The number is not wrong. It is measuring the wrong thing, and its existence is the obstacle.
Retire it. Report work product instead: pages corrected, changes merged, questions answered.
4. Put helping into promotion, not into badges
19 percent feel their effort to learn is noticed by anyone who matters. It is the lowest number in the survey.
61 percent would answer colleagues' questions if that were visible and credited. Three times as many.
That is the cheapest capacity in the whole program, and it is unlocked by recognition rather than by hiring.
One warning from the research. Rewards tied to finishing something reduce the motivation people already had, and the effect lasts after the reward stops. Do not pay for completion. Credit the contribution, in the system where credit actually counts, which is promotion.
The other six
5. Have senior people ask basic questions in public
Half of staff have avoided asking a question they needed answered, because of how it would look. That is worse among people with long service.
Asking is learned from watching. If the most senior person at a site asks something basic in writing every week, the cost of asking drops for everyone below them. It costs nothing.
6. Keep a wrongness log
57 percent have had the assistant give them an answer that was wrong and sounded certain. Two thirds of developers surveyed elsewhere report the same thing.
Right now each of those is a private event that teaches one person to stop using the tool.
Collect them in one open page. It turns individual burns into shared knowledge, and it teaches the one skill nobody has taught: telling which answers need checking. 22 percent can do that today.
7. Write it down before you decide it
GitLab runs a public handbook on one rule: if it is not written down, it is not real. A change is documented before it is announced.
That is a norm, not a course. It makes writing load-bearing, which is the only thing that keeps documentation alive.
8. Look it up before you ask
The other half of the same practice. It sounds unfriendly and does the opposite: it means the written answer has to be good enough to find, so the effort moves to where it helps everyone.
Pair it with rule 5 or it will read as a rebuke.
9. Give the local experts a title and hours
Five of 9 sites have nobody who can answer a question during the working day. Those sites are 17 days slower to a first merged change.
Name a person at each, and give them a fifth of their week. Unpaid goodwill decays, and then the fix looks like it failed.
10. Write for the second language by default
People working in English as a second language report confidence 14 points lower, after allowing for which site they are at.
Prose pitched higher than it needs to be is a tax on them and on everyone else. This is a writing standard, applied to what already exists. It is not a program and it does not need a budget.
What this costs
Six of the ten are decisions. They can be made in an afternoon and cost nothing: the role promise, canceling an initiative, retiring the completion metric, senior people asking in public, the two GitLab norms.
Two are cheap and ongoing: the wrongness log and the writing standard.
Two need real money: the local expert time, and changing promotion criteria, which needs the people team rather than a purchase.
Not one of them is training, and together they move more than the training will.
Project plan: closing the capability gap
9 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
For Halden Systems, across 9 sites.
Bottom line. 4 phases over 18 months, each ending at a gate with a named decider and the evidence they decide on. Phase 1 is already running and costs nothing. No phase begins before the gate ahead of it closes, and 2 of the 4 gates can stop the program.
The phases
| Phase | Months | What it produces | Gate | Who decides |
|---|---|---|---|---|
| 1. Pilot and instrument | 0 to 3 | Dublin result, survey, baseline escalation rate | Is the gap real and measurable | Director of Enablement |
| 2. Prove at 2 sites | 4 to 8 | The unit run twice by people who did not design it | Does it work without its author | COO |
| 3. Procure and build | 9 to 13 | Vendor signed, content and assessment in place | Is the platform earning its cost | COO with Procurement |
| 4. Scale to 9 sites | 14 to 18 | Local capability at every site | Is this now a permanent function | COO with the site directors |
Phase 1 is running. Dublin has already produced its result, and the survey is in the field, so the first gate is about reading evidence rather than waiting for it.
The gates, and what closes them
A gate is not a review meeting. Each one names a number, and the number is available before the meeting.
Gate 1, is the gap real. Closes on the survey showing a completion-to-understanding gap above 30 points, and on the baseline escalation rate being measurable from calendars and ticket queues rather than estimated. Opens if both hold. Stops the program if the gap is under 15 points, because at that size the cost of the program exceeds the cost of the problem.
Gate 2, does it work without its author. Closes on 2 sites running the unit with no involvement from the person who designed it, and on median days to a first merged change falling below 20 at both. Stops the program if either site needs the designer to run it, because a program that cannot be handed over is a consulting engagement with an internal invoice.
Gate 3, is the platform earning its cost. Closes on the vendor meeting the acceptance criteria in the statement of work. Does not stop the program: a failure here changes the vendor or reverts to building, and the capability work continues either way.
Gate 4, is this permanent. Closes on escalations per person per week falling below 1.2 across the inner population, sustained for 2 quarters. A failure here does not stop the program; it stops the expansion and returns the function to maintaining what works.
Dependencies, including the ones we do not control
| Dependency | Owner | Outside our control |
|---|---|---|
| Survey response rate above 40% | Director of Enablement | Partly. Site directors decide whether to encourage it |
| Baseline escalation data from calendars and tickets | VP Engineering | No |
| A named local expert at each of the 4 sites that have one | Site directors | Yes. They are volunteering their people |
| Vendor response window | Procurement | Yes. 3 to 6 weeks and it does not compress |
| Legal review of the data-processing terms | General Counsel | Yes. Queue depth is not ours to set |
| The job commitment being published | Chief People Officer | Yes. This is the one that can quietly not happen |
The last row is the dependency most likely to be missed, because nothing breaks visibly when it slips. It is on the risk register for that reason.
Risks, with owners
| Risk | Owner | What we do about it |
|---|---|---|
| The program is not needed. The gap closes on its own as people learn by exposure | Director of Enablement | Gate 1 measures it. Under 15 points we stop and say so |
| Site directors read this as a budget grab | COO | Regional budgets stay regional through year 1, stated in the strategy |
| The job commitment is never published | Chief People Officer | Gate 2 does not close without it |
| The vendor is late or fails acceptance | Procurement | Payment is tied to acceptance, not dates. Build-against-buy is revisited rather than absorbed |
| Engineering resents the time cost | VP Engineering | The program removes engineer hours rather than asking for them. Measured monthly and reported |
| Assistance rollout outruns capability again | Director of Enablement | No new tool reaches the outer population without a capability check first |
| The designer leaves | COO | Gate 2 exists to prove this does not matter |
The first risk is the one most plans omit. A program whose own plan cannot describe the condition under which it should not exist is asking for trust rather than a decision.
What runs in parallel, and what cannot
Content authoring, assessment design and the local-expert network all run from month 4 and do not wait for procurement. That is deliberate: the parts that depend on a vendor are the parts we should be least attached to.
Procurement cannot start before gate 1, because an evaluation matrix written before the survey would be weighted against a problem we had guessed at.
Scaling cannot start before gate 2. Running the unit at 9 sites when it has worked at 1 is how a program that worked becomes a program nobody trusts.
Statement of work: learning platform and content services
10 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Between Halden Systems and the selected supplier.
Bottom line. 4 deliverables, each with acceptance criteria a third party could apply without us in the room. Payment follows acceptance, never a date. What is out of scope is listed at the same length as what is in, because that is the half every dispute turns on.
Deliverables and acceptance criteria
Acceptance is a test somebody outside both organizations could run. Where a criterion needs judgment, it names who exercises it and against what.
| # | Deliverable | Accepted when |
|---|---|---|
| D1 | Platform configured for 9 sites | 3 named staff at each site complete a scripted task list unaided, in their own timezone, on their own device |
| D2 | Content migration of 140 existing units | 100% render without manual repair, and a random sample of 20 passes our accessibility standard |
| D3 | Assessment engine wired to our identity provider | A revoked account loses access within 15 minutes, verified by our security team |
| D4 | Reporting against the 4 measurement levels | Every figure in the strategy document can be produced from the platform without a spreadsheet |
D4 is the one vendors negotiate hardest and the one we hold. A platform that cannot report level 2 leaves us measuring completion again, which is the failure the program exists to correct.
Out of scope
Listed at length on purpose. A vendor who has read this and signed cannot later present any of it as a change order, and we cannot later expect it for free.
Content authoring. The vendor migrates what exists and builds none of it. Our staff write the material, because the material is the capability and outsourcing it defeats the program.
Anything touching production engineering systems. The platform reads our identity provider and nothing else. No repository access, no pipeline integration, no write path into an engineering tool. This is a security boundary rather than a preference.
Translation and localization. 4 continents and we are still delivering in English. That is a known gap, it is recorded in the strategy, and it is not being solved by this contract.
Change management and internal communications. The vendor does not talk to our staff. Adoption is our problem and buying it produces a rollout with no owner inside the company.
Success metrics beyond the 4 levels. No engagement scores, no leaderboards, no completion dashboards presented as outcomes. We already have too much of that number.
Data migration out. Export is in scope. Anything the vendor would do to help a successor is not, and we have priced our own exit accordingly.
Payment against acceptance
| Milestone | Share | Released on |
|---|---|---|
| Contract signature | 10% | Signature |
| D1 accepted | 25% | Acceptance test passed at all 9 sites |
| D2 and D3 accepted | 30% | Both, not either |
| D4 accepted | 25% | Acceptance, plus 30 days of reporting we did not have to correct |
| Retention | 10% | 90 days after D4, held against defects |
No milestone is released on a date. A vendor who is late is late, and a vendor who is late and paid has been told that the schedule is ours to worry about.
The 30-day tail on D4 exists because reporting is the deliverable most likely to pass a demo and fail in use.
When acceptance fails
First failure. The vendor has 15 working days to remedy, at their cost. The clock on dependent milestones stops. This is expected at least once and is not a dispute.
Second failure on the same deliverable. We take a 15% reduction on that milestone and either accept the remedy or move the deliverable out of scope with a matching reduction. Which of those happens is our choice, not theirs.
Third failure, or a failure on D3. Termination for cause, retention forfeit, and the export obligation triggers immediately. D3 is singled out because a security boundary that fails twice is not a quality problem.
If we cause the failure. Where acceptance fails because we did not supply access, data or people on time, the remedy period pauses and the vendor is entitled to a schedule extension of the same length. Stating this protects the relationship: a buyer who never admits a cause is a buyer vendors price defensively.
Governance after signature
Monthly against the scorecard in the budget and vendor document. The relationship after signature is where most of the money is actually lost, and a statement of work that ends at acceptance has described the cheapest part of the contract.
Budget model and vendor evaluation
11 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. $2.1M over 3 years against a gap costing Halden Systems $6.1M a year. We recommend buying the platform and building the content, because the content is the capability and the platform is not. The evaluation weights below were fixed before any vendor was approached.
Total cost of ownership, 3 years
Internal time is in the model. A buy that removes a license cost and adds 0.5 of a person has moved the money rather than saved it, and a model that leaves internal effort out is the standard way that gets hidden.
| Year 1 | Year 2 | Year 3 | Total | |
|---|---|---|---|---|
| Platform licences | $180,000 | $195,000 | $210,000 | $585,000 |
| Implementation and migration | $145,000 | $145,000 | ||
| Content authoring, internal | $310,000 | $180,000 | $150,000 | $640,000 |
| Program staff, 2.5 FTE | $220,000 | $230,000 | $240,000 | $690,000 |
| Local expert time, 9 sites | $0 | $0 | $0 | $0 |
| Total | $855,000 | $605,000 | $600,000 | $2,060,000 |
The zero row is deliberate and is the number a reviewer should push on. Local expert time is real and it is not new money: those people already answer these questions, at 760 engineer hours a week. The program moves that time rather than adding it, and if it does not, the model is wrong and gate 4 will show it.
Consolidating the $410,000 of duplicate regional spend covers most of year 2 and year 3 on its own. We have not netted it off above, because a saving counted inside the cost of the thing producing it is how a budget stops being checkable.
Build against buy
| Build | Buy | |
|---|---|---|
| 3-year cost | $1.74M | $2.06M |
| Time to first site live | 11 months | 4 months |
| Ongoing engineering ownership | 1.5 FTE, permanently | none |
| Accessibility conformance | ours to prove | contractual, and testable |
| Exit cost | none | 1 quarter of export and re-platforming |
We recommend buying. Building is cheaper on paper and the paper is wrong in 2 places. It assumes the 1.5 FTE of permanent engineering ownership is available, which it is not while those same engineers are absorbing 760 hours a week. It also puts an internal team between every content change and the people who need it, which is the exact shape of the problem we are trying to fix.
The content is the opposite case and we build it. It encodes how our systems actually work, it is where the capability lives, and a vendor writing it would produce material that teaches a generic tool rather than ours.
Evaluation matrix, weights fixed 2026-09-07
These weights were set before any vendor was approached, and this file's history shows it. Weights chosen after seeing submissions are a justification rather than an evaluation, and a procurement reader can tell the difference immediately.
| Criterion | Weight | Why it carries that weight |
|---|---|---|
| Level 2 reporting: can it show understanding, not completion | 25 | The whole program turns on this measure |
| Accessibility conformance, evidenced | 20 | Non-negotiable, and cheaper to demand than to retrofit |
| Asynchronous delivery across 17 hours | 15 | No shared working hour exists |
| Identity integration and revocation | 15 | A security boundary rather than a convenience |
| Export and exit cost | 10 | The clause nobody reads until they need it |
| Content migration effort | 10 | Real, and one-off |
| License cost | 5 | The number everybody optimizes and the smallest line |
License cost is weighted last on purpose. It is 28% of the model and the easiest thing to negotiate after selection, and weighting it heavily selects for the vendor best at discounting rather than the one best at level 2.
Scorecard after signature
Reviewed monthly by the Director of Enablement, reported quarterly to the COO. This is where the money is actually lost.
| Measure | Target | Trigger |
|---|---|---|
| Support tickets we raise, resolved within SLA | 90% | 2 consecutive months below triggers escalation |
| Platform availability during any site's working hours | 99.5% | Any month below triggers a service credit |
| Reporting accuracy: figures we did not have to correct | 100% | Any correction is logged and reviewed |
| Roadmap items delivered against commitment | 70% | 2 quarters below opens the exit clause |
| Accessibility regressions | 0 | Any regression pauses the next payment |
The last row has no tolerance because a regression there breaks a commitment we made to our own staff, and a target with a tolerance is a target that will be spent.
Scaling proposal
12 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
How the program reaches all 9 Halden Systems sites.
Bottom line. This has to run without the person who designed it. One unit is 1 local expert supporting 25 people for 6 weeks, at about $4,100 of time nobody is currently accounting for. 9 sites need 46 units over 18 months. If the unit cannot be run by somebody who did not design it, we stop, and the kill criteria below say when.
What one unit is
The unit is a person at a site, not a course in a catalog. That follows from the evidence: a 22-point confidence gap between sites with a local expert and sites without one is wider than the gap between job families, so proximity to help predicts capability better than role does.
| Participants | 25, drawn from the inner population |
| Duration | 6 weeks |
| Local expert time | 4 hours a week |
| Participant time | 2 hours a week, in their own timezone |
| New content required | none after month 6 |
| Direct cost | $0 |
| Cost in time, at loaded rates | about $4,100 |
The direct cost is genuinely zero and the time cost is genuinely $4,100. Reporting only the first is how programs get approved and then quietly starve.
Dublin ran this once already and produced median days to a first merged change of 9 against a company median of 41.
The train-the-trainer path
A unit is run by somebody who was in the previous unit. That is the mechanism, and it is why the program compounds rather than scaling linearly with the founder's calendar.
| Stage | What the new expert does | Who supports them |
|---|---|---|
| 1 | Completes a unit as a participant | The site's existing expert |
| 2 | Co-runs the next unit, taking half the sessions | The same expert |
| 3 | Runs a unit alone, with a weekly 30-minute check-in | The enablement function |
| 4 | Supports somebody at stage 2 | Nobody |
A site is self-sufficient when it has 2 people at stage 4. 4 sites have somebody at stage 1 today, which is where the 22-point gap comes from, and 5 have nobody at all.
The material a new expert needs is a runbook, a session outline and the assessment. All 3 exist already, which is what makes stage 3 a real handover rather than an aspiration.
What breaks at 10 times the size
Named specifically, because a caveat that says "scaling brings challenges" tells a reader nothing and costs whoever wrote it nothing.
The check-in does not scale and fails silently. At stage 3 the enablement function spends 30 minutes a week per new expert. At 46 concurrent units that is 23 hours a week against a function of 2.5 people. It breaks somewhere around 20 concurrent units, and the failure mode is not a crisis, it is check-ins quietly being skipped. Fix: stage 4 experts take the check-ins from unit 15 onward, which is why stage 4 exists.
Content goes stale faster than 2.5 people can revise it. 140 units against 4 engineering systems that ship continuously. At the current change rate, roughly 12% of the content is wrong at any moment. Fix: content ownership moves to the teams that own each system, with the enablement function holding the standard rather than the text.
Assessment integrity. At 25 people a facilitator knows whether somebody understood. At 1,150 they do not, and the assessment becomes the only signal at exactly the point where it starts being gamed. Fix: assessment against a real change in a real repository, which cannot be answered from a recording.
The 5 sites with nobody to ask stay last. Every scaling plan reaches the sites that already have an expert first, because those go faster, and the 22-point gap is at the other 5. Fix: the sequence is fixed in the project plan and starts at 2 of the 5, which costs us speed in year 1 on purpose.
Kill criteria
The conditions under which we stop. A scaling proposal with no such condition is a budget request wearing a plan's clothes.
| We stop when | Measured | Why this one |
|---|---|---|
| A unit cannot be run by a stage 3 expert without the designer | Gate 2, month 8 | The program is a consulting engagement with an internal invoice |
| Escalations per person do not fall below 1.8 within 2 quarters of a site going live | Monthly, from calendars | The mechanism does not work here, and more of it will not help |
| Completion rises while understanding does not | Every unit | We have bought more of the metric that was already full |
| Fewer than 2 sites reach stage 4 by month 18 | Gate 4 | It does not propagate, so it is a service and should be costed as one |
| The gap closes without us | Annual survey | The program is not needed, which is a good outcome and still a stop |
The last row is the one that matters. A program that cannot describe the world in which it should not exist is asking for trust rather than a decision, and the person being asked can tell.
Stopping is not failure in 2 of these 5 cases. It is the finding.
Executive readout: the capability gap
13 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. $2.1M over three years, against a capability gap costing $6.1M a year. The decision asked for today is funding, this quarter.
Two kinds of number appear below. Figures about Halden Systems are invented, and come from
data/scenario.yamlanddata/survey.yaml. Figures credited to DORA, Stack Overflow, McKinsey or PwC are real, were read on 2026-09-07, and are recorded indata/literature.yaml.
9 slides for the COO, around day 65. The strategy document is the argument; this is the version that survives a thirty-minute meeting with a finance review after it. Everything that is not a decision is in the appendix, and the appendix is a separate file on purpose. A deck that contains the strategy document is a strategy document nobody read.
Slide 1. The ask
We are asking for $2.1M over three years to close a capability gap that costs us $6.1M a year.
The decision today: fund it this quarter, or tell us to keep absorbing the cost.
Procurement is already in flight. No vendor contract is signed, and none will be inside these 90 days.
Speaker note: the amount and the decision go on the first slide because an executive who has to wait 6 slides to learn what is being asked spends those 6 slides guessing.
Slide 2. What is actually broken
78 percent can follow the steps. 31 percent can say what the steps do.
Four years ago we moved most non-engineering work into engineering systems. Every move was run as a tooling rollout. The tooling arrived and the capability did not.
Our dashboard reports the 78 and calls it success.
Slide 3. What it costs to leave alone
| Engineer hours absorbed each week | 760 |
| At $145 fully loaded | $5.7M a year |
| Duplicate regional training spend | $410,000 a year |
| Tickets a month whose answer already exists | 1,900 |
Every figure is already in a calendar, a budget or a queue. None of it needs a new measurement, and none of it is a projection.
Slide 4. Where the value is, in order of size
Revenue. Six solution-engineer days per enterprise customer, 240 customers a year. 95 partner implementations that need certified partner staff.
Cost. $5.7M in engineer time. $410,000 in duplicate spend.
Efficiency. 41 median days to a first merged change, and 58 at the sites with nobody to ask. The pilot runs at nine.
Speaker note: revenue opens the room because it is the largest number. Say the next line out loud, because somebody in the finance review will say it first if we do not.
Slide 5. The number we will not defend
Revenue attribution for an internal program is the weakest evidence we have.
We lead on it because it is the largest, and we rest the case on cost and efficiency because those are the ones a skeptic can check without our help.
Every learning function that has overclaimed revenue has been caught doing it. We would rather be the one that said so first.
Slide 6. Why more training will not fix it
Four findings, each ruling something out.
- One program for everyone measurably degrades the experts in it.
- Proximity to help predicts capability better than role does. Our own gap between sites with and without a local expert is 22 points, wider than the gap between job families.
- Recordings teach the steps and skip the thinking, which is exactly the population we have.
- Paying people to finish courses buys finished courses. We are at 78 percent completion already.
So the incentive is built on contribution, not completion.
Slide 7. AI is a chapter, not a second program
71 percent have used the assistant. 38 percent would sign their name to its output. 22 percent can tell which answers need checking.
That last figure is the skill, and it is the one nobody measured before rollout.
McKinsey: 88 percent of organizations have adopted AI, 6 percent capture real value, and the gap did not move between 2025 and 2026.
A strategy about a tool expires with the tool. This one absorbs the next arrival instead of being rewritten for it.
Slide 8. How you will know it worked
| Level | What we watch | Reachable |
|---|---|---|
| 1. Reaction | Can people use what we give them | Now |
| 2. Learning | Can somebody now do the thing | Now, and it is the point |
| 3. Behavior | Escalations per person, days to first change | Now |
| 4. Results | Engineer hours, tickets, duplicate spend | Cost and efficiency, yes. Revenue, no. |
And how you will know it failed. Escalations do not fall within two quarters of a site going live. Failure also looks like completion rising while understanding does not, which means we bought more of the metric that is already full.
Slide 9. The decision
Fund $2.1M over three years, this quarter.
What you get in 90 days: the funding decision made, procurement in flight, the pilot running, and a named owner on the job commitment we owe people.
What you do not get: a signed vendor contract. That is deliberate.
If it cannot be run by somebody who did not design it, we will tell you and stop it. The kill criteria are in the scaling proposal, and they are ours rather than yours to enforce.
Appendix, which is not presented
Kept separate so the deck stays 9 slides. Reach for these only if asked.
- A1. Survey instrument, sample and response rates, including the items inconvenient for the
program.
data/survey.yaml. - A2. The 22-point site confidence gap, by site.
- A3. Three-year total cost of ownership, and the build-against-buy analysis.
- A4. The evaluation matrix, with weights fixed before any vendor was named.
- A5. The literature review, and the one claim we refuse to cite.
- A6. What we are deliberately not solving, and why each was cut.
What this set cost to produce
14 of 14 · read this one on its own page
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. What the set represents, and what assistance changed. About 168 working hours unassisted and 69 assisted, a factor of 2.4. The calendar barely moves: 133 elapsed days against 100, a factor of 1.3, because most of the calendar is waiting for other people.
The set is the 11 documents describing Halden Systems, listed in the table below.
Every figure below is an estimate, not a measurement. An estimate presented as a measurement is the same failure as an unverified statistic.
The headline, with the qualifier attached
| Unassisted | Assisted | Factor | |
|---|---|---|---|
| Working hours | 168 | 69 | 2.4 |
| Elapsed days | 133 | 100 | 1.3 |
Those two rows describe the same work, and the gap between them is the point.
Per artifact, which is where the pattern is visible:
| Artifact | Hours | Hours, assisted | Days | Days, assisted |
|---|---|---|---|---|
| Literature review and executive summary | 24 | 8 | 10 | 4 |
| Survey design, fielding and analysis | 40 | 22 | 35 | 31 |
| Personas derived from survey themes | 12 | 4 | 4 | 2 |
| Strategy document | 20 | 7 | 12 | 6 |
| Project plan | 14 | 5 | 8 | 5 |
| Statement of work | 10 | 4 | 6 | 4 |
| Budget model and vendor evaluation | 24 | 10 | 45 | 40 |
| Scaling proposal | 14 | 5 | 8 | 5 |
| Executive presentation | 10 | 4 | 5 | 3 |
| Total | 168 | 69 |
Assistance compresses the hours somebody spends. It barely moves the calendar, because most of the calendar is waiting: 3 weeks of a survey in the field, a vendor response window, references who answer when they answer.
Reporting the first number as delivery speed is how this claim is usually overstated. Anyone who has run a procurement will catch it in the room.
The honest version is more useful anyway. It says where to put the assistance and where not to bother.
Where the compression is, and where it is not
The survey is the clearest case in the set: 45 percent fewer working hours and 11 percent less clock time. The instrument can be drafted in an afternoon. The 3 weeks in the field do not move, and the pilot before it does not move either.
The budget and vendor work has the longest elapsed time of anything here and is the least affected. 45 days becomes forty. The response window is the response window.
The literature review compresses hardest on the finding-and-summarizing and least on the part that matters. Assistance is good at locating sources and stating what they claim. Deciding which are load-bearing, and which are a vendor citing its own survey back to itself, requires having read them. That reading is most of the 8 hours, and skipping it is how unverified figures enter a document and stay there for a decade.
What did not compress at all
Every artifact in data/effort.yaml carries a line naming the part assistance did not touch, and
they have a shape in common. Each one is a decision rather than a draft.
Knowing that revenue is the largest number in the case and the weakest evidence in it. Choosing to say so on the slide, rather than hoping nobody asks.
Deciding what each persona rules out. That is the difference between six plausible people and an instrument that refutes a bad proposal.
Writing the out-of-scope section of a contract, which is specific to what this buyer will assume. Writing kill criteria for your own program. Cutting the deck.
That distinction is the argument this portfolio is making about AI, and it is the same argument the program in the scenario makes to 3,000 people: assistance compresses drafting and synthesis, and does not compress deciding. A leader who cannot tell those apart will either refuse the tool or over-trust it, and both failures are visible in the survey data.
Why this belongs in a summary a recruiter reads
Two reasons, and the second is the real one.
It sets the scale. A reader can tell in one line whether they are looking at a weekend's work or a month's, which is otherwise guessed from polish.
It also demonstrates the competence being hired for. A director of learning will be asked what these tools do to their function's capacity, by an executive who has heard a tenfold claim and does not believe it. The answer here is 2.4 on hours, 1.3 on the calendar, with the boundary named. That answer is worth more in the room than a larger number would be, because it can survive the follow-up question.