Capability strategy
Written by a person. Last read by a person on 2026-09-07, 1 day ago. Its facts were checked by the eval suite on 2026-09-07.
You are deciding whether to fund this, or you want the spine of the argument in one document.
Scenario-based sample. Halden Systems is invented, and so is every figure about it.
Bottom line. The 1,150 non-technical staff who work in our engineering systems can follow the steps and cannot read the system, which costs about $6.1M a year in absorbed engineer time and duplicated spend. We are asking for $2.1M over three years to close it, decided this quarter.
Two kinds of number appear below. Figures about Halden Systems and its staff come from
data/scenario.yamlanddata/survey.yaml, and are invented for a fictional company. Figures credited to DORA, Stack Overflow, McKinsey or PwC are real, were read on 2026-09-07, and are recorded indata/literature.yaml. They are never mixed unlabeled.
For the Chief Operating Officer. Prepared in the first 90 days, and the decision it asks for is a funding commitment, not an approval to plan.
The gap, and what a year of it costs
Over 4 years we moved most of our non-engineering work into engineering systems, and we ran every one of those moves as a tooling rollout rather than as a capability problem. The tooling arrived. The capability did not.
78 percent of the people who work in those systems can follow the documented steps. 31 percent can say what those steps do. That 47-point gap is the diagnosis, and our completion dashboard reports it as a success because completion is the only thing it measures.
A person who can follow steps without a model of what the steps do is a person who escalates whenever the steps stop matching what is on the screen. That happens 2.3 times a week per person, across roughly 1,150 people, and each escalation lands on an engineer.
The arithmetic is deliberately plain, because every term in it is already recorded somewhere we can check.
| Engineer hours absorbed each week | 760 |
| Fully loaded cost per hour | $145 |
| Annual cost of the interruptions alone | about $5.7M |
| Duplicate regional training spend | $410,000 a year |
| Annual cost of the gap | about $6.1M |
| Internal tickets a month whose answer already exists | about 1,900 |
Nothing in that table requires a new measurement. The interruptions are in calendars, the duplicate spend is in nine regional budgets, and the tickets are in a queue with their answers already written down somewhere else.
That is the cost of the year we just had. It is also the cost of the year ahead, because none of the four moves that produced it is being reversed and a fifth has already started.
What we are not solving
A strategy that addresses everything is a wish list, and an executive reading one correctly concludes that nobody has chosen. This program does not attempt the following, and each omission is a decision rather than an oversight.
We are not retraining the engineering organization. Their capability is not what is failing, and a program that arrives asking for their time on top of the interruptions will be read as a second tax.
We are not building a learning platform. Section 5 reaches a buy recommendation. We have no advantage in that market and no reason to acquire one.
We are not centralizing regional training budgets in year one. Nobody who bought duplicate training was wrong to: waiting for a central answer costs a site director more than buying locally. Taking that authority away before we have earned it converts our most useful constituency into an opposition, and the regional-sponsor persona refutes the plan that tries.
We are not promising a revenue number we can defend. See the ordering below.
We are not fixing the fifteen years of undocumented internal knowledge that sits behind those 1,900 tickets. That is a real content project and it is not this one.
How we will know, before what we will do
The measurement plan comes first on purpose. A reader who sees the solution before the measure has no way to tell what would count as failure, and by then the answer is designed to fit.
We report against the four levels our learning function already uses, and we say plainly which of them we can reach today.
Level 1, reaction. Whether people can use what we give them. Reachable now, from the same route counters the documentation already carries.
Level 2, learning. Whether somebody can now do the thing. This is the level the whole program turns on, because it is the level where the 47-point gap lives, and it is the one most programs skip because completion is cheaper to collect. Each unit publishes criteria you can observe, and a unit is admitted only when somebody who worked through it alone can meet them.
Level 3, behavior. Whether the work changes. Measured as escalations per person per week and as median days to a first merged change, both of which we are already able to count.
Level 4, results. What it costs the business. Engineer hours returned, tickets deflected, duplicate spend consolidated. Reachable for cost and efficiency, and honestly not reachable for revenue, which is why revenue is ordered where it is.
The failure conditions are stated with the same numbers. If escalations per person do not fall within two quarters of a site going live, the design is wrong rather than the rollout. If the completion rate rises while the understanding rate does not, we have bought more of the metric we already had too much of.
The case, in the order it will be heard
Revenue leads because it is the largest number. Cost and efficiency come second and carry the argument, because they are the ones we can prove.
We say that out loud rather than hoping nobody notices the ordering, and the reason is not modesty. Revenue attribution for an internal enablement program is the weakest evidence any learning function offers, and every function that has overclaimed it has been caught doing so by somebody in a finance review. A reader who spots the ordering before we name it stops trusting the rest of the document.
Revenue, and the weakest evidence. Solution engineers spend about 6 days per new enterprise customer teaching material a certification program would carry, across 240 new enterprise customers a year. 95 partner implementations a year depend on partner staff being certified, and each is delivery capacity we do not have to hire for.
This is the part of learning that earns rather than spends. It is also the part where attribution is an argument rather than a measurement, and we will be asked which it is.
Cost, and the easiest to verify. The $5.7M of absorbed engineer time is the largest single line and the one a skeptic can check without our help. The $410,000 of duplicate regional spend is the fastest saving available, because consolidating it needs an owner rather than a behavior change.
Efficiency, and the one that compounds. Median days to a first merged change is 41 company-wide and 58 at the 5 sites with nobody to ask. The pilot runs at nine. We are not targeting zero: a first change that takes an hour would mean the task was not real.
It compounds because the person who reaches a first merged change stops being a source of escalations and starts being somebody a colleague can ask. That is the same mechanism as the 22-point site gap, run forwards instead of backwards, and it is why the number to watch is the 5 sites without a local expert rather than the company median.
What the evidence says about the design
Four findings from the literature review shape the program, and each one rules something out.
One program for everyone is worse than none. Support that helps a beginner measurably degrades an expert's performance. The program is therefore tiered by what somebody can already do, and a single mandatory curriculum is refuted before it is proposed.
Sitting near somebody who knows beats being trained. Proximity to help predicts capability better than role does, and our own 22-point confidence gap between sites with and without a local expert is larger than the gap between job families. The unit of the program is therefore a person at a site, not a course in a catalog.
Recordings teach the steps and skip the thinking, which is the half that transfers. That is precisely how we produced a population at 78 percent completion and 31 percent understanding.
Paying people to finish a course buys finished courses. Completion is already at 78 percent. Buying more of a metric that is nearly full is the definition of spending without effect, so the incentive is built on contribution rather than completion. Somebody who answers a colleague's question, fixes a runbook, or lands a review is doing the thing we want more of, and all three are countable without a new system.
There is external evidence for the shape of this, and one figure carries it. DORA measured how much each technical capability lifts organizational performance, then split teams by the quality of their internal documentation. Version control returned 278 percent above average and 27 percent below it. The same practice, adopted by both groups, paid back roughly ten times more where the writing was good.
The commitment we owe people, and who owns it
Half our staff think this tool may cut their jobs, and 2 in five think that getting good at it would prove they can be replaced. 1 in eight would say either thing to a manager, so every conversation leadership has had about this has been with the 1 in eight.
We have already tried reassurance. We told people nobody is losing their job, and under a quarter believe we have been straight with them. Those are the same fact: a promise with no date, no scope and no notice period cannot be kept, because it was never specific enough to be broken.
The commitment is therefore a program component with a named owner, rather than a courtesy offered in a town hall. The Chief People Officer owns it, it states a period, the roles and sites it covers, and the notice we will give if the position changes, and it is published where anybody can hold us to it.
This is not a soft item. PwC surveyed 50,000 workers and found that daily AI users feel more job-secure than occasional ones, at 58 percent against 36. Fear makes people retreat to what they already know, and use appears to reduce the fear, so the two findings are a loop rather than a contradiction. Our job is to break into the loop safely and early, not to argue with it.
AI, which is a chapter and not a program
The assistants went to everybody, into the same workflows, with the same rollout playbook that produced the gap we are already paying for.
71 percent have used the assistant. 38 percent would sign their name to its output. 22 percent say they can tell which answers need checking, and that last figure is the actual skill. 57 percent have had an answer that was wrong and sounded certain, which is the best predictor we have of somebody having quietly stopped.
The market is not ahead of us here. McKinsey found 88 percent of organizations have adopted AI and 6 percent are capturing real value, and the gap did not close between their 2025 and 2026 surveys. Two years of a flat gap is a structural finding rather than an early-days one.
Stack Overflow found developer use rising from 76 to 84 percent while distrust of accuracy rose from 31 to 46 in the same population of 49,000 people. More adoption is not producing better judgment.
None of that argues for a separate AI program. It argues that the assistant is the fourth thing we have handed this population without a model of what it does, and it will not be the last. A strategy whose subject is a tool expires with the tool. This one is about the capability underneath, so the next arrival is a chapter rather than a rewrite.
Who owns it
The Chief Operating Officer funds it, because the problem crosses every function and no single department can pay for it out of its own budget.
It is not owned by engineering, and that needs saying because the VP of Engineering is its loudest advocate. It is their engineers absorbing the 760 hours, so their enthusiasm is rational. A program owned by engineering also gets built for people who think like engineers, which is the population that is not failing.
The nine regional site directors are the constituency that decides whether this works, because they are the ones currently paying for training out of their own budgets and coping without a shared owner. They are consulted on the design before it is funded, not briefed on it afterwards.
What we are asking for
$2.1M over three years, against an annual gap cost of $6.1M, decided this quarter. The model is in the budget document and the first year is recovered inside the first year, which is a lower claim than these programs usually make and one we can hold.
Procurement is already in flight and no vendor contract is signed inside the 90 days. A plan that signs inside the window tells anybody who has run a procurement that we have not.
The scaling proposal in section 6 carries the kill criteria. If the program cannot be run by somebody who did not design it, it is a consulting engagement with an internal invoice, and we should stop it rather than fund a second year.