Phil Wagner

Staff Learning Designer and Technical Writer
Technical Lead of AI Enablement Education

The situation at Halden Systems

Written by a person. Last read by a person on 2026-09-07, 21 days ago. Its facts were checked by the eval suite on 2026-09-28.

You are starting the set, or you want the situation without the argument built on it.

Scenario-based sample. Halden Systems is invented, and so is every figure about it.

Bottom line. 78 percent of our people reached the end of the training for our engineering systems. 31 percent can predict what a step in those systems will do. That 47-point gap costs about $6.1M a year, and no better tool will close it. Another approach is needed.

They can follow the steps. They cannot say what the steps do.

That is not a complaint about depth of understanding. It is the difference between running a procedure and being able to predict its result, which is what somebody needs in order to tell when the screen has stopped matching the instructions, to judge whether an action can be undone, and to recover without asking. Lacking it, the only safe move is to stop and escalate, which is what 2.3 escalations per person per week are.

The organization

Halden Systems builds and sells engineering software.

Staff 5,000
Sites, across 4 continents 9
Hours between the first site's morning and the last site's evening 17
Technical staff 2,000
Non-technical staff 3,000

The timezone spread is the figure that decides the delivery model, and the split between the last two rows is the one that decides who the program is for.

2 groups inside those 3,000 non-technical staff need naming, because the program treats them differently.

The 1,150 who work in engineering systems daily. Writers, designers, researchers, program managers and the senior half of support. Their work lives in repositories, issue trackers and build pipelines. They are the group whose capability gap costs money today.

The other 1,850, who were given an AI assistant anyway. Their work does not touch repositories, and the assistant went to every employee without regard to whether it suited the job. They are counted here because a tool nobody assessed them for is now in their hands.

The 4 earlier migrations at least moved people whose work had to move. This one reached 1,850 people with no case for reaching them, and it repeats the same error more widely: capability was never measured before the tool arrived.

5 of the 9 sites have nobody who can answer a question about repositories, pipelines or issue trackers during the local working day.

What is failing

Over 4 years Halden moved most of its non-engineering work into engineering systems. Documentation became docs-as-code. Design tokens went into a repository. Program management moved into the issue tracker. Support began authoring runbooks in markdown.

Every one of those moves was run as a tooling migration: a cutover, a demonstration, a recording, and a channel for questions. None was run as a capability problem, and two measurements show what that produced.

78 percent of the 1,150 reached the end of the training. 31 percent can predict what a step will do. Our completion dashboard reports the seventy-eight and calls the program finished.

A person who follows steps without understanding them is fine until the screen stops matching the instructions. Then they escalate. That happens a median 2.3 times a week each, which across the 1,150 comes to roughly 760 engineer hours absorbed every week.

64 percent say they are afraid of breaking something that cannot be undone. Given that they cannot predict what any step will do, that fear is accurate rather than timid, and it is the reason the escalation rate is a floor rather than a peak.

How we measured the 47-point gap

Bottom line. The 78 and the 31 come from two different instruments, and we say so before anybody asks. The 78 is a completion count we already had. The 31 is a scored check of whether somebody can predict what one of our systems will do.

Both known errors inflate the same side of the subtraction, so 47 points is a floor, not an estimate. The check becomes the level 2 measure (whether somebody can now do the thing), quarterly, with a stated condition under which we throw it out.

The 78 The 31
What it records Reaching the end of a module Predicting what an action will do, and why
Source Learning platform, 4 migration curricula, 2022 to 2026 Capability check, 6 scored items, fielded over 2 weeks
Population Census of all 1,150 712 respondents, a 62 percent response rate
Known error Inflated. Mandatory, and clicking through is free Inflated. People who expect to do badly do not sit it
What it is good for Operations Capability

Three things a reader should take from this and nothing more.

  1. Most of this population finishes training it cannot act on. That is the only claim the comparison supports, and it is enough.
  2. 47 is a floor. Non-response flatters the 31 and mandatory completion flatters the 78. If the measurement is wrong, it is wrong in the direction of making the problem look smaller.
  3. The check is repeatable and the program is accountable to it. Same item bank, same scoring, quarterly, reported by site before it is reported company-wide.

The condition under which we drop it. If the pass rate rises while escalations per person per week do not fall and median days to a first merged change does not move, we are measuring test-taking rather than capability, and the instrument is wrong rather than vindicated. Both of those behavioral measures are already counted and neither is self-reported.

The instrument, its scoring, the distribution and the cadence are in the capability survey.

The AI assistant is the fifth rollout, not a new problem

The assistants went to the same people, through the same workflows, on the same playbook.

The survey asked all 3,000 non-technical staff four questions about it.

Share
Have used the assistant 71%
Would put their name on its output 38%
Can tell which answers need checking 22%
Have been given an answer that was wrong and sounded confident 57%

The third row is the one that matters, because an assistant is only safe for somebody who can judge when it is wrong. On that measure fewer than a quarter of them are equipped to use it, and more than half have already met the failure it produces.

Somebody who cannot yet judge the output learns the wrong lesson from a confident wrong answer, which is that they cannot tell good answers from bad ones. The reasonable response is to stop using the assistant, and to say nothing about having stopped, because the rollout was announced as a success.

That is why usage will keep looking healthy while the benefit does not arrive. We are counting who opened the tool, not who can use it safely.

What was already tried, and what it produced

1 site ran a pilot at its own cost. Dublin, 24 people from the inner population, 6 weeks, a written runbook and one facilitated change each, with a named person to ask.

Median days to a first merged change fell from 41 to 9 for that group. The pilot cost nothing beyond existing staff time, which matters for one reason: it answers the objection that nothing can begin before the funding decision. Something already did.

What we already built, and what it teaches

Four years of rollouts produced a library of 140 training modules. It is the largest asset the program inherits and the clearest evidence about what went wrong.

Reading all 140 against one question, whether somebody who works through it comes out able to predict what the system will do, 96 fail it.

What the audit found Modules
Teaches the steps and not the understanding, mostly screen recordings 41
Covers a tool the company no longer uses 24
A regional duplicate of another module 19
Compliance content that belongs with the policy 12
Still correct and still needed 22
Right subject, wrong form 22

Three of those rows are the diagnosis restated as an asset register. The 41 are how 78 percent learned to follow steps they cannot explain. The 19 duplicates are the $410,000 of regional spend, showing up as content rather than as invoices. The 24 are 4 years of tool churn nobody was funded to clean up.

That is why the answer is not a bigger library. The content plan keeps 44 of the 140, adds 35 short pieces, and sizes the result at 79.

What the scale changes

At 1,200 people in one building, one determined person holds a program together by force of will. At 5,000 across 9 sites, 3 constraints bind that would not otherwise.

1. No hour of the working day is shared by every site. 17 hours separate the first site's morning from the last site's evening, and 18 percent say they would attend an optional live session. Any plan that rests on live attendance is refuted by our own survey before anyone costs it.

2. Proximity to help predicts capability better than job title does. The confidence gap between sites with a local expert and sites without one is 22 points, wider than the gap between job families. Median days to a first merged change is 41 company-wide and 58 at the 5 sites with nobody to ask.

3. 9 sites are each buying their own training, at about $410,000 a year of overlap. No site director was wrong to do that. Waiting for a central answer costs them more than buying locally, and the duplication is a symptom of having no shared owner rather than of poor judgment.

Where the money is

Three lines, and they do not come from the same population.

Revenue, which comes from customers rather than staff. Halden sells engineering software, and its customers need the same kind of teaching its own people do. Solution engineers currently spend about 6 days per new enterprise customer delivering that teaching by hand, across 240 new enterprise customers a year. 95 partner implementations a year depend on partner staff holding a certification we do not currently offer.

The connection to the internal problem is the machinery, not the audience. The content, the assessment and the certification that would close the internal gap are the same assets a customer program sells, which is why one function should own both rather than two functions building them twice.

Cost, which comes from the 1,150 and their colleagues. Roughly 760 engineer hours a week at $145 fully loaded, plus the $410,000 of duplicated regional spend, plus about 1,900 internal tickets a month whose answer already exists in writing.

Efficiency, which comes from time. 41 median days to a first merged change, 58 at the sites without an expert, and nine in the Dublin pilot.

Who decides

The Chief Operating Officer funds it. The problem crosses every function, and no single department can pay for it from its own budget.

The VP of Engineering is its loudest advocate, because these are their engineers absorbing 760 hours a week. They are also the wrong owner. A program owned by engineering gets built for people who think like engineers, and those people are not the ones failing.

The nine regional site directors decide whether it works. They hold the budgets that currently pay for the duplication, they employ the people the program serves, and they are the ones who have been coping without help. A plan that takes their budgets before it has earned their trust turns the constituency it needs into the opposition that stops it.