I have been building digital learning since 1994.
In that time the way it gets delivered has been rebuilt four times. Every rebuild was real. Each one made the work faster, cheaper and better.
In three decades, the number we report at the end has not changed once.
The disc
In the mid-nineties, courses shipped on CD-ROM. Designed on one continent, built on two others.
In 1995 CBT Systems floated on Nasdaq. First Irish technology company to do it. First e-learning company in the world to do it.
What did we tell a customer about their people? That they had completed the course.
The browser
By the end of the nineties everything was moving onto the web.
No shipping. No sending out corrected discs. A mistake could be fixed on Tuesday and everyone had the fix on Wednesday.
What did we tell a customer about their people? That they had completed the course.
The platform
Through the noughties I was running a production company, and the learning platform became standard. At last everything could be tracked centrally. Records, reports, dashboards.
The measure got slightly richer. Completed, plus a mark out of a hundred.
That was the last time the measure changed.
The regulated years
Later I worked in pharmaceutical manufacturing. Inspectors read the training records there. A wrong one can hold a batch.
This is the most seriously regulated training in the world.
And the record said the operator had completed the training.
In the one setting where it truly mattered whether someone could do the job, the answer was still that they had reached the end of the material.
Now
These days I build certification programmes for professional bodies.
Microlearning. Mobile. Pathways. AI. All of it is better than what I was making in 1996, and I do not say that grudgingly.
Every board report I see counts completions.
But our programme has an assessment at the end
It probably does. Almost every one I have built does, and they are not decoration. They measure something real.
Three things go wrong anyway.
Most assessments are written from the content. So they test whether somebody took in the material. That is a genuine result and it is worth having. It is not the same as whether they can do the job. Knowing the answer in a quiet room is not the same as making the right call at four o'clock on a Tuesday when a client is waiting.
The result comes out as a total. Picture two members who both score seventy-two. One was steadily average throughout. The other was excellent everywhere except the single answer that would put them in front of a disciplinary panel, and their strong answers quietly absorbed it. The total cannot tell them apart, and neither can you. I have written about why the scoring method matters more than the questions separately.
And then almost all of it is thrown away.
Your assessment may know exactly which decision went wrong. It may know the member hesitated, changed their mind, and picked the option that breaches confidentiality.
None of that reaches your records. What reaches your records is one word and a number.
The assessment understood. The organisation never found out.
Why that happens
Most learning platforms still record results using rules written in 2001. Under those rules, a result can only be one of six words.
Passed. Completed. Failed. Incomplete. Browsed. Not attempted.
There is no seventh word.
Rustici Software published figures from their own platform in February 2026. Across the previous year, ninety-two percent of course launches still ran on those rules, and the 2001 version was the most common of all.
A better way has existed since 2013. It is called xAPI, and it can carry what a person actually did, choice by choice. It has been sitting there, finished, for thirteen years, and hardly anybody uses it. When the Learning Guild asked organisations who had not adopted it why not, the leading answer was not cost. It was that they did not know how.
People have asked this before
It would be easy to write all this as though nobody ever tried.
They did. Medicine has been asking whether people can actually do the job since the 1970s, and it built a proper answer. Candidates move between stations, each with a real task and a trained examiner watching. Aviation did the same with simulators. Both work, because both measure what somebody does rather than what they can recall.
So why has none of it reached the rest of us?
Look at the price. Published estimates for one of these examinations run from about twenty dollars a candidate to over a thousand. Cheap versions come in at around fifty to seventy dollars a head. One German university put the cost of a single neurology sitting at eighty-six euro per student.
That is per person, per sitting. And almost all of it is human time. Somebody has to watch, and somebody has to write down what they saw.
A professional body with twenty thousand members cannot spend eighty-six euro a head on a CPD module. Nobody can, outside the professions where a mistake kills someone.
Medicine knew that too. Which is why it built a cheaper version that scales, where the candidate faces written situations from a real working week and picks what they would do. That is a situational judgement test, and I have written about how it should be scored. Outside medicine and a few regulated professions, almost nobody uses that either.
So it was never neglect. It was arithmetic.
The assessment that tells you whether somebody can do the job has always existed. It cost too much per person.
The plumbing that could carry a rich result has existed since 2013. Nobody built the layer that turned it into something a person would read.
Both are the same problem in different clothes. Collecting the detail was never hard. Interpreting it was, because interpreting meant a qualified human sitting down with everything a member did and working out what it meant.
That is the expensive part. It always was.
What has actually changed
Machines can now do that reading and writing.
Not the judging. What counts as a good decision still has to be set down in advance by people who know the profession, and I would not hand that to a machine. But the part where somebody works through what a member did and writes it up in language they will understand is no longer a person's afternoon. It costs a few cents per member rather than tens of euro.
That makes two things possible which were not, three years ago.
A member who fails can be handed a plain account of what they chose, where it went wrong, and why that particular choice matters. Not a score. An explanation of their own decisions.
And a board can be handed a picture of where its membership is actually weak. Not how many finished. Which duty people are getting wrong, how often, and whether it is better or worse than last year.
Both come out of an assessment that already knew all of this and had nowhere to put it.
What it does not fix
It does not replace an examiner watching somebody perform a procedure. Hands-on skill still needs a human in the room, and it always will.
And it does not rescue a badly built assessment. If your questions test whether people remember the content, better reporting will only tell you in more detail that you tested the wrong thing. The design has to be right first. The reporting is what lets you see it.
Where that leaves us
Four rebuilds of how this work gets delivered. None at all in what we can say about the people who did it.
That was never because the question was uninteresting. Every board wants to know whether its members can do the job. It was because answering it properly cost more than anyone would pay, so everybody settled for the number that came free.
The price has changed. Most of the sector has not caught up. Completion is still what gets reported, because it is still what the systems hand over.
So the next time somebody shows you a completion rate, do not ask whether the number is good. Ask why it is the only number they have.
Want to know what your own programme could tell you that it currently does not? The Programme Design Diagnostic is a fixed-fee, independent read on a single programme: a board-ready diagnosis and a costed build scope, in two to three weeks.
Sources for the figures above: Rustici Software's published SCORM Cloud usage data, February 2026; the Learning Guild's survey work on xAPI adoption and the barriers reported by non-adopters; and the published cost literature on objective structured clinical examinations, including the range of roughly twenty to over one thousand dollars per examinee, shoestring estimates of fifty to seventy dollars, and the University of Ulm neurology costing of eighty-six euro per student. All retrieved 24 August 2026.
For how a judgement-based result should be scored once you have one, see one test, five ways of scoring it. For why decision-based formats became affordable at all, see the best digital learning experience I know was built in 2013. For the design failure underneath all of it, see assessment-bolted-on and the four kinds of credential drift. See more insights from LearnFrame.