Three colleagues in a bright modern glass-walled meeting room with a city skyline behind them, discussing printed pages spread across a pale wood table
← Back to Insights

One test, five ways of scoring it. Only one of them worked.

19 August 2026 · 7 min read · By Paul Robinson · LearnFrame Insights

Every certification body knows a quiz cannot measure judgement. Fewer know that the scoring method matters more than the questions.

In 2024 a group of researchers at the California University of Science and Medicine did something unusual. They took 142 fourth-year medical students, gave them all the same test, and then scored the same answers five different ways.

Not five different tests. One test. One set of responses. Five methods of turning those responses into a number.

The methods ranged from simple to less simple. Rank the options in order. Mark each one right or wrong. Mark each one right or wrong and subtract for wrong ones. Measure how far each answer sat from the expert's preferred answer. And finally, measure that distance and square it, so that being badly wrong counted for much more than being slightly wrong.

Then they checked how reliable the test was under each method. Reliability here means something specific and unglamorous: does the instrument hold together, or is it producing noise?

The answers ranged from 0.39 to 0.85.

For context, anything below about 0.7 is generally treated as too unreliable to make a decision on. So the same students, giving the same answers to the same questions, produced a test that was unusable under some scoring methods and defensible under one of them. The only method that cleared the bar was the last one, distance from the expert answer, squared.

Nothing about the questions changed. Only the arithmetic applied afterwards.

Why that should worry anyone who awards a credential

If you run education at a professional body, you have almost certainly commissioned a scenario-based module at some point. Someone faces a realistic situation. They choose what to do. They get feedback. It is a good format and members generally like it.

Now ask a harder question about the one you have. What is the score actually made of?

In most commercial learning products, the honest answer is that there is no score. There are branches and there is feedback, and at the end there is a record that a person reached the end. In the ones that do produce a number, that number is almost always a plain points total. Right answers add points. Wrong ones do not.

Given what those 142 students demonstrated, a plain points total is close to the weakest option available. It is the arithmetic most likely to produce a number that looks meaningful and is not.

That matters more for you than for almost anyone else, because you are not running a training course. You are putting your name on a statement about a member's competence.

Two mature fields that have almost nothing to do with each other

There is a reason this gap has stayed open, and it is structural rather than anyone's fault.

Situational judgement tests are a serious instrument with about fifty years of methods literature behind them. The founding paper dates to 1990. They are used at national scale to select doctors, and they are sold commercially by a handful of substantial assessment firms. In that world, the scoring key is the product. A panel of experienced practitioners works through every option on every scenario and agrees how good each one is. That agreed key is what the learner's answers are measured against.

Scenario-based learning, meanwhile, is completely standard in corporate training. Every mid-sized supplier offers it.

The two fields use nearly identical vocabulary. Both talk about scenarios, decisions, workplace judgement, consequences. From either side, it looks as though the other side is doing the same thing.

They are not. The difference is not the scenario, and it is not the branching. The difference is whether there is an expert-validated key sitting behind the response options, and whether the output measures something or simply records that something happened.

Without a key, a scenario module is teaching wearing the clothes of measurement.

We wrote in May about the four ways a credential drifts when assessment is designed last. One of those was capability drift: an assessment that tests recall when the curriculum was building judgement. That piece named the problem. This one is about what you do instead.

What the CPD system currently measures

It is worth being blunt about the baseline, because it explains why nobody has pushed on this.

Continuing professional development, in every profession and jurisdiction I have looked at, is counted in hours, units and points. One recertification framework for pharmacy specialists, for example, requires eighty units across a seven-year cycle with a minimum of two a year. Accreditation bodies in this space largely accredit the provider against quality criteria rather than assessing what any individual learner can now do.

The sector says the quiet part out loud in its own marketing. Providers routinely distinguish a CPD certificate from "a certificate of attendance", and then define the CPD certificate as confirmation that you completed a specific course.

Completing a course is attendance with extra steps.

There is nothing dishonest about this. Hours are easy to count, easy to audit and easy to defend. But it does mean that the entire apparatus, at no point, measures whether a professional can actually exercise judgement in a live situation. That is the gap, and it has been sitting there in plain sight.

The thing a points total cannot do

There is a second problem with points totals, separate from reliability, and it is the one that should bother a regulator most.

A points total is compensatory. It lets a good score in one place make up for a bad score in another. That is fine when you are measuring knowledge, because knowing more about topic A genuinely does offset knowing less about topic B.

It is not fine when you are measuring professional conduct.

Suppose a member works through five workplace decisions and handles four of them well. On the fifth, they upload a client's confidential file into a public AI tool. Under a points total, a strong performance elsewhere can absorb that. They pass. They receive your certificate.

No professional body on earth would defend that outcome if it were put to them directly. Yet it is the arithmetic almost every scored module quietly runs.

The fix is not complicated, but it has to be a deliberate design decision. You identify the small number of actions that are not a matter of degree, the ones where doing it once is disqualifying regardless of everything else, and you make those non-compensatory. No total buys them back. Everything else is scored on the distance-from-the-expert-key method, where being slightly off costs a little and being badly off costs a lot.

That is two design decisions. Together they are the difference between a number that means something and a number that looks like it does.

The honest limit, which almost nobody states

Here is the part that most suppliers will not tell you, and it is the question worth asking first.

If a module offers you a profile of a member across, say, five professional duties, ask how many scored decisions sit behind each duty. If the answer is one, the profile is not a measurement. A single item has no internal consistency, no error estimate, and enormous sensitivity to how that one scenario happened to be worded. Change a few words in the scenario and the person's apparent strength on that duty moves.

The literature would want at least three scored decisions per duty before reporting on it separately, and would prefer more. That is fifteen decisions minimum for a five-duty profile, which is more than fits in a twenty-minute module.

So a per-duty profile is a programme-level output, not a module-level one. Anyone showing you a rich five-domain profile off the back of a short module is overclaiming, and it is a fair question to put to them in a procurement conversation.

I will say plainly that this applies to our own work. The first module of the programme we publish carries five scored decisions. It reports a whole-instrument result, and the per-duty view is labelled as indicative rather than measured, because one item per duty cannot support anything stronger. We rebuilt the scoring itself on the evidence above, moving off a points total to distance from a declared expert key, with a small set of non-compensatory rules that no score can recover. That change required no new scenarios at all. It was arithmetic applied to responses we were already capturing.

One more thing worth knowing

There is a common assumption that the more advanced route here is an adaptive system, one where AI generates or adjusts the questions in response to the learner.

The evidence does not support that for this job. Fixed-response tests, where every option is written and rated in advance, have been shown to be the more effective format for formative assessment, precisely because a pre-defined list of options is what lets you measure a person's grasp of the standards and expected behaviours. You cannot measure someone against a standard if the options were invented on the fly.

That has a second consequence. Under the EU AI Act, AI systems used to evaluate learning outcomes fall into the high-risk category, with the associated obligations. A fixed, pre-authored instrument does not. So the design that is better at the formative job also happens to be the one that carries the lighter regulatory load.

That is an unusual position to be in, and it is worth understanding before anyone sells you the adaptive version as the modern option.

Where this leaves you

If you are responsible for a credential, the question is not whether your programme has scenarios. Most do now.

The question is whether anything behind those scenarios was ever validated by people who know the work, and whether the arithmetic turning answers into a result was chosen deliberately or inherited by default.

If nobody can tell you what the scoring key is or who agreed it, you do not have a measurement. You have a very good piece of teaching with a number attached to it. Those are both worth having. They are not the same thing, and only one of them is safe to put a credential on.

A concrete next step

Wondering what is actually behind the score on one of your own programmes? The Programme Design Diagnostic is a fixed-fee, independent read on a single programme: a board-ready diagnosis and a costed build scope, in two to three weeks.

The Diagnostic →

Building member or staff CPD on AI?

AI for Professional Practice is open now, and it is built to be licensed by professional bodies, training providers and corporate academies, branded as yours and re-verified quarterly. Early conversations shape the licensing round.

Licence the Programme → Arrange a Conversation →

Sources for the figures above: the scoring-method comparison pilot at the California University of Science and Medicine; Whetzel, Sullivan and McCloy (2020) on situational judgement test development; Motowidlo, Dunnette and Carter (1990) as the origin paper; Sahota and colleagues in Medical Education (2026) on judgement testing and later professionalism; and published guidance on fixed-response formats in formative assessment. All retrieved 28 July 2026.

For the problem this piece answers, see assessment-bolted-on and the four kinds of credential drift. For why decision-based formats became affordable at all, see the best digital learning experience I know was built in 2013. See more insights from LearnFrame.

About the author

Paul Robinson is the founder of LearnFrame, which designs and builds custom eLearning programmes for professional certification bodies, regulated training providers, and corporate academies. He has worked in digital learning for three decades, since the 1995 Nasdaq IPO of CBT Systems. Connect on LinkedIn.