Immersive training pilots rarely fail on the technology. They fail at the budget meeting eighteen months later, when someone asks what the return was and the answer is a satisfaction score and a slide of people wearing headsets looking engaged.
Engagement is real, and it is not an argument. Every budget committee has seen an engaging pilot that changed nothing. Any AR VR development company that hands over a program without an evaluation design has handed over a project that will be cancelled at its first serious review, however good the simulation is.
The metrics that do not survive scrutiny
Three appear in almost every deck and none of them will carry a funding decision.
- Satisfaction scores. Learners enjoy VR because it is novel. Novelty decays, and enjoyment was never correlated with competence anyway.
- Completion rates. These measure whether people finished, which mostly measures whether it was mandatory.
- Time-in-headset. A pure activity metric. More minutes is not more learning, and a finance director will say so.
The common flaw is that all three measure the training rather than the work. Nobody funds training. They fund outcomes that training is supposed to produce.
Start from the cost of the error you are preventing
Immersive training earns its premium in exactly one situation: when the procedure is expensive to get wrong and expensive to practise. If either half is missing, a video and a checklist will beat VR on cost every time, and you should say so rather than sell a headset.
So the analysis begins before the build, by quantifying the error. How often does it occur, and what does one occurrence cost - in downtime, rework, scrapped material, safety incident, or delayed commissioning? That number is the ceiling on what the program can possibly be worth, and it is worth knowing before committing to build anything.
It also frequently ends the conversation early, which is a good outcome. A procedure performed twice a year with a €400 failure cost does not justify a simulation, no matter how good the demo looked.
Four measures that hold up
- Time to competence. How long from first exposure to independently signed off, compared to the existing method? This is usually the largest and most defensible saving, because trainer hours and trainee wages are already in a budget line somebody owns.
- Error rate in live work. Track the specific errors the simulation targets, for trained versus untrained cohorts, over the first 90 days on the job. This is the number that persuades operations.
- Equipment and downtime avoided. Training on a real production line has an opportunity cost that operations can usually quantify precisely. Simulation removes it.
- Retention at 90 days. Re-test competence a quarter later. Immersive training's genuine advantage over classroom delivery tends to show up here rather than immediately after, and skipping this measurement forfeits the strongest result.
If you cannot name the specific error you are preventing and what one instance costs, you are not measuring ROI. You are collecting reassurance.
Design the comparison before you build
The most common evaluation mistake is measuring only the trained group. Their error rate drops, and it means nothing without a comparison - the process may have improved, the intake cohort may have been stronger, the season may differ.
You rarely need a randomised trial. A staged rollout is usually enough and is easier to get approved: one site or shift adopts the simulation while a comparable one continues as before, then they swap. You get a comparison group without denying anyone the training permanently, which is normally the objection that kills a control group.
Capture the baseline before the first headset arrives. Retrospective baselines are always contested, and the argument is unwinnable because the data was not collected for that purpose. Two weeks of measurement beforehand protects the entire investment.
Instrument the simulation itself
A simulation can record what a classroom cannot: not only whether someone completed a procedure, but where they hesitated, which step they repeated, what order they worked in, and which mistakes they made before self-correcting.
This is worth building in from the start, for two reasons. It gives you competence data far richer than a pass mark, and it tells you which parts of the procedure are genuinely hard - which frequently improves the procedure itself, or the physical design of the equipment. Several programs we have delivered produced more value from that finding than from the training.
Tie every recorded event to a named training objective. Telemetry that is not mapped to an objective becomes a large dataset nobody analyses.
What a defensible business case contains
One page: the specific error, its frequency and unit cost, the baseline time-to-competence, the projected improvement with its evidence, the total program cost including hardware refresh and scenario maintenance, and the payback period. Then the measurement plan that will confirm or refute it, with dates.
Include the ongoing costs honestly. Headsets get replaced, procedures change, and scenarios need maintaining. Programs that omit this look excellent in year one and get cancelled in year three when the true running cost surfaces at a review nobody prepared for.




