· Kai Roer

Security awareness is not a checkbox. Here's how to measure if it works.

Security awareness is not a checkbox. Completion rates measure activity, not effect. Here is how to measure whether awareness actually changes employee behaviour.

Security awareness is not a checkbox. Here's how to measure if it works.

Most organisations treat security awareness as a compliance activity. Run the training, record the completion, report the number, move on. The programme exists. The checkbox is checked. The board receives a slide showing 90 per cent completion, nods, and moves to the next agenda item.

That approach satisfies the process requirement. It does not satisfy the effectiveness requirement that GDPR, NIS2, and DORA now impose. The gap between those two things — between “we did it” and “it worked” — is where most security awareness programmes are stuck.

The proxy problem

The security industry has converged on phishing simulations as the primary measurement tool for awareness effectiveness. The logic is understandable. Phishing is the most common attack vector. Simulations test whether employees fall for it. If click rates go down over time, training is working.

The problem is that what employees do in a simulation does not reliably predict what they do at their desks. A simulation is a controlled test in an artificial environment. The employee may recognise it as a test. They may be primed by the training they just completed. They may simply be more cautious in the days immediately following a training session.

The peer-reviewed evidence is consistent on this point. Lain, Kostiainen and Capkun (IEEE S&P 2022) found that click-time training did not make employees more resilient to phishing in a study of more than 14,000 employees over 15 months. Hillman, Harel and Toch (Computers & Security 2023) found no significant effect of training timing on click rates across 5,000 employees at an Israeli financial institution. Hydari et al. (arXiv 2026, under peer review) found that observed improvement in simulation click rates is largely explained by stable individual traits — who someone is — rather than by the training effect.

Simulations measure personality, not learning. Declining click rates may reflect that the employees who were never going to click continue not clicking, while the employees who were always going to click continue clicking regardless of training. That is a selection effect, not a training effect.

What actual measurement requires

Measuring whether a security awareness programme works requires tracing the connection between an intervention and a change in real behaviour. Not simulated behaviour. The actual thing employees do in their actual work environment — whether they enable MFA, whether they follow data handling procedures, whether they report suspicious activity through the proper channel.

This requires three things.

First, a baseline. What did behaviour look like before the intervention? Without a baseline, there is no “before” to compare the “after” against. An improvement claim without a baseline is not a measurement. It is a guess.

Second, an event. What was the intervention, and when exactly did it happen? The event is the independent variable. If you cannot identify it precisely — this training was delivered to this population on this date — you cannot attribute any subsequent change to it.

Third, a before-and-after comparison. What did the relevant behavioural indicators look like before the event, and what did they look like after? Did MFA adoption increase? Did reporting rates change? Did the specific behaviour the training targeted actually shift? And if it shifted, by how much, and for how long?

The time dimension

A one-week improvement that reverts by week four is not effectiveness. It is a short-term performance effect — the kind of boost you see immediately after any intervention, before people revert to their baseline behaviour.

Real measurement requires following the signal across time. What happened in the first week? What about week four? Month three? Six months later? A training that produces a seven-day improvement followed by full reversion has not produced evidence of effectiveness. It has produced evidence that a temporary performance spike does not translate into lasting change. That is useful information, but it is not the information most organisations report.

The effectiveness requirement in the EU frameworks is not satisfied by showing a temporary blip. It requires evidence that the measure produced a sustained effect. That means measurement over months, not days.

The accountability question

Before attributing a security gap to employee behaviour, examine what the organisation has actually put in place.

If MFA has not been enabled, that is an organisational failure, not a behavioural one. The employee cannot use a control that does not exist. If MFA has been enabled but not enforced — if it is optional, and the organisation relies on employees choosing to adopt it — then the organisation has chosen recommendation over enforcement. Recommendation is a structurally weaker driver than enforcement. An organisation that chooses the weaker approach and then blames employees for the predictable outcome has misdirected its accountability.

This principle applies broadly. If data classification tools are available but not configured, that is an infrastructure decision. If password policies are set but not enforced, that is a governance decision. Accountability flows to the lowest level where the failure occurred. In many cases, that level is not the employee. It is the organisation that chose not to enforce the control it could have enforced.

The infrastructure now exists

The measurement problem that made this difficult for most of the industry’s history was an infrastructure problem. You could not measure real employee behaviour at scale because the data did not exist in an accessible form. Organisations could track training completion and simulation clicks because those were events within systems they controlled. They could not track what employees actually did in their daily work.

That constraint has changed. Microsoft 365 emits continuous behavioural signals from real work — authentication events, MFA usage, data handling patterns, configuration states. The signals exist. They are accessible through standard APIs. The measurement problem is now a solved infrastructure question.

The remaining question is whether organisations are building the process to use those signals — to establish baselines, tie interventions to dates, and track whether the intervention produced a sustained change in the thing it was designed to change. The data is there. The process to turn data into evidence is what most organisations have not built yet.

For the full regulatory argument and enforcement precedents, read the white paper: Measuring Security Control Effectiveness: From Attestation to Evidence

Ready to see your employees' security behaviors?

Connect your Microsoft 365 and see months of behavioral data in 15 minutes. Free 30-day trial — no credit card, no sales call.

Start Free Trial