· Thea Mannix

Receipts or Results – Part 3: Measuring Security Culture

Part 3 of the Receipts or Results series explores why one-size-fits-all security culture scoring fails, and how baseline-first measurement reveals meaningful behavioral change.

Receipts or Results – Part 3: Measuring Security Culture

Hand a three-year-old an IQ test built for adults and they will fail. Every question. Confidently, cheerfully, completely. If you didn’t know any better, you’d write “significant cognitive impairment” in the file.

Nothing is wrong with the child. Something is wrong with the test. This is precisely why intelligence tests come in different versions for different age groups in the first place - a test built for a 30 year old and a test built for a 3 year old are measuring against entirely different baselines of what “typical” even looks like at that stage. Psychology solved this problem decades ago. You calibrate the instrument to the population being tested, or the result is noise wearing a lab coat.

Security culture measurement hasn’t caught up.

One score, every organisation

Most culture surveys and maturity models are built the same way. Ten questions, five-point scale, deployed identically whether you’re a forty-person startup or a forty-thousand-person bank. The bank scores a 4.2. The startup scores a 2.8. Somebody writes “culture problem” in the startup’s file and recommends more training.

Nobody asks whether the instrument was built with the startup in mind at all. We’ve talked before about how population-wide benchmarking smooths out exactly the context that makes a number meaningful. Comparing yourself to an average tells you very little about your own environment. A universal culture score does the same thing, just with better production values. It looks precise, but without being calibrated to your baseline it isn’t measuring what it claims to.

What the same score can hide

Take a regulated bank with mandatory annual training, tracked completion, and a policy library everyone can recite. It scores well on most culture instruments, because most culture instruments are quietly measuring training completion and policy recall dressed up as “culture.”

Now take a 40 person fintech with no formal training programme, but where every engineer flags anomalies fast, escalates without waiting for permission, and treats a weird login as everyone’s problem rather than the security team’s. That organisation may well score lower on the same instrument, because the instrument was never built to see judgment and instinct, only compliance and recitation.

One of these organisations is describing behaviour. The other is describing paperwork. The test can’t tell them apart, and it isn’t the startup that has the problem.

Calibrating the instrument, not the excuse

Fixing this doesn’t mean giving up on measuring culture. It means being honest about what actually needs calibrating before you measure it: industry risk profile, organisational size and structure, regulatory exposure, workforce composition, the maturity of the controls already in place. A 12 person team and a global enterprise are not the same population, and no amount of survey polish changes that.

In practice, that looks like benchmarking against comparable organisations rather than a global average, and against your own baseline over time rather than someone else’s snapshot. A score that’s meaningfully improved from where you were six months ago tells you something. A score that’s lower than a bank’s tells you almost nothing, because you were never the bank’s population to begin with.

It also means treating the output as a signal to investigate, not a verdict to file. A low score should prompt the question “what is this actually picking up on here,” the same way a clinician would ask before writing anything down about that 3 year old. We wouldn’t hand a toddler the adult matrices and diagnose them on the result. Yet we do the organizational equivalent every time we run one culture survey across every environment and call whatever comes out the truth.

Ultimately, we need a baseline-first approach to measurement. We need to standardise the method before we work on tests.

Ready to see your employees' security behaviors?

Connect your Microsoft 365 and see months of behavioral data in 15 minutes. Free 30-day trial — no credit card, no sales call.

Start Free Trial