You could build this. Here's what it actually takes.
The data is in your tenant, the Microsoft Graph API is documented, and a competent engineer can pull a first report in a week. That part is true. This page is the honest version of what happens after the first report — where a build usually stops, and when it is the right call anyway.
What you get in week one
A real, useful answer. With read-only Graph access, an engineer can surface how many people have MFA registered, who is sharing files externally, how mail-forwarding rules are being set, and how sign-in risk is distributed. Pulled into a BI layer, that is genuinely more than most organizations can see today, and it is worth having. If your question is "what does our tenant look like right now?", a week of engineering answers it well.
Where it stops
Microsoft's retention window
Behavioral history is retained for a limited period, so what you did not capture is gone permanently. Baselines cannot be reconstructed backwards — the day you start building is the earliest your trend can begin.
Normalisation
A raw event count is not a behavior indicator. Turning signals into something comparable across teams of different sizes and over time — so a rise means a real change, not a headcount change — is the actual work, and it is most of it.
Statistical validity
A before-and-after comparison needs a method that survives an auditor asking how you controlled for seasonality, headcount change and sampling. Without that, the number is suggestive, not defensible.
Privacy architecture
Employee behavioral data has works-council, GDPR and internal-trust implications. The design decisions — what is aggregated, what is never stored, who can see what — are hard to reverse once the pipeline is live.
Maintenance
Graph schemas change. The report is not the product; the ongoing correctness is. Every change upstream is a maintenance event, and the cost is continuous, not one-off.
Nobody owns it
The engineer who built it moves teams. Without a clear owner, a home-built measure quietly rots — and it tends to fail at the moment someone finally asks it for a trend.
Peer-reviewed field research
29–55%
of the variation in phishing susceptibility is attributable to organizational-level factors, not individual ones.
Measured across 83,269 employees in 510 organizations, using their real responses to phishing campaigns.
Petrič, G., & Roer, K. (2022). The impact of formal and informal organizational norms on susceptibility to phishing. Telematics and Informatics, 67, 101766. Licensed under CC BY 4.0.
Organizational-level measurement is a research problem before it is an engineering one. Reading the signal is the easy part; measuring it in a way that holds up is the hard part.
The honest cost comparison
We won't put fabricated numbers on your build — you know your own rates. Weigh the engineering days to build it, the ongoing days to maintain it, and the opportunity cost of the team not doing something else, against a published per-employee price. Then put your own figures into the model and see where the line falls for your organization.
Model your own build-vs-buy numbers in the ROI calculatorWhen building is the right call
Sometimes it genuinely is, and pretending otherwise would make the rest of this page dishonest. Build it yourself when:
- you have a one-off question, not an ongoing reporting requirement;
- your data source is unusual and no product covers it well;
- you have a data-engineering team with genuine spare capacity and an owner who will stay with it;
- you have no regulatory obligation to produce a defensible, maintained effectiveness trend.
If several of those are true, a build is a reasonable choice. If none of them are, the ongoing cost usually points the other way.
A query is not a method
Evidence you built yourself still has to answer an assessor's question about method. The clause does not care who wrote the query — it cares whether the measurement is sound and documented.
| What your reports show | What the regulation asks for | What closes the gap |
|---|---|---|
| A dashboard built from your own Graph extraction | NIS2 CIR Annex §7 — the effectiveness of risk-management measures is evaluated with a defensible method | A documented, repeatable measurement method that survives review |
| Raw event counts trended over time | GDPR Art 32(1)(d) — a process for regularly testing, assessing and evaluating effectiveness | Normalisation and controls for headcount and seasonality, documented |
| A report one engineer maintains | DORA Art 13(4) and 13(6) — ongoing monitoring of effectiveness over time | Continuity that does not depend on one person staying in the role |
Building it is a legitimate choice. But the assessor's question is about method and documentation — and that is the part that outlasts the first report.
See what an assessor actually asks you to evidenceQuestions engineers ask
Can’t a competent engineer just build this from Graph?
What does Microsoft’s retention window mean for a build?
When is building it yourself the right call?
Does evidence I built myself satisfy NIS2, DORA or GDPR effectiveness requirements?
Compare it against your own build
Connect Microsoft 365 in 15 minutes and see the behavioral baseline a build would take months to reach — then decide.
Start your free 30-day trialNo credit card. No commitment. Results in 15 minutes, or don't continue.
See the price — published, so you can weigh it against engineering days.
Read what an assessor requires — method and documentation, not who wrote the query.
Related comparisons
Already measure this in Microsoft?
Defender, Secure Score and Purview report on your tenant. See which layer answers which question — and why Graph availability isn't the same as measurement.
Compare every platform category
See what each kind of platform measures, and what none of them measure, in one place.