Human Risk Score: How to Quantify Employee Security Risk
A practical guide to building a human risk score: which signals to use, how to weight them, and how to turn employee risk into a metric boards understand.

Security teams can tell you their mean time to patch, their EDR coverage, and the CVSS score of every vulnerability in the backlog. Ask them to quantify the risk sitting in row 4 of the finance department, and most have nothing — even though the Verizon 2026 Data Breach Investigations Report attributes 62% of breaches to the human element.
That asymmetry is the case for a human risk score: a single, continuously updated number that does for people what CVSS does for software — makes risk comparable, trackable, and actionable. This guide covers what belongs in the score, how to build one that survives contact with reality, and the mistakes that turn a useful metric into a compliance ornament.
Why quantify human risk at all?
Three reasons, in ascending order of importance.
Prioritization. No security team can coach everyone. The IEEE S&P 2025 study of 19,500 employees showed that blanket annual training barely moves behavior; what works is targeted, timely intervention. Targeting requires knowing who — and a score is how you know.
Communication. Boards do not read phishing click reports. They read trend lines. A score that moves from 68 to 41 over two quarters is a story a CISO can tell in one slide, and it converts "awareness" from a cost center into a measured risk-reduction program — the core argument in our breakdown of security awareness training ROI.
Strategy. Gartner predicts that by 2027, half of large-enterprise CISOs will have adopted human-centric security design practices, and NIST SP 800-50r1 reframes awareness as an ongoing, measured program rather than an annual event. Both shifts presuppose measurement. You cannot run a human-centric program on gut feel.
A human risk score is not a grade you give employees. It is a triage system you give your security team.
What goes into the score
A defensible score combines two dimensions: behavior (what people actually do) and exposure (what happens if they get it wrong). Someone who clicks every lure but has no access to anything sensitive is a smaller problem than a domain admin who clicks occasionally.
| Signal category | Example inputs | Why it matters |
|---|---|---|
| Simulation behavior | Phishing simulation failure rate, report rate, time-to-report | Directly observed behavior under realistic conditions |
| Real-world behavior | Reported real phish, malware events, policy violations, risky browsing | What simulations predict, reality confirms |
| Credential hygiene | Password reuse, appearance in breach dumps, MFA adoption | Credentials are the attacker's favorite door |
| Training engagement | Completion, recency, knowledge-check performance | Weak alone, meaningful in combination |
| Exposure & privilege | Admin rights, payment authority, data access, public profile | Converts likelihood into potential impact |
| Targeting | Volume and sophistication of real attacks reaching the person | Attackers already run their own targeting model |
Two design notes on this table. First, difficulty matters: a click on a crude lure and a click on a pixel-perfect spear phish are not the same signal. The NIST Phish Scale exists precisely to rate how hard each phishing email is to detect — weight failures (and reports) by it. Second, reporting is a positive signal, and it should visibly lower risk. A team that clicks occasionally but reports within minutes is operationally safer than a team that never clicks and never reports, because reporting is what gives your SOC a chance to contain the messages that got through.
How to build one: a five-step method
1. Define the unit and the scale. Score individuals, but act on aggregates — teams, departments, locations. A 0–100 scale with named bands (low / moderate / elevated / critical) is easier to communicate than raw probabilities. Decide up front whether higher means riskier and never flip it.
2. Start with signals you already have. Most organizations can compute a first version from phishing simulation results, training records, MFA enrollment, and breach-exposure monitoring alone. Do not wait for perfect telemetry; a score built on four honest signals beats a roadmap slide with fourteen.
3. Weight behavior over attitudes, and recency over history. Observed actions (clicked, reported, reused a password) should dominate self-reported survey answers. Apply time decay: research covered in our ROI analysis shows behavior change fades within four to six months, so a failure from last week should weigh far more than one from last year — and improvements should show up quickly enough that people can feel the score respond to their effort.
4. Calibrate against reality. A score is a prediction; test it like one. Do high-scoring users actually generate more real incidents, more malware detections, more genuine phish reports against them? If the top decile of your score does not account for a disproportionate share of real events after a couple of quarters, revisit the weights. This is the difference between a risk model and a vanity metric.
5. Wire it to interventions, not punishments. The score's output should be actions: elevated-risk users automatically enter more frequent, harder simulations and short targeted training; critical-risk users with privileged access get conditional-access policies tightened. What it should never trigger is discipline. The moment a score is used punitively, employees stop reporting, your data quality collapses, and the score starts lying to you.
Pitfalls that sink scoring programs
The single-signal trap — using phishing click rate alone — produces a score that one lucky or unlucky template can swing by 30 points. The black-box trap — a score nobody can explain — dies in the first works-council meeting; every user-facing score should decompose into its contributing factors on demand. The league-table trap — publishing ranked lists of "riskiest employees" — converts a security tool into a morale problem. Share team-level trends widely and individual detail narrowly.
Privacy deserves explicit design, not an afterthought. Under GDPR, employee scoring is defensible when it is transparent, proportionate, and not used for automated decisions with significant effects — which aligns exactly with the "interventions, not punishments" rule above. Document the legitimate-interest assessment and involve the DPO early.
From score to program
A human risk score is the connective tissue of a human risk management program: simulations and real-world telemetry feed the score, the score drives adaptive training and controls, and the resulting behavior change flows back into the score. Run that loop continuously and you get what static awareness programs never deliver — a falsifiable, board-ready answer to "is our human risk going down?"
If you want to see what a continuously computed score looks like across simulation behavior, credential hygiene, and training signals, that is exactly what the NOUSEC Human Risk Score was built to do.
Frequently asked questions
What is a human risk score?
A human risk score is a single, continuously updated metric that quantifies how likely a person, team, or organization is to cause a security incident. It combines behavioral signals — phishing simulation results, report rates, credential hygiene, training engagement — with exposure factors such as access privileges and attack frequency, so security teams can prioritize interventions the way they already prioritize vulnerabilities.
Which signals should feed a human risk score?
Start with signals you can measure objectively and repeatedly: simulated phishing failure and report rates, real phishing reports, credential exposure in breach dumps, MFA adoption, and training completion and recency. Then weight in exposure: privileged access, public-facing roles, and how often the person is actually targeted. Self-reported survey data is the weakest input and should never dominate the score.
Is scoring employees legal under GDPR?
Generally yes, if done correctly. Scoring for security purposes is typically based on legitimate interest, but you must be transparent about what is measured, minimize the data collected, avoid fully automated decisions with significant effects on employees, and involve your works council or data protection officer where required. The score should drive support and training — not disciplinary action — both for legal safety and because punitive use destroys the reporting culture you need.
How often should a human risk score be updated?
Continuously, or as close to it as your data allows. Research shows security behavior decays within four to six months of training, and attack targeting shifts week to week. A score refreshed quarterly is a report card; a score refreshed continuously is an operational metric you can actually act on — triggering adaptive training the moment risk rises rather than months later.
NOUSEC simulates attacks across 8 channels and turns the results into one number your board can read.
Book a demo