Phishing Simulation Best Practices: How to Test Employees Without Breaking Trust
Evidence-based phishing simulation best practices: cadence, difficulty, metrics beyond click rate, and a 90-day rollout plan for security teams.

Sixty percent of confirmed breaches still involve a human action, according to the Verizon Data Breach Investigations Report. That single number explains why phishing simulations have become a standard control in every serious security program. But it hides a second truth: a badly run simulation program can do real damage — eroding trust, suppressing incident reports, and teaching employees that the security team is out to trick them.
The difference between the two outcomes is not the tooling. It is the design of the program. This guide covers the practices that separate simulations that change behavior from simulations that generate resentment, based on what the current data from Verizon, KnowBe4, and NIST actually supports.
Why simulations work — when they're done right
The evidence for well-designed simulation programs is strong:
- KnowBe4's 2026 Phishing by Industry benchmark, drawn from 42 million simulated phishing emails across 14.8 million users at 64,000 organizations, found that untrained organizations start with a baseline Phish-Prone Percentage of 33.2% — roughly one in three employees interacts with a baseline test. That falls to 20.1% after ninety days and 4.2% after twelve months of combined training and simulation. We quote the published rates rather than a headline reduction percentage, because KnowBe4's own summaries of the same data have given different reduction figures.
- The Verizon DBIR found that employees who had received recent security awareness training were about four times more likely to report a phishing attempt than untrained peers.
- The 2026 DBIR adds a new warning: click rates on mobile devices are running roughly 40% higher than on desktop, which means simulation programs that only test corporate email on laptops are measuring the easy half of the problem.
Reporting is the metric that matters most in that list. A phishing email that gets reported in two minutes is an incident-response head start; a phishing email that gets silently deleted — or silently clicked — is not. Everything in your program design should push in one direction: more reports, faster.
The goal of a phishing simulation is not to catch employees failing. It is to make reporting a suspicious email as automatic as locking your screen.
Eight best practices for phishing simulations
1. Define the goal before the first send
"Reduce clicks" is not a goal; it is a side effect. Write down what you are actually trying to change: time-to-first-report, the percentage of employees who report rather than ignore, repeat-clicker counts in high-risk teams. Get executive sign-off on those goals and on the ground rules (no punitive consequences, no public shaming), so the program has cover when someone senior clicks — and someone senior will click.
2. Establish an honest baseline
Run your first campaign before any announcement or training, using a mid-difficulty template, across the whole organization. This is the only unbiased measurement you will ever get, and every future improvement claim depends on it. Resist the temptation to soften the first test to make the numbers look better.
3. Test monthly, randomize everything
Behavior decays and headcount churns. A monthly cadence with randomized send times, rotating templates, and staggered target groups prevents the two classic failure modes: the quarterly "phishing season" everyone learns to anticipate, and the single mass-send that gets outed in the company chat within minutes.
4. Control difficulty deliberately with the Phish Scale
If every simulation is an easy catch, your metrics flatter you. If every simulation is a spear-phish masterpiece, employees learn helplessness. The NIST Phish Scale gives you a repeatable way to rate each template's detection difficulty based on its cues and contextual alignment, so you can mix difficulty levels intentionally and interpret click rates fairly — a 12% click rate on a hard template is a better result than 4% on an obvious one.
5. Never punish a click
This is the hill to die on. Programs that dock bonuses, name-and-shame, or threaten termination for simulation failures consistently see the same result: reporting collapses, employees tip each other off, and the security team loses the relationship it needs most during a real incident. Coach individuals, escalate only on sustained patterns, and say all of this out loud in your program charter.
6. Measure reporting, not just clicking
Click rate is the vanity metric of phishing simulation. The numbers that predict real-world resilience are the reporting rate, the report-to-click ratio, and the median time between delivery and first report. A team with an 8% click rate and a 70% report rate is far safer than a team with a 3% click rate where nobody reports anything.
| Metric | What it tells you | Target direction |
|---|---|---|
| Click rate | Susceptibility to the lure | Down over time |
| Reporting rate | Willingness + ability to escalate | Up — this is the headline metric |
| Report-to-click ratio | Culture of reporting vs. silent failure | Above 1.0, then climbing |
| Time to first report | Incident-response head start | Under 5 minutes |
| Repeat-clicker rate | Where coaching should focus | Shrinking cohort |
| Credential submission rate | Worst-case severity | As close to zero as possible |
7. Go beyond email
Real attackers already have. Voice phishing, smishing, QR-code lures, and MFA-fatigue prompts are now routine initial-access techniques — and with mobile click rates running well above desktop, an email-only program tests yesterday's threat model. Fold SMS and QR scenarios into your rotation, and pair simulations with scenario training for the channels you cannot safely simulate. Our social engineering statistics roundup tracks how quickly the channel mix is shifting.
8. Segment by risk, not just by org chart
Finance approves payments, executive assistants control calendars and travel, IT admins hold privileged access — each faces different lures and deserves different scenarios. Mature programs assign each employee a risk level based on role, access, and past simulation behavior, then tune scenario difficulty and frequency accordingly. This is exactly the problem a Human Risk Score is built to solve: turning scattered simulation results into a per-person risk signal you can act on.
A 90-day rollout plan
Days 1–15: Foundation. Write the program charter (goals, ground rules, no-punishment policy), get executive sign-off, deploy and publicize the report-phishing button. If employees have no one-click way to report, fix that before anything else.
Days 16–30: Baseline. Run the unannounced baseline campaign with a Phish Scale–rated mid-difficulty template. Record click rate, report rate, and time-to-first-report. Brief leadership on the results without naming individuals.
Days 31–60: First training cycle. Announce the ongoing program. Deliver short, role-relevant training — the DBIR data suggests recency matters, so favor frequent micro-sessions over an annual marathon. Run the second campaign with rotated templates and randomized sends.
Days 61–90: Operationalize. Move to the monthly rhythm. Add one non-email scenario. Start the repeat-clicker coaching loop. Publish a program dashboard showing trend lines for click and report rates — transparency about aggregate results builds the trust that punitive programs destroy.
From day 91 onward, the program becomes a flywheel: simulate, measure, coach, increase difficulty, repeat. Platforms like NOUSEC's simulation engine automate the mechanics — template rotation, difficulty scaling, per-user scheduling, instant just-in-time coaching moments — so a small security team can run an enterprise-grade program.
Common mistakes that sink programs
The same failure patterns appear again and again: testing only at quarterly intervals that everyone anticipates; judging the program on click rate alone; using deeply emotional lures (fake bonuses, fake layoffs) that generate backlash instead of learning; excluding executives from testing; and treating the simulation as the product rather than the measurement. Phishing, spoofing, and their variants remain the most-reported cybercrime category in the FBI's IC3 data, with overall cybercrime losses passing $16 billion in the latest report — the threat is real enough without manufacturing outrage internally.
A phishing simulation program is ultimately a trust-building exercise conducted at scale. Run it with clear goals, honest measurement, and zero punishment, and within a year the data says you can expect clicks to fall by more than 80% while reports multiply — which is exactly the trade you want when the real thing lands in an inbox.
Frequently asked questions
How often should we run phishing simulations?
At least monthly for most organizations. Quarterly campaigns leave gaps of 60+ days in which behavior decays and new joiners are never tested. A monthly cadence with rotating templates and randomized send windows keeps data fresh without overwhelming employees, and high-risk groups such as finance or executive assistants can be tested more frequently.
Should employees be punished for clicking a simulated phish?
No. Punitive consequences reliably backfire: they suppress incident reporting, push employees to warn each other about live campaigns, and destroy the trust a security team depends on. Treat every click as a coaching moment with immediate, short, relevant training — and reserve escalation for patterns, not single events.
What is a good phishing simulation click rate?
Untrained organizations typically start around 33% on baseline tests according to KnowBe4's 2025 benchmark, while Verizon's DBIR reports a median click-through of roughly 1.5% in mature simulation programs. Rather than chasing a single number, track the trend in your click rate, your reporting rate, and your report-to-click ratio over time.
Are phishing simulations still worth it now that attackers use AI?
Yes — arguably more than ever. AI-generated lures raise the quality of real attacks, so employees need regular, realistic practice at spotting harder phish. Simulations remain the only safe way to rehearse that skill and to measure whether reporting behavior is improving before a real campaign hits.
NOUSEC simulates attacks across 8 channels and turns the results into one number your board can read.
Book a demo