Security Awareness Training

The industry practice of teaching employees to resist phishing and social engineering — and, as of the mid-2020s, one of the best-measured and most uncomfortable evidence bases in security: at organizational scale, current training approaches barely work, and the field’s own measurements explain why vendors can claim otherwise.

What it is

“Security awareness training” (SAT) covers a spectrum: annual compliance e-learning (lecture videos + quiz), simulated phishing campaigns with embedded training (click a lure, get immediately routed to a training page), interactive/gamified exercises, just-in-time warnings, and — at the far end — behavioral-science consultancies that diagnose why employees fail (knowledge, opportunity, motivation) rather than re-lecturing them. Legal and standards mandates drive most of it: HIPAA requires SAT in healthcare, PCI DSS and GDPR encourage or mandate it, ISO 27001 calls for awareness programs.1

The evidence base (2022–2026)

Three large-scale field studies anchor the current picture:

  • Lain, Kostiainen, Čapkun (IEEE S&P 2022) — 15 months, 14,773 employees of one company. Voluntary embedded training not only failed to improve resilience; employees routed to the training page were more likely to click again (repeat-clickers: 801 with training vs 647 without). Email warnings, by contrast, worked well, and crowd-sourced reporting proved practical at scale — fast campaign detection with acceptable operational load.2
  • Ho et al. (IEEE S&P 2025) — 19,500+ healthcare employees; minimal training impact (via Rozema & Davis’s review).3
  • Rozema & Davis (arXiv 2025/2026, USENIX Security 2026 artifact) — 12,511 fintech employees, randomized lecture vs lecture+interactive vs control. Overall click rate 10.4%; control 9.8%, trained 10.5%. No significant training effect on clicks (p=0.450) or reporting (p=0.417); all effect sizes η² < 0.01.4

The same group’s CCS 2024 follow-up (Content, Nudges and Incentives) isolates the mechanism: embedded training’s modest effect comes from the nudge — the periodic reminder of the threat — not from the content, which employees rarely consume. “Phishing is an attention problem, rather than a knowledge one, even for the most susceptible employees.”5

The NIST Phish Scale — and compliance theater

Rozema & Davis provide the first enterprise-scale validation of the NIST Phish Scale, which rates lure difficulty on two dimensions: phishing cues (observable errors — spelling, suspicious URLs) and premise alignment (fit to the recipient’s normal mail flow). Difficulty strongly predicts clicks: 7.0% (easy) → 8.7% (medium) → 15.0% (hard).6

That validation cuts twice. If lure difficulty determines outcomes more than training does, then whoever picks the templates controls the metrics. Vendor libraries ship without standardized difficulty ratings, so an organization can — accidentally or deliberately — run easy lures and report year-over-year “improvement.” Rozema & Davis name this as a concrete mechanism for compliance theater: optimizing metrics that satisfy regulators rather than reducing risk.7

What works instead

The convergent recommendation is not “stop training” but demote it within defense-in-depth:

  1. Technical controls first — filtering, attachment analysis, URL protection, and phishing-resistant authentication (passkeys / FIDO2) remove the burden from human judgment.
  2. Targeted training — role- or scenario-specific, not generic company-wide.
  3. Reporting culture as collective defense — Lain et al. showed crowd-sourced detection works; Rozema & Davis’s Organizational Inoculation Index found 36–55% of campaigns had a report precede the first click, independent of training condition. Median time-to-report: 21 minutes.
  4. Warnings and design interventions — banner warnings measurably reduce repeat clicks.89

The behavioral-science turn (market context)

The market’s response to weak SAT outcomes has been to rebrand upward: “human risk management,” culture measurement, and behavioral consultancies that treat security failure as an organizational-behavior problem. Nathan’s commissioned market research (2026-08-05, conversation-sourced) found a structural split: in the Netherlands, behavioral-security expertise lives in consultancies (Awareways, Northwave, Infosequre, BV Cyber — Inge Wetzer’s school); in North America it lives in platforms (Living Security, CybSafe, Hoxhunt) with consulting as an afterthought. Whether the behavioral turn fixes the evidence problem or just reprices it is an open question — the studies above measured conventional training, not the consultancy model, which mostly lacks comparable independent evaluation.

Open questions

  • Does consultancy-grade behavioral programming (culture scans, motivation/opportunity diagnosis) outperform lecture-and-simulate under the same randomized measurement? Nobody has run it at scale.
  • Do GenAI-crafted lures break the NIST Phish Scale itself (fewer cues to rate), and do they make any training hopeless? Rozema & Davis flag this explicitly — one author failed his own organization’s GenAI-lure test.
  • Will regulators (HIPAA, PCI DSS, cyber-insurance underwriting) adapt to the evidence, or does mandate-driven demand persist regardless of efficacy?

Connections

  • nudges-vs-prices — the policy-economics twin: behavioral interventions evaluated only relative to alternatives. Security training is the enterprise instance of the same question, with the same answer shape (small effects, cheap at the margin, displaced at scale)
  • passkeys — the phishing-resistant technical control the evidence points toward
  • institutionally-constrained-technology-adoption — compliance theater as metric-optimization over outcomes; same shape at the governance layer
  • bullshit-jobs — annual mandatory training nobody believes in, sustained by mandate, is Graeber-adjacent organizational ritual

Sources

Footnotes

  1. 2026 — Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale

  2. 2022 — Phishing in Organizations: Findings from a Large-Scale and Long-Term Study

  3. 2026 — Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale

  4. 2026 — Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale

  5. 2022 — Phishing in Organizations: Findings from a Large-Scale and Long-Term Study

  6. 2026 — Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale

  7. 2026 — Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale

  8. 2022 — Phishing in Organizations: Findings from a Large-Scale and Long-Term Study

  9. 2026 — Anti-Phishing Training (Still) Does Not Work: A Large-Scale Reproduction of Phishing Training Inefficacy Grounded in the NIST Phish Scale