Skip to content

Note

Generated by python -m gauntlet bench into results/RESULTS.md and copied here. Explore it interactively in the coverage explorer.

Benchmarks and results

ATT&CK Enterprise v19.2 - SigmaHQ r2026-07-01 - 96 OTRF Security-Datasets Windows recordings covering 54 techniques. Generated by python -m gauntlet bench.

Detection coverage on real recorded attack telemetry

Brackets are 95% Wilson intervals (54 techniques / 96 recordings are small samples).

Rule set Rules Technique coverage Exact-ID coverage Coverage w/o OTRF-citing rules Recordings detected Off-target rules / recording Off-target alerts / 10k events
legacy 6 5.6% [1.9, 15.1] 3.7% n/a 3.1% [1.1, 8.8] 0.15 0.23
sigma-core 1365 66.7% [53.4, 77.8] 57.4% 61.1% (-19 rules) 53.1% [43.2, 62.8] 1.55 11.18
sigma-all 2519 81.5% [69.2, 89.6] 66.7% 77.8% (-41 rules) 69.8% [60.0, 78.1] 3.49 53.68

Threat-weighted coverage per CTI profile

Profile ATT&CK groups legacy sigma-core sigma-all
ransomware 18 7.7% 67.8% 81.2%
espionage 57 8.3% 67.4% 81.6%
financial 24 8.6% 68.1% 82.7%
cloud 4 6.5% 68.8% 83.9%

Coverage sprint (demo scenario 5)

Hand-written baseline rules: 5.6% technique coverage (7.7% ransomware-weighted). Adding only the 10 cheapest-win SigmaHQ rules: 35.2% (44.6% ransomware-weighted).

Cheapest wins (greedy, ransomware-weighted, sigma-core)

# Rule New techniques Cumulative weighted coverage
1 CobaltStrike Service Installations - Security T1021, T1021.002 5.4%
2 Dumping of Sensitive Hives Via Reg.EXE T1003.002, T1003.004 10.3%
3 PCRE.NET Package Image Load T1059 13.6%
4 DPAPI Domain Backup Key Extraction T1003 16.7%
5 T1047 Wmiprvse Wbemcomn DLL Hijack T1047 19.5%
6 Meterpreter or Cobalt Strike Getsystem Service Installation - Security T1134.001, T1134.002 22.4%
7 LSASS Dump Keyword In CommandLine T1003.001 25.1%
8 NetNTLM Downgrade Attack - Registry T1112 27.6%
9 Successful Overpass the Hash Attempt T1550.002 30.1%
10 Suspicious Service Path Modification T1543.003 32.6%

Telemetry ablation (sigma-core): techniques lost if a channel is not collected

Channel Techniques lost Coverage without
microsoft-windows-sysmon/operational 5 57.4%
security 3 61.1%
microsoft-windows-powershell/operational 1 64.8%
windows powershell 0 66.7%
system 0 66.7%

Figures

Per-tactic coverage

Prioritization recall curves, ransomware

Per-tactic coverage (sigma-core)

Tactic Covered / recorded techniques
collection 0/1
credential-access 6/9
defense-impairment 1/3
discovery 1/7
execution 6/7
lateral-movement 5/6
persistence 5/6
privilege-escalation 5/7
stealth 7/8

Prioritization: CTI-ranked vs breadth-first emulation (leave-one-group-out)

Recall of a held-out actor's techniques after k emulations, universe = Windows techniques with an Atomic Red Team test. Lower steps to 80% is better.

Profile Strategy recall@10 recall@25 recall@50 steps to 50% steps to 80% AUC
ransomware (n=18) cti 0.070 0.201 0.376 64.3 118.9 0.718
ransomware (n=18) relevance 0.066 0.200 0.418 62.3 126.6 0.710
ransomware (n=18) prevalence 0.080 0.195 0.381 66.8 126.9 0.710
ransomware (n=18) breadth 0.061 0.140 0.255 99.8 194.4 0.568
ransomware (n=18) random 0.041 0.095 0.186 131.5 211.9 0.502
espionage (n=57) cti 0.114 0.309 0.516 52.2 103.2 0.761
espionage (n=57) relevance 0.109 0.305 0.515 53.2 106.0 0.755
espionage (n=57) prevalence 0.127 0.303 0.511 52.1 103.2 0.761
espionage (n=57) breadth 0.045 0.125 0.244 93.5 196.8 0.580
espionage (n=57) random 0.042 0.096 0.189 131.5 211.2 0.503
financial (n=24) cti 0.080 0.214 0.443 59.5 123.7 0.727
financial (n=24) relevance 0.068 0.213 0.427 60.2 134.8 0.710
financial (n=24) prevalence 0.091 0.238 0.436 58.8 124.5 0.727
financial (n=24) breadth 0.041 0.117 0.236 95.5 201.9 0.569
financial (n=24) random 0.042 0.096 0.188 131.8 211.8 0.502
cloud (n=4) cti 0.109 0.238 0.448 61.2 147.2 0.692
cloud (n=4) relevance 0.090 0.222 0.405 73.0 178.8 0.640
cloud (n=4) prevalence 0.118 0.253 0.437 61.8 123.0 0.728
cloud (n=4) breadth 0.062 0.129 0.265 94.5 196.2 0.577
cloud (n=4) random 0.040 0.095 0.184 130.8 210.7 0.502

Paired comparison over held-out groups (mean difference, 95% bootstrap CI; negative steps = CTI needs fewer emulations):

Profile Comparison Δ steps to 80% Δ AUC
ransomware cti minus breadth -75.5 [-85.0, -65.8] +0.150 [+0.136, +0.162]
ransomware cti minus prevalence -8.0 [-12.2, -3.9] +0.008 [+0.003, +0.013]
ransomware cti minus random -93.0 [-97.8, -88.4] +0.215 [+0.205, +0.226]
espionage cti minus breadth -93.6 [-102.2, -85.0] +0.181 [+0.165, +0.199]
espionage cti minus prevalence +0.0 [-1.0, +0.9] +0.000 [-0.001, +0.001]
espionage cti minus random -108.0 [-115.8, -100.6] +0.258 [+0.240, +0.276]
financial cti minus breadth -78.2 [-88.3, -68.4] +0.158 [+0.143, +0.174]
financial cti minus prevalence -0.8 [-3.6, +1.9] -0.000 [-0.002, +0.002]
financial cti minus random -88.1 [-93.8, -82.0] +0.224 [+0.212, +0.236]
cloud cti minus breadth -49.0 [-74.5, -23.0] +0.114 [+0.062, +0.167]
cloud cti minus prevalence +24.2 [+9.0, +44.8] -0.036 [-0.058, -0.018]
cloud cti minus random -63.4 [-94.8, -25.3] +0.190 [+0.134, +0.245]

Next-technique prediction (co-occurrence vs popularity, leave-one-group-out)

Level Model recall@5 recall@10 recall@20 MRR
sub_technique_level (n=165) cooccurrence 0.131 0.235 0.366 0.897
sub_technique_level (n=165) popularity 0.115 0.191 0.311 0.865
technique_level (n=168) cooccurrence 0.172 0.311 0.495 0.909
technique_level (n=168) popularity 0.168 0.278 0.456 0.926
Level co-occurrence − popularity recall@10 (95% CI) recall@10 mean ± sd over 5 hide-split seeds
sub_technique_level +0.044 [+0.026, +0.064] cooccurrence 0.237 ± 0.003; popularity 0.194 ± 0.002
technique_level +0.034 [+0.021, +0.048] cooccurrence 0.311 ± 0.005; popularity 0.277 ± 0.002

Sigma evaluator support

2519 Windows rules evaluated, 35 skipped as unsupported. Top reasons: service appxdeployment-server (9), service msexchange-management (8), category file_access (6), service iis-configuration (3), service dns-server-analytic (1), service smbclient-connectivity (1), service appxpackaging-om (1), service certificateservicesclient-lifecycle-system (1).