Note
Generated by python -m gauntlet bench into results/RESULTS.md and copied here. Explore it interactively in the coverage explorer.
Benchmarks and results
ATT&CK Enterprise v19.2 - SigmaHQ r2026-07-01 - 96 OTRF Security-Datasets Windows recordings covering 54 techniques. Generated by python -m gauntlet bench.
Detection coverage on real recorded attack telemetry
Brackets are 95% Wilson intervals (54 techniques / 96 recordings are small samples).
| Rule set | Rules | Technique coverage | Exact-ID coverage | Coverage w/o OTRF-citing rules | Recordings detected | Off-target rules / recording | Off-target alerts / 10k events |
|---|---|---|---|---|---|---|---|
| legacy | 6 | 5.6% [1.9, 15.1] | 3.7% | n/a | 3.1% [1.1, 8.8] | 0.15 | 0.23 |
| sigma-core | 1365 | 66.7% [53.4, 77.8] | 57.4% | 61.1% (-19 rules) | 53.1% [43.2, 62.8] | 1.55 | 11.18 |
| sigma-all | 2519 | 81.5% [69.2, 89.6] | 66.7% | 77.8% (-41 rules) | 69.8% [60.0, 78.1] | 3.49 | 53.68 |
Threat-weighted coverage per CTI profile
| Profile | ATT&CK groups | legacy | sigma-core | sigma-all |
|---|---|---|---|---|
| ransomware | 18 | 7.7% | 67.8% | 81.2% |
| espionage | 57 | 8.3% | 67.4% | 81.6% |
| financial | 24 | 8.6% | 68.1% | 82.7% |
| cloud | 4 | 6.5% | 68.8% | 83.9% |
Coverage sprint (demo scenario 5)
Hand-written baseline rules: 5.6% technique coverage (7.7% ransomware-weighted). Adding only the 10 cheapest-win SigmaHQ rules: 35.2% (44.6% ransomware-weighted).
Cheapest wins (greedy, ransomware-weighted, sigma-core)
| # | Rule | New techniques | Cumulative weighted coverage |
|---|---|---|---|
| 1 | CobaltStrike Service Installations - Security | T1021, T1021.002 | 5.4% |
| 2 | Dumping of Sensitive Hives Via Reg.EXE | T1003.002, T1003.004 | 10.3% |
| 3 | PCRE.NET Package Image Load | T1059 | 13.6% |
| 4 | DPAPI Domain Backup Key Extraction | T1003 | 16.7% |
| 5 | T1047 Wmiprvse Wbemcomn DLL Hijack | T1047 | 19.5% |
| 6 | Meterpreter or Cobalt Strike Getsystem Service Installation - Security | T1134.001, T1134.002 | 22.4% |
| 7 | LSASS Dump Keyword In CommandLine | T1003.001 | 25.1% |
| 8 | NetNTLM Downgrade Attack - Registry | T1112 | 27.6% |
| 9 | Successful Overpass the Hash Attempt | T1550.002 | 30.1% |
| 10 | Suspicious Service Path Modification | T1543.003 | 32.6% |
Telemetry ablation (sigma-core): techniques lost if a channel is not collected
| Channel | Techniques lost | Coverage without |
|---|---|---|
| microsoft-windows-sysmon/operational | 5 | 57.4% |
| security | 3 | 61.1% |
| microsoft-windows-powershell/operational | 1 | 64.8% |
| windows powershell | 0 | 66.7% |
| system | 0 | 66.7% |
Figures


Per-tactic coverage (sigma-core)
| Tactic | Covered / recorded techniques |
|---|---|
| collection | 0/1 |
| credential-access | 6/9 |
| defense-impairment | 1/3 |
| discovery | 1/7 |
| execution | 6/7 |
| lateral-movement | 5/6 |
| persistence | 5/6 |
| privilege-escalation | 5/7 |
| stealth | 7/8 |
Prioritization: CTI-ranked vs breadth-first emulation (leave-one-group-out)
Recall of a held-out actor's techniques after k emulations, universe = Windows techniques with an Atomic Red Team test. Lower steps to 80% is better.
| Profile | Strategy | recall@10 | recall@25 | recall@50 | steps to 50% | steps to 80% | AUC |
|---|---|---|---|---|---|---|---|
| ransomware (n=18) | cti | 0.070 | 0.201 | 0.376 | 64.3 | 118.9 | 0.718 |
| ransomware (n=18) | relevance | 0.066 | 0.200 | 0.418 | 62.3 | 126.6 | 0.710 |
| ransomware (n=18) | prevalence | 0.080 | 0.195 | 0.381 | 66.8 | 126.9 | 0.710 |
| ransomware (n=18) | breadth | 0.061 | 0.140 | 0.255 | 99.8 | 194.4 | 0.568 |
| ransomware (n=18) | random | 0.041 | 0.095 | 0.186 | 131.5 | 211.9 | 0.502 |
| espionage (n=57) | cti | 0.114 | 0.309 | 0.516 | 52.2 | 103.2 | 0.761 |
| espionage (n=57) | relevance | 0.109 | 0.305 | 0.515 | 53.2 | 106.0 | 0.755 |
| espionage (n=57) | prevalence | 0.127 | 0.303 | 0.511 | 52.1 | 103.2 | 0.761 |
| espionage (n=57) | breadth | 0.045 | 0.125 | 0.244 | 93.5 | 196.8 | 0.580 |
| espionage (n=57) | random | 0.042 | 0.096 | 0.189 | 131.5 | 211.2 | 0.503 |
| financial (n=24) | cti | 0.080 | 0.214 | 0.443 | 59.5 | 123.7 | 0.727 |
| financial (n=24) | relevance | 0.068 | 0.213 | 0.427 | 60.2 | 134.8 | 0.710 |
| financial (n=24) | prevalence | 0.091 | 0.238 | 0.436 | 58.8 | 124.5 | 0.727 |
| financial (n=24) | breadth | 0.041 | 0.117 | 0.236 | 95.5 | 201.9 | 0.569 |
| financial (n=24) | random | 0.042 | 0.096 | 0.188 | 131.8 | 211.8 | 0.502 |
| cloud (n=4) | cti | 0.109 | 0.238 | 0.448 | 61.2 | 147.2 | 0.692 |
| cloud (n=4) | relevance | 0.090 | 0.222 | 0.405 | 73.0 | 178.8 | 0.640 |
| cloud (n=4) | prevalence | 0.118 | 0.253 | 0.437 | 61.8 | 123.0 | 0.728 |
| cloud (n=4) | breadth | 0.062 | 0.129 | 0.265 | 94.5 | 196.2 | 0.577 |
| cloud (n=4) | random | 0.040 | 0.095 | 0.184 | 130.8 | 210.7 | 0.502 |
Paired comparison over held-out groups (mean difference, 95% bootstrap CI; negative steps = CTI needs fewer emulations):
| Profile | Comparison | Δ steps to 80% | Δ AUC |
|---|---|---|---|
| ransomware | cti minus breadth | -75.5 [-85.0, -65.8] | +0.150 [+0.136, +0.162] |
| ransomware | cti minus prevalence | -8.0 [-12.2, -3.9] | +0.008 [+0.003, +0.013] |
| ransomware | cti minus random | -93.0 [-97.8, -88.4] | +0.215 [+0.205, +0.226] |
| espionage | cti minus breadth | -93.6 [-102.2, -85.0] | +0.181 [+0.165, +0.199] |
| espionage | cti minus prevalence | +0.0 [-1.0, +0.9] | +0.000 [-0.001, +0.001] |
| espionage | cti minus random | -108.0 [-115.8, -100.6] | +0.258 [+0.240, +0.276] |
| financial | cti minus breadth | -78.2 [-88.3, -68.4] | +0.158 [+0.143, +0.174] |
| financial | cti minus prevalence | -0.8 [-3.6, +1.9] | -0.000 [-0.002, +0.002] |
| financial | cti minus random | -88.1 [-93.8, -82.0] | +0.224 [+0.212, +0.236] |
| cloud | cti minus breadth | -49.0 [-74.5, -23.0] | +0.114 [+0.062, +0.167] |
| cloud | cti minus prevalence | +24.2 [+9.0, +44.8] | -0.036 [-0.058, -0.018] |
| cloud | cti minus random | -63.4 [-94.8, -25.3] | +0.190 [+0.134, +0.245] |
Next-technique prediction (co-occurrence vs popularity, leave-one-group-out)
| Level | Model | recall@5 | recall@10 | recall@20 | MRR |
|---|---|---|---|---|---|
| sub_technique_level (n=165) | cooccurrence | 0.131 | 0.235 | 0.366 | 0.897 |
| sub_technique_level (n=165) | popularity | 0.115 | 0.191 | 0.311 | 0.865 |
| technique_level (n=168) | cooccurrence | 0.172 | 0.311 | 0.495 | 0.909 |
| technique_level (n=168) | popularity | 0.168 | 0.278 | 0.456 | 0.926 |
| Level | co-occurrence − popularity recall@10 (95% CI) | recall@10 mean ± sd over 5 hide-split seeds |
|---|---|---|
| sub_technique_level | +0.044 [+0.026, +0.064] | cooccurrence 0.237 ± 0.003; popularity 0.194 ± 0.002 |
| technique_level | +0.034 [+0.021, +0.048] | cooccurrence 0.311 ± 0.005; popularity 0.277 ± 0.002 |
Sigma evaluator support
2519 Windows rules evaluated, 35 skipped as unsupported. Top reasons: service appxdeployment-server (9), service msexchange-management (8), category file_access (6), service iis-configuration (3), service dns-server-analytic (1), service smbclient-connectivity (1), service appxpackaging-om (1), service certificateservicesclient-lifecycle-system (1).