Skip to content

CLI and API reference

CLI

python -m gauntlet profiles                         # ATT&CK groups per threat profile
python -m gauntlet plan --profile ransomware --top 15
python -m gauntlet manifest --profile ransomware --top 10 --out plan.json   # DRY RUN
python -m gauntlet predict --observed T1566.001,T1059.001 -k 10
python -m gauntlet replay --profile ransomware --top 15 --ruleset sigma-core        --navigator layer.json --json report.json [--baseline old.json]      # exit 2 on regression
python -m gauntlet bench                            # full benchmark -> results/
python -m gauntlet kb                               # rebuild the shipped ATT&CK KB
python -m gauntlet sim --profile ransomware         # offline simulation demo

Run python -m gauntlet <command> --help for all options.

Python API

gauntlet.prioritize

CTI prioritizer on real ATT&CK data.

Threat profiles are sets of real ATT&CK groups. By default a profile is selected reproducibly from group descriptions (e.g. every group whose ATT&CK description mentions ransomware), or explicitly by group ids/names.

relevance(T)  = share of profile groups with a documented procedure for T
prevalence(T) = share of *all* ATT&CK groups with a documented procedure for T
score(T)      = relevance(T) * prevalence(T) / max(prevalence)     ("cti")

Alternative strategies (for the research question "does CTI-prioritized emulation reach relevant coverage faster than breadth-first?"): breadth (ATT&CK id order), random, prevalence (global only), relevance (profile only), cti (product, default).

Leave-one-group-out evaluation of prioritization strategies.

For each group g in the profile: relevance is computed from the other profile groups, prevalence from all groups except g, the candidate list is universe (techniques GAUNTLET can emulate), and we measure how quickly each ordering covers g's techniques inside the universe.

profile_groups(kb, profile)

Resolve a named profile or a comma list of group ids / names / aliases.

rank(candidates, rel, prev, strategy='cti', seed=0)

Order candidate technique ids by a strategy (ties broken by id, deterministic).

gauntlet.coverage

Coverage scoring on replayed real telemetry, gap ranking and ATT&CK Navigator export.

Ground truth is the dataset's ATT&CK mapping. A rule firing on a recording is on-target when one of its ATT&CK tags is in the same technique family as one of the recording's techniques (T1059 / T1059.001 / T1059.003 are one family); a stricter exact-id mode is also available. Rules that fire but are not on-target are counted as off-target alerts: the recordings contain lab background activity and other attack steps, so this is an upper bound on the false-positive burden, not an exact FP rate.

CoverageSummary dataclass

technique_coverage property

Share of techniques with >=1 on-target detection (partial counts as covered).

channel_ablation(results, rules, channels, kb=None)

Techniques lost if a telemetry channel were not collected (value of each log source).

greedy_rule_selection(results, rules, weights=None, top=10, kb=None)

'Cheapest wins': the few rules that buy the most (weighted) technique coverage.

navigator_layer(summary, attack_version='19', name=None, weights=None)

ATT&CK Navigator layer (format 4.5): green detected / orange partial / red missed.

score(ruleset, results, rules, kb=None, drop_channels=(), exact=False)

Score replay results. With kb, revoked ATT&CK ids in dataset labels and rule tags (e.g. T1086 -> T1059.001) are mapped to their current replacement first.

gauntlet.sigma

A dependency-light Sigma rule evaluator for Windows event logs.

This is not a SIEM backend: it evaluates SigmaHQ YAML rules directly against JSON event records (e.g. OTRF Security-Datasets / Mordor recordings), which is what a coverage benchmark needs. Supported:

  • detection items: field maps (AND of keys, OR of list values), lists of maps (OR), keyword lists (match anywhere in the event), null values
  • value modifiers: contains, startswith, endswith, all, windash, re (+ i/m/s), base64, base64offset, utf16le/utf16be/utf16/wide, cased, cidr, exists, gt/gte/lt/lte, fieldref; plain values support */? wildcards
  • conditions: and / or / not / parentheses, 1 of X*, all of X*, 1 of them, all of them
  • logsources: Sysmon categories, Security 4688 process creation (with field mapping), PowerShell script/module/classic logs and a set of Windows services

Not supported (the rule is reported as unsupported rather than silently passing): aggregations (| count()), correlation rules, expand placeholders, non-Windows products. See docs/adr/0002-sigma-evaluator.md.

SigmaRule dataclass

cites(pattern)

True if the rule's references/description match pattern (case-insensitive).

load_rules(source, subdir_prefix='')

Load (and memoise per process) Sigma rules; see :func:_load_rules.

gauntlet.replay

Replay recorded telemetry through a rule set (the real-data 'OBSERVE + DETECT' stage).

Instead of executing techniques, GAUNTLET replays public recordings of them (OTRF Security-Datasets) through Sigma rules. Rules are indexed by (channel, EventID) so each event is only tested against rules whose logsource can apply to it.

load_ruleset(spec)

Rule-set spec: sigma-all:<zip-or-dir>, sigma-core:<zip-or-dir> or legacy:<dir>.

sigma-core keeps rules with status stable/test and level high/critical -- roughly what a SOC would page on.

replay_many(spec, datasets, workers=None, cache=None, progress=True)

Replay many datasets in parallel; results cached to cache (JSON) if given.

gauntlet.predict

Technique co-occurrence model: "given what we've seen, what comes next?"

Item-item collaborative filtering over ATT&CK group -> technique sets. For an observed set S, a candidate technique t scores

score(t | S) = sum_{s in S} cooc(s, t) / sqrt(n(s) * n(t))   (cosine)

plus a tiny popularity prior to break ties. Evaluated with leave-one-group-out: hide half of a group's techniques, predict them from the other half using a model trained on every other group, compare with a popularity baseline.

evaluate(sets, ks=(5, 10, 20), hide=0.5, seed=0, min_size=10)

Leave-one-group-out recall@k / MRR, co-occurrence vs popularity.

evaluate_seeds(sets, seeds=range(5), **kw)

Repeat :func:evaluate over several hide-split seeds; report seed 0 plus mean/sd across seeds.

gauntlet.stats

Small, dependency-free uncertainty helpers used by the benchmark.

  • :func:wilson -- 95% Wilson score interval for a proportion (coverage, recall).
  • :func:bootstrap_ci -- percentile bootstrap CI of the mean of per-unit values (e.g. one value per held-out ATT&CK group).
  • :func:paired_bootstrap_ci -- CI of the mean difference between two strategies measured on the same units (like-for-like comparison).

All resampling uses a fixed seed so results are reproducible.

paired_bootstrap_ci(a, b, n_boot=2000, alpha=0.05, seed=0)

Mean of a - b with a bootstrap CI; a and b are aligned per unit.