Limitations & roadmap¶
REVENANT is a research prototype. These are the honest boundaries; several are inherent to the available public ground truth or need hardware/data that this project deliberately does not require.
Limitations¶
- B1 covers process lineage and process→action edges only. Cross-entity edges (dropped file executed, logon session) have no public ground truth, so they are untested.
- B2 uses REVENANT's own heuristics to decide "found", so it measures ranking and grouping, not detection quality. The gain over a flat suspicion-sorted list is small.
- B3 precision is low, and the labels come from file names. Service-crash style log suppression is not modelled.
.evtxparsing through python-evtx is slow (about 21 records/s on the benchmark machine). Converting to JSON first withevtx_dumpis much faster.- No published numbers exist for these exact tasks on these corpora. The comparisons are against baselines only.
- Anti-forensics checks on MFT against
$LogFile/$UsnJrnlare limited to plaso's$SI/$FNfields. There is no raw NTFS parser. - The LLM report-drafting assistant from the spec is intentionally not built. No claim comes from a model (ADR 0001).
Roadmap¶
- [ ]
evtx_dump/ Hayabusa JSON ingest for fast.evtxhandling - [ ] Event-log service-crash and MRU-deletion tamper indicators (the B3 misses)
- [ ] Labelled cross-entity edges from a self-captured lab scenario
- [ ] Timesketch importer/exporter
- [ ] PostgreSQL custody backend
Design stance on AI¶
The core uses rules and a graph on purpose. Courts distrust black boxes, so explainability is a design constraint. Every edge names the rule that produced it, and every report line cites an event id and hash. The only learned part is the per-rule calibration table, fitted from public ground truth and shipped as readable JSON.