
Nimbus Vestige — Updated!
A forensic reconstruction engine for cloud and identity incident response.
Nimbus Vestige (NV)
A forensic reconstruction engine for cloud and identity incident response.
Given fragmented cloud, SaaS, and identity telemetry — control-plane logs, sign-in events, token and consent activity — NV reconstructs how an intrusion most likely moved across accounts and services, enumerates the other most-likely paths it could have taken, and reports every step with calibrated, explainable confidence. It refuses to assert what the evidence cannot support.
Detection tells you that something happened. Nimbus Vestige tells you how — and what else.
Why this exists
Modern intrusions do not "break in." They log in. Identity is now the primary attack vehicle, implicated in the large majority of cloud IR investigations, and most intrusions span multiple surfaces — identity plus cloud plus SaaS plus endpoint. The telemetry that records them is fragmented and inconsistent, which already forces responders to reconstruct the story by hand from incomplete data.
That manual reconstruction is slow, and it fails in a predictable way: the responder anchors on the first plausible narrative and misses the real one. NV automates the reconstruction and directly attacks that failure mode by always presenting the ranked space of plausible paths, not a single story.
Crucially, NV does this with honesty as the product. The market's stated enemy is black-box confidence — a tool that asserts a conclusion without showing its work. Every number NV produces is traceable to the specific evidence that earned it, and anything it cannot support is withheld rather than guessed.
What it is — and is not
NV is not a detector, a scanner, an auditor, or an attack-runner. Those tools tell you that something happened and measure it. NV works backward — abductive reconstruction of the mechanism from patterns, across a fragmented identity estate.
- Not a detector — it doesn't fire alerts on activity; it explains how activity fits together.
- Not a SIEM — it enters through a narrow wedge (post-incident narrative reconstruction), not as a log platform.
- Not a rule engine — when nothing matches a known pattern, it still reconstructs, degrading gracefully instead of going blind.
Why it is valid long-term
-
Blue team regenerates; red team commoditizes. Offensive tooling finds a finite, patchable set of holes and is folding into automated CI/CD safety pipelines. Reconstructing how an intrusion happened never resolves — attackers keep inventing, so the need is permanent and self-renewing.
-
The engine is substrate-independent. The core logic — reconstruct the mechanism from patterns, with calibrated confidence and a stop rail — is committed to cloud/identity first, but ports to network, endpoint, and OT later. The target can change without rewriting the thesis.
-
Reconstruction enables hardening. Once you know how they got in — and which other doors were open — you build the defenses. The untaken-but-plausible paths are a hardening backlog, often worth more than the reconstruction itself, because most breaches exploit preventable exposure, not novel tradecraft.
-
It degrades gracefully on the novel intrusion — the very incident that matters most. A signature/rule engine goes blind on a zero-day; NV's abductive core still produces a most-likely path, honestly flagged as lower-confidence.
-
Honesty is a moat. Calibrated, case-based confidence with ranked alternatives is exactly what the market says it wants and what black-box competitors structurally cannot offer without redesign.
Architecture
Raw provider logs → normalized event/entity graph → reconstruction engine → JSON → GUI.
O365 / Entra audit logs nv_extract_identity_events.py
AWS CloudTrail (IAM/STS/S3) nv_extract_cloudtrail_events.py (ingest + normalize)
▼
normalized identity events (JSONL) ← one event model, any substrate
│ nv/graph.py (typed per-actor timelines)
▼
reconstruction engine
├─ nv/patterns.py known layer: ATT&CK identity pattern library
├─ nv/providers.py provider packs: per-substrate op→ATT&CK vocab (Entra + AWS)
├─ nv/validation.py Phase 3: provenance + integrity gate on every pattern
├─ nv/feeds.py live ATT&CK STIX / TAXII / Sigma clients + scheduler
├─ nv/confidence.py evidence-corroboration scorer (calibratable weights)
├─ nv/calibration.py fit + measure confidence against labeled ground truth
├─ nv/reconstruct.py most-likely chain + ranked competing paths + trust floor
└─ nv/scope.py authorized-scope gate (refuses unauthorized tenants/accounts)
│ run_nv.py
▼
reconstruction.json ──► nv_gui.html (embedding canvas + two-column view + trust slider)
The GUI is a pure view layer — it renders engine output and lets the analyst move the trust floor. It contains no reconstruction logic of its own.
The confidence model (the credibility of the whole product)
Every step carries a confidence in [0, 1], built by evidence corroboration:
confidence = per-technique base rate
+ bonus for each independent corroborating signal
(source IP, device, successful outcome, temporal adjacency, broad-consent flags)
− penalty for missing signals (e.g. no source IP to corroborate origin)
- verified — confidence ≥ 0.70 and two or more independent signals agree.
- pattern-only — above the trust floor but weakly or singly supported.
- withheld — below the trust floor; shown greyed, never silently guessed or hidden.
Path scores compose their links by geometric mean, so one weak link honestly drags a path down rather than being averaged away.
Every weight above — the per-technique base rates, the per-signal bonuses, the missing
penalties — lives in a single PARAMS table in nv/confidence.py. Those are NV's v1
priors, and they are calibratable: nv/calibration.py can refit them against labeled
ground truth and the engine will load the result, falling back to the v1 priors when no
calibration is supplied (see below).
The trust floor
The withhold rule, baked into the engine's output contract — not just the UI. Below the floor, a step is withheld. Dragging the floor is the honesty-versus-coverage tradeoff made physical: up for strict/high-trust, down for permissive/high-coverage. The default floor is an editorial decision about where NV sits before the analyst touches it.
Calibrating confidence against ground truth
The v1 weights are defensible, but they are priors. nv/calibration.py tunes them
against labeled ground truth — events where we know which ones were part of the
intrusion and which were benign — and, crucially, measures whether tuning actually
helped, so a calibration is adopted only if it lowers calibration error rather than just
moving numbers around.
Calibrated confidence means one thing: NV's confidence for a step should equal the probability that the step really was on the attack path. So each labeled event's confidence is treated as a predicted probability, and the harness fits:
- per-operation base rates — the empirical hit rate for that operation, Bayesian- shrunk toward the v1 prior so small samples don't overfit; and
- per-signal bonuses and the two penalties — by bounded coordinate descent that minimizes Brier score with an L2 pull back toward the v1 values.
The verified/withheld thresholds are policy, not calibration, and are left untouched.
It reports Brier score, log-loss, ECE, AUC, a reliability table, and a trust-floor sweep (true-step recall vs benign false-positive rate at each floor), before and after — so the honesty-versus-coverage tradeoff is legible rather than a single number.