
A forensic reconstruction engine for cloud and identity incident response.
A forensic reconstruction engine for cloud and identity incident response.
Given fragmented cloud, SaaS, and identity telemetry — control-plane logs, sign-in events, token and consent activity — NV reconstructs how an intrusion most likely moved across accounts and services, enumerates the other most-likely paths it could have taken, and reports every step with calibrated, explainable confidence. It refuses to assert what the evidence cannot support.
Detection tells you that something happened. Nimbus Vestige tells you how — and what else.
Modern intrusions do not "break in." They log in. Identity is now the primary attack vehicle, implicated in the large majority of cloud IR investigations, and most intrusions span multiple surfaces — identity plus cloud plus SaaS plus endpoint. The telemetry that records them is fragmented and inconsistent, which already forces responders to reconstruct the story by hand from incomplete data.
That manual reconstruction is slow, and it fails in a predictable way: the responder anchors on the first plausible narrative and misses the real one. NV automates the reconstruction and directly attacks that failure mode by always presenting the ranked space of plausible paths, not a single story.
Crucially, NV does this with honesty as the product. The market's stated enemy is black-box confidence — a tool that asserts a conclusion without showing its work. Every number NV produces is traceable to the specific evidence that earned it, and anything it cannot support is withheld rather than guessed.
NV is not a detector, a scanner, an auditor, or an attack-runner. Those tools tell you that something happened and measure it. NV works backward — abductive reconstruction of the mechanism from patterns, across a fragmented identity estate.
Blue team regenerates; red team commoditizes. Offensive tooling finds a finite, patchable set of holes and is folding into automated CI/CD safety pipelines. Reconstructing how an intrusion happened never resolves — attackers keep inventing, so the need is permanent and self-renewing.
The engine is substrate-independent. The core logic — reconstruct the mechanism from patterns, with calibrated confidence and a stop rail — is committed to cloud/identity first, but ports to network, endpoint, and OT later. The target can change without rewriting the thesis.
Reconstruction enables hardening. Once you know how they got in — and which other doors were open — you build the defenses. The untaken-but-plausible paths are a hardening backlog, often worth more than the reconstruction itself, because most breaches exploit preventable exposure, not novel tradecraft.
It degrades gracefully on the novel intrusion — the very incident that matters most. A signature/rule engine goes blind on a zero-day; NV's abductive core still produces a most-likely path, honestly flagged as lower-confidence.
Honesty is a moat. Calibrated, case-based confidence with ranked alternatives is exactly what the market says it wants and what black-box competitors structurally cannot offer without redesign.
Raw provider logs → normalized event/entity graph → reconstruction engine → JSON → GUI.
O365 / Entra audit logs nv_extract_identity_events.py
AWS CloudTrail (IAM/STS/S3) nv_extract_cloudtrail_events.py (ingest + normalize)
▼
normalized identity events (JSONL) ← one event model, any substrate
│ nv/graph.py (typed per-actor timelines)
▼
reconstruction engine
├─ nv/patterns.py known layer: ATT&CK identity pattern library
├─ nv/providers.py provider packs: per-substrate op→ATT&CK vocab (Entra + AWS)
├─ nv/validation.py Phase 3: provenance + integrity gate on every pattern
├─ nv/feeds.py live ATT&CK STIX / TAXII / Sigma clients + scheduler
├─ nv/confidence.py evidence-corroboration scorer (calibratable weights)
├─ nv/calibration.py fit + measure confidence against labeled ground truth
├─ nv/reconstruct.py most-likely chain + ranked competing paths + trust floor
└─ nv/scope.py authorized-scope gate (refuses unauthorized tenants/accounts)
│ run_nv.py
▼
reconstruction.json ──► nv_gui.html (embedding canvas + two-column view + trust slider)
The GUI is a pure view layer — it renders engine output and lets the analyst move the trust floor. It contains no reconstruction logic of its own.
Every step carries a confidence in [0, 1], built by evidence corroboration:
confidence = per-technique base rate
+ bonus for each independent corroborating signal
(source IP, device, successful outcome, temporal adjacency, broad-consent flags)
− penalty for missing signals (e.g. no source IP to corroborate origin)
Path scores compose their links by geometric mean, so one weak link honestly drags a path down rather than being averaged away.
Every weight above — the per-technique base rates, the per-signal bonuses, the missing
penalties — lives in a single PARAMS table in nv/confidence.py. Those are NV's v1
priors, and they are calibratable: nv/calibration.py can refit them against labeled
ground truth and the engine will load the result, falling back to the v1 priors when no
calibration is supplied (see below).
The withhold rule, baked into the engine's output contract — not just the UI. Below the floor, a step is withheld. Dragging the floor is the honesty-versus-coverage tradeoff made physical: up for strict/high-trust, down for permissive/high-coverage. The default floor is an editorial decision about where NV sits before the analyst touches it.
The v1 weights are defensible, but they are priors. nv/calibration.py tunes them
against labeled ground truth — events where we know which ones were part of the
intrusion and which were benign — and, crucially, measures whether tuning actually
helped, so a calibration is adopted only if it lowers calibration error rather than just
moving numbers around.
Calibrated confidence means one thing: NV's confidence for a step should equal the probability that the step really was on the attack path. So each labeled event's confidence is treated as a predicted probability, and the harness fits:
The verified/withheld thresholds are policy, not calibration, and are left untouched.
It reports Brier score, log-loss, ECE, AUC, a reliability table, and a trust-floor sweep (true-step recall vs benign false-positive rate at each floor), before and after — so the honesty-versus-coverage tradeoff is legible rather than a single number.