
HiddenSteps — a local-first personal workflow intelligence platform
This is the honest, current-state complement to docs/design/02-system-architecture.md's target module map. It says what's actually built, what's verified against a real backend vs. a mock, and what's still genuinely missing — not what's planned (that's docs/roadmap/01-implementation-roadmap.md).
Run cargo build --workspace && cargo test --workspace && cargo clippy --workspace --all-targets -- -D warnings from the repo root. As of this writing: 12 crates, 193 passing tests, zero clippy warnings, cargo fmt --check clean — 183 across the 11 crates that need no display or external service, plus 10 in hiddensteps-observation that need a live X11 display (verified where one exists; see that row). Four tests are #[ignore]d by design (see below) and are not counted as failures or as part of the 193.
| Crate | Implements | Verified how |
|---|---|---|
hiddensteps-domain | Core types: PrivacyLevel/PrivacyState, EventSummary/SignalType, Pattern, Recommendation, AuditEntry, and CapturedSignal — a type that structurally cannot be persisted (no Serialize), enforcing ADR-0006's raw-data rule at the type level | Unit tests: level round-tripping/ordering, Deep-mode TTL gating |
hiddensteps-security | SecretStore (ADR-0008): real OS-vault (KeyringSecretStore) + in-memory (test) implementations; CSPRNG master-key generation (returned in a zeroize::Zeroizing wrapper so the key is wiped on drop rather than lingering in freed memory); Argon2id passphrase derivation for Portable Mode (PassphraseKey zeroizes its derived key on drop, keeping the non-secret salt). hiddensteps-event-store likewise holds the key-bearing PRAGMA key/rekey SQL text in Zeroizing | Unit tests against the in-memory store and the KDF; the real-vault round trip is #[ignore]d (see below) |
hiddensteps-event-store | SqlCipherEventStore (ADR-0003): the full schema from docs/design/07-database-schema.md, CRUD for privacy state, events, audit log, patterns, pattern↔event links, pattern embeddings (see note below), recommendations, LLM provider config, and generic settings, plus delete_all_data (transactional; also rekeys for a "delete everything" that survives a relaunch)/export_data/count_rows (diagnostics)/delete_expired_events (the Deep-mode TTL sweep, called from apps/desktop/src-tauri's periodic recommendation loop — ttl_expires_at was persisted since v0.1.0 but nothing deleted a row past it before this); foreign-key enforcement (PRAGMA foreign_keys = ON) so schema.sql's ON DELETE CASCADE on pattern↔event links actually runs | 33 tests against a real SQLCipher file: wrong key fails to open, same key reopens correctly, delete-all clears every table including the newest ones, rekey round-trips, TTL sweep leaves non-expired events alone, cascade delete leaves no orphaned pattern↔event links |
hiddensteps-redaction | The Redaction Engine (docs/design/05-privacy-model.md §4): regex+Luhn detectors for API keys/tokens/PEM keys/emails/SSNs/credit cards, an entropy-based ambiguous-secret detector, and the drop-on-uncertainty policy | 30 tests, including deliberately adversarial inputs (secrets embedded in prose, near-miss non-secrets like git SHAs, dashless/spaced SSNs, digit-padded card numbers, all-one-case high-entropy tokens) |
hiddensteps-pipeline | The Event Pipeline (ADR-0006): Classify → Redact → Summarize, privacy-level gating per signal type, Deep-mode TTL assignment | 8 tests covering redaction-triggered drops, level-gating drops, and successful summarization |
hiddensteps-observation | ObservationSource (ADR-0005) + Linux: ActiveWindowSource (X11 GetInputFocus), FileOperationSource (inotify via notify), ClipboardMetadataSource (X11 selection, metadata-only), GlobalShortcutSource (X11 XGrabKey). Plus macOS/Windows source files (see below) | 10 of 11 tests run against real backends in this environment — a live X11 display (WSLg's DISPLAY=:0) and real inotify, not mocks. 1 test (GlobalShortcutSource's real grab) is #[ignore]d by design |
hiddensteps-llm-provider | LlmProvider (ADR-0004): Ollama client (with a think: Option<bool> request field for hybrid-reasoning models), an OpenAI-wire-compatible client (covers OpenAI/Azure/OpenRouter/Together/Groq/DeepSeek/LocalAI), an Anthropic Messages client, and local-runtime auto-detection. Every client sets a request timeout (build_http_client) so a hung remote can't block a call forever; Ollama forwards max_tokens as its nested options.num_predict | 19 tests against wiremock mock servers (including a real timeout-fires check and that Ollama actually sends num_predict), plus 2 real-Ollama integration tests (tests/ollama_live.rs, #[ignore]d — see below) that found and fixed a real problem: the same prompt took over two minutes against a real local hybrid-thinking model with think left at its default, and a few seconds with think: Some(false) |
hiddensteps-patterns | Pattern Detection (sliding-window n-gram sequence matching) + Workflow Graph (transition graph with edge weights) — ADR-0010's Layer 1 | 16 tests, including a direct analog of PROMPT.md's own "observed 31 times" example and a regression test that overlapping windows over a continuous repeat aren't double-counted |
hiddensteps-recommendations | The Recommendation Engine's Layer 2 (ADR-0010): LLM synthesis with a structured-JSON prompt contract, a narrative-contradiction validator, and a retry loop — critically, the numeric fields (estimated_time_saved_minutes) are never parsed from the LLM's output at all, only computed from Layer 1 | 23 tests, including malformed-JSON retry, narrative-contradiction retry (covering spelled-out numbers and every LLM-controlled field, not just why), and string-aware JSON extraction, against a scripted test provider |
hiddensteps-privacy-engine | The cloud-dispatch gate (docs/design/03-data-flow-diagrams.md §5) and consent versioning (docs/design/05-privacy-model.md §5); PrivacyGatedProvider wraps any LlmProvider so the gate can't be bypassed by the normal call path | 13 tests, including that Level-4 content is blocked even with every consent granted |
hiddensteps-plugin-host | The WASM Plugin Host (ADR-0009): closed capability enumeration, manifest validation, a wasmtime-backed sandbox that links in only granted capabilities' host functions, plus fuel metering and a memory ResourceLimiter bounding a plugin instance's CPU/memory regardless of what capabilities it holds — the two axes docs/research/06-threat-model.md's Denial-of-Service section names that capability enforcement alone can't address (a capability-free module can still loop or grow memory forever). instantiate_from_manifest is the safe entry point: it forces manifest validation (the Level-4-required-for-screenshot rule) and rejects granting anything the manifest didn't declare, before any capability reaches the linker — the plain instantiate capability slice has no such link to a manifest at all | 20 tests, including real capability-escape attempts: hand-written WAT modules compiled at test time, proving an ungranted capability's import is genuinely unresolved (instantiation fails), not merely unused; plus a real infinite-loop module and an unbounded-memory.grow module that trap instead of hanging/exhausting memory |