Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
sift-kg — Turn any collection of documents into a knowledge graph. Extract entities and relationships via LLM, deduplicate with your approval. Map domains, find hidden connections, spot patterns across documents — knowledge that persists and compounds, for you and your AI agents. All from the CLI. | Kitploit
Tools/GitHubGitHub/juanceresa/sift-kg
OSINT (Open Source Intelligence)ForensicsInformation GatheringData RecoveryDigital ForensicsPapers & ResearchLearning & Education
GitHubjuanceresa/sift-kg

sift-kg

View Repository
66758264 months agoReviewed by Kitploit

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →

About

Turn any collection of documents into a knowledge graph. Extract entities and relationships via LLM, deduplicate with your approval. Map domains, find hidden connections, spot patterns across documents — knowledge that persists and compounds, for you and your AI agents. All from the CLI.

Share

sift-kg

Turn any collection of documents into a knowledge graph.

No code, no database, no infrastructure — just a CLI and your documents. Drop in PDFs, papers, articles, or records — get a browsable knowledge graph that shows how everything connects, in minutes. sift-kg extracts entities and relationships via LLM, deduplicates with your approval, and generates an interactive viewer you can explore in your browser. Concept maps for anything, at your fingertips.

The same graph that powers your visualizations also works as an AI second brain. Everyone's spending months building knowledge bases in Notion and Obsidian. Who has time for that? sift-kg is the structured memory you build in 2 minutes instead of 2 years. Just point at your docs and your AI has a structured understanding of how everything connects.

Live demos → graphs generated entirely by sift-kg

pip install sift-kg

sift init                           # create sift.yaml + .env.example
sift extract ./documents/           # extract entities & relations
sift build                          # build knowledge graph
sift resolve                        # find duplicate entities
sift review                         # approve/reject merges interactively
sift apply-merges                   # apply your decisions
sift narrate                        # generate narrative summary
sift view                           # interactive graph in your browser
sift export graphml                 # export to Gephi, yEd, Cytoscape, SQLite, etc.

How It Works

Documents (PDF, DOCX, text, HTML, and 75+ formats)
       ↓
  Text Extraction (Kreuzberg, local) — with optional OCR (Tesseract, EasyOCR, PaddleOCR, or Google Cloud Vision)
       ↓
  Schema Discovery (LLM designs entity/relation types from your data — or use a predefined domain)
       ↓
  Entity & Relation Extraction (LLM, using discovered or predefined schema)
       ↓
  Knowledge Graph (NetworkX, JSON)
       ↓
  Entity Resolution (LLM proposes → you review)
       ↓
  Narrative Generation (LLM)
       ↓
  Interactive Viewer (browser) / Export (GraphML, GEXF, CSV, SQLite)

Every entity and relation links back to the source document and passage. You control what gets merged. The graph is yours.

Features

  • Zero-config start — point at a folder, get a knowledge graph. Or drop a sift.yaml in your project for persistent settings
  • Any LLM provider — OpenAI, Anthropic, Mistral, Ollama (local/private), or any LiteLLM-compatible provider
  • Schema-free by default — one LLM call samples your documents and designs a schema tailored to the corpus, saved as discovered_domain.yaml for reuse and editing. Or use a structured domain (general, osint, academic) for fixed schemas, or define your own in YAML
  • Human-in-the-loop — sift proposes entity merges, you approve or reject in an interactive terminal UI
  • CLI search — sift search "SBF" finds entities by name or alias, with optional relation and description output
  • Interactive viewer — explore your graph in-browser with community regions (colored zones showing graph structure), hover preview, focus mode (double-click to isolate neighborhoods), keyboard navigation (arrow keys to step through connections), trail breadcrumb (persistent path that tracks your exploration — trace back through every node you visited), search, type/community/relation toggles, source document filter, and degree filtering. Pre-filter with CLI flags: --neighborhood, --top, --community, --source-doc, --min-confidence
  • Export anywhere — GraphML (yEd, Cytoscape), GEXF (Gephi), SQLite, CSV, or native JSON for advanced analysis
  • Narrative generation — prose reports with relationship chains, timelines, and community-grouped entity profiles
  • Source provenance — every extraction links to the document and passage it came from
  • Multilingual — extracts from documents in any language, outputs a unified English knowledge graph. Proper names stay as-is, non-Latin scripts are romanized automatically
  • 75+ document formats — PDF, DOCX, XLSX, PPTX, HTML, EPUB, images, and more via Kreuzberg extraction engine
  • OCR for scanned PDFs — local OCR via Tesseract (default), EasyOCR, or PaddleOCR (--ocr flag), with optional Google Cloud Vision fallback (--ocr-backend gcv)
  • Budget controls — set --max-cost to cap LLM spending
  • Runs locally — your documents stay on your machine

Use Cases

  • Research & education — map how theories, methods, and findings connect across a body of literature. Generate concept maps for courses, literature reviews, or self-study
  • Business intelligence — drop in competitor whitepapers, market reports, or internal docs and see the landscape
  • Investigative work — analyze FOIA releases, court filings, public records, and document leaks
  • Legal review — extract and connect entities across document collections
  • Genealogy — trace family relationships across vital records

AI Knowledge Base

sift-kg generates structured knowledge that AI agents can operate from directly.

Point sift at your documents, notes, or project files. The output — a JSON knowledge graph — gives any AI agent a persistent, structured understanding of how everything in your world connects. No manual organization, no tagging, no wiki links. The structure emerges from the content.

sift extract ./my-stuff/
sift build
sift topology          # structural overview (JSON, for agents)
sift query "topic"     # entity neighborhood subgraph (JSON, for agents)
sift search "X" --json # entity lookup (JSON, for agents)
sift info --json       # project stats (JSON, for agents)

The graph persists across sessions and grows incrementally — extract new documents into the same output directory and rebuild. Entity deduplication ensures the graph stays coherent as it grows.

What this gives your agent:

  • Structure — not just text chunks, but entities, relationships, communities, and how they connect
  • Topology — which knowledge clusters exist, what bridges them, what's isolated
  • Durability — the graph survives context window resets. Your agent stops starting from zero every session

Bundled agent skill: sift-kg ships with a skill at .agents/skills/sift-kg/SKILL.md that teaches agents how to use the knowledge graph as persistent memory — session orientation, entity exploration, link-knowledge-islands reasoning, and grounded suggestion generation.

Bundled Domains

sift-kg ships with specialized domains you can use out of the box:

sift domains                              # list available domains
sift extract ./docs/ --domain-name osint  # use a bundled domain

Set a domain in sift.yaml so you don't need the flag every time:

domain: academic

Works with bundled names (schema-free, general, osint, academic) or a path to a custom YAML file.

DomainFocusKey Entity TypesKey Relation Types
schema-freeAuto-discovered from your data (default)(LLM designs per corpus)(LLM designs per corpus)
generalGeneral document analysisPERSON, ORGANIZATION, LOCATION, EVENT, DOCUMENTASSOCIATED_WITH, MEMBER_OF, LOCATED_IN
osintInvestigations & FOIASHELL_COMPANY, FINANCIAL_ACCOUNTBENEFICIAL_OWNER_OF, TRANSACTED_WITH, SIGNATORY_OF
academicLiterature review & topic mappingCONCEPT, THEORY, METHOD, SYSTEM, FINDING, PHENOMENON, RESEARCHER, PUBLICATION, FIELD, DATASETSUPPORTS, CONTRADICTS, EXTENDS, IMPLEMENTS, EXPLAINS, PROPOSED_BY, USES_METHOD, APPLIED_TO, INVESTIGATES
Download Tool