
Turn any collection of documents into a knowledge graph. Extract entities and relationships via LLM, deduplicate with your approval. Map domains, find hidden connections, spot patterns across documents — knowledge that persists and compounds, for you and your AI agents. All from the CLI.
Turn any collection of documents into a knowledge graph.
No code, no database, no infrastructure — just a CLI and your documents. Drop in PDFs, papers, articles, or records — get a browsable knowledge graph that shows how everything connects, in minutes. sift-kg extracts entities and relationships via LLM, deduplicates with your approval, and generates an interactive viewer you can explore in your browser. Concept maps for anything, at your fingertips.
The same graph that powers your visualizations also works as an AI second brain. Everyone's spending months building knowledge bases in Notion and Obsidian. Who has time for that? sift-kg is the structured memory you build in 2 minutes instead of 2 years. Just point at your docs and your AI has a structured understanding of how everything connects.
Live demos → graphs generated entirely by sift-kg
pip install sift-kg
sift init # create sift.yaml + .env.example
sift extract ./documents/ # extract entities & relations
sift build # build knowledge graph
sift resolve # find duplicate entities
sift review # approve/reject merges interactively
sift apply-merges # apply your decisions
sift narrate # generate narrative summary
sift view # interactive graph in your browser
sift export graphml # export to Gephi, yEd, Cytoscape, SQLite, etc.
Documents (PDF, DOCX, text, HTML, and 75+ formats)
↓
Text Extraction (Kreuzberg, local) — with optional OCR (Tesseract, EasyOCR, PaddleOCR, or Google Cloud Vision)
↓
Schema Discovery (LLM designs entity/relation types from your data — or use a predefined domain)
↓
Entity & Relation Extraction (LLM, using discovered or predefined schema)
↓
Knowledge Graph (NetworkX, JSON)
↓
Entity Resolution (LLM proposes → you review)
↓
Narrative Generation (LLM)
↓
Interactive Viewer (browser) / Export (GraphML, GEXF, CSV, SQLite)
Every entity and relation links back to the source document and passage. You control what gets merged. The graph is yours.
sift.yaml in your project for persistent settingsdiscovered_domain.yaml for reuse and editing. Or use a structured domain (general, osint, academic) for fixed schemas, or define your own in YAMLsift search "SBF" finds entities by name or alias, with optional relation and description output--neighborhood, --top, --community, --source-doc, --min-confidence--ocr flag), with optional Google Cloud Vision fallback (--ocr-backend gcv)--max-cost to cap LLM spendingsift-kg generates structured knowledge that AI agents can operate from directly.
Point sift at your documents, notes, or project files. The output — a JSON knowledge graph — gives any AI agent a persistent, structured understanding of how everything in your world connects. No manual organization, no tagging, no wiki links. The structure emerges from the content.
sift extract ./my-stuff/
sift build
sift topology # structural overview (JSON, for agents)
sift query "topic" # entity neighborhood subgraph (JSON, for agents)
sift search "X" --json # entity lookup (JSON, for agents)
sift info --json # project stats (JSON, for agents)
The graph persists across sessions and grows incrementally — extract new documents into the same output directory and rebuild. Entity deduplication ensures the graph stays coherent as it grows.
What this gives your agent:
Bundled agent skill: sift-kg ships with a skill at .agents/skills/sift-kg/SKILL.md that teaches agents how to use the knowledge graph as persistent memory — session orientation, entity exploration, link-knowledge-islands reasoning, and grounded suggestion generation.
sift-kg ships with specialized domains you can use out of the box:
sift domains # list available domains
sift extract ./docs/ --domain-name osint # use a bundled domain
Set a domain in sift.yaml so you don't need the flag every time:
domain: academic
Works with bundled names (schema-free, general, osint, academic) or a path to a custom YAML file.
| Domain | Focus | Key Entity Types | Key Relation Types |
|---|---|---|---|
schema-free | Auto-discovered from your data (default) | (LLM designs per corpus) | (LLM designs per corpus) |
general | General document analysis | PERSON, ORGANIZATION, LOCATION, EVENT, DOCUMENT | ASSOCIATED_WITH, MEMBER_OF, LOCATED_IN |
osint | Investigations & FOIA | SHELL_COMPANY, FINANCIAL_ACCOUNT | BENEFICIAL_OWNER_OF, TRANSACTED_WITH, SIGNATORY_OF |
academic | Literature review & topic mapping | CONCEPT, THEORY, METHOD, SYSTEM, FINDING, PHENOMENON, RESEARCHER, PUBLICATION, FIELD, DATASET | SUPPORTS, CONTRADICTS, EXTENDS, IMPLEMENTS, EXPLAINS, PROPOSED_BY, USES_METHOD, APPLIED_TO, INVESTIGATES |