将任何文档集合转化为知识图谱。
无需代码,无需数据库,无需基础设施——只需要一个命令行工具和你的文档。放入PDF、论文、文章或记录——几分钟内即可获得一个可浏览的知识图谱,展示所有内容之间的关联。sift-kg通过LLM提取实体和关系,经你确认后去重,并生成一个交互式查看器,你可以在浏览器中探索。任何内容的概念图,唾手可得。
支撑你可视化的同一个图谱也可以作为AI第二大脑。每个人都在Notion和Obsidian上花费数月构建知识库。谁有那个时间?sift-kg是你用2分钟而不是2年构建的结构化记忆。只需指向你的文档,你的AI就能结构化地理解所有内容之间的关联。
现场演示 → 完全由sift-kg生成的图谱```bash pip install sift-kg
sift init # create sift.yaml + .env.example sift extract ./documents/ # extract entities & relations sift build # build knowledge graph sift resolve # find duplicate entities sift review # approve/reject merges interactively sift apply-merges # apply your decisions sift narrate # generate narrative summary sift view # interactive graph in your browser sift export graphml # export to Gephi, yEd, Cytoscape, SQLite, etc.
## 工作原理```
Documents (PDF, DOCX, text, HTML, and 75+ formats)
↓
Text Extraction (Kreuzberg, local) — with optional OCR (Tesseract, EasyOCR, PaddleOCR, or Google Cloud Vision)
↓
Schema Discovery (LLM designs entity/relation types from your data — or use a predefined domain)
↓
Entity & Relation Extraction (LLM, using discovered or predefined schema)
↓
Knowledge Graph (NetworkX, JSON)
↓
Entity Resolution (LLM proposes → you review)
↓
Narrative Generation (LLM)
↓
Interactive Viewer (browser) / Export (GraphML, GEXF, CSV, SQLite)
每个实体和关系都链接回源文档和段落。你控制哪些内容被合并。这张图是你的。
sift.yaml 以持久化设置discovered_domain.yaml 供复用和编辑。或者使用结构化领域(general、osint、academic)获得固定模式,也可在YAML中自行定义sift search "SBF" 按名称或别名查找实体,可选择输出关系和描述--neighborhood、--top、--community、--source-doc、--min-confidence--ocr 标志),可选 Google Cloud Vision 作为备用(--ocr-backend gcv)--max-cost 以限制LLM花费sift-kg 生成结构化知识,AI代理可以直接基于此进行操作。
将 sift 指向你的文档、笔记或项目文件。输出——一个JSON知识图谱——为任何AI代理提供持久化、结构化的理解,展示你世界中一切事物之间的联系。无需手动组织、无需标签、无需维基链接。结构从内容中自然涌现。```bash sift extract ./my-stuff/ sift build sift topology # structural overview (JSON, for agents) sift query "topic" # entity neighborhood subgraph (JSON, for agents) sift search "X" --json # entity lookup (JSON, for agents) sift info --json # project stats (JSON, for agents)
图结构跨会话持久保存,并会增量扩展——将新文档提取到相同的输出目录并重建。实体去重确保图在扩展过程中保持连贯。
**这为您的智能体带来了什么:**
- **结构** —— 不仅仅是文本块,而是实体、关系、社区及其连接方式
- **拓扑** —— 存在哪些知识集群,哪些桥接它们,哪些是孤立的
- **持久性** —— 图在上下文窗口重置后仍然存在。您的智能体不再从零开始每个会话
**附带的智能体技能:** sift-kg 附带一个位于 `.agents/skills/sift-kg/SKILL.md` 的技能文件,该文件教导智能体如何将知识图用作持久记忆——会话定位、实体探索、连接知识孤岛推理以及基于基础信息的建议生成。
## 附带的领域
sift-kg 附带了一些开箱即用的专用领域:```bash
sift domains # list available domains
sift extract ./docs/ --domain-name osint # use a bundled domain
在 sift.yaml 中设置一个域名,这样就不需要每次都使用该标志:```yaml
domain: academic
Works with bundled names (`schema-free`, `general`, `osint`, `academic`) or a path to a custom YAML file.
| Domain | Focus | Key Entity Types | Key Relation Types |
|--------|-------|------------------|--------------------|
| `schema-free` | Auto-discovered from your data (default) | *(LLM designs per corpus)* | *(LLM designs per corpus)* |
| `general` | General document analysis | PERSON, ORGANIZATION, LOCATION, EVENT, DOCUMENT | ASSOCIATED_WITH, MEMBER_OF, LOCATED_IN |
| `osint` | Investigations & FOIA | SHELL_COMPANY, FINANCIAL_ACCOUNT | BENEFICIAL_OWNER_OF, TRANSACTED_WITH, SIGNATORY_OF |
| `academic` | Literature review & topic mapping | CONCEPT, THEORY, METHOD, SYSTEM, FINDING, PHENOMENON, RESEARCHER, PUBLICATION, FIELD, DATASET | SUPPORTS, CONTRADICTS, EXTENDS, IMPLEMENTS, EXPLAINS, PROPOSED_BY, USES_METHOD, APPLIED_TO, INVESTIGATES |
**academic** 域映射研究领域的知识图谱——输入论文,即可获得理论、方法、系统、发现和概念如何相互连接的图形。区分抽象思想(THEORY、METHOD)和具体成果(SYSTEM——例如 GPT-2、BERT、GLUE)。专为文献综述、主题映射以及理解观点在何处一致、矛盾或相互构建而设计。
**schema-free** 域(默认)在提取前执行**模式发现**步骤——一次 LLM 调用会抽样您的文档,并为该语料库设计实体和关系类型。发现的模式会保存到 `output/discovered_domain.yaml` 中,并在后续运行中重复使用,因此类型在所有分块和文档之间保持一致。您可以检查、手动编辑或复制该文件,作为自定义域的起点。使用 `--force` 重新发现。它不会将关系强制纳入预定义类别(如 ASSOCIATED_WITH),而是生成特定类型(如 FUNDED、TESTIFIED_AGAINST 或 ENROLLED_AT)。当您想要预先定义的固定模式时,请使用结构化域(如 `general` 或 `osint`)。
**general** 域提供了一个固定模式,包含 PERSON、ORGANIZATION、LOCATION、EVENT 和 DOCUMENT 实体类型以及常见关系类型。当您希望在文档中获得可预测的一致类型时非常有用。
**osint** 域增加了空壳公司、金融账户和离岸司法管辖区的实体类型,以及用于追踪实益所有权和资金流动的关系类型。
未经您的批准,不会合并任何内容——LLM 提出建议,您进行验证。每次提取都会链接回源文档和段落。
See [`examples/transformers/`](https://github.com/juanceresa/sift-kg/blob/main/examples/transformers) for 12 foundational AI papers mapped as a concept graph (425 entities, ~$0.72), and [`examples/ftx/`](https://github.com/juanceresa/sift-kg/blob/main/examples/ftx) for the FTX collapse (431 entities from 9 articles). [**Explore the live demos**](https://juanceresa.github.io/sift-kg/) — no install, no API key.
## Civic Table
Looking for a hosted platform with forensic legal analysis and analyst verification?
[**Civic Table**](https://github.com/juanceresa/forensic_analysis_platform) is a forensic intelligence platform built on the sift-kg pipeline. It adds a 4-tier verification system where analysts and JDs validate AI-extracted facts before they're treated as evidence, LaTeX dossier generation for legal submissions, and a web interface for sharing results with clients and families. Built for property restitution, investigative journalism, and any context where documentary provenance matters.
sift-kg is the open-source CLI. Civic Table is the full platform — and where the output gets vetted by analysts and JDs before it carries evidentiary weight.
## Installation
Requires Python 3.11+.```bash
pip install sift-kg
对于OCR支持(扫描的PDF、图片):```bash
brew install tesseract # macOS sudo apt install tesseract-ocr # Ubuntu/Debian
对于 Google Cloud Vision OCR 作为替代后端(可选):```bash
pip install sift-kg[ocr]
# Then use: sift extract ./docs/ --ocr --ocr-backend gcv
对于实体解析过程中的语义聚类(可选,PyTorch约需2GB):```bash pip install sift-kg[embeddings]
用于开发:```bash
git clone https://github.com/juanceresa/sift-kg.git
cd sift-kg
pip install -e ".[dev]"