使用AI追踪840多个社交媒体账户
--correlate)--recurse-depth N)<username>.{com,io,net,…} 是否已注册且可访问?(--domains)--watch 6h --notify <url>)--resume file.jsonl)aliens_eye tui,可选额外组件)aliens_eye serve,可选额外组件)--proxy socks5://... 或直接 --tor--site github,reddit、--exclude-site、--no-nsfw,以及即插即用的 sites.d/ 插件站点映射aliens_eye selfcheck 报告每个站点的精确率 / 召回率 / F1 / FPRaliens_eye corpus record / selfcheck --corpus)aliens_eye eval ablate 使用 bootstrap 置信区间对检测器配置评分;eval external 在相同的存储响应上与 Sherlock / Maigret / WhatsMyName 规则进行比较aliens_eye train 重新训练,或使用 aliens_eye label 手动标记不确定的命中pip install aliens-eye
可选额外组件:
pip install "aliens-eye[browser]" # Playwright fallback for hard pages
python -m playwright install chromium
pip install "aliens-eye[train]" # scikit-learn, for retraining the ML model
pip install "aliens-eye[correlate]" # Pillow, for avatar-image matching in --correlate
pip install "aliens-eye[pdf]" # reportlab, for --format pdf
pip install "aliens-eye[tui]" # textual, for the interactive `tui` browser
pip install "aliens-eye[serve]" # mcp, for the `serve` MCP server
或使用 Docker:
docker build -t aliens-eye .
docker run --rm -it aliens-eye username
从源码安装:
git clone https://github.com/arxhr007/Aliens_eye.git
cd Aliens_eye
pip install -e .
# Interactive prompts
aliens_eye
# Single username
aliens_eye username
# Multiple usernames
aliens_eye username1 username2
# Advanced scan level (prefix/suffix variations)
aliens_eye username -l advanced
# Only scan specific sites
aliens_eye username --site github,reddit,gitlab
# Skip NSFW sites
aliens_eye username --no-nsfw
# Route through Tor (needs a local Tor daemon)
aliens_eye username --tor
# Any HTTP or SOCKS proxy
aliens_eye username --proxy socks5://127.0.0.1:1080
# Export everything
aliens_eye username --format all --output results
# Heuristics only, no ML
aliens_eye username --no-ml
# Non-interactive preset: quick / full / aggressive
aliens_eye username --profile quick
# Plain output for scripts and CI (no colors/progress)
aliens_eye username --plain
# View results from a previous scan
aliens_eye -r results/username_advanced_20260611_120000.json
# Correlate hits into "likely same person" clusters + check domains
aliens_eye username --correlate --domains
# Follow linked usernames out of found bios and re-scan them
aliens_eye username --recurse-depth 1
# Export a graph of the results (import into Gephi / Maltego / Mermaid)
aliens_eye username --correlate --format gexf,mermaid,maltego
# Investigator PDF with embedded avatars
aliens_eye username --format pdf
# Watch for changes every 6 hours and POST them to a webhook
aliens_eye username --watch 6h --notify https://hooks.example/aliens
# Resume an interrupted scan
aliens_eye username --resume scan.jsonl
# Compare two saved reports
aliens_eye diff results/old.json results/new.json
# Validate detection accuracy (precision / recall / F1 per site)
aliens_eye selfcheck --negatives 2 --report json
# Record a frozen response corpus, then evaluate against it reproducibly
aliens_eye corpus record --out corpus/v1 --split all --negatives 4
aliens_eye corpus stats corpus/v1
aliens_eye selfcheck --split holdout --corpus corpus/v1 --report json
# Compare detector configurations over that corpus, with confidence intervals
aliens_eye eval ablate --corpus corpus/v1 --split holdout
# Compare against Sherlock / Maigret / WhatsMyName rules (fetch their data yourself)
aliens_eye eval external --corpus corpus/v1 --sherlock data.json --whatsmyname wmn-data.json
# Rebuild the ground-truth splits from those projects account lists
aliens_eye eval groundtruth --sherlock data.json --whatsmyname wmn-data.json
# Interactively label uncertain hits into a training set
aliens_eye label results/username_basic_20260611_120000.json --out labeled.csv
# Interactive terminal browser (needs [tui])
aliens_eye tui username
# Run the MCP server for LLM agents (needs [serve])
aliens_eye serve
自定义平台:将 { "site_name": "https://site/{}" } JSON 文件放入
./sites.d/(或用户配置目录的 sites.d/),它会自动合并;
--sites-dir DIR 可添加另一个位置。
每个响应都会被转换为一个 30 维特征向量:HTTP 状态分桶、用户名位置(路径/标题/meta/canonical)、错误与资料关键词、DOM 结构(图片、表单、资料/错误 CSS 类)、结构化数据信号(og:type、JSON-LD Person)、响应时间、重定向次数,以及从先前扫描中学到的每站点指纹匹配。
然后两个评判器进行投票:
混合概率映射为 Found / Maybe / Not Found,并附带置信度百分比。加载的模型同时提供混合权重和阈值 — 随附模型使用 0.9 * ml + 0.1 * heuristic,Found 高于 0.620,Not Found 低于 0.360。如果模型文件缺失或无效,扫描器会静默回退到 core/detector.py 中的默认启发式(0.4 ML 权重,0.6 / 0.35 阈值)。完整表格见 WORKING.md。
将 Found/Maybe 视为需要验证的线索,而非发现。 随附模型在来自 286 个平台的 1,455 个样本上拟合,并在它从未见过的 142 个平台上评估:精确率 0.65,召回率 0.51,假阳性率 9%(F1 0.57)。它倾向于减少错误线索而非捕获每个账户,并且在留出平台上,没有任何配置——包括此配置——在 F1 上能与简单的 HTTP 状态检查在统计上区分开来。
pip install "aliens-eye[train]"
# 1. Scan ground-truth accounts + random non-existent usernames to build a dataset
# (reads the train split only; the eval holdout is never touched)
aliens_eye train collect --out dataset.csv --negatives 4
# 2. Fit and export the model
aliens_eye train fit --data dataset.csv --out model.json
# 3. Score it on platforms it never trained on
aliens_eye selfcheck --split holdout --model model.json --report json
# 4. Use it
aliens_eye username --model model.json
Ground truth 按 站点不相交 拆分为 data/selfcheck.json(训练,30
个站点)和 data/eval_holdout.json(留出,13 个站点)。对 --split train
评分衡量的是拟合程度,而非泛化能力,会偏高。
aliens_eye.api 是在 Aliens Eye 之上构建其他工具的稳定接口。它
返回形状类似 JSON 报告的普通字典,默认不打印任何内容。
import asyncio
from aliens_eye import api