
Reddit OSINT by username. Discover indexed posts, comments, deleted and live content, activity patterns, exposed identifiers, and an overall exposure score.
Reddit OSINT by username. Give it a username and it returns everything the public Reddit archives have indexed for that account: posts, comments, what was deleted and what is still live, activity patterns, leaked identifiers, and an exposure score.
No dependencies, no API keys, no login. Just python3.

python3 reddit_osint.py someuser
The same analysis is also available as MCP tools for LLM clients, see MCP server.
pip install -r requirements.txt && python mcp_server.py
Reddit deletes things, but third-party data archives keep the copy that existed before the
deletion (or at least the [removed] / [deleted] marker). This tool queries those archives and
reconstructs the activity profile of an account.
That is the premise behind the h3kz-reddint web tool, which is the reference for this script:
it reimplements that pipeline (two sources, merge, PII regexes, graph, score) as a standalone
CLI, with tests and optional proxy support.
| Source | What it is | Role |
|---|---|---|
Arctic Shift (arctic-shift.photon-reddit.com) | Free Reddit archive, community-run successor to Pushshift. Covers 2022→today and retains the original scrape (selftext, removed_by_category, media_metadata, flair). No API key, per-IP rate limit. | Primary source. Provides 100% of the useful data. |
PullPush (api.pullpush.io) | Another index of the same Reddit data (historical Pushshift dumps). | Fallback. Queried in parallel and merged by id. Its subreddit API returns 404, so it only serves posts/comments. |
reddit.com | — | Never queried. Only used to build permalink URLs in the report. The official Reddit API is not used (no login), which is why this works on suspended accounts. |
Measured on a real sample account: Arctic Shift returned 39 posts / 139 comments, PullPush
returned 38 posts with overlapping IDs. Arctic Shift wins; --source arctic is faster and
sufficient.
pip install. Standard library only.pip install pysocks (optional, only needed for socks5://)pip install -r requirements.txt (FastMCP). The CLI does not need it.python3 reddit_osint.py someuser
Prints the report as JSON on stdout (summary only, without the per-item dump) and logs on
stderr.
python3 reddit_osint.py someuser --json report.json
python3 reddit_osint.py someuser --json report.json --quiet > /dev/null
report.json includes the all_items key with every post and comment (text, subreddit,
score, date, status, permalink). The on-screen summary omits it so it does not dump 200 KB.
python3 reddit_osint.py someuser --subreddit dotnet # one subreddit only
python3 reddit_osint.py someuser --keywords "kubernetes" # posts mentioning X
python3 reddit_osint.py someuser --over18 true # NSFW only
python3 reddit_osint.py someuser --source arctic # single source (faster)
python3 reddit_osint.py someuser --max-pages 10 # cap pagination (100 items/page)
python3 reddit_osint.py someuser --delay 1.0 # slower, better with proxies
python3 reddit_osint.py someuser --rate-limit-retries 0 # fail fast when throttled
python3 reddit_osint.py someuser --proxies proxies.txt
proxies.txt, one proxy per line. Accepted formats:
http://1.2.3.4:8080
https://1.2.3.4:8443
socks5://1.2.3.4:1080
socks5://user:[email protected]:1080
1.2.3.4:8080 # no scheme -> assumed http://
# comments and blank lines are ignored
Duplicates are removed. If the file contains no valid proxy, the script exits with code 2 and makes no requests.
Rotation. Proxies are walked round-robin. When a request comes back with a blocking/rate-limit
code, that proxy is put on cooldown for --cooldown seconds and the script moves to the next one:
| Flag | Default | What it does |
|---|---|---|
--proxies FILE | — | Proxy list. Without it: direct connection. |
--cooldown SECONDS | 30 | How long a banned proxy stays sidelined. |
--max-attempts N | (proxies+1) × 3 | Attempt ceiling per request before giving up. |
--no-direct-fallback | off | Wait instead of connecting directly when all proxies are cooling down. |
--shuffle-proxies | off | Shuffle the list at startup (spreads load from the first request). |
--proxy-stats | off | Adds proxy_stats to the report: ok / banned / errors per proxy. |
--rate-limit-retries N | 2 | Retries of the whole page with exponential backoff (30 s, 60 s, capped at 120 s) when the API answers 429/422. |
Codes treated as a ban (they rotate the proxy): 401 403 407 422 429 500 502 503 504 520 521 522 523 524, plus any 400/409 whose body contains slow down / rate limit /
too many requests / timeout. Arctic Shift uses 422 "Timeout. Maybe slow down a bit" for
rate limiting, so it is deliberately handled as a ban.
A 404 does not rotate: that is a bad request, not a bad proxy, and retrying elsewhere will
not fix it.
Backoff. After a ban it waits min(cooldown, 10 s) — this applies to direct connections too,
which is where it hurts most. If a page exhausts its attempts, fetch_all retries the page with
exponential backoff (--rate-limit-retries, 2 attempts by default = 30 s and 60 s). That is what
stops the script from silently losing comments when it trips a rate limit mid-run.
Sample log with rotation:
[proxy] http://1.2.3.4:8080 blocked (429), backing off 10s -> rotating
[proxy] http://5.6.7.8:1080 failed (Connection refused), rotating
[proxy] direct connection blocked (422), backing off 10s -> rotating
[!] arctic/comments rate limited, backing off 30s (retry 1/2)
When every proxy is cooling down, the script waits out the longest remaining cooldown and retries
(or uses a direct connection unless --no-direct-fallback was passed).
Warning: a proxy is the egress path your IP takes to a third party. If you do not trust the proxy vendor, do not use one with this tool.
| Variable | Purpose |
|---|---|
REDDIT_OSINT_ARCTIC | Override the Arctic Shift URL (mirrors, tests). |
REDDIT_OSINT_PULLPUSH | Override the PullPush URL. |
id, title, selftext, body, author, subreddit, created_utc, score, permalink,
url, over_18, num_comments, link_flair_text, preview.images, media_metadata,
removed_by_category.
| Status | Detection |
|---|---|
removed | Text = [removed] or [ Removed by Reddit ], or removed_by_category present (posts) → a moderator removed it. |
deleted | Text = [deleted] or author == "[deleted]" → the author deleted it. |
live | Everything else. |
This is the whole point of the tool: the archive captures the fact of the removal even when it has no content, and that alone is useful metadata.
deleted_pct.total_karma, post_karma, comment_karma, num_posts,
num_comments, earliest_post_at, last_comment_at (from /api/users/search).