Skip to content
KitploitKITPLOIT
ToolsExploitsBlog
Log in
Submit
ToolsExploitsBlog
Submit

Hacking, PenTest, and Cybersecurity Tools for Your Security Arsenal!

Kitploit is a directory of hacking, cybersecurity, and pentesting tools. Discover the latest project updates to find vulnerabilities, analyze systems, automate testing, and strengthen your security.

FeedsContactPrivacy© 2026 Kitploit

Tool Directory

Categories

View all categories
Loading categories
Tools/GitHubGitHub/albertrg99/redditosint
OSINT (Open Source Intelligence)ReconnaissanceOSINT for Social EngineeringScripting & AutomationData ExfiltrationInformation GatheringPrivacySocial EngineeringUtilities & FrameworksThreat Intelligence
GitHub
41161 day agoNot yet reviewed

Most Popular

View all →

Discover the most used tools by our community.

Explore all tools

Browse our collection of tools

View all tools →
Share
albertrg99/redditosint

RedditOSINT

Reddit OSINT by username. Discover indexed posts, comments, deleted and live content, activity patterns, exposed identifiers, and an overall exposure score.

View Repository

RedditOSINT

Reddit OSINT by username. Give it a username and it returns everything the public Reddit archives have indexed for that account: posts, comments, what was deleted and what is still live, activity patterns, leaked identifiers, and an exposure score.

No dependencies, no API keys, no login. Just python3.

image

python3 reddit_osint.py someuser

The same analysis is also available as MCP tools for LLM clients, see MCP server.

pip install -r requirements.txt && python mcp_server.py

What it does and why it exists

Reddit deletes things, but third-party data archives keep the copy that existed before the deletion (or at least the [removed] / [deleted] marker). This tool queries those archives and reconstructs the activity profile of an account.

That is the premise behind the h3kz-reddint web tool, which is the reference for this script: it reimplements that pipeline (two sources, merge, PII regexes, graph, score) as a standalone CLI, with tests and optional proxy support.

The sources

SourceWhat it isRole
Arctic Shift (arctic-shift.photon-reddit.com)Free Reddit archive, community-run successor to Pushshift. Covers 2022→today and retains the original scrape (selftext, removed_by_category, media_metadata, flair). No API key, per-IP rate limit.Primary source. Provides 100% of the useful data.
PullPush (api.pullpush.io)Another index of the same Reddit data (historical Pushshift dumps).Fallback. Queried in parallel and merged by id. Its subreddit API returns 404, so it only serves posts/comments.
reddit.com—Never queried. Only used to build permalink URLs in the report. The official Reddit API is not used (no login), which is why this works on suspended accounts.

Measured on a real sample account: Arctic Shift returned 39 posts / 139 comments, PullPush returned 38 posts with overlapping IDs. Arctic Shift wins; --source arctic is faster and sufficient.


Requirements

  • Python 3.9+ (tested on 3.13)
  • No pip install. Standard library only.
  • For SOCKS proxies: pip install pysocks (optional, only needed for socks5://)
  • For the MCP server only: pip install -r requirements.txt (FastMCP). The CLI does not need it.

Usage

Basic

python3 reddit_osint.py someuser

Prints the report as JSON on stdout (summary only, without the per-item dump) and logs on stderr.

Saving the full report

python3 reddit_osint.py someuser --json report.json
python3 reddit_osint.py someuser --json report.json --quiet > /dev/null

report.json includes the all_items key with every post and comment (text, subreddit, score, date, status, permalink). The on-screen summary omits it so it does not dump 200 KB.

Filters

python3 reddit_osint.py someuser --subreddit dotnet       # one subreddit only
python3 reddit_osint.py someuser --keywords "kubernetes"  # posts mentioning X
python3 reddit_osint.py someuser --over18 true            # NSFW only
python3 reddit_osint.py someuser --source arctic          # single source (faster)
python3 reddit_osint.py someuser --max-pages 10           # cap pagination (100 items/page)
python3 reddit_osint.py someuser --delay 1.0             # slower, better with proxies
python3 reddit_osint.py someuser --rate-limit-retries 0   # fail fast when throttled

Proxies

python3 reddit_osint.py someuser --proxies proxies.txt

proxies.txt, one proxy per line. Accepted formats:

http://1.2.3.4:8080
https://1.2.3.4:8443
socks5://1.2.3.4:1080
socks5://user:[email protected]:1080
1.2.3.4:8080              # no scheme -> assumed http://
# comments and blank lines are ignored

Duplicates are removed. If the file contains no valid proxy, the script exits with code 2 and makes no requests.

Rotation. Proxies are walked round-robin. When a request comes back with a blocking/rate-limit code, that proxy is put on cooldown for --cooldown seconds and the script moves to the next one:

FlagDefaultWhat it does
--proxies FILE—Proxy list. Without it: direct connection.
--cooldown SECONDS30How long a banned proxy stays sidelined.
--max-attempts N(proxies+1) × 3Attempt ceiling per request before giving up.
--no-direct-fallbackoffWait instead of connecting directly when all proxies are cooling down.
--shuffle-proxiesoffShuffle the list at startup (spreads load from the first request).
--proxy-statsoffAdds proxy_stats to the report: ok / banned / errors per proxy.
--rate-limit-retries N2Retries of the whole page with exponential backoff (30 s, 60 s, capped at 120 s) when the API answers 429/422.

Codes treated as a ban (they rotate the proxy): 401 403 407 422 429 500 502 503 504 520 521 522 523 524, plus any 400/409 whose body contains slow down / rate limit / too many requests / timeout. Arctic Shift uses 422 "Timeout. Maybe slow down a bit" for rate limiting, so it is deliberately handled as a ban.

A 404 does not rotate: that is a bad request, not a bad proxy, and retrying elsewhere will not fix it.

Backoff. After a ban it waits min(cooldown, 10 s) — this applies to direct connections too, which is where it hurts most. If a page exhausts its attempts, fetch_all retries the page with exponential backoff (--rate-limit-retries, 2 attempts by default = 30 s and 60 s). That is what stops the script from silently losing comments when it trips a rate limit mid-run.

Sample log with rotation:

[proxy] http://1.2.3.4:8080 blocked (429), backing off 10s -> rotating
[proxy] http://5.6.7.8:1080 failed (Connection refused), rotating
[proxy] direct connection blocked (422), backing off 10s -> rotating
[!] arctic/comments rate limited, backing off 30s (retry 1/2)

When every proxy is cooling down, the script waits out the longest remaining cooldown and retries (or uses a direct connection unless --no-direct-fallback was passed).

Warning: a proxy is the egress path your IP takes to a third party. If you do not trust the proxy vendor, do not use one with this tool.

Environment variables

VariablePurpose
REDDIT_OSINT_ARCTICOverride the Arctic Shift URL (mirrors, tests).
REDDIT_OSINT_PULLPUSHOverride the PullPush URL.

What information it extracts

Raw fields from the APIs

id, title, selftext, body, author, subreddit, created_utc, score, permalink, url, over_18, num_comments, link_flair_text, preview.images, media_metadata, removed_by_category.

Status classification

StatusDetection
removedText = [removed] or [ Removed by Reddit ], or removed_by_category present (posts) → a moderator removed it.
deletedText = [deleted] or author == "[deleted]" → the author deleted it.
liveEverything else.

This is the whole point of the tool: the archive captures the fact of the removal even when it has no content, and that alone is useful metadata.

Profile and metrics

  • Totals for posts, comments, items, plus a live/removed/deleted breakdown with deleted_pct.
  • Real first and last activity (derived from items, not from archive metadata).
  • Archive metadata: total_karma, post_karma, comment_karma, num_posts, num_comments, earliest_post_at, last_comment_at (from /api/users/search).

Temporal patterns

Download Tool