
PhishCollector is a research framework for collecting, analysing, and tracking phishing sites.
PhishCollector is a research framework for collecting, analysing, and tracking phishing sites. It is intentionally designed as a starting point — the detection rules, technology signatures, wordlists, and plugins are all plain data structures that researchers are expected to read, extend, and adapt to their own threat landscape.
Submit a suspicious URL and PhishCollector will:
All results are accessible via a REST API, a web dashboard, and a CLI.


cp .env.example .env # configure (see below)
docker compose up --build # starts db + app + frontend
| Service | URL |
|---|---|
| GUI | http://localhost:3000 |
| API docs | http://localhost:8000/docs |
| DB | localhost:5432 |
All settings are environment variables with the PHISH_ prefix. Copy .env.example to .env and adjust.
| Variable | Default | Description |
|---|---|---|
PHISH_DATABASE_URL | postgres://… | PostgreSQL DSN |
PHISH_API_KEY | (empty) | If set, all requests require X-API-Key: <value> |
PHISH_DATA_DIR | /data | Storage root for screenshots, HTML, assets |
PHISH_BROWSER_TIMEOUT | 30000 | Page-load timeout in ms |
PHISH_REQUEST_TIMEOUT | 15 | HTTP sub-request timeout in seconds |
PHISH_MAX_SPIDER_PAGES | 50 | Max URLs the spider visits per job |
PHISH_MAX_ASSET_SIZE | 10485760 | Max JS/CSS file size to store (bytes) |
PHISH_PROXY_URL | (empty) | Outbound proxy — see below |
PHISH_PROXY_SSL_VERIFY | true | Set false for intercepting proxies — see below |
PHISH_URLHAUS_ENABLED | false | Enable URLhaus reputation check |
PHISH_VIRUSTOTAL_API_KEY | (empty) | VirusTotal v3 API key (leave empty to disable) |
Routing all outbound traffic through a proxy keeps your analyst IP hidden from the phishing server.
PHISH_PROXY_URL=socks5://127.0.0.1:9050
PHISH_PROXY_SSL_VERIFY=true # Tor does not intercept TLS
Burp acts as a TLS man-in-the-middle and presents its own CA certificate for every HTTPS connection. Without disabling SSL verification every HTTPS request through the proxy will fail.
PHISH_PROXY_URL=http://127.0.0.1:8080
PHISH_PROXY_SSL_VERIFY=false # required for Burp / intercepting proxies
Note:
PHISH_PROXY_SSL_VERIFY=falseonly affects outbound HTTPS connections made by the Python backend (plugins, fingerprinter, spider). The Playwright browser already operates withignore_https_errors=trueregardless of this setting.
Warning: Never set
PHISH_PROXY_SSL_VERIFY=falsewithout a proxy configured — it would disable certificate validation for all external API calls (URLhaus, VirusTotal).
Base path: /api/v1
| Method | Path | Description |
|---|---|---|
POST | /collections | Submit a URL for collection |
GET | /collections | List all collections |
GET | /collections/{id} | Full detail + fingerprint |
GET | /collections/{id}/screenshot | Full-page PNG |
GET | /collections/{id}/html | Captured HTML (downloaded as plain-text) |
GET | /collections/{id}/requests | Network request log |
GET | /collections/{id}/spider | Spider results |
GET | /collections/{id}/plugins | Threat-intel plugin results |
POST | /collections/{id}/plugins/refresh | Re-run plugins (e.g. fetch pending VT result) |
POST | /collections/{id}/rescan | Re-collect the same URL (original is preserved) |
PATCH | /collections/{id} | Update tags and notes |
GET | /collections/{id}/export?format=json|csv | Export collection data |
DELETE | /collections/{id} | Delete a collection and all its artifacts |
GET | /search | Search fingerprints by IP, favicon hash, technology, country, title |
Full interactive docs at /docs (Swagger UI).
curl -X POST http://localhost:8000/api/v1/collections \
-H 'Content-Type: application/json' \
-d '{"url": "https://suspicious-site.example.com", "use_wordlist": true}'
# Install (inside container or local venv with requirements.txt)
pip install -e .
# Submit a URL and wait for completion
phishcollector collect https://target.example.com --wait
# With wordlist fuzzing
phishcollector collect https://target.example.com --wordlist --wait
# List recent jobs
phishcollector list
# View full detail
phishcollector detail <job-id>
# Download screenshot
phishcollector screenshot <job-id> -o capture.png
# Search by tech stack / favicon hash / country
phishcollector search --tech WordPress --country RU
phishcollector search --favicon-hash -1234567890
Requires a free Auth-Key from auth.abuse.ch.
PHISH_URLHAUS_ENABLED=true
PHISH_URLHAUS_API_KEY=<your-auth-key>
Requires a free or paid API key from virustotal.com.
PHISH_VIRUSTOTAL_API_KEY=<your-key>
When a URL has not yet been analysed by VT, PhishCollector submits it for scanning and automatically re-fetches the result every 30 seconds until it resolves.
Each plugin is a single file in phishcollector/plugins/ that exposes one async function:
# phishcollector/plugins/myplugin.py
from . import CheckResult
async def check(url: str, proxy_url=None, ssl_verify=True) -> CheckResult:
# query your feed / API here
return CheckResult(
plugin_name="myplugin",
status="malicious", # malicious | suspicious | clean | unknown | error
score=0.95, # 0.0–1.0, or None
result={"raw": ...}, # stored as JSONB, displayed in the GUI
)
Then register it in phishcollector/plugins/runner.py:
from .myplugin import check as myplugin_check
tasks.append(myplugin_check(url, proxy_url=settings.proxy_url, ssl_verify=settings.proxy_ssl_verify))
No other changes are needed — the result is automatically stored, displayed in the dashboard, and factored into the threat score.
The detection engine is intentionally kept as plain, readable data so researchers can tune it to the kits and campaigns they are tracking. Everything lives in one file:
phishcollector/collector/fingerprint.py
PHISHING_PATTERNS — regex rules scanned against rendered HTML + JSEach entry is a (regex, human_readable_label) tuple grouped into categories. A match in any category is surfaced in the Indicators tab and counts toward the threat score.
PHISHING_PATTERNS: dict[str, list[tuple[str, str]]] = {