
WWWGrep は、HTML 要素をタイプ別に検査し、フォーカスされた(単一)、複数(ファイルベースの URL)、再帰的(ルートドメインに対する有無)検索を実行できる高速な「グレップ」メカニズムです。ヘッダー名と値も同様に再帰的に検索できます。WWWGrep は、検査対象のコードベースを迅速に調査するために、ブレーカーとビルダーの両方を支援するように設計されています。以下にいくつかのユースケースと例を示します。
インストール
git clone
pip3 install -r requirements.txt
python3 wwwgrep.py <arguments and parameters>
依存関係 (pip3 install -r requirements.txt)
- Python 3.5+
- BeautifulSoup 4
- UrlLib.parse
- requests_html
- argparse
- requests
- re
- os.path
wwwgrep.py [target/file] [search_string] [search params/criteria/recursion etc]
Search Inputs
search_string Specify the string to search for or alternatively “”
for all objects of type specified in search parameters
-t --target Specify a single URL as a target for the search
-f --file Specify a file containing a list of URLs to search
Recursion
-rr --recurse-root Limits URL recursion to the domain provided in the target
-ra --recurse-any Allows recursion to extend beyond the domain of the target
Matching Criteria
-i --ignore-case Performs case insensitive matching (default is to respect case)
-d --dedupe Allow duplicate findings per page (default is to de-duplicate findings)
-r --no-redirects Do not allow redirects (default is to allow redirects)
-b --no-base-url Omit the URL of the match from the output (default is to include the URL)
-x --regex Allows the use of RegEX matches (search_string is treated as a RegEX, default is off)
-e --separator Specify and output specifier (default is : )
-j --java-render Turns on JavaScript rendering of page objects and text (default is off)
-p --linked-js-on Turns on searching of linked (script src tags) Java Script (default is off)
Request Parameters
-ps --https-proxy Specify a proxy for the HTTPS protocol in https://<ip>:<port> format
-pp --http-proxy Specify a proxy for the HTTP protocol in http://<ip>:<port> format
-hu --user-agent Specify a string to use as the user agent in the request
-ha --auth-header Specify a bearer token or other auth string to use in the request header
Search Parameters
-s --all Search all page HTML and scripts for terms that match the search specification
-sr --relative Search page links that match the search specification as relative URLs
-sa --absolute Search page links that match the search specification as absolute URLs
-si --input-fields Search page input fields that match the search specification
-ss --scripts Search scripts tags that match the search specification
-st --text Search visible text on the page that matches the search specification
-sc --comments Search comments on the page that match the search specification
-sm --meta Search in page metadata for matches to the search specification
-sf --hidden Search in hidden fields for specific matches to the search specification
-sh --header-name Search response headers for specific matches to the search specification
-sv --header-value Search response header values for specific matches to the search specification
大文字小文字を区別せず、ルートドメイン内で再帰的にサイト上の 'login' という名前のすべての入力フィールドを見つける
wwwgrep.py -t https://www.target.com -i -si “login” -rr
サイト上のすべてのページで「to do」を含むすべてのコメントを見つける
wwwgrep.py -t https://www.target.com -i -sc “to do” -rr
特定のWebページ上のすべてのコメントを見つける
wwwgrep.py -t https://www.target.com/some_page -i -sc “”
ファイル input.txt に含まれるWebアプリケーションのリスト内のすべての隠しフィールドをサイト再帰を使用して見つける
wwwgrep.py -f input.txt -sf “” -rr