
सरल, तेज़ वेब क्रॉलर जो वेब एप्लिकेशन के भीतर एंडपॉइंट्स और संपत्तियों की आसान, त्वरित खोज के लिए डिज़ाइन किया गया है।
URL और JavaScript फ़ाइल स्थानों को एकत्रित करने के लिए तेज़ गोलैंग वेब क्रॉलर। यह मूलतः शानदार Gocolly लाइब्रेरी का एक सरल कार्यान्वयन है।
एकल URL:
echo https://google.com | hakrawler
एकाधिक URL:
cat urls.txt | hakrawler
प्रत्येक stdin लाइन के लिए 5 सेकंड के बाद टाइमआउट:
cat urls.txt | hakrawler -timeout 5
सभी अनुरोधों को प्रॉक्सी के माध्यम से भेजें:
cat urls.txt | hakrawler -proxy http://localhost:8080
उपडोमेन शामिल करें:
echo https://google.com | hakrawler -subs
नोट: एक सामान्य समस्या यह है कि टूल कोई URL नहीं लौटाता। यह आमतौर पर तब होता है जब कोई डोमेन निर्दिष्ट किया जाता है (https://example.com), लेकिन यह किसी उपडोमेन (https://www.example.com) पर रीडायरेक्ट करता है। उपडोमेन दायरे में शामिल नहीं होता, इसलिए कोई URL प्रिंट नहीं होता। इसे हल करने के लिए, या तो रीडायरेक्ट श्रृंखला में अंतिम URL निर्दिष्ट करें या उपडोमेन शामिल करने के लिए
-subsविकल्प का उपयोग करें।
Google के सभी उपडोमेन प्राप्त करें, उन्हें खोजें जो http(s) पर प्रतिक्रिया देते हैं, और उन सभी को क्रॉल करें।
echo google.com | haktrails subdomains | httpx | hakrawler
पहले, आपको go इंस्टॉल करना होगा।
फिर hakrawler को डाउनलोड और कंपाइल करने के लिए यह कमांड चलाएँ:
go install github.com/hakluke/hakrawler@latest
अब आप ~/go/bin/hakrawler चला सकते हैं। यदि आप बिना पूर्ण पथ के केवल hakrawler चलाना चाहते हैं, तो आपको export PATH="~/go/bin/:$PATH" करना होगा। यदि आप चाहते हैं कि यह स्थायी रहे, तो आप इस पंक्ति को अपनी ~/.bashrc फ़ाइल में भी जोड़ सकते हैं।
echo https://www.google.com | docker run --rm -i hakluke/hakrawler:v2 -subs
ऊपर दी गई डॉकरहब विधि का उपयोग करना बहुत आसान है, लेकिन यदि आप इसे स्थानीय रूप से चलाना पसंद करते हैं:
git clone https://github.com/hakluke/hakrawler
cd hakrawler
sudo docker build -t hakluke/hakrawler .
sudo docker run --rm -i hakluke/hakrawler --help
नोट: यह सभी सुविधाओं के बिना hakrawler का एक पुराना संस्करण स्थापित करेगा, और यह बगी हो सकता है। मैं अन्य तरीकों में से किसी एक का उपयोग करने की सलाह देता हूँ।
sudo apt install hakrawler
फिर, hakrawler चलाने के लिए:
echo https://www.google.com | docker run --rm -i hakluke/hakrawler -subs
Usage of hakrawler:
-d int
Depth to crawl. (default 2)
-dr
Disable following HTTP redirects.
-h string
Custom headers separated by two semi-colons. E.g. -h "Cookie: foo=bar;;Referer: http://example.com/"
-i Only crawl inside path
-insecure
Disable TLS verification.
-json
Output as JSON.
-proxy string
Proxy URL. E.g. -proxy http://127.0.0.1:8080
-s Show the source of URL based on where it was found. E.g. href, form, script, etc.
-size int
Page size limit, in KB. (default -1)
-subs
Include subdomains for crawling.
-t int
Number of threads to utilise. (default 8)
-timeout int
Maximum time to crawl each URL from stdin, in seconds. (default -1)
-u Show only unique urls.
-w Show at which link the URL is found.