Skip to content
KitploitKITPLOIT
도구블로그
제출
도구블로그
제출

해킹, 침투 테스트 및 사이버 보안 도구를 당신의 보안 무기고에!

Kitploit은 해킹, 사이버 보안 및 침투 테스트 도구 디렉토리입니다. 최신 프로젝트 업데이트를 발견하여 취약점을 찾고, 시스템을 분석하고, 테스트를 자동화하고, 보안을 강화하세요.

··피드·문의·개인정보·© 2026 Kitploit

도구 디렉토리

카테고리

모든 카테고리 보기
Loading categories
Crawlector — Crawlector는 악성 객체를 검사하기 위해 웹사이트를 스캔하도록 설계된 위협 헌팅 프레임워크입니다. | Kitploit
도구/GitHubGitHub/mfmokbel/crawlector
OSINT (Open Source Intelligence)Vulnerability ScannersThreat Feeds & AggregatorsInformation GatheringWeb SecurityMalware AnalysisThreat IntelligenceCrawler
GitHubmfmokbel/crawlector

Crawlector

Crawlector는 악성 객체를 검사하기 위해 웹사이트를 스캔하도록 설계된 위협 헌팅 프레임워크입니다.

저장소 보기
123108개월 전Kitploit 검토 완료

인기

모두 보기 →

커뮤니티에서 가장 많이 사용되는 도구를 찾아보세요.

모든 도구 탐색

도구 컬렉션을 둘러보세요

모든 도구 보기 →
공유
웹사이트

Crawlector

Crawlector (이름 Crawlector는 Crawler & Detector의 조합입니다)는 웹사이트에서 악성 객체를 스캔하기 위해 설계된 위협 헌팅 프레임워크입니다.

참고-1: 이 프레임워크는 2022년 10월 22일 이탈리아 베르가모에서 열린 No Hat 컨퍼런스에서 처음 발표되었습니다 (슬라이드, YouTube 녹화). 또한, 2022년 12월 2일 싱가포르에서 열린 AVAR 컨퍼런스에서 두 번째로 발표되었습니다.

참고-2: 발표에서 언급된 동반 도구 EKFiddle2Yara (EKFiddle 규칙을 가져와 Yara 규칙으로 변환하는 도구)도 두 컨퍼런스에서 함께 공개되었습니다.

참고-3: 버전 2.0 (Photoid Build:180923)은 2023년 9월 18일에 출시된 중요한 릴리스입니다.

참고-4: 버전 2.1 (Universe-647 Build:031023)은 2023년 10월 3일에 출시되었습니다. 주요 추가 사항은 Slack 알림 기능입니다.

참고-5: 버전 2.2 (Hallstatt Build:051123)은 2023년 11월 5일에 출시되었습니다. 주요 추가 사항은 Slack 원격 제어 기능입니다.

참고-6: 버전 2.3 (Munich Build:241123)은 2023년 11월 24일에 출시되었습니다. 주요 추가 사항은 DNS 네임서버 기능입니다.

참고-6: 버전 2.3.1 {Nero Build:131225}은 2025년 12월 13일에 출시되었습니다. 유지보수 릴리스입니다.

기능

  • 웹사이트 스파이더링을 지원하여 추가 링크 찾기 (최대 2레벨 깊이)
  • Yara를 백엔드 엔진으로 통합하여 규칙 스캔
  • 온라인 및 오프라인 스캔 지원
  • 도메인/사이트 디지털 인증서 크롤링 지원
  • URLhaus 쿼리를 통해 페이지 내 악성 URL 탐지 지원
  • 심층 객체 추출 (DOE)
  • Slack 알림
  • HTTP 리디렉션에 대한 파라미터화된 지원
  • Whois 정보 검색
  • 페이지 콘텐츠 해싱 지원: TLSH (Trend Micro Locality Sensitive Hash) 및 md5, sha1, sha256, ripemd128 등 기타 표준 암호화 해시 함수
    • TLSH는 페이지 크기가 50바이트 미만이거나 데이터에 충분한 무작위성이 없으면 값을 반환하지 않음
  • 모든 URL의 평점 및 카테고리 조회 지원
  • 주어진 사이트에 대해 동일한 도메인의 사용 가능한 모든 TLD 및/또는 서브도메인을 찾아 확장 지원
    • 이 기능은 Omnisint Labs API (이 사이트는 2023년 3월 10일 기준 다운됨) 및 RapidAPI API 사용
    • TLD 확장 구현은 네이티브
    • 이 기능은 평점 및 카테고리화와 함께 원본 도메인에 대한 스캠/피싱/악성 도메인을 찾는 기능을 제공
  • 도메인 해석 지원 (IPv4 및 IPv6)
  • 스캔된 웹사이트 페이지를 나중에 스캔할 수 있도록 저장 (zip 압축 가능)
  • 프레임워크의 모든 설정은 단일 사용자 정의 설정 파일로 제어
  • 모든 스캔 세션은 잘 구조화된 CSV 파일에 저장되며, 스캔된 웹사이트에 대한 풍부한 정보와 트리거된 Yara 규칙에 대한 정보 포함
  • 기타 여러 기능...
  • 모든 HTTP(S) 통신은 프록시 인식
  • 단일 실행 파일
  • C++로 작성됨

URLHaus 스캔 및 API 통합

이는 스캔 중인 모든 페이지에 대해 악성 URL을 확인하는 기능입니다. 프레임워크는 URLHaus 서버 (설정: url_list_web)에서 악성 URL 목록을 쿼리하거나 디스크의 파일 (설정: url_list_file)에서 쿼리할 수 있으며, 후자가 지정되면 전자보다 우선합니다.

작동 방식: 모든 페이지의 콘텐츠를 url_list_web 또는 url_list_file의 모든 URL 항목에 대해 검색하여 모든 발생을 확인합니다. 또한, 일치하는 경우 check_url_api 설정 옵션이 true로 설정되어 있으면 Crawlector는 url_api 설정 옵션에 설정된 API URL로 POST 요청을 보내며, 일치하는 URL에 대한 추가 정보가 포함된 JSON 객체를 반환합니다. 이러한 정보에는 urlh_status (예: online, offline, unknown), urlh_threat (예: malware_download), urlh_tags (예: elf, Mozi), urlh_reference (예: https://urlhaus.abuse.ch/url/1116455/)가 포함됩니다. 이 정보는 check_url_api가 true로 설정된 경우에만 로그 파일 cl_mlog_<current_date><current_time><(pm|am)>.csv (아래 참조)에 포함됩니다. 그렇지 않으면 로그 파일에는 urlh_url (일치하는 악성 URL 목록) 및 urlh_hit (일치하는 각 악성 URL의 발생 횟수) 열이 포함되며, 이는 check_url이 true로 설정된 경우에만 해당됩니다.

URLHaus 기능은 설정 옵션 check_url을 false로 설정하여 완전히 비활성화할 수 있습니다.

이 기능은 확인해야 하는 방대한 악성 URL (~ 현재 약 1억 3천만 개 항목)과 URLHaus 서버에서 추가 정보를 가져오는 시간(옵션 check_url_api가 true로 설정된 경우)으로 인해 스캔 속도가 느려질 수 있습니다.

파일 및 폴더 구조

  1. \cl_sites
    • 방문하거나 크롤링할 사이트 목록이 저장되는 위치.
    • 여러 파일 및 디렉토리 지원.
  2. \crawled
    • 모든 크롤링/스파이더링된 URL이 텍스트 파일로 저장되는 위치.
  3. \certs
    • 모든 도메인/사이트 디지털 인증서가 저장되는 위치 (.der 형식).
  4. \results
    • 방문한 웹사이트가 저장되는 위치. 옵션 results_dir을 통해 설정 가능.
  5. \pg_cache
    • 스파이더 기능의 일부가 아닌 사이트에 대한 프로그램 캐시. 옵션 cache_dir, 섹션 [default] 을 통해 설정 가능.
  6. \cl_cache
    • 스파이더 기능의 일부인 사이트에 대한 크롤러 캐시. 옵션 cache_dir, 섹션 [spider] 를 통해 설정 가능.
  7. \yara_rules
    • 모든 Yara 규칙이 저장되는 위치. 이 디렉토리에 있는 모든 규칙은 엔진에 의해 로드되고, 구문 분석, 유효성 검사 및 실행 전 평가됨.
  8. cl_config.ini
    • 프레임워크의 동작에 영향을 주기 위해 조정할 수 있는 모든 구성 매개변수가 포함된 파일.
  9. cl_mlog_<current_date><current_time><(pm|am)>.csv
    • 방문한 웹사이트에 대한 풍부한 정보가 포함된 로그 파일
    • 날짜, 시간, Yara 스캔 상태, 각 일치 항목의 오프셋 및 길이와 함께 트리거된 Yara 규칙 목록, ID, URL, HTTP 상태 코드, 연결 상태, HTTP 헤더, 페이지 크기, 디스크에 저장된 페이지 경로, URLHaus 결과 관련 기타 열 포함.
    • 파일 이름은 세션별로 고유함.
  10. cl_offl_mlog_<current_date><current_time><(pm|am)>.csv
    • 오프라인으로 스캔된 파일에 대한 정보가 포함된 로그 파일.
    • 일치 항목의 오프셋 및 길이와 함께 트리거된 Yara 규칙 목록, 디스크에 저장된 페이지 경로 포함.
    • 파일 이름은 세션별로 고유함.
  11. cl_certs_<current_date><current_time><(pm|am)>.csv
    • 발견된 디지털 인증서에 대한 풍부한 정보가 포함된 로그 파일.
  12. \expanded\exp_subdomain_<pm|am>.txt
    • 발견된 서브도메인 포함 ([site] 섹션의 일부)
  13. \expanded\exp_tld_<pm|am>.txt
    • 발견된 도메인 포함 ([site] 섹션의 일부)

설정 파일 (cl_config.ini)

세션을 실행하기 전에 반드시 설정 파일 cl_config.ini에 익숙해져야 합니다. 모든 섹션과 매개변수는 설정 파일 자체에 문서화되어 있습니다.

Yara 오프라인 스캔 기능은 독립 실행형 옵션입니다. 즉, 활성화되면 Crawlector는 다른 활성화된 기능과 관계없이 이 기능만 실행합니다. 도메인/사이트 디지털 인증서 크롤링 기능도 마찬가지입니다. 어느 쪽이든 설정 파일에서 사용하지 않는 모든 기능을 비활성화하는 것이 좋습니다.

  • 설정(log_to_file 또는 log_to_cons)에 따라 Yara 규칙이 모듈 속성(예: PE, ELF, Hash 등)만 참조하는 경우, Crawlector는 일치 시 규칙 이름만 표시하고 오프셋 및 길이 데이터는 제외합니다.

참고: 경로를 사용하는 모든 옵션에는 항상 절대 경로를 제공하십시오.

사이트 형식 패턴

웹사이트를 방문/스캔하려면 URL 목록이 텍스트 파일에 저장되어야 하며, 디렉토리 "cl_sites"에 있어야 합니다.

Crawlector는 세 가지 유형의 URL을 허용합니다:

  1. 유형 1: 한 줄에 하나의 URL
    • Crawlector는 각 URL에 URL 호스트 이름에서 파생된 고유 이름을 할당합니다.
  2. 유형 2: 한 줄에 하나의 URL, 고유 이름과 함께 [a-zA-Z0-9_-]{1,128} = <url>
  3. 유형 3: 스파이더 기능을 위해 고유한 형식이 사용됩니다. 한 줄에 하나의 URL은 다음과 같습니다:

<id>[depth:<0|1>-><\d+>,total:<\d+>,sleep:<\d+>] = <url>

예를 들어,

mfmokbel[depth:1->3,total:10,sleep:0] = https://www.mfmokbel.com

이는 다음과 동일합니다: mfmokbel[d:1->3,t:10,s:0] = https://www.mfmokbel.com

여기서, <id> := [a-zA-Z0-9_-]{1,128}

depth, total 및 sleep can also be replaced with their shortened versions d, t and s, respectively.

  • depth: 스파이더는 추가 URL을 찾기 위해 두 레벨 깊이로 이동을 지원합니다 (이는 설계 결정입니다).
  • 0 값은 레벨 1 깊이를 나타내며, "->" 뒤의 값은 무시됩니다.
  • 레벨 1 깊이는 total 매개변수에 의해 제어됩니다. 따라서 스파이더는 먼저 지정된 URL에서 가능한 많은 추가 URL을 찾으려고 시도합니다.
  • "->" 뒤의 값은 total 매개변수 값에 따라 발견된 각 URL에 대해 스파이더링할 최대 URL 수를 나타냅니다.
  • 1 값은 레벨 2 깊이를 나타내며, "->" 뒤의 값은 total 매개변수에 따라 발견된 각 URL에 대해 찾을 최대 URL 수를 나타냅니다. 명확히 하자면, 위의 예에서와 같이 먼저 스파이더는 total 매개변수에 지정된 대로 10개의 URL을 찾고, 발견된 각 URL은 최대 3개의 URL까지 스파이더링됩니다. 따라서 최상의 시나리오에서 우리는 40 (10 + (10*3))개의 URL을 얻게 됩니다.
  • sleep 매개변수는 각 HTTP 요청 사이에 대기할 밀리초 단위의 정수 값을 사용합니다.

참고 1: 유형 3 URL은 설정 파일의 spider 섹션에서 설정 매개변수 live_crawler를 false로 설정하여 유형 1 URL로 전환할 수 있습니다.

참고 2: 빈 줄과 ";" , "#" 또는 "//"로 시작하는 줄은 무시됩니다.

스파이더 기능

스파이더 기능은 Crawlector에 대상 페이지에서 추가 링크를 찾는 기능을 제공합니다. 스파이더는 다음 기능을 지원합니다:

  • 스파이더 기능이 작동하려면 도메인이 유형 3이어야 합니다.
  • exclude_url 설정 옵션을 통해 스파이더링에서 일치하는 URL을 제외하기 위해 와일드카드 패턴 목록(파이프로 구분)을 지정할 수 있습니다. 예를 들어, *.zip|*.exe|*.rar|*.zip|*.7z|*.pdf|.*bat|*.db
  • include_url 설정 옵션을 통해 패턴과 일치하는 URL만 스파이더링하도록 와일드카드 패턴 목록(파이프로 구분)을 지정할 수 있습니다. 예를 들어, */checkout/*|*/products/*
  • exclude_https 설정 옵션을 통해 HTTPS URL을 제외할 수 있습니다.
  • add_ext_links 설정 옵션을 통해 기본 페이지만의 아웃바운드/외부 링크를 고려할 수 있습니다. 이 기능은 exclude_url 및 include_url 설정 옵션을 따릅니다.
  • ext_links_only 설정 옵션을 통해 기본 페이지만의 아웃바운드/외부 링크를 고려하고 다른 모든 URL은 제외할 수 있습니다. 이 기능은 exclude_url 및 include_url 설정 옵션을 따릅니다.

ID 유형

릴리스 2.0에서 ID는 ID 자체에 다음 유형 중 하나를 추가하여 명시적으로 유형이 할당됩니다:

각 ID가 유형을 함께 전달하면 결과를 탐색하고 필터링하기가 더 쉽습니다. 또한, 이는 내부적으로 다양한 목적으로 사용됩니다.

사이트 순위 기능

  • 웹사이트 순위를 확인하는 기능입니다.
  • 순위가 있는 웹사이트 목록이 포함된 파일을 CSV 파일 형식으로 제공합니다.
  • 웹사이트 순위 목록을 제공하는 서비스에는 Alexa top-1m (2022년 5월 중단), Cisco Umbrella, Majestic, Quantcast, Farsight 및 Tranco 등이 있습니다.
  • CSV 파일 형식 (2개 열만): 첫 번째 열은 순위, 두 번째 열은 도메인 이름입니다.
  • 셀에 따옴표로 묶인 데이터가 포함된 경우 자동으로 따옴표가 제거됩니다.
  • 따옴표로 묶인 텍스트에는 줄 바꿈이 허용되지 않습니다.
  • 셀에서 읽을 때 앞뒤 공백이 제거됩니다.
  • 빈 줄과 주석 줄은 건너뜁니다.
  • 설정 파일의 site_ranking 섹션은 CSV 파일을 읽는 방식을 변경하는 몇 가지 옵션을 제공합니다.
  • 이 쿼리의 성능은 CSV 파일의 레코드 수에 따라 달라집니다.
  • Crawlector는 CSV 파일의 모든 항목을 조사 중인 도메인과 비교하며, 그 반대는 아닙니다.
  • 등록된/지불 수준 도메인만 비교됩니다.

TLD 및 서브도메인 찾기 - [site] 섹션

  • site 섹션은 동일한 도메인에 대해 사용 가능한 모든 최상위 도메인(TLD) 및/또는 서브도메인을 찾아 주어진 사이트를 확장하는 기능을 제공합니다. 발견된 새 TLD/서브도메인은 다른 도메인과 마찬가지로 확인됩니다.
  • 이 기능은 Omnisint Labs (https://omnisint.io/) 및 RapidAPI API를 사용합니다.
  • Omnisint Labs API는 서브도메인과 TLD를 반환하는 반면, RapidAPI는 서브도메인만 반환합니다 (Omnisint Labs API는 2023년 3월 10일 기준 다운되었지만, 사이트가 복구되는 경우를 대비해 구현은 여전히 유지됩니다).
  • RapidAPI의 경우 RapidAPI에서 요청할 수 있는 유효한 "Domains records" API 키가 필요하며, 설정 파일의 rapid_api_key 키에 입력합니다.
  • find_tlds가 활성화된 경우, Omnisint Labs API의 tld 결과 외에도 프레임워크는 tlds_file 또는 tlds_url의 모든 tld 항목을 검토하여 다른 활성/등록된 도메인을 찾으려고 시도합니다.
  • tlds_url이 설정된 경우, 각 줄에 하나씩 tld를 호스팅하는 URL을 가리켜야 합니다 (';', '#' 또는 '//' 문자로 시작하는 줄은 무시됩니다).
  • tlds_file은 tld 목록이 포함된 파일 이름을 보유합니다 (tlds_url과 동일, '.'을 제외한 tld만, 예: "com", "org").
  • tlds_file이 설정되면 tlds_url보다 우선합니다.
  • tld_dl_time_out은 확인 중인 도메인이 해석되는지 여부를 확인할 때 dnslookup 함수의 최대 시간 제한을 설정합니다.

리디렉션 기능

이전 릴리스의 URL 리디렉션 기능은 손상되었습니다. 이 릴리스는 리디렉션 기능의 완전한 재작성을 제공하며, 작동 제어를 위한 높은 수준의 매개변수화를 제공합니다. 릴리스 버전 2.0에서 리디렉션은 설정 파일에 [redirect] 라는 전용 섹션이 있습니다. 리디렉션 기능 전체는 [default] 섹션의 follow_redir 옵션을 통해 켜고 끌 수 있습니다.

리디렉션 함수는 HTTP 응답 상태 코드 301, 302, 303, 307 및 308을 확인합니다. 일치하는 경우 Crawlector는 Location 헤더를 구문 분석하여 리디렉션 대상 URL을 확인하며, 절대 및 상대 리디렉션 URL을 모두 처리합니다. Crawlector의 리디렉션 기능은 성능과 민첩성을 위해 설계되었습니다. [redirect] 섹션은 다음 옵션 목록을 제공합니다:

[redirect]

  • depth = all ; (t: string)
  • max_redirect = 200 ; (t: uint16_t)
  • visit = true ; (t: bool)
  • skip_similar = true ; (t: bool)

depth 옵션은 값 last 또는 all 중 하나를 사용합니다. visit 옵션이 활성화되었는지 여부에 따라 발견된 리디렉션 URL 중 방문할 대상을 제어합니다. all 은 발견된 모든 리디렉션 URL을 방문하는 것입니다. last 는 마지막 리디렉션 URL만 방문하는 것입니다. 해당 URL의 방문은 동일/현재 세션에서 이루어집니다. depth 값에 관계없이 Crawlector는 발견된 모든 리디렉션 URL 목록을 절대 형식으로 총 개수와 함께 기록합니다. 이는 cl_mlog CSV 파일의 열 redirect_urls 및 redirect_total 아래에 기록됩니다.

max_redirect 옵션은 발견할 URL 리디렉션의 총 개수에 상한을 설정합니다.

skip_similar 옵션은 다음 예를 통해 가장 잘 설명됩니다:

Crawlector에 크롤링할 원래 URL이 "https://www.mfa.gov.law"이고 발견된 redirect_urls 중 하나가 "https://mfa.gov.law/"라고 가정합니다. 보시다시피 유일한 차이점은 URL 끝의 슬래시입니다. 이 두 URL은 동일하며 서버는 동일한 페이지로 응답합니다. visit 옵션이 true로 설정되면 Crawlector는 두 URL을 모두 크롤링하여 리소스를 낭비하고 동일한 작업을 두 번 수행합니다. 이는 1~2개의 URL에는 문제가 되지 않을 수 있지만, 크롤링하려는 URL이 1000개 이상이고 visit 옵션이 활성화된 경우 절반 이상이 그러한 발견된 URL을 가질 가능성이 매우 높아져 이는 중요한 문제가 됩니다. 따라서 skip_similar 옵션을 true로 설정하면 유사한 URL 방문을 건너뛰어 이 문제를 해결하는 데 도움이 됩니다. skip_similar 옵션은 forward_slash 시나리오 외에도 다음 두 가지 시나리오를 처리합니다: 리디렉션 URL이 접두사 "https://" 및 "www." 중 하나 또는 둘 모두만 다른 경우.

심층 객체 추출 (DOE)

릴리스 2.0의 주요 추가 사항 중 하나는 페이지에서 다양한 유형의 객체를 추출하여 디스크에 저장하고 Yara 및 URLHaus로 스캔한 후 결과를 CSV 파일에 저장하는 기능입니다. 이 기능을 활성화하려면 [page] 섹션의 extract_obj 옵션을 true로 설정하십시오.The implementation of the deep object extraction feature works by creating an MHT web archive file from the webpage, including external scripts, images and CSS files. All embedded files will be extracted into the path specified by the option obj_dir (path: obj_dir/objects/), where each file will be scanned. The implementation is not to be confused with headless browser functionality. DOE is different and doesn't involve loading the page to retrieve all dynamically queried URLs. Therefore, it has its limitations.

All of the extracted objects will have some of their metadata written into the CSV file. Things to keep in mind when reading the CSV file: the ID of the domain with the extracted object has a unique format, as follows, <domain_id>_<type>_p_obj_<counter> (for example, _mfa_gov_cef40bc5-ba6a-41_t1_p_obj_0_). And, the url will have the following format, <url>__<object_filename> (for example, https://www.mfa.gov.law\_\_bilmur.min.js).

If the option delete_obj is set to true, then all extracted objects that aren't being detected by Yara are deleted from disk. If the option log_all_objs is set to true, then log all extracted objects metadata to the same cl_mlog CSV file. If the option check_urlhaus under the [page] section is set to true, then every extracted object will be URLHaus-scanned. Note that this option's options are inherited from the section [urlhaus].

Note: if the domain being crawled redirects to another domain, then the last redirect to URL has to be passed to DOE to work. Moreover, the domain has to start with "HTTP(S)://" for DOE to work.

Slack Alert Notification

Sometimes, you might want to run Crawlector sessions that might take days to complete, for example, by crawling the top 1-million Alexa websites, and for such a scenario, you need a way to monitor the framework's operation and progress remotely. Therefore, in release 2.1, I've added the Slack alert notification feature to provide a mechanism to monitor the execution of Crawlector in real-time, by sending Yara's alerts, std::exit() events, and process warnings and errors, to a Slack channel of your choosing. In addition to that, Crawlector installs a console handler in an attempt to monitor certain event types, including ctrl_c, ctrl_close, ctrl_break, ctrl_logoff and ctrl_shutdown. It is important to keep in mind that Crawlector doesn't change/alter the default handler's behaviour; it merely reports to the Slack channel the receipt of any of the listed events. This could be extended in the future to account for other types of events.

This feature uses Slack REST API, and for authentication with the server, it uses OAuth 2.0. You'll need a Slack API token to use it, and a channel configured with the right permissions. This feature only posts messages to the Slack channel and doesn't receive or process any incoming messages.

The [slack_alert] section provides the following list of options:

[slack_alert]

  • alert = true ; (t: bool)
  • api_token = ; (t: string)
  • channel = ; (t: string)
  • sleep = ; (t: uint32_t) in milliseconds

To disable or enable this feature, simply set the option alert to true or false. Moreover, you need to specify the api_token, with a channel name.

Note-1: In the initialization phase of Crawlector, it tests whether the provided authentication token is valid or not, or if the channel is set, and in case of failure, this feature is disabled automatically.

All alerts reported to the Slack channel are reported under the user's name Crawlector v<version_number>, for example, Crawlector v2.1. The user has the icon of a spider web. Additionally, all alerts are threaded, meaning all subsequent alerts after the first starting message, are posted as replies. This was a design decision and helps in case you're running multiple sessions at the same time, all reporting to the same channel. Some alerts use the markdown markup language for formatting.

When the process has finished successfully and is about to exit, it posts the following message:

Crawlector has finished and is shutting down successfully

Note-2: Slack rate limit on the post message API is one message per second, with leeway for some bursts. Crawlector does not queue messages to account for more posts per second. This might change in the future if required; however, the option sleep allows for the process to sleep for a specified amount of time after every successfully posted message.

Slack Remote Control

With release 2.2 (code-named Hallstatt), I'm introducing the capability to remotely control Crawlector via a selected set of specially designed control commands. The reason for introducing this functionality is to monitor and control certain behaviours of sessions that are supposed to run for hours or days. For example, you might want to turn on/off the Slack alert functionality, terminate Crawlector, and upload a configuration file, among others.

This feature uses Slack REST API, and for authentication with the server, it uses OAuth 2.0. You'll need a Slack API token to use it, and a channel configured with the right permissions. The API token is the same as that used in the [slack_alert] section, option api_token.

The [slack_alert] section provides the following additional list of options for the remote control functionality:

[slack_alert] (control options)

  • control = true ; (t: bool)
  • ctrl_channel = ; (t: string) it has to be the channel ID and not the channel name
  • ctrl_sleep = ; (t: uint32_t) in milliseconds

To disable or enable this feature, simply set the option control to true or false. The ctrl_channel name has to be the channel ID name and not the channel name. You can get it by right-clicking on the channel name -> View channel details -> Scroll down to the bottom of the window, and you'll see the Channel ID: <channel_id> field.

The option ctrl_sleep determines the frequency of calling out to the control channel specified in the ctrl_channel option for retrieving control commands. You could also update this option via the control command cl_update_delay <time_in_ms>.

The list of supported control commands is the following:

Note-1: In the initialization phase of Crawlector, it tests whether the provided authentication token is valid or not, or if the channel is set, and in case of failure, this feature is disabled automatically.

If this functionality is enabled, and once it passes API token validation, Crawlector sends the message "Crawlector is ready for receiving control commands. Type the command cl_help for a list of supported control commands." to the designated ctrl_channel.

All responses to a given control command are threaded. Moreover, control commands are read on a session-by-session basis, from the time a session is started.

Note-2: Slack rate limit on the retrieval (conversation history) message API is one request per second, with leeway for some bursts. So, if the ctrl_sleep option is set to a value less than a second or greater than a second, Crawlector does queue messages to account for more control commands per second, and execute them in the order received.

DNS Nameservers

With release 2.3 (code-named Munich), the capability to specify a list of DNS nameservers for all DNS queries and DNS-to-IP resolutions attempted by Crawlector is introduced with a high level of control. This is important in case you're crawling blocked or malicious websites. This feature applies to every function in Crawlector where a DNS query or DNS-to-IP request is made. More importantly, it provides the capability to perform DNS over TLS for every nameserver that supports it.

The [dns_ns] section provides the following list of options for administering this functionality:

[dns_ns]

  • enable = false ; (t: bool)
  • name_servers = 8.8.8.8(e_tls),12.13.14.15(d_tls) ; (t: string)
  • dns_tls = yes ; (t: string) (yes, no or force)
  • keep_default = false ; (t: bool)
  • conn_time_out = 3000 ; (t: uint32_t) in milliseconds (0 to wait indefinitely)

The option name_servers takes a parametrized list of DNS name servers to use, comma-separated. The value of this option has the format: <IPv4_address>(<tls_option>) where <tls_option> takes either of the values "d_tls" or "e_tls". The options "d_tls" or "e_tls" indicate whether the nameserver in question supports DNS over TLS or not, respectively. This option will be enforced depending on the value set for the option dns_tls. For example, the entry 8.8.8.8(e_tls) indicates to use the Google DNS server 8.8.8.8 with TLS support, whereas the entry 12.13.14.15(d_tls) indicates to use the DNS server 12.13.14.15 with no TLS support.

The option dns_tls specifies the required level of TLS enforcement. This option takes either of the values, "yes" "no" or "force".

  • yes
    • DNS nameservers with TLS support will be attempted first, and if no NS with TLS support is found, then UDP/TCP resolution will be attempted instead.
  • no
    • no TLS support (do not use DNS nameservers with TLS support)
  • force
    • only DNS nameservers with TLS support will be used, and any DNS query initiated by Crawlector will use DoT (DNS over TLS).

The option keep_default is for whether to add the default nameserver(s) to the list of nameservers. A default nameserver is assumed not to support TLS.

The option conn_time_out specifies the time in milliseconds to wait for an answer to a DNS query.

The option enable turns this functionality on or off.

Miscellaneous Improvements in Version 2.0

  • Added the command line options "-v" and "-c". The option "-v" is for printing version info to the console. The option "-c" is for reading a different configuration file, other than the default "cl_config.ini".
  • Various code optimizations and minor improvements
  • For every read domain (not with a subdomain), Crawlector will prepend "www." to every read site entry, if it doesn't exist already. For example, in the case of RapidAPI subdomain enumeration query, the domain being queried has to start with "www.".
  • Added the options clear_dns and upg_2_https to the [default] section. The former clears the hostname to the IP DNS cache, and the latter upgrades every site to HTTPS by prepending "https://" to it. Similarly, the option tld_upgrd_2_https has been added to the section [site], for upgrading active domains with different TLDs to https.
  • Added the options rapid_api_weeks and rapid_api_limit to the [site] section for configuring the API request to RapidAPI. Both options are optional. The former specifies the number of weeks to query from the DB, while the latter specifies the number of subdomains to return per site.
  • In release 2.0, you can specify a different cache directory for the [spider] and [default] sections, via the option cache_dir.
  • Many options were added to the page section in release 2.0, including:
  • The whois_info option retreives whois domain information including, registrar, registered_on, expires_on and updated_on. This data is pulled from https://www.whois.com/whois/. The data is saved to the cl_mlog CSV file.
  • The page_title option saves the the page title to the cl_mlog CSV file.
  • The option results_dir was added to provide the capability to save pages under a different path. If not set, the folder "results" will be created in the same directory as Crawlector.
  • Upgraded Yara engine to 4.3.2

Design Considerations

  • A URL page is retrieved by sending a GET request to the server, reading the server response body, and passing it to the Yara engine for detection.
  • Some of the GET request attributes are defined in the [default] section in the configuration file, including the User-Agent and Referer headers, and connection timeout, among other options.
  • Although Crawlector logs a session's data to a CSV file, converting it to an SQL file is recommended for better performance, manipulation and retrieval of the data. This becomes evident when you’re crawling thousands of domains.
  • Repeated domains/URLs in the cl_sites are allowed.

Limitations

  • Single-threaded
  • Static detection (no dynamic evaluation of a given page's content). Please check the DOE feature instead.
  • No headless browser support, yet!

Third-party libraries used

  • Chilkat: library for website spidering, HTTP communications, hashing, JSON parsing and file compression (ZIP), among others
  • Yara: for rule scanning (v4.5.4)
  • CrossGuid: for generating GUID/UUID
  • Inih: for parsing configuration file
  • Rapidcsv: for parsing CSV files
  • Color Console: for console coloring
  • TLSH (Trend Micro Locality Sensitive Hash) (v4.8.2)

Contributing

Open for pull requests and issues. Comments and suggestions are greatly appreciated.

Author

Mohamad Mokbel (@MFMokbel)

도구 다운로드
id_postfix (유형)설명
_t1_pID가 없는 일반 유형 1
_sd서브도메인에 대한 하위 유형
_tldTLD에 대한 하위 유형
_t2_pID가 있는 일반 유형 2
_t3_s스파이더링된 도메인 유형 3
_t3_sc하위 노드가 있는 스파이더링된 도메인 유형 3
_t3_ss유형 3 (_t3_s) URL이 유형 1 URL로 전환된 경우 유형 3
_t3_s_e스파이더링된 도메인 유형 3의 외부 링크
_obj_심층 스캔 및 객체 추출용
_t4_ru리디렉션 URL용 (모든 유형)
  • tld_use_connect는 옵션 tlds_connect_ports에 정의된 포트 목록을 통해 해당 도메인에 연결하는 기능을 활성화합니다.
  • 옵션 tlds_connect_ports는 쉼표로 구분된 포트 목록 또는 25-40,90-100,80,443,8443과 같은 범위 목록(범위 시작과 끝 포함)을 허용합니다.
    • tld_con_time_out은 연결 함수의 최대 시간 제한을 설정합니다.
  • tld_con_use_ssl은 도메인에 연결할 때 SSL 사용을 활성화/비활성화합니다.
  • save_to_file_subd가 true로 설정되면 발견된 서브도메인이 "\expanded\exp_subdomain_<pm|am>.txt"에 저장됩니다.
  • save_to_file_tld가 true로 설정되면 발견된 도메인이 "\expanded\exp_tld_<pm|am>.txt"에 저장됩니다.
  • exit_here가 true로 설정되면 Crawlector는 이 [site] 함수를 실행한 후 다른 활성화된 옵션에 관계없이 종료됩니다. 즉, 발견된 사이트는 크롤링/스파이더링되지 않습니다.
  • Control CommandDescription
    cl_get_dateRetrieves the date and time Crawlector was started and the current date and time.
    cl_pingSends back the message "Pong...". This is to check that the C&C channel is working.
    cl_get_configUploads the currently used configuration (e.g., cl_config.ini) file as a text file.
    cl_update_delay <integer_in_milliseconds>Updates the check-in time between every pull request for control commands.

    - Changes the value (ctrl_sleep) for the current session only.

    cl_turn_off_slack_alertTurns off Slack alert feature for the currently active session.
    cl_turn_on_slack_alertTurns on Slack alert feature for the currently active session
    cl_helpLists this help message.
    cl_exitTerminates Crawlector, forcefully.