
Hashcat 규칙 분석기 및 인터프리터 - hashcat 규칙 구문을 해독하기 위한 로제타석
| | | | __ _ | | ___ __ | || _ \ ___ ___ | || | __ _
| || |/ / __| '_ \ / __/ _ | __| |) / _ / __|/ _ \ __| / ` |
| _ | (| _ \ | | | (| (| | || _ < () __ \ / || || (| |
|| ||_,|/| ||____,|_|| __/|/___|_|__,|
비밀번호 크래킹 규칙의 로제타 스톤을 해독하세요
# HashcatRosetta
hashcat 디버그 모드 4 및 모드 5 출력 파일을 분석하여 가장 효율적인 규칙을 식별하고 비밀번호 크래킹 공격 중 사용된 기본 단어 빈도 패턴을 추적하도록 설계된 Python 프로젝트입니다.
## 기능
- **hashcat 디버그 파일 파싱** (`--debug-mode 4` 및 `--debug-mode 5`) - 기본 단어 및 규칙 자동 추출
- **후보를 소스 단어 목록에 귀속** (모드 5) - 단어 목록별 통계
- **여러 지표로 규칙 효율성 추적**:
- 적용 빈도 (가장 일반적으로 적용되는 규칙)
- 기본 단어 분포 (가장 많은 고유 기본 단어에 적용된 규칙)
- 후보 생성 (가장 많은 고유 후보를 생성하는 규칙)
- **기본 단어 패턴 모니터링** - 상세 발생 로그 및 통계 포함
- **상세 보고서 생성** - 규칙 및 기본 단어 분석 포함
- **분석 내보내기** - 추가 처리를 위해 JSON 또는 CSV 형식으로 내보내기
- **명령줄 인터페이스** - 쉬운 분석 및 보고
## 설치
### uv 사용 (권장)
[uv](https://github.com/astral-sh/uv)를 사용하는 경우 설치 없이 실행할 수 있습니다:
```bash
# 저장소 복제
git clone https://github.com/bandrel/HashcatRosetta.git
cd HashcatRosetta
# 모듈로 실행 (권장)
uv run python -m hashcat_rosetta --help
# 또는 설치된 명령 사용
uv run hashcat-rosetta --help
uv tool install git+https://github.com/bandrel/HashcatRosetta.git
개발 도구는 dev 의존성 그룹에 있습니다.
# uv 사용 (기본적으로 dev 그룹 설치)
uv sync
# pip 사용 (25.1+)
pip install -e . --group dev
hashcat 디버그 파일 분석 (기본적으로 요약 표시):
hashcat-rosetta debug_output.txt
빈도별 상위 규칙 표시:
hashcat-rosetta debug_output.txt --rules --top 10 --metric frequency
다른 지표별 상위 규칙 표시:
hashcat-rosetta debug_output.txt --rules --metric basewords
hashcat-rosetta debug_output.txt --rules --metric candidates
여러 번 나타나는 기본 단어 표시:
hashcat-rosetta debug_output.txt --basewords --top 10
상위 단어 목록 표시 (디버그 모드 5 전용):
hashcat-rosetta debug_output.txt --wordlists --top 10
상세한 단어 목록별 통계 표시 (고유 기본 단어, 후보, 규칙):
hashcat-rosetta debug_output.txt --wordlists --top 10 --detail
--wordlists 출력은 --rules와 유사합니다: Top N Wordlists 헤더 다음에
번호가 매겨진 Wordlist: <name> (<count>) 줄이 이어집니다. 모드 5 파일을
출력 플래그 없이 분석하면 기본 요약에 Wordlist Statistics 섹션도 포함됩니다.
자동 감지 대신 특정 디버그 모드를 강제 지정:
hashcat-rosetta debug_output.txt --debug-mode 5 --wordlists
상세한 기본 단어 분석 표시:
hashcat-rosetta debug_output.txt --basewords --top 10 --detail --min-occurrences 2
전체 분석 보고서 내보내기:
hashcat-rosetta debug_output.txt --export report.json --format json
hashcat-rosetta debug_output.txt --export report.csv --format csv
hashcat 규칙이 수행하는 작업을 단계별로 설명:
hashcat-rosetta --explain "c$1" --baseword admin
hashcat-rosetta --explain "u$!" --baseword myword
로컬 LLM을 사용하여 영어 설명에서 hashcat 마스크 생성:
hashcat-rosetta --mask "The word 'Summer' followed by six digits."
출력:
Mask Suggestions for: 'The word 'Summer' followed by six digits.'
======================================================================
1. Summer?d?d?d?d?d?d
literal "Summer", then 6 × digit → 1,000,000 candidates
Why: matches the literal word followed by a 6-digit number
생성된 마스크를 파일에 저장:
hashcat-rosetta --mask "The word 'Summer' followed by six digits." -o masks.hcmask
다른 설명에서 마스크 생성:
hashcat-rosetta --mask "a capitalized season, two digits, and a special char"
hashcat-rosetta --mask "year 2020-2025 followed by exclamation or question mark"
마스크 생성 기능은 OpenAI 호환 채팅 엔드포인트를 실행하는 로컬 Ollama 서버를
사용합니다. 기본적으로 http://localhost:11434에 연결하고 gemma3:27b 모델을
사용합니다 (아래 참조). 환경 변수 또는 CLI 플래그로 구성할 수 있습니다:
# 환경 변수 사용
OLLAMA_HOST=http://192.168.1.100:11434 OLLAMA_MODEL=llama2:70b \
hashcat-rosetta --mask "your description here"
# CLI 플래그 사용 (환경 변수 재정의)
hashcat-rosetta --mask "your description" --ollama-host http://custom.host:11434 --model llama2
보안 참고: 마스크 설명은 구성한 Ollama 엔드포인트(기본적으로 localhost,
또는 --ollama-host/OLLAMA_HOST가 가리키는 곳)로만 전송되며, 클라우드
제공자에게는 절대 전송되지 않습니다. OpenAI SDK는 해당 엔드포인트에 대한 HTTP
클라이언트로만 사용되며, 데이터나 API 키는 api.openai.com으로 전송되지
않습니다.
gemma3:27b인가?기본값은 scripts/benchmark_mask_models.py에 의해 선택되며, 이 스크립트는
14개의 고정된 --mask 스타일 프롬프트(사용자 정의 문자셋 역참조 및 카테고리
회상 프롬프트(성경 책, 성경 구절 참조, 유럽 수도) 포함)를 로컬에 설치된 모든
후보 모델에 대해 실행하고 각 응답을 세 가지 방식으로 평가합니다:
mp64(maskprocessor)가 설치된 경우 이를 통해 독립적으로 검증됩니다. 이는
프로덕션에서 generate_masks() 자체가 모든 제안에 대해 실행하는 것과 동일한
검사이므로, 여기서 벤치마크 하드 실패가 발생하면 실제 --mask 사용에서도
거부되었을 것입니다.gemma3:12b, 자기 평가
편향을 피하기 위해 후보가 아닌 모델로 특별히 선택됨)이 모든 응답을 원래
요청을 얼마나 잘 충족하는지 1-5점으로 평가합니다. 요청은 사고 기능이
활성화된 상태로 전송됩니다. 느린 모델은 어차피 왕복 시간이 발생하기
때문입니다.권장 기본값은 하드 실패가 0이고 평균 심사 점수가 ≥ 4인 가장 작은 모델입니다. 이전 라운드에서 이월된 세 가지 최종 후보를 전체 14개 프롬프트 세트에 대해 다시 실행한 결과:
| model | size | hard fails | mean score | time |
|---|---|---|---|---|
gemma3:27b | 16.2 GB | 0 | 4.1 | 180s |
dengcao/Qwen3-30B-A3B-Instruct-2507:latest | 17.4 GB | 1 | 4.5 | 176s |
laguna-xs-2.1:latest | 18.9 GB | 2 | 4.7 | 730s |
gemma3:27b는 세 모델 중 유일하게 하드 실패가 0이므로, 가장 높은 원점수를
받지 못했음에도 선택되었습니다. dengcao와 laguna-xs-2.1:latest는 더 높은
점수를 받았지만 각각 최소 하나의 프롬프트에서 완전히 실패했으며
(laguna-xs-2.1:latest는 또한 훨씬 느린 730s 대 ~180s로, 자체적으로 매우 큰
네이티브 컨텍스트 윈도우 때문입니다).
이전 스윕 라운드의 여러 Qwen3 계열 모델(qwen3:8b/30b/32b, qwen3.5:9b/27b)은
뚜렷한 실패 모드를 보였습니다: 숨겨진 "사고" 토큰과 거대한 네이티브 컨텍스트
윈도우가 --mask 규모 하드웨어에서 수 분에서 30분까지 멈추는 현상을
유발했습니다. 이전 기본값인 qwen3.6:35b-a3b도 같은 이유로 교체되었습니다
(CHANGELOG.md 참조).
uv run python scripts/benchmark_mask_models.py로 직접 스윕을 다시 실행할 수
있습니다. 누락된 후보를 가져오고 업데이트된 권장 사항을 출력합니다.
from hashcat_rosetta import DebugAnalyzer
analyzer = DebugAnalyzer()
# Analyze a debug file
result = analyzer.analyze_debug_file('debug_output.txt')
print(f"Total entries: {result['total_entries']}")
print(f"Unique rules: {result['unique_rules']}")
print(f"Unique basewords: {result['unique_basewords']}")
# Get top rules by frequency
top_rules = analyzer.get_top_rules_by_frequency(10)
for rule, count in top_rules:
print(f"Rule: {rule}, Applications: {count}")
# Get top basewords
top_basewords = analyzer.get_top_basewords_by_frequency(10)
for baseword, count in top_basewords:
print(f"Baseword: {baseword}, Occurrences: {count}")
# Get basewords appearing multiple times
frequent_basewords = analyzer.get_basewords_with_min_occurrences(2)
print(f"Basewords appearing 2+ times: {len(frequent_basewords)}")
# Get detailed information about a specific baseword
detail = analyzer.get_baseword_detail('password')
print(f"Rules applied to 'password': {detail['unique_rules']}")
print(f"Occurrences: {len(detail['occurrences'])}")
# Export complete analysis
export = analyzer.export_to_dict()
분석기는 두 가지 hashcat 디버그 출력 형식을 자동으로 감지하고 지원합니다:
baseword:rule:candidate
COMPUTER:} } } } t:retupmoc
EXAMPLE:sa@ se3 so0:3x@mpl3
admin:$1 $5 c ^@:@Admin15
각 줄은 세 개의 콜론으로 구분된 필드를 포함합니다:
hashcat은 항상 이 형식을 출력해 왔습니다 (src/debugfile.c가 orig, :, rule, :, mod를 기록).
baseword rule candidate
password c P@ssword
password u PASSWORD
admin l admin
letmein [ etmein
각 줄은 위와 동일한 의미를 가진 세 개의 공백으로 구분된 필드를 포함합니다. 이것은 이 파서도 허용하는 오래된 레거시 형식입니다.
참고: 분석기는 파일이 사용하는 형식을 자동으로 감지합니다. 수동 구성이 필요하지 않습니다!
hashcat --debug-mode 5는 각 콜론 구분 줄에 후행 wordlist 필드를 추가합니다: