| | | | __ _ | | ___ __ | || _ \ ___ ___ | || | __ _
| || |/ / __| '_ \ / __/ _ | __| |) / _ / __|/ _ \ __| / ` |
| _ | (| _ \ | | | (| (| | || _ < () __ \ / || || (| |
|| ||_,|/| ||____,|_|| __/|/___|_|__,|
解码密码破解规则的罗塞塔石碑
# HashcatRosetta
一个 Python 项目,旨在分析 hashcat 调试模式 4 和模式 5 的输出文件,以识别最高效的规则并跟踪密码破解攻击中使用的基词频率模式。
## 功能特性
- **解析 hashcat 调试文件**(`--debug-mode 4` 和 `--debug-mode 5`),自动提取基词和规则
- **将候选密码归因到源词表**(模式 5),以获取每个词表的统计数据
- **通过多种指标跟踪规则效率**:
- 应用频率(最常被应用的规则)
- 基词覆盖范围(应用于最多不同基词的规则)
- 候选密码生成量(生成最多不同候选密码的规则)
- **监控基词模式**,提供详细的出现日志和统计数据
- **生成详细报告**,包含规则和基词分析
- **导出分析结果**为 JSON 或 CSV 格式,以便进一步处理
- **命令行界面**,便于分析和报告
## 安装
### 使用 uv(推荐)
如果你使用 [uv](https://github.com/astral-sh/uv),无需安装即可运行:
```bash
# 克隆仓库
git clone https://github.com/bandrel/HashcatRosetta.git
cd HashcatRosetta
# 作为模块运行(推荐)
uv run python -m hashcat_rosetta --help
# 或使用已安装的命令
uv run hashcat-rosetta --help
uv tool install git+https://github.com/bandrel/HashcatRosetta.git
开发工具位于 dev 依赖组 中。
# 使用 uv(默认安装 dev 组)
uv sync
# 使用 pip(25.1+)
pip install -e . --group dev
分析 hashcat 调试文件(默认显示摘要):
hashcat-rosetta debug_output.txt
按频率显示排名靠前的规则:
hashcat-rosetta debug_output.txt --rules --top 10 --metric frequency
按其他指标显示排名靠前的规则:
hashcat-rosetta debug_output.txt --rules --metric basewords
hashcat-rosetta debug_output.txt --rules --metric candidates
显示多次出现的基词:
hashcat-rosetta debug_output.txt --basewords --top 10
显示排名靠前的词表(仅调试模式 5):
hashcat-rosetta debug_output.txt --wordlists --top 10
显示每个词表的详细统计信息(唯一基词、候选密码和规则):
hashcat-rosetta debug_output.txt --wordlists --top 10 --detail
--wordlists 的输出与 --rules 类似:先是 Top N Wordlists 标题,随后是
带编号的 Wordlist: <name> (<count>) 行。当分析模式 5 文件且未指定任何输出标志时,
默认摘要还会包含一个 词表统计 部分。
强制指定特定的调试模式,而不是自动检测:
hashcat-rosetta debug_output.txt --debug-mode 5 --wordlists
显示详细的基词分析:
hashcat-rosetta debug_output.txt --basewords --top 10 --detail --min-occurrences 2
导出完整的分析报告:
hashcat-rosetta debug_output.txt --export report.json --format json
hashcat-rosetta debug_output.txt --export report.csv --format csv
逐步解释 hashcat 规则的作用:
hashcat-rosetta --explain "c$1" --baseword admin
hashcat-rosetta --explain "u$!" --baseword myword
使用本地 LLM 从英文描述生成 hashcat 掩码:
hashcat-rosetta --mask "The word 'Summer' followed by six digits."
输出:
Mask Suggestions for: 'The word 'Summer' followed by six digits.'
======================================================================
1. Summer?d?d?d?d?d?d
literal "Summer", then 6 × digit → 1,000,000 candidates
Why: matches the literal word followed by a 6-digit number
将生成的掩码保存到文件:
hashcat-rosetta --mask "The word 'Summer' followed by six digits." -o masks.hcmask
从其他描述生成掩码:
hashcat-rosetta --mask "a capitalized season, two digits, and a special char"
hashcat-rosetta --mask "year 2020-2025 followed by exclamation or question mark"
掩码生成功能使用运行 OpenAI 兼容聊天端点的本地 Ollama 服务器。
默认情况下,它连接到 http://localhost:11434 并使用模型
gemma3:27b(见下文)。这些可以通过环境变量
或 CLI 标志进行配置:
# 使用环境变量
OLLAMA_HOST=http://192.168.1.100:11434 OLLAMA_MODEL=llama2:70b \
hashcat-rosetta --mask "your description here"
# 使用 CLI 标志(覆盖环境变量)
hashcat-rosetta --mask "your description" --ollama-host http://custom.host:11434 --model llama2
安全说明: 掩码描述仅发送到你配置的 Ollama 端点
(默认为 localhost,或 --ollama-host/OLLAMA_HOST 指向的位置)——绝不会发送到
云服务提供商。OpenAI SDK 纯粹作为针对该端点的 HTTP 客户端使用;
任何数据或 API 密钥都不会传输到 api.openai.com。
gemma3:27b?默认值由 scripts/benchmark_mask_models.py 选定,该脚本运行一组固定的
14 个 --mask 风格提示——包括自定义字符集反向引用和类别回忆
提示(圣经书卷、圣经经文引用、欧洲首都城市)——针对每个
本地安装的候选模型,并以三种方式对每个响应进行评分:
mp64(maskprocessor)进行验证(如果已安装)——这与
generate_masks() 本身在生产环境中对每个建议运行的检查相同,因此
此处的基准硬失败也意味着实际的 --mask 使用会拒绝它。gemma3:12b,专门
选择因为它本身不是候选模型——避免自我评分偏差)对
每个响应按 1-5 分评分,评估其满足原始请求的程度。请求发送时
启用了思考功能,因为慢速模型无论如何已经消耗了往返时间。推荐的默认值是零硬失败且平均评判分数 ≥ 4 的最小模型。 从早期轮次中保留下来、针对完整的 14 个提示集重新运行的三个决赛选手:
| 模型 | 大小 | 硬失败 | 平均分 | 时间 |
|---|---|---|---|---|
gemma3:27b | 16.2 GB | 0 | 4.1 | 180s |
dengcao/Qwen3-30B-A3B-Instruct-2507:latest | 17.4 GB | 1 | 4.5 | 176s |
laguna-xs-2.1:latest | 18.9 GB | 2 | 4.7 | 730s |
gemma3:27b 是三者中唯一零硬失败的,因此尽管原始分数不是最高,
它仍是首选——dengcao 和 laguna-xs-2.1:latest 得分更高,但
各自至少有一个提示完全失败(而且 laguna-xs-2.1:latest 也慢得多,
730s 对约 180s,这是由其自身非常大的原生上下文窗口导致的)。
早期扫描轮次中的几个 Qwen3 系列模型(qwen3:8b/30b/32b、qwen3.5:9b/27b)
表现出一种明显的失败模式:隐藏的“思考”令牌加上巨大的原生上下文窗口
导致在 --mask 规模的硬件上出现数分钟到 30 分钟的挂起。之前的默认值
qwen3.6:35b-a3b 因同样原因被替换(见 CHANGELOG.md)。
你可以使用 uv run python scripts/benchmark_mask_models.py 自行重新运行扫描——它会拉取
任何缺失的候选模型并打印更新的推荐。
from hashcat_rosetta import DebugAnalyzer
analyzer = DebugAnalyzer()
# 分析调试文件
result = analyzer.analyze_debug_file('debug_output.txt')
print(f"Total entries: {result['total_entries']}")
print(f"Unique rules: {result['unique_rules']}")
print(f"Unique basewords: {result['unique_basewords']}")
# 按频率获取排名靠前的规则
top_rules = analyzer.get_top_rules_by_frequency(10)
for rule, count in top_rules:
print(f"Rule: {rule}, Applications: {count}")
# 获取排名靠前的基词
top_basewords = analyzer.get_top_basewords_by_frequency(10)
for baseword, count in top_basewords:
print(f"Baseword: {baseword}, Occurrences: {count}")
# 获取多次出现的基词
frequent_basewords = analyzer.get_basewords_with_min_occurrences(2)
print(f"Basewords appearing 2+ times: {len(frequent_basewords)}")
# 获取特定基词的详细信息
detail = analyzer.get_baseword_detail('password')
print(f"Rules applied to 'password': {detail['unique_rules']}")
print(f"Occurrences: {len(detail['occurrences'])}")
# 导出完整分析
export = analyzer.export_to_dict()
分析器自动检测并支持两种 hashcat 调试输出格式:
baseword:rule:candidate
COMPUTER:} } } } t:retupmoc
EXAMPLE:sa@ se3 so0:3x@mpl3
admin:$1 $5 c ^@:@Admin15
每行包含三个冒号分隔的字段:
hashcat 一直输出这种格式(src/debugfile.c 写入 orig、:、rule、:、mod)。
baseword rule candidate
password c P@ssword
password u PASSWORD
admin l admin
letmein [ etmein
每行包含三个空格分隔的字段,含义与上述相同。这是一种较旧的遗留格式,此解析器也接受。
注意:分析器会自动检测你的文件使用哪种格式。无需手动配置!
Hashcat --debug-mode 5 在每个冒号分隔行末尾添加一个词表字段:
baseword:rule:candidate:wordlist
password:c:P@ssword:/opt/wordlists/rockyou.txt
admin:l:admin:/opt/wordlists/rockyou.txt
letmein:[:etmein:<stdin>
每行包含四个字段:
<stdin>、<generic>、<none>)模式 5 解锁了 --wordlists 输出以及默认摘要中的词表统计部分。
模式 4 分析保持不变。
默认情况下,分析器通过计数字段来自动检测模式。使用 --debug-mode
强制指定解释方式:
hashcat-rosetta debug.txt --debug-mode auto # 默认:根据字段数检测
hashcat-rosetta debug.txt --debug-mode 4 # 强制模式 4
hashcat-rosetta debug.txt --debug-mode 5 # 强制模式 5
--debug-mode 仅适用于调试文件分析(不适用于 --analyze-rules)。
Windows 路径限制:Hashcat 不转义冒号,因此基词、候选密码
和 Windows 词表路径(例如 C:\wordlists\rockyou.txt)可能包含 :。解析器
假设末尾的词表字段不包含冒号,这对 Linux 路径和
哨兵值成立,但对 Windows 驱动器号路径不成立。强制 --debug-mode 5 通过
将候选密码之后的所有内容视为词表字段来缓解此问题。
使用 hashcat 的 --debug-mode 4 或 --debug-mode 5 生成调试输出: