Evidence-focused malware reverse engineering with deep PE/.NET inspection, Ghidra reconstruction, AI cross-checks, YARA, and ELF debugging
Evidence-focused reverse-engineering CLI. Andrey Pautov develops the Python analyst interface, offline triage, string intelligence, PE inspection and reporting workflows. Capstone, Ghidra, GDB and optional AI providers are integrations, not original tools authored by this project.
Role relevance: Malware triage, reverse engineering tooling, safe AI-assisted analysis, Python delivery.
Run the safe local demo and review validation scope · Recorded local validation
AIDebug is an evidence-focused malware reverse-engineering CLI and terminal UI. It combines deterministic offline triage, whole-file hex inspection, deep PE structure analysis, Capstone disassembly, Ghidra reconstruction, optional LLM cross-checks, local ELF debugging, compiled learning exercises, and analyst-review reporting.
Current source version: AIDebug 3.1.0. See the 3.1.0 release notes.
The latest immutable published release remains AIDebug v3.0.0, available as
1200km-aidebug, until the version-matched 3.1.0 tag and GitHub release complete the verified publishing workflow.
Install the stable package from PyPI:
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip
python -m pip install 1200km-aidebug==3.0.0
aidebug --version
Install optional capabilities as needed:
# Remote/local LLM providers and validated YARA generation
python -m pip install "1200km-aidebug[ai]==3.0.0"
# Frida dynamic instrumentation
python -m pip install "1200km-aidebug[dynamic]==3.0.0"
# All optional Python integrations
python -m pip install "1200km-aidebug[all]==3.0.0"
For development:
git clone https://github.com/anpa1200/AIDebug.git
cd AIDebug
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[dev,dynamic]"
Ghidra, GDB, Bubblewrap, a C compiler, and Frida target components are external tools used only by the workflows that require them.
Open a PE or ELF sample in the main terminal interface:
aidebug --binary /path/to/sample.exe --offline
Run deterministic analysis without the full-screen UI and export evidence:
aidebug --binary /path/to/sample.exe \
--offline --no-tui --report --json-export --yara \
--out-dir reports/
Use Ghidra reconstruction:
aidebug --binary /path/to/sample.exe --offline --no-tui --decompile
aidebug --binary /path/to/sample.exe --offline --no-tui \
--decompile-all reports/sample-reconstruction.c
Analyze one C translation unit through a temporary, non-executed ELF artifact:
aidebug --source /path/to/example.c --offline --no-tui
Identify an arbitrary file independently of its filename extension:
aidebug --identify /path/to/renamed-or-unknown-file --offline
--identify reports structured JSON with the declared type, MIME type, common
extensions, confidence, method, evidence, SHA-256, and size. Deterministic
coverage includes common executable and bytecode formats, archives and disk
images, Office/OpenDocument/EPUB containers, documents, images, audio/video,
packet captures, databases, registry/event-log artifacts, scripts, and text.
ZIP-based formats are inspected by bounded member names and small metadata
reads; files are never executed or extracted.
Install python-magic plus the operating system's libmagic database for
additional signatures known to the local platform:
python -m pip install python-magic
When no deterministic signature, structure, or text rule matches, a configured
AI provider may infer a candidate from bounded metadata: the extension, size,
SHA-256, up to 96 header bytes, 32 tail bytes, sample entropy, and NUL ratio.
The file body, extracted strings, and filesystem path are not sent. AI-only
results are labeled ai-inference, capped at 60% confidence, and require
analyst validation. Use --offline to disable the fallback completely; an
unresolved type is reported as Unknown with exit status 2.
Press S in the main terminal interface, or start directly in the workspace:
aidebug --binary /path/to/sample.exe --offline --strings
The workspace preserves file offsets, mapped addresses when available, encoding, byte and character lengths, duplicate-occurrence information, section context, confidence, triage score, and the deterministic reasons for each classification. Filters cover minimum length, encoding, category, and free-text search; column sorting and pagination keep large inventories usable. Every selected encoding scans the complete size-bounded artifact. The retained inventory is capped at 25,000 records and 4,096 displayed characters per value; exact candidate/omission counts and full-byte coverage make either cap visible. Each record retains at most 32 DLL/API annotations and 4,096 description characters; adversarial overflows are reported in the record reasons.