
Reverse-Engineering-Framework in Python
Miasm ist ein kostenloses und quelloffenes (GPLv2) Reverse-Engineering-Framework. Miasm zielt darauf ab, Binärprogramme zu analysieren / zu modifizieren / zu generieren. Hier ist eine unvollständige Liste der Funktionen:
Siehe den offiziellen Blog für weitere Beispiele und Demos.
Importiere die Miasm x86-Architektur:```pycon
from miasm.arch.x86.arch import mn_x86 from miasm.core.locationdb import LocationDB
Standort-Datenbank abrufen:```pycon
>>> loc_db = LocationDB()
Stelle eine Zeile zusammen:```pycon
l = mn_x86.fromstring('XOR ECX, ECX', loc_db, 32) print(l) XOR ECX, ECX mn_x86.asm(l) ['1\xc9', '3\xc9', 'g1\xc9', 'g3\xc9']
Operand ändern:```pycon
>>> l.args[0] = mn_x86.regs.EAX
>>> print(l)
XOR EAX, ECX
>>> a = mn_x86.asm(l)
>>> print(a)
['1\xc8', '3\xc1', 'g1\xc8', 'g3\xc1']
Disassemblieren Sie das Ergebnis:```pycon
print(mn_x86.dis(a[0], 32)) XOR EAX, ECX
Verwendung der `Machine`-Abstraktion:```pycon
>>> from miasm.analysis.machine import Machine
>>> mn = Machine('x86_32').mn
>>> print(mn.dis('\x33\x30', 32))
XOR ESI, DWORD PTR [EAX]
Für MIPS:```pycon
mn = Machine('mips32b').mn print(mn.dis(b'\x97\xa3\x00 ', "b")) LHU V1, 0x20(SP)
Zwischendarstellung
---------------------------
Erstelle eine Anweisung:```pycon
>>> machine = Machine('arml')
>>> instr = machine.mn.dis('\x00 \x88\xe0', 'l')
>>> print(instr)
ADD R2, R8, R0
Erstellen Sie ein Objekt der Zwischendarstellung:```pycon
lifter = machine.lifter_model_call(loc_db)
Erstelle eine leere ircfg:```pycon
>>> ircfg = lifter.new_ircfg()
Anweisung zum Pool hinzufügen:```pycon
lifter.add_instr_to_ircfg(instr, ircfg)
Aktuellen Pool ausgeben:```pycon
>>> for lbl, irblock in ircfg.blocks.items():
... print(irblock)
loc_0:
R2 = R8 + R0
IRDst = loc_4
Arbeiten mit IR, zum Beispiel durch das Erhalten von Nebeneffekten:```pycon
for lbl, irblock in ircfg.blocks.items(): ... for assignblk in irblock: ... rw = assignblk.get_rw() ... for dst, reads in rw.items(): ... print('read: ', [str(x) for x in reads]) ... print('written:', dst) ... print() ... read: ['R8', 'R0'] written: R2
read: [] written: IRDst
Weitere Informationen zur Miasm IR finden Sie im [zugehörigen Jupyter Notebook](https://github.com/cea-sec/miasm/blob/master/doc/expression/expression.ipynb).
Emulation
---------
Angenommen, ein Shellcode:```pycon
00000000 8d4904 lea ecx, [ecx+0x4]
00000003 8d5b01 lea ebx, [ebx+0x1]
00000006 80f901 cmp cl, 0x1
00000009 7405 jz 0x10
0000000b 8d5bff lea ebx, [ebx-1]
0000000e eb03 jmp 0x13
00000010 8d5b01 lea ebx, [ebx+0x1]
00000013 89d8 mov eax, ebx
00000015 c3 ret
>>> s = b'\x8dI\x04\x8d[\x01\x80\xf9\x01t\x05\x8d[\xff\xeb\x03\x8d[\x01\x89\xd8\xc3'
Importiere den Shellcode dank der Container-Abstraktion:```pycon
from miasm.analysis.binary import Container c = Container.from_string(s, loc_db) c <miasm.analysis.binary.ContainerUnknown object at 0x7f34cefe6090>
Disassemblieren des Shellcodes an Adresse `0`:```pycon
>>> from miasm.analysis.machine import Machine
>>> machine = Machine('x86_32')
>>> mdis = machine.dis_engine(c.bin_stream, loc_db=loc_db)
>>> asmcfg = mdis.dis_multiblock(0)
>>> for block in asmcfg.blocks:
... print(block)
...
loc_0
LEA ECX, DWORD PTR [ECX + 0x4]
LEA EBX, DWORD PTR [EBX + 0x1]
CMP CL, 0x1
JZ loc_10
-> c_next:loc_b c_to:loc_10
loc_10
LEA EBX, DWORD PTR [EBX + 0x1]
-> c_next:loc_13
loc_b
LEA EBX, DWORD PTR [EBX + 0xFFFFFFFF]
JMP loc_13
-> c_to:loc_13
loc_13
MOV EAX, EBX
RET
Initialisieren der JIT-Engine mit einem Stack:```pycon
jitter = machine.jitter(loc_db, jit_type='python') jitter.init_stack()
Fügen Sie den Shellcode an einer beliebigen Speicherstelle hinzu:```pycon
>>> run_addr = 0x40000000
>>> from miasm.jitter.csts import PAGE_READ, PAGE_WRITE
>>> jitter.vm.add_memory_page(run_addr, PAGE_READ | PAGE_WRITE, s)
Erstellen Sie einen Wächter, um die Rückkehr des Shellcodes abzufangen:```Python def code_sentinelle(jitter): jitter.running = False jitter.pc = 0 return True
jitter.add_breakpoint(0x1337beef, code_sentinelle) jitter.push_uint32_t(0x1337beef)
Aktive Logs:```pycon
>>> jitter.set_trace_log()
An beliebiger Adresse ausführen:```pycon
jitter.init_run(run_addr) jitter.continue_run() RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000000 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 RIP 0000000040000000 40000000 LEA ECX, DWORD PTR [ECX+0x4] RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000004 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 .... 4000000e JMP loc_0000000040000013:0x40000013 RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000004 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 RIP 0000000040000013 40000013 MOV EAX, EBX RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000004 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 RIP 0000000040000013 40000015 RET
Interaktion mit dem Jitter:```pycon
>>> jitter.vm
ad 1230000 size 10000 RW_ hpad 0x2854b40
ad 40000000 size 16 RW_ hpad 0x25e0ed0
>>> hex(jitter.cpu.EAX)
'0x0L'
>>> jitter.cpu.ESI = 12
Initialisieren des IR-Pools:```pycon
lifter = machine.lifter_model_call(loc_db) ircfg = lifter.new_ircfg_from_asmcfg(asmcfg)
Initialisierung der Engine mit standardmäßigen symbolischen Werten:```pycon
>>> from miasm.ir.symbexec import SymbolicExecutionEngine
>>> sb = SymbolicExecutionEngine(lifter)
Starten der Ausführung:```pycon
symbolic_pc = sb.run_at(ircfg, 0) print(symbolic_pc) ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10)
Gleich, mit Schrittprotokollen (nur Änderungen werden angezeigt):```pycon
>>> sb = SymbolicExecutionEngine(lifter, machine.mn.regs.regs_init)
>>> symbolic_pc = sb.run_at(ircfg, 0, step=True)
Instr LEA ECX, DWORD PTR [ECX + 0x4]
Assignblk:
ECX = ECX + 0x4
________________________________________________________________________________
ECX = ECX + 0x4
________________________________________________________________________________
Instr LEA EBX, DWORD PTR [EBX + 0x1]
Assignblk:
EBX = EBX + 0x1
________________________________________________________________________________
EBX = EBX + 0x1
ECX = ECX + 0x4
________________________________________________________________________________
Instr CMP CL, 0x1
Assignblk:
zf = (ECX[0:8] + -0x1)?(0x0,0x1)
nf = (ECX[0:8] + -0x1)[7:8]
pf = parity((ECX[0:8] + -0x1) & 0xFF)
of = ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1))[7:8]
cf = (((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1)) ^ ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1)))[7:8]
af = ((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1))[4:5]
________________________________________________________________________________
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5]
pf = parity((ECX + 0x4)[0:8] + 0xFF)
zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1)
ECX = ECX + 0x4
of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8]
nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8]
cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8]
EBX = EBX + 0x1
________________________________________________________________________________
Instr JZ loc_key_1
Assignblk:
IRDst = zf?(loc_key_1,loc_key_2)
EIP = zf?(loc_key_1,loc_key_2)
________________________________________________________________________________
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5]
EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10)
pf = parity((ECX + 0x4)[0:8] + 0xFF)
IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10)
zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1)
ECX = ECX + 0x4
of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8]
nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8]
cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8]
EBX = EBX + 0x1
________________________________________________________________________________
>>>
Wiederholen Sie die Ausführung mit einem konkreten ECX. Hier erreicht die symbolic / concolic execution das Ende des Shellcodes:```pycon
from miasm.expression.expression import ExprInt sb.symbols[machine.mn.regs.ECX] = ExprInt(-3, 32) symbolic_pc = sb.run_at(ircfg, 0, step=True) Instr LEA ECX, DWORD PTR [ECX + 0x4] Assignblk: ECX = ECX + 0x4
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5] EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) pf = parity((ECX + 0x4)[0:8] + 0xFF) IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1) ECX = 0x1 of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8] nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8] cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8] EBX = EBX + 0x1
Instr LEA EBX, DWORD PTR [EBX + 0x1] Assignblk: EBX = EBX + 0x1
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5] EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) pf = parity((ECX + 0x4)[0:8] + 0xFF) IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1) ECX = 0x1 of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8] nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8] cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8] EBX = EBX + 0x2
Instr CMP CL, 0x1 Assignblk: zf = (ECX[0:8] + -0x1)?(0x0,0x1) nf = (ECX[0:8] + -0x1)[7:8] pf = parity((ECX[0:8] + -0x1) & 0xFF) of = ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1))[7:8] cf = (((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1)) ^ ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1)))[7:8] af = ((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1))[4:5]
af = 0x0 EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) pf = 0x1 IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x2
Instr JZ loc_key_1 Assignblk: IRDst = zf?(loc_key_1,loc_key_2) EIP = zf?(loc_key_1,loc_key_2)
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x10 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x2
Instr LEA EBX, DWORD PTR [EBX + 0x1] Assignblk: EBX = EBX + 0x1
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x10 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3
Instr LEA EBX, DWORD PTR [EBX + 0x1] Assignblk: IRDst = loc_key_3
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x13 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3
Instr MOV EAX, EBX Assignblk: EAX = EBX
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x13 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3 EAX = EBX + 0x3
Instr RET Assignblk: IRDst = @32[ESP[0:32]] ESP = {ESP[0:32] + 0x4 0 32} EIP = @32[ESP[0:32]]
af = 0x0 EIP = @32[ESP] pf = 0x1 IRDst = @32[ESP] zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3 ESP = ESP + 0x4 EAX = EBX + 0x3
Wie funktioniert es?
=================
Miasm enthält einen eigenen Disassembler, eine Zwischensprache und Anweisungssemantik. Es ist in Python geschrieben.
Um Code zu emulieren, verwendet es LLVM, GCC, Clang oder Python, um die Zwischendarstellung per JIT zu kompilieren. Es kann Shellcodes und ganze oder Teile von Binärdateien emulieren. Python-Callbacks können ausgeführt werden, um mit der Ausführung zu interagieren, zum Beispiel um Effekte von Bibliotheksfunktionen zu emulieren.
Dokumentation
=============
Einige Dokumentationsressourcen sind im Ordner [doc](https://github.com/cea-sec/miasm/blob/HEAD/doc) verfügbar.
Eine automatisch generierte Dokumentation ist verfügbar:
* [Doxygen](http://miasm.re/miasm_doxygen)
* [pdoc](http://miasm.re/miasm_pdoc)
Miasm beziehen
===============
* Repository klonen: [Miasm auf GitHub](https://github.com/cea-sec/miasm/)
* Eines der Docker-Images unter [Docker Hub](https://registry.hub.docker.com/u/miasm/) beziehen
Softwareanforderungen
---------------------
Miasm verwendet:
* python-pyparsing
* python-dev
* optional python-pycparser (Version >= 2.17)
Um Code-JIT zu aktivieren, ist eines der folgenden Module erforderlich:
* GCC
* Clang
* LLVM mit Numba llvmlite, siehe unten
'Optional' kann Miasm auch verwenden:
* Z3, der [Theorembeweiser](https://github.com/Z3Prover/z3)
Konfiguration
-------------
Um den Jitter zu verwenden, wird GCC oder LLVM empfohlen
* GCC (beliebige Version)
* Clang (beliebige Version)
* LLVM
* Debian (testing/unstable): Nicht getestet
* Debian stable/Ubuntu/Kali/whatever: `pip install llvmlite` oder Installation von [llvmlite](https://github.com/numba/llvmlite)
* Windows: Nicht getestet
* Miasm bauen und installieren:```pycon
$ cd miasm_directory
$ python setup.py build
$ sudo python setup.py install
Wenn während der Kompilierung der Jitter-Module etwas schiefgeht, überspringt Miasm den Fehler und deaktiviert das entsprechende Modul (siehe Kompilierungsausgabe).
Die meisten IDA-Plugins von Miasm verwenden eine Teilmenge der Miasm-Funktionalität. Eine schnelle Möglichkeit, sie zum Laufen zu bringen, ist das Hinzufügen von:
pyparsing.py nach C:\...\IDA\python\ oder pip install pyparsingmiasm/miasm-Verzeichnis nach C:\...\IDA\python\Alle Funktionen außer denen, die mit dem JITter zusammenhängen, sind verfügbar. Für eine vollständige Installation siehe die obigen Absätze.
Miasm wird mit einer Reihe von Regressionstests geliefert. Um alle auszuführen:```pycon cd miasm_directory/test
python test_all.py
python -m unittest test_all.py # sequential, requires 'unittest' python -m pytest test_all.py # sequential, requires 'pytest' python -m pytest -n auto test_all.py # parallel, requires 'pytest' and 'pytest-xdist'
Einige Optionen können angegeben werden:
* Mono-Threading: `-m`
* Codeabdeckungsinstrumentierung: `-c`
* Nur schnelle Tests: `-t long` (schließt die langen Tests aus)
Sie verwenden bereits Miasm
===========================
Werkzeuge
---------
* [Sibyl](https://github.com/cea-sec/Sibyl): Ein Tool zur Funktionsdivination
* [R2M2](https://github.com/guedou/r2m2): Verwenden Sie miasm als radare2-Plugin
* [CGrex](https://github.com/mechaphish/cgrex): Gezielter Patcher für CGC-Binärdateien
* [ethRE](https://github.com/jbcayrou/ethRE): Reverse-Engineering-Tool für Ethereum EVM (mit entsprechender Miasm2-Architektur)
Blogbeiträge / Artikel / Konferenzen
-------------------------------------
* [Deobfuscation: recovering an OLLVM-protected program](http://blog.quarkslab.com/deobfuscation-recovering-an-ollvm-protected-program.html)
* [Taming a Wild Nanomite-protected MIPS Binary With Symbolic Execution: No Such Crackme](https://doar-e.github.io/blog/2014/10/11/taiming-a-wild-nanomite-protected-mips-binary-with-symbolic-execution-no-such-crackme/)
* [Génération rapide de DGA avec Miasm](https://www.lexsi.com/securityhub/generation-rapide-de-dga-avec-miasm/): Schnelle Berechnung von DGA (französischer Artikel)
* [Enabling Client-Side Crash-Resistance to Overcome Diversification and Information Hiding](https://www.internetsociety.org/sites/default/files/blogs-media/enabling-client-side-crash-resistance-overcome-diversification-information-hiding.pdf): Erkennen potenzieller Argumente bei undirekten Aufrufen
* [Miasm: Framework de reverse engineering](https://www.sstic.org/2012/presentation/miasm_framework_de_reverse_engineering/) (Französisch)
* [Tutorial miasm](https://www.sstic.org/2014/presentation/Tutorial_miasm/) (Französisches Video)
* [Graphes de dépendances : Petit Poucet style](https://www.sstic.org/2016/presentation/graphes_de_dpendances__petit_poucet_style/): DepGraph (Französisch)
Bücher
------
* [Practical Reverse Engineering: X86, X64, Arm, Windows Kernel, Reversing Tools, and Obfuscation](http://eu.wiley.com/WileyCDA/WileyTitle/productCd-1118787315,subjectCd-CSJ0.html): Einführung in Miasm (Kapitel 5 "Obfuscation")
* [BlackHat Python - Appendix](https://github.com/oreilly-japan/black-hat-python-jp-support/tree/master/appendix-A): Beispiele des japanischen Sicherheitsbuchs