
Marco de ingeniería inversa en Python
Miasm es un framework de ingeniería inversa gratuito y de código abierto (GPLv2). Miasm tiene como objetivo analizar / modificar / generar programas binarios. Aquí hay una lista no exhaustiva de características:
Consulte el blog oficial para más ejemplos y demostraciones.
Importar arquitectura x86 de Miasm:```pycon
from miasm.arch.x86.arch import mn_x86 from miasm.core.locationdb import LocationDB
Obtén una base de datos de ubicación:```pycon
>>> loc_db = LocationDB()
Montar una línea:```pycon
l = mn_x86.fromstring('XOR ECX, ECX', loc_db, 32) print(l) XOR ECX, ECX mn_x86.asm(l) ['1\xc9', '3\xc9', 'g1\xc9', 'g3\xc9']
Modificar un operando:```pycon
>>> l.args[0] = mn_x86.regs.EAX
>>> print(l)
XOR EAX, ECX
>>> a = mn_x86.asm(l)
>>> print(a)
['1\xc8', '3\xc1', 'g1\xc8', 'g3\xc1']
Desensamble el resultado:```pycon
print(mn_x86.dis(a[0], 32)) XOR EAX, ECX
Usando la abstracción `Machine`:```pycon
>>> from miasm.analysis.machine import Machine
>>> mn = Machine('x86_32').mn
>>> print(mn.dis('\x33\x30', 32))
XOR ESI, DWORD PTR [EAX]
Para MIPS:```pycon
mn = Machine('mips32b').mn print(mn.dis(b'\x97\xa3\x00 ', "b")) LHU V1, 0x20(SP)
Representación intermedia
---------------------------
Crear una instrucción:```pycon
>>> machine = Machine('arml')
>>> instr = machine.mn.dis('\x00 \x88\xe0', 'l')
>>> print(instr)
ADD R2, R8, R0
Crear un objeto de representación intermedia:```pycon
lifter = machine.lifter_model_call(loc_db)
Crear un ircfg vacío:```pycon
>>> ircfg = lifter.new_ircfg()
Añadir instrucción al pool:```pycon
lifter.add_instr_to_ircfg(instr, ircfg)
Imprimir pool actual:```pycon
>>> for lbl, irblock in ircfg.blocks.items():
... print(irblock)
loc_0:
R2 = R8 + R0
IRDst = loc_4
Trabajando con IR, por ejemplo, obteniendo efectos secundarios:```pycon
for lbl, irblock in ircfg.blocks.items(): ... for assignblk in irblock: ... rw = assignblk.get_rw() ... for dst, reads in rw.items(): ... print('read: ', [str(x) for x in reads]) ... print('written:', dst) ... print() ... read: ['R8', 'R0'] written: R2
read: [] written: IRDst
Más información sobre el IR de Miasm está en el [Jupyter Notebook correspondiente](https://github.com/cea-sec/miasm/blob/master/doc/expression/expression.ipynb).
Emulación
---------
Dado un shellcode:```pycon
00000000 8d4904 lea ecx, [ecx+0x4]
00000003 8d5b01 lea ebx, [ebx+0x1]
00000006 80f901 cmp cl, 0x1
00000009 7405 jz 0x10
0000000b 8d5bff lea ebx, [ebx-1]
0000000e eb03 jmp 0x13
00000010 8d5b01 lea ebx, [ebx+0x1]
00000013 89d8 mov eax, ebx
00000015 c3 ret
>>> s = b'\x8dI\x04\x8d[\x01\x80\xf9\x01t\x05\x8d[\xff\xeb\x03\x8d[\x01\x89\xd8\xc3'
Importe el shellcode gracias a la abstracción Container:```pycon
from miasm.analysis.binary import Container c = Container.from_string(s, loc_db) c <miasm.analysis.binary.ContainerUnknown object at 0x7f34cefe6090>
Desensamblando el shellcode en la dirección `0`:```pycon
>>> from miasm.analysis.machine import Machine
>>> machine = Machine('x86_32')
>>> mdis = machine.dis_engine(c.bin_stream, loc_db=loc_db)
>>> asmcfg = mdis.dis_multiblock(0)
>>> for block in asmcfg.blocks:
... print(block)
...
loc_0
LEA ECX, DWORD PTR [ECX + 0x4]
LEA EBX, DWORD PTR [EBX + 0x1]
CMP CL, 0x1
JZ loc_10
-> c_next:loc_b c_to:loc_10
loc_10
LEA EBX, DWORD PTR [EBX + 0x1]
-> c_next:loc_13
loc_b
LEA EBX, DWORD PTR [EBX + 0xFFFFFFFF]
JMP loc_13
-> c_to:loc_13
loc_13
MOV EAX, EBX
RET
Inicializando el motor JIT con una pila:```pycon
jitter = machine.jitter(loc_db, jit_type='python') jitter.init_stack()
Añade el shellcode en una ubicación de memoria arbitraria:```pycon
>>> run_addr = 0x40000000
>>> from miasm.jitter.csts import PAGE_READ, PAGE_WRITE
>>> jitter.vm.add_memory_page(run_addr, PAGE_READ | PAGE_WRITE, s)
Crear un centinela para capturar el retorno del shellcode:```Python def code_sentinelle(jitter): jitter.running = False jitter.pc = 0 return True
jitter.add_breakpoint(0x1337beef, code_sentinelle) jitter.push_uint32_t(0x1337beef)
Registros activos:```pycon
>>> jitter.set_trace_log()
Ejecutar en dirección arbitraria:```pycon
jitter.init_run(run_addr) jitter.continue_run() RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000000 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 RIP 0000000040000000 40000000 LEA ECX, DWORD PTR [ECX+0x4] RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000004 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 .... 4000000e JMP loc_0000000040000013:0x40000013 RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000004 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 RIP 0000000040000013 40000013 MOV EAX, EBX RAX 0000000000000000 RBX 0000000000000000 RCX 0000000000000004 RDX 0000000000000000 RSI 0000000000000000 RDI 0000000000000000 RSP 000000000123FFF8 RBP 0000000000000000 zf 0000000000000000 nf 0000000000000000 of 0000000000000000 cf 0000000000000000 RIP 0000000040000013 40000015 RET
Interactuando con el jitter:```pycon
>>> jitter.vm
ad 1230000 size 10000 RW_ hpad 0x2854b40
ad 40000000 size 16 RW_ hpad 0x25e0ed0
>>> hex(jitter.cpu.EAX)
'0x0L'
>>> jitter.cpu.ESI = 12
Inicializando el pool de IR:```pycon
lifter = machine.lifter_model_call(loc_db) ircfg = lifter.new_ircfg_from_asmcfg(asmcfg)
Inicializando el motor con valores simbólicos predeterminados:```pycon
>>> from miasm.ir.symbexec import SymbolicExecutionEngine
>>> sb = SymbolicExecutionEngine(lifter)
Iniciando la ejecución:```pycon
symbolic_pc = sb.run_at(ircfg, 0) print(symbolic_pc) ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10)
Igual, con registros de pasos (solo se muestran los cambios):```pycon
>>> sb = SymbolicExecutionEngine(lifter, machine.mn.regs.regs_init)
>>> symbolic_pc = sb.run_at(ircfg, 0, step=True)
Instr LEA ECX, DWORD PTR [ECX + 0x4]
Assignblk:
ECX = ECX + 0x4
________________________________________________________________________________
ECX = ECX + 0x4
________________________________________________________________________________
Instr LEA EBX, DWORD PTR [EBX + 0x1]
Assignblk:
EBX = EBX + 0x1
________________________________________________________________________________
EBX = EBX + 0x1
ECX = ECX + 0x4
________________________________________________________________________________
Instr CMP CL, 0x1
Assignblk:
zf = (ECX[0:8] + -0x1)?(0x0,0x1)
nf = (ECX[0:8] + -0x1)[7:8]
pf = parity((ECX[0:8] + -0x1) & 0xFF)
of = ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1))[7:8]
cf = (((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1)) ^ ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1)))[7:8]
af = ((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1))[4:5]
________________________________________________________________________________
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5]
pf = parity((ECX + 0x4)[0:8] + 0xFF)
zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1)
ECX = ECX + 0x4
of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8]
nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8]
cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8]
EBX = EBX + 0x1
________________________________________________________________________________
Instr JZ loc_key_1
Assignblk:
IRDst = zf?(loc_key_1,loc_key_2)
EIP = zf?(loc_key_1,loc_key_2)
________________________________________________________________________________
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5]
EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10)
pf = parity((ECX + 0x4)[0:8] + 0xFF)
IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10)
zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1)
ECX = ECX + 0x4
of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8]
nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8]
cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8]
EBX = EBX + 0x1
________________________________________________________________________________
>>>
Reintentar ejecución con un ECX concreto. Aquí, la ejecución simbólica / concolic alcanza el final del shellcode:```pycon
from miasm.expression.expression import ExprInt sb.symbols[machine.mn.regs.ECX] = ExprInt(-3, 32) symbolic_pc = sb.run_at(ircfg, 0, step=True) Instr LEA ECX, DWORD PTR [ECX + 0x4] Assignblk: ECX = ECX + 0x4
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5] EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) pf = parity((ECX + 0x4)[0:8] + 0xFF) IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1) ECX = 0x1 of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8] nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8] cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8] EBX = EBX + 0x1
Instr LEA EBX, DWORD PTR [EBX + 0x1] Assignblk: EBX = EBX + 0x1
af = (((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[4:5] EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) pf = parity((ECX + 0x4)[0:8] + 0xFF) IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) zf = ((ECX + 0x4)[0:8] + 0xFF)?(0x0,0x1) ECX = 0x1 of = ((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1))[7:8] nf = ((ECX + 0x4)[0:8] + 0xFF)[7:8] cf = (((((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8]) & ((ECX + 0x4)[0:8] ^ 0x1)) ^ ((ECX + 0x4)[0:8] + 0xFF) ^ (ECX + 0x4)[0:8] ^ 0x1)[7:8] EBX = EBX + 0x2
Instr CMP CL, 0x1 Assignblk: zf = (ECX[0:8] + -0x1)?(0x0,0x1) nf = (ECX[0:8] + -0x1)[7:8] pf = parity((ECX[0:8] + -0x1) & 0xFF) of = ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1))[7:8] cf = (((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1)) ^ ((ECX[0:8] ^ (ECX[0:8] + -0x1)) & (ECX[0:8] ^ 0x1)))[7:8] af = ((ECX[0:8] ^ 0x1) ^ (ECX[0:8] + -0x1))[4:5]
af = 0x0 EIP = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) pf = 0x1 IRDst = ((ECX + 0x4)[0:8] + 0xFF)?(0xB,0x10) zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x2
Instr JZ loc_key_1 Assignblk: IRDst = zf?(loc_key_1,loc_key_2) EIP = zf?(loc_key_1,loc_key_2)
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x10 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x2
Instr LEA EBX, DWORD PTR [EBX + 0x1] Assignblk: EBX = EBX + 0x1
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x10 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3
Instr LEA EBX, DWORD PTR [EBX + 0x1] Assignblk: IRDst = loc_key_3
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x13 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3
Instr MOV EAX, EBX Assignblk: EAX = EBX
af = 0x0 EIP = 0x10 pf = 0x1 IRDst = 0x13 zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3 EAX = EBX + 0x3
Instr RET Assignblk: IRDst = @32[ESP[0:32]] ESP = {ESP[0:32] + 0x4 0 32} EIP = @32[ESP[0:32]]
af = 0x0 EIP = @32[ESP] pf = 0x1 IRDst = @32[ESP] zf = 0x1 ECX = 0x1 of = 0x0 nf = 0x0 cf = 0x0 EBX = EBX + 0x3 ESP = ESP + 0x4 EAX = EBX + 0x3
¿Cómo funciona?
=================
Miasm incluye su propio desensamblador, lenguaje intermedio y semántica de instrucciones. Está escrito en Python.
Para emular código, utiliza LLVM, GCC, Clang o Python para JIT de la representación intermedia. Puede emular shellcodes y partes completas o parciales de binarios. Se pueden ejecutar callbacks de Python para interactuar con la ejecución, por ejemplo, para emular los efectos de funciones de biblioteca.
Documentación
=============
Algunos recursos de documentación están disponibles en la carpeta [doc](https://github.com/cea-sec/miasm/blob/HEAD/doc).
Una documentación autogenerada está disponible:
* [Doxygen](http://miasm.re/miasm_doxygen)
* [pdoc](http://miasm.re/miasm_pdoc)
Obteniendo Miasm
===============
* Clona el repositorio: [Miasm en GitHub](https://github.com/cea-sec/miasm/)
* Obtén una de las imágenes Docker en [Docker Hub](https://registry.hub.docker.com/u/miasm/)
Requisitos de software
---------------------
Miasm utiliza:
* python-pyparsing
* python-dev
* opcionalmente python-pycparser (versión >= 2.17)
Para habilitar JIT de código, uno de los siguientes módulos es obligatorio:
* GCC
* Clang
* LLVM con Numba llvmlite, ver más abajo
'Opcionalmente' Miasm también puede utilizar:
* Z3, el [Theorem Prover](https://github.com/Z3Prover/z3)
Configuración
-------------
Para usar el jitter, se recomienda GCC o LLVM
* GCC (cualquier versión)
* Clang (cualquier versión)
* LLVM
* Debian (testing/unstable): No probado
* Debian stable/Ubuntu/Kali/lo que sea: `pip install llvmlite` o instalar desde [llvmlite](https://github.com/numba/llvmlite)
* Windows: No probado
* Construir e instalar Miasm:```pycon
$ cd miasm_directory
$ python setup.py build
$ sudo python setup.py install
Si algo sale mal durante la compilación de uno de los módulos jitter, Miasm omitirá el error y deshabilitará el módulo correspondiente (consulte la salida de compilación).
La mayoría de los complementos IDA de Miasm utilizan un subconjunto de la funcionalidad de Miasm. Una forma rápida de hacerlos funcionar es agregar:
pyparsing.py a C:\...\IDA\python\ o pip install pyparsingmiasm/miasm a C:\...\IDA\python\Todas las funciones, excepto las relacionadas con JITter, estarán disponibles. Para una instalación más completa, consulte los párrafos anteriores.
Miasm viene con un conjunto de pruebas de regresión. Para ejecutarlas todas:```pycon cd miasm_directory/test
python test_all.py
python -m unittest test_all.py # sequential, requires 'unittest' python -m pytest test_all.py # sequential, requires 'pytest' python -m pytest -n auto test_all.py # parallel, requires 'pytest' and 'pytest-xdist'
Algunas opciones pueden especificarse:
* Subprocesamiento único: `-m`
* Instrumentación de cobertura de código: `-c`
* Solo pruebas rápidas: `-t long` (excluye las pruebas largas)
Ya usan Miasm
=====================
Herramientas
-----
* [Sibyl](https://github.com/cea-sec/Sibyl): Una herramienta de adivinación de funciones
* [R2M2](https://github.com/guedou/r2m2): Usa miasm como plugin de radare2
* [CGrex](https://github.com/mechaphish/cgrex): Parcheador dirigido para binarios CGC
* [ethRE](https://github.com/jbcayrou/ethRE): Herramienta de reversión para Ethereum EVM (con la correspondiente arquitectura Miasm2)
Publicaciones de blog / artículos / conferencias
---------------------------------
* [Deobfuscation: recovering an OLLVM-protected program](http://blog.quarkslab.com/deobfuscation-recovering-an-ollvm-protected-program.html)
* [Taming a Wild Nanomite-protected MIPS Binary With Symbolic Execution: No Such Crackme](https://doar-e.github.io/blog/2014/10/11/taiming-a-wild-nanomite-protected-mips-binary-with-symbolic-execution-no-such-crackme/)
* [Génération rapide de DGA avec Miasm](https://www.lexsi.com/securityhub/generation-rapide-de-dga-avec-miasm/): Cálculo rápido de DGA (Artículo en francés)
* [Enabling Client-Side Crash-Resistance to Overcome Diversification and Information Hiding](https://www.internetsociety.org/sites/default/files/blogs-media/enabling-client-side-crash-resistance-overcome-diversification-information-hiding.pdf): Detecta posibles argumentos de llamadas no dirigidas
* [Miasm: Framework de reverse engineering](https://www.sstic.org/2012/presentation/miasm_framework_de_reverse_engineering/) (Francés)
* [Tutorial miasm](https://www.sstic.org/2014/presentation/Tutorial_miasm/) (Video en francés)
* [Graphes de dépendances : Petit Poucet style](https://www.sstic.org/2016/presentation/graphes_de_dpendances__petit_poucet_style/): DepGraph (Francés)
Libros
-----
* [Practical Reverse Engineering: X86, X64, Arm, Windows Kernel, Reversing Tools, and Obfuscation](http://eu.wiley.com/WileyCDA/WileyTitle/productCd-1118787315,subjectCd-CSJ0.html): Introducción a Miasm (Capítulo 5 "Ofuscación")
* [BlackHat Python - Appendix](https://github.com/oreilly-japan/black-hat-python-jp-support/tree/master/appendix-A): Muestras del libro de seguridad de Japón