
Playing with the VMProtect software protection. Automatic deobfuscation of pure functions using symbolic execution and LLVM.
An experimental dynamic approach to devirtualize pure functions protected by VMProtect 3.x
I am sharing some notes about a dynamic approach to devirtualize pure functions protected by VMProtect. This approach has shown very good results if the virtualized function only contains one basic block (regardless of its size). This is a common scenario when binaries protect arithmetic operations. However, this approach is a bit more experimental when the target function contains more than one basic block. Nevertheless, we managed to devirtualize and reconstruct the binary code from samples that contain 2 basic blocks which suggests that it is possible to fully devirtualize small functions dynamically.
VMProtect is a software protection that protects code by running it through a virtual machine with non-standard architecture. This protection is a great playground for asm lovers [0, 1, 2, 3, 4, 5, 6, 11]. Also, there are already numerous tools that attack this protection [7, 8, 9, 12, 13]. In 2016 we took a look at the Tigress software protection solution and managed to defeat its virtualization using symbolic execution and LLVM. This approach has been presented at DIMVA 2018 [10] and I wanted to test it on VMProtect. Note that there is no magic solution that works on every binaries, there are always tradeoffs depending on the target and your goals. This modest contribution aims to provide an example of a dynamic attack against pure functions that are virtualized by VMProtect. The main advantage of a dynamic attack is that it defeats by design some VMProtect's static protections like self modifying code, key and operands encryption etc.
We consider a pure function a function with a finite number of paths and that does not have side effects. There can be several inputs but only one output. Below is an example of a pure function:
int secret(int x, int y) {
int r = x ^ y;
return r;
}
We rely on the key intuition that an obfuscated trace T' (from the obfuscated code P') combines original instructions from the original code P (the trace T corresponding to T' in the original code) and instructions of the virtual machine VM such that T' = T + VM(T). If we are able to distinguish between these two subsequences of instructions T and VM(T), we then are able to reconstruct one path of the original program P from a trace T'. By repeating this operation to cover all paths of the virtualized program, we will be able to reconstruct the original program P. In our practical example, the original code has a finite number of executable paths, which is the case in many situations involving intellectual property protection. To do so, we proceed with the following steps:
Let's take as a first example the following function: it takes two inputs and returns x ^ y which is protected by VMProtect.
int secret(int x, int y) {
VMProtectBegin("secret");
int r = x ^ y;
VMProtectEnd();
return r;
}
We start by identifying where functions are using VMProtect and how many arguments they have. For our example we may have something like below:
Just by reading the code we know that the function starts at the address 0x4011c0, have two 32-bit arguments (edi and esi)
and returns at 0x4011ef. That's all the reverse-engineering we need. Next parts will be automatic. Now we
have to generate an trace execution of this virtualized function. To do so we use a Pintool.
It only needs a start and an end address (for our example, 0x4011c0 and 0x4011ef) which represents the range
of the instrumentation. Note that any kind of DBI or emulator could do this job.
$ ./pin/pin -t ./pin/source/tools/VMP_Trace/obj-intel64/VMP_Trace.so -start 4198848 -end 4198895 -- ./vmp_binaries/binaries/sample2.vmp.bin 1 2 &> ./vmp_traces/sample2.vmp.trace
You can see the result here. The trace format uses three kind of operations: mr, r and i.
mr is a memory read access done by the instruction i, and r are the CPU registers. For example:
mr:0x7ffda459d718:8:0x227db4f8
r:0x40200a:0x0:0x7ffda459f571:0x2:0x40200a:0x0:0x0:0x7ffda459d688:0x0:0x0:0x7feee9b80ac0:0x7feee9b8000f:0xad1c3e:0x0:0x0:0x0
i:0x89173e:8:488BB42490000000
We have a memory read that loads an 8 bytes constant 0x227db4f8 from the address 0x7ffda459d718.
The instruction is executed at the address 0x89173e and its 8-bytes long opcode is 488BB42490000000 which is a
mov rsi, qword ptr [rsp + 0x90].
The register state before the execution is the following:
(1) RAX = 0x40200a (9) R8 = 0
(2) RBX = 0 (10) R9 = 0
(3) RCX = 0x7ffda459f571 (11) R10 = 0x7feee9b80ac0
(4) RDX = 0x2 (12) R11 = 0x7feee9b8000f
(5) RDI = 0x40200a (13) R12 = 0xad1c3e
(6) RSI = 0 (14) R13 = 0
(7) RBP = 0 (15) R14 = 0
(8) RSP = 0x7ffda459d688 (16) R15 = 0
Once the VMP trace has been generated, we replay it using the attack_vmp.py script. This script uses Triton to build the path predicate of the trace. Note that all expressions which involve symbolic variables (inputs of the function) are kept symbolic while all non related input expressions are concretized. In other words, our symbolic expressions do not contain any operation related to the virtual machine (the machinery itself does not depend on the user) but only operations related to the original program.
For example, below is an example of a concretization. On the left we have an AST that contains subexpressions which do not involve
symbolic variable (1 + 2 and 6 ^ 3). So these branches are concretized and replaced by constants 3 and 5 which leads to
the AST on the right. This is how we devirtualize code.