
Dynamic branch-divergence finder for native code -- traces two Frida executions and finds the exact instruction where they diverge.
════════════════════════════════════════════════════════════════ VERIDIFF the exact instruction where two runs of the same code split ════════════════════════════════════════════════════════════════
Two traces go in. One address comes out: the exact basic block where your
two runs stopped agreeing, disassembled, with the branch that split them
marked — cmp/jne on x86, subs/b.ne on ARM64, a jump-table dispatcher
if the code is flattened. (When a compiler decides a condition without
branching at all — csel, cmov — the comparison itself sits in a block
both runs executed, and what you get is the block where the paths actually
forked. That distinction is demonstrated, not glossed over: see
Obfuscated control flow below.)
That's the whole product. Not a framework. Not a platform. An engine.
Current release: v0.4.0. Fixes silent trace truncation: Frida's default
Stalker queue holds 16384 events and drops the rest without an error, so
every trace longer than that was quietly incomplete — measured at up to 98%
of blocks lost while reporting success. The engine now sizes the queue itself
(and lets you size it), and every Trace carries queue_saturated so
truncation can never be silent again. Everything in this README is now
live-verified on both x86-64 and ARM64/Android, including control-flow
flattening and a 1.5-million-block trace. Breaking: trace_call gained an
option field; see Quickstart. Full history in
CHANGELOG.md.
Contents — Quickstart · Proof, Not Promises · How it works · Warm-up mode · Obfuscated control flow · Scale · Branch classification · Field notes · The Law
Reverse engineering a license check, an anti-cheat heuristic, or an OLLVM-flattened dispatcher usually comes down to the same manual loop: run it once with good input, run it again with bad input, and stare at two disassembly listings trying to spot where they stopped agreeing with each other. That loop doesn't scale past a few hundred basic blocks, and it doesn't survive control-flow flattening at all — every path already goes through the same dispatcher block, so staring at addresses tells you nothing.
Veridiff automates the diffing, not the reversing. Attach Frida's Stalker to a thread, run it twice, and get back the one thing that actually matters: the last block both runs agreed on, the first block they didn't, and the disassembly of the instruction that split them. Millions of basic blocks in, one address out.
It ships as two things, on purpose:
python/main.py — one file, pip install frida, done.rust/ — a real library crate (src/lib.rs) plus a thin
demo binary (src/main.rs), for when Python's call overhead is the
bottleneck instead of Frida's own instrumentation cost.Both expose the same four calls (trace_call, disassemble_block,
find_first_divergence, find_divergence_regions) and produce
byte-identical results against the same target. Neither one has a
single line of CLI, GUI, or presentation logic in it. That's not an
oversight — see The Law below.
Python — one file, one dependency:
pip install frida
python3 python/main.py <executable> <hex-address-of-target-fn> <arg-a> <arg-b>
Rust — a real library, importable by path/git from any other crate:
# your Cargo.toml
[dependencies]
veridiff = { path = "../Veridiff/rust" }
let options = TraceCallOptions { warm_up: true };
let trace = engine.trace_call(&mut script, addr, &args, "int", None, options)?;
let d = VeridiffEngine::find_first_divergence(&trace_a, &trace_b);
src/main.rs in this repo is nothing but a thin CLI wrapper over that
same API — read it as a usage example, not as the product.
New in v0.2.0, both languages: trace_call(..., warm_up=True) / a
TraceCallOptions { warm_up: true } argument (see Warm-up mode), and
instruction.is_conditional_branch / insn.is_conditional_branch() on
every Instruction returned by disassemble_block (see Cross-architecture
branch classification). Both are additive on the Python side; the Rust
trace_call signature gained a required trailing parameter, which is a
breaking change for existing callers on this pre-1.0 crate -- pass
TraceCallOptions::default() for the old behavior.
Pre-1.0, so these come as minor bumps. Both are small.
v0.4.0 added stalker_queue_capacity to TraceCallOptions, which breaks
Rust struct-literal construction. Use ..Default::default() and future knobs
won't break you either:
let options = TraceCallOptions { warm_up: true, ..Default::default() };
Trace also gained event_count and queue_capacity, and with them
queue_saturated() / .queue_saturated — check it on any trace long
enough to matter, because a saturated trace is silently truncated (see
Field notes). Python's trace_call takes the capacity as a keyword
argument, so nothing there breaks.
v0.3.0 made DivergenceRegion's branch_a/branch_b Option<u64> in
Rust (Optional[int] in Python), so "one trace ran out while the other kept
going" can be reported instead of silently dropped.
The agent and host also exchange a per-call id and a queue capacity, so a host
and an AGENT_SOURCE from different versions must not be mixed. Within a
release they always match, since each engine embeds its own copy.
No benchmarks against products we've never touched, no "battle-tested
against Denuvo" fantasy claims. Here is a small, deliberately
unremarkable C function — examples/licensecheck.c —
compiled with nothing that makes Frida's job easier, traced twice with
a valid and an invalid key, on this machine, moments before this
paragraph was written:
$ gcc -O0 -no-pie -fno-pie -fno-inline -o licensecheck examples/licensecheck.c
$ nm licensecheck | grep check_license
0000000000401146 T check_license
$ python3 python/main.py ./licensecheck 0x401146 VALID-KEY-123 WRONG-KEY
trace A ('VALID-KEY-123'): 9 blocks, returned 1
trace B ('WRONG-KEY'): 6 blocks, returned 0
first divergence at trace index 1
last common block : 0x114e
run A took block : 0x1164
run B took block : 0x116a
disassembly of the deciding block:
0x40114e mov qword ptr [rbp - 0x18], rdi
0x401152 mov dword ptr [rbp - 4], 0
0x401159 mov rax, qword ptr [rbp - 0x18]
0x40115d movzx eax, byte ptr [rax]
0x401160 cmp al, 0x56
0x401162 jne 0x40116a <-- decides here
(Trace A is 9 blocks here, not 11 -- v0.1.0's README showed this same run
before warm-up mode existed. See Warm-up mode below for what changed
and why. The <-- decides here marker is v0.2.0's conditional-branch
classification, covered in Cross-architecture branch classification.)
0x56 is 'V'. The engine never saw the source. It never saw
"VALID-KEY-123" as a string to search for — it doesn't know what a
license key is. It ran two traces, found where they stopped matching,
and handed back the exact comparison, because that's the first byte
the two arguments actually disagree on. That's the entire trick, and
it's the only trick: turn "where did these diverge" into a linear
scan instead of a stare-at-two-listings exercise.
The Rust engine, same binary, same two keys:
$ cd rust && cargo run --release -- ../licensecheck 0x401146 VALID-KEY-123 WRONG-KEY
trace A ("VALID-KEY-123"): 9 blocks, returned Some("1")
trace B ("WRONG-KEY"): 6 blocks, returned Some("0")