
Benchmark dataset e artefatti di sintesi per QSynth, contenenti funzioni C offuscate con Tigress, binari x86_64, tracce di esecuzione, ground truth e risultati AST sintetizzati.
Questi artefatti forniscono i seguenti quattro dataset:
Ogni directory di dataset contiene i seguenti file:
original.c: file sorgente contenente tutte le funzioni
obfuscated.c: file sorgente offuscato generato da Tigress
run_tigress.sh: script Shell per rigenerare il file obfuscated.c
ground_truth.json: un file JSON contenente la ground truth. Per ogni funzione contiene l'espressione originale (in original.c) e la sua controparte offuscata (in obfuscated.c)
obfuscated: binario precompilato come x86_64 (no PIE)
trace.db: la traccia di esecuzione come database SQLite (probabilmente il file più interessante)
results:
Il file qsynth.json contiene per ogni funzione le seguenti voci, come ad esempio:
{
"target_1": {
"fun_name": "target_1",
"start": 164,
"stop": 238,
"orig_ast": "~(b*d ^ d*d) - b",
"orig_node": 10,
"orig_depth": 5,
"obfu_ast": "(((b & d)*(b | d) + (b & ~d)*(~b & d) & (d & d)*(d | d) + (d & ~d)*(~d & d)) - ((b & d)*(b | d) + (b & ~d)*(~b & d) | (d & d)*(d | d) + (d & ~d)*(~d & d)) - 1 ^ b) - ((~(((b & d)*(b | d) + (b & ~d)*(~b & d) & (d & d)*(d | d) + (d & ~d)*(~d & d)) - ((b & d)*(b | d) + (b & ~d)*(~b & d) | (d & d)*(d | d) + (d & ~d)*(~d & d)) - 1) & b) + (~(((b & d)*(b | d) + (b & ~d)*(~b & d) & (d & d)*(d | d) + (d & ~d)*(~d & d)) - ((b & d)*(b | d) + (b & ~d)*(~b & d) | (d & d)*(d | d) + (d & ~d)*(~d & d)) - 1) & b))",
"obfu_node": 229,
"obfu_depth": 12,
"triton_ast": "(((b & d)*(b | d) + (~b & d)*(~d & b) & d*d) - (d*d | (b & d)*(b | d) + (~b & d)*(~d & b)) - 1 ^ b) - (((d*d ^ (b & d)*(b | d) + (~b & d)*(~d & b)) & b) + ((d*d ^ (b & d)*(b | d) + (~b & d)*(~d & b)) & b))",
"triton_node": 95,
"triton_depth": 10,
"synthesized_ast": "~((d*d ^ d*b) + b)",
"synth_node": 10,
"synth_depth": 5,
"dse_t": 0.10445356369018555,
"synthesis_t": 0.03491616249084473,
"sem_orig_obf": "UNK",
"sem_obf_trit": "UNK",
"sem_orig_synth": "OK",
"is_simplified": true,
"is_fully_synthesized": true
}
}
Start e stop sono gli offset nella traccia di esecuzione (quindi l'id nel DB). Poi per ogni
espressione l'espressione stessa e la sua dimensione in nodi e profondità. Poi dse_t, synthesis_t forniscono
rispettivamente il tempo di esecuzione simbolica e il tempo di sintesi. is_simplified e is_fully_synthesized
indicano se la funzione è stata semplificata e, in tal caso, se lo è stata completamente. sem_orig_obf, sem_obf_trit
e sem_orig_synth indicano se la semantica è preservata tra, ad esempio, originale e sintetizzata
(sem_orig_synth). "UNK" indica che non è stato verificato o che ha prodotto un timeout.
Le tabelle utilizzate per i benchmark sono disponibili qui: https://ret2libc.com/static/various/lts_15/ Sono oggetti pickle di Python. Le espressioni sono codificate con una notazione polacca inversa (RPN) simile a quella di Syntia.