
Per il nostro articolo CCS24 🏆 "ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped Binaries" di Danning Xie, Zhuo Zhang, Nan Jiang, Xiangzhe Xu, Lin Tan e Xiangyu Zhang. 🏆 Vincitore dell'ACM SIGSAC Distinguished Paper Award
Questo repository fornisce gli artefatti per l'articolo "ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped Binaries" (CCS 2024).
🏆 Vincitore dell'ACM SIGSAC Distinguished Paper Award
Nota: Stiamo attivamente mantenendo/aggiornando i nostri artefatti. Assicurati di utilizzare la versione più recente.
process_data. Genera dati di addestramento con informazioni sui simboli ground truth. Lo script è avviabile con un semplice comando e le istruzioni per l'uso sono fornite nella cartella.ReSym_rawdata). Questi includono i file binari grezzi e il corrispondente codice decompilato che abbiamo utilizzato in questo progetto:
bin/: Contiene i file binari non strippati grezzi con informazioni di debug.decompiled/: Codice decompilato da binari completamente strippati.metadata.json: Metadati per i binari, incluse le informazioni sul progetto.training_src per i modelli VarDecoder e FieldDecoder.ReSym_data). Questi includono: dati di addestramento, dati di test e risultati di predizione per FieldDecoder e VarDecoder.posterior_reasoning. I dettagli e le istruzioni si trovano nella cartella.training_src.@inproceedings{10.1145/3658644.3670340,
author = {Xie, Danning and Zhang, Zhuo and Jiang, Nan and Xu, Xiangzhe and Tan, Lin and Zhang, Xiangyu},
title = {ReSym: Harnessing LLMs to Recover Variable and Data Structure Symbols from Stripped Binaries},
year = {2024},
isbn = {9798400706363},
publisher = {Association for Computing Machinery},
address = {New York, NY, USA},
url = {https://doi.org/10.1145/3658644.3670340},
doi = {10.1145/3658644.3670340},
abstract = {Decompilation aims to recover a binary executable to the source code form and hence has a wide range of applications in cyber security, such as malware analysis and legacy code hardening. A prominent challenge is to recover variable symbols, including both primitive and complex types such as user-defined data structures, along with their symbol information such as names and types. Existing efforts focus on solving parts of the problem, e.g., recovering only types (without names) or only local variables (without user-defined structures). In this paper, we propose ReSym, a novel hybrid technique that combines Large Language Models (LLMs) and program analysis to recover both names and types for local variables and user-defined data structures. Our method encompasses fine-tuning two LLMs to handle local variables and structures, respectively. To overcome the token limitations inherent in current LLMs, we devise a novel Prolog-based algorithm to aggregate and cross-check results from multiple LLM queries, suppressing uncertainty and hallucinations. Our experiments show that ReSym is effective in recovering variable information and user-defined data structures, substantially outperforming the state-of-the-art methods.},
booktitle = {Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security},
pages = {4554–4568},
numpages = {15},
keywords = {large language models, program analysis, reverse engineering},
location = {Salt Lake City, UT, USA},
series = {CCS '24}
}