A Hybrid Transformer–LLM Pipeline for Function Name Recovery in Stripped Binaries
Remus M. Petrache, Camelia Lemnaru · 2025
Recovering meaningful function names from stripped executables is a difficult challenge in reverse engineer- ing and security analysis. Without debug symbols, functions receive generic names, complicating both manual and automated analysis. Building on AsmDepictor [1], we propose architectural refinements to reduce repetitive and ambiguous predictions, along with API-based integration of a Large Language Model for enhanced name suggestions. We evaluate our approach on a dataset of 26 million functions, extracted from roughly 6800 open-source Windows binaries, compiled with MSVC $2017-2022$ for both X86 and X86_64 architectures. Preliminary evaluations incorporating the LLM model show notable reductions in repeti- tive naming errors. The overall approach highlights the potential to restore relevant semantic information in stripped binaries, promoting more efficient reverse engineering, malware analysis, and software composition analysis.