On the Vulnerability of Large Corpora Source Code
Joseph R. Barr, Tyler Thatcher · 2022
This paper is a part of a continual effort to score functions in source code for vulnerability. For practical reasons we've restricted our attention to the C and C++ programming languages. We demonstrate an auto-encoder network and techniques to embed source code into a low-dimensional Euclidean space and some of the issues encountered where dealing with a very large code base. We also describe a process of developing ‘code smell’ features and a classifier when data is extremely unbalanced. Finally we explore how the workflow may generalize to other projects and programming languages.