Investigating influence of data storage organization on structured code search performance

Daniel Bernau, Olga Mordvinova, Jan Karstens, Susan Hickl · 2011

Code search in an industrial environment is driven by the programmers wish to scan huge source code repositories with high precision in a very short time. Given a challenging scenario of a huge software repository, the question for an efficient code search backend is relevant. This paper discusses the question of an appropriate data storage model for a structured code search engine applied in an industrial development scenario, where a search on large software repositories is common. To investigate this, a search engine approach with integrated Abstract Syntax Trees is adapted. Using the capabilities of a hybrid in-memory database, we stored a big amount of structured data obtained from the source code repository into column-, row-, and a hybrid store layout and performed a set of typical queries using an SQL interface on them. The results have shown the superiority of the column-oriented approach for the investigated scenario.

Read the paper · More papers on PaperTik