Analyzing ultra-large-scale code corpus with boa

ROBERT L. DYER, Hoan Anh Nguyen, Hridesh Rajan, Tien N. Nguyen · 2012

Analyzing the wealth of information contained in software repositories requires significant expertise in mining techniques as well as a large infrastructure. In order to make this information more reachable for non-experts, we present the Boa language and infrastructure. Using Boa, these mining tasks are much simpler to write as the details are abstracted away. Boa programs also run on a distributed cluster to automatically provide massive parallelization to users and return results in minutes instead of potentially days.

Read the paper · More papers on PaperTik