RABID -- A General Distributed R Processing Framework Targeting Large Data-Set Problems
Hao Lin, Shuo Yang, Samuel Pratt Midkiff · 2013
Large-scale data mining and deep data analysis are in high demand in modern enterprises. This work describes the RABID (R Analytics for BIg Data) framework to provide a highly parallel R. We achieve the goal of providing data analysts with an easy-to-use R interface to effectively perform deep data analysis on clusters by integrating R and a MapReduce-like platform. By leveraging a distributed runtime system, our framework enables R, the single-threaded language, to efficiently perfrom parallel analysis of data that cannot fit into a single shared memory machine in parallel. Experiments of data mining benchmarks on our framework show promising results.