Parallel computing framework for big data
LU Ke-zhong, Rui Mao, GuoLiang CHEN · Chinese Science Bulletin (Chinese Version) · 2015
Big data has received a great deal of attention with respect to its use in research and application in information technology. However, most current efforts focus on systems and applications instead of the theoretical foundation. Based on computational complexity theory, according to the volume, velocity, and variety challenges of big data, we study the computability and computational principles of big data. First, various types of big data can be abstracted into metric space for universal representation to handle the variety challenge. Big data can then be partitioned in metric space according to distance. Finally, NC-class computing theory can be applied to solve big data problems in parallel and handle the volume and velocity challenges. Last of all, from a wider perspective, we propose a processing strategy according to the challenges and innovation of big data research methodology.