The Use of Locality Information on Data Intensive Parallel File Systems
Ricardo Ryoiti Sugawara, Líria Matsumoto Sato · 2013
Many recent data intensive parallel systems builds with cost effective hardware and combine compute and storage facilities. Since bandwidth-bisecting networks are the norm, distributing jobs near data provides significant performance improvements. However, the data locality information is not easily available to the programmer. It requires interaction with file system internals, or the adoption of a custom programming and run-time frameworks that provide locality-aware job scheduling, such as Mapreduce and Hadoop. In this paper, we present a parallel file system implementation combined with virtual files presented as text that can be queried for locality data or written to control the placement of new data blocks. This simplifies how software that do not adhere to Mapreduce's model can benefit from computing near the data. To evaluate the proposed approach, a number of tests were run on an initial implementation using fast disks, with locality-aware cases showing from 2 to 9 times faster reads and higher processor utilization.