Locality analysis
Jacob Brock, Hao Luo, Chen Ding · ACM SIGMETRICS Performance Evaluation Review · 2014
The rise of social media and cloud computing, paired with ever-growing storage capacity are bringing big data into the limelight, and rightly so. Data, it seems, can be found everywhere; It is harvested from our cars, our pockets, and soon even from our eyeglasses. While researchers in machine learning are developing new techniques to analyze vast quantities of sometimes unstructured data, there is another, not-so-new, form of big data analysis that has been quietly laying the architectural foundations of efficient data usage for decades. Every time a piece of data goes through a processor, it must get there through the memory hierarchy. Since retrieving the data from the main memory takes hundreds of times longer than accessing it from the cache, a robust theory of data usage can lay the groundwork for all efficient caching. Since everything touched by the CPU is first touched by the cache, the cache traces produced by the analysis of big data will invariably be bigger than big. In this paper we first summarize the locality problem and its history, and then we give a view of the present state of the field as it adapts to the industry standards of multicore CPUs and multithreaded programs before exploring ideas for expanding the theory to other big data domains.