Towards Intelligent Distributed Data Systems for Scalable Efficient and Accurate Analytics
Peter Triantafillou · 2018
Large analytics tasks are currently executed over Big Data Analytics Stacks (BDASs) which comprise a number of distributed systems as layers for back-end storage management, resource management, distributed/parallel execution frameworks, etc. In the current state of the art, the processing of analytical queries is too expensive, accessing large numbers of data server nodes where data is stored, crunching and transferring large volumes of data and thus consuming too many system resources, taking too much time, and failing scalability desiderata. With this vision paper, we wish to push the research envelope, offering a drastically different view of analytics processing in the big data era. The radical new idea is to process analytics tasks employing learned models of data and queries, instead of accessing any base data - we call this data-less big data analytics processing. We put forward the basic principles for designing the next generation intelligent data system infrastructures realizing this new analytics-processing paradigm and present a number of specific research challenges that will take us closer to realizing the vision, which are based on the harmonic symbiosis of statistical and machine learning models with traditional system techniques. We offer a plausible research program that can address said challenges and offers preliminary ideas towards their solution. En route, we describe initial successes we have had recently with achieving scalability, efficiency, and accuracy for specific analytical tasks, substantiating the potential of the new paradigm.