AstroSpark
Mariem Brahem, Stéphane Lopes, Laurent Yeh, Karine Zeitouni · 2016
Large amounts of astronomical data are continuously collected. As a result, support of scalable and high performance query processing of such data has become increasingly necessary. Apache Spark has been widely adopted as a successor to Apache Hadoop MapReduce to analyze Big Data in distributed frameworks. Despite its rich features, this framework can not be directly exploited towards processing astronomical data. In this work, we present AstroSpark, a distributed data server for astronomical data. AstroSpark extends Spark, a distributed in-memory computing framework, to analyze and query huge volume of astronomical data. It supports astronomical operations such as cone search, cross-match and histogram. AstroSpark introduces data partitioning and optimization techniques to achieve high performance query execution.