IDEA -- An API for Parallel Computing with Large Spatial Datasets
Baoqiang Yan, Philip J. Rhodes · 2011
We describe IDEA, an API designed specifically for the parallel processing of large spatial datasets on a cluster. Because such datasets present special challenges for efficient I/O and communication, it is especially valuable to provide an API that frees the user from the burden of partitioning the data among the processors. IDEA allows the user to address a communication to neighboring blocks of data, rather than processes or nodes. In addition to being very natural for the user, this data-centric view allows communication to a data block before it has been assigned a process. This is a key ability when handling data sets larger than the aggregate memory capacity of the cluster, since the dataset must be processed in a piecewise fashion.