Multi-threading and performance tuning a hydrologic model: a case study
Jean‐Michel Perraud, Jamie Vleeshouwer · 2009
The computer industry has evolved from single-core to many-core architectures to keep offering increasing processing power. Parallel programming is not new but in practice much remains to do to take advantage of these many-core architectures. Most software is still designed for serial execution, and the level of parallel execution is often limited to running the user interface and engine in separate threads to keep an application responsive. This paper focuses on task parallelization in an environmental modelling framework, The Invisible Modelling Environmental framework (TIME). The case study arises from a project requiring the calibration of rainfall-runoff models in numerous unimpaired gauged catchments located in Northern Australia. The hydrologic model structure is spatially distributed, such that each catchment model can have hundreds of input time series at a daily time step. These models are run on multi-core computers in a cluster, with the cumulated memory requirements possibly surpassing the available memory, slowing the computation to unacceptable levels due to virtual memory swapping. We thus tailor the number of catchment calibration tasks per compute node such that the cumulated memory footprint fits in the random access memory of that node. In order to still maximize the use of processing power on these multi-core nodes, we need to be able to parallelize these calibration tasks. Three parallelization strategies are considered, characterized mainly by different granularities for the tasks considered for parallelization: (1) parallelizing the calibration algorithm, (2) parallelizing the model along the spatial dimension, and (3) parallelizing the model along spatial and temporal dimensions. Assessing each against several criteria notably runtime performance gain, technological know-how, the amount of code changes and architectural impacts, solution (2) with multi-threading is preferred as the best compromise between these criteria. The architectural changes required by the parallelization and concomitant performance tuning are described along with the technical characteristics of .NET based solutions for multi-threading. We find that the performance tuning process necessary to make the parallelization a net benefit proves to be the bulk of the work. Two main performance bottlenecks are, not unexpectedly, identified: the use of software reflection (a.k.a. software introspection), and the access to input and output time series using date-time as an index. We use dynamic code generation to overcome the former, and introduce new interfaces for time series access for the latter, while avoiding pervasive changes to the framework. We find a satisfying speedup of roughly 80% of a theoretical linear speedup for a dual-threaded catchment model with 256 spatial grid cells. While additional threads improve the overall runtime, the runtime for eight threads is a somewhat disappointing 250% of the 'ideal' runtime, only partly explained by the expected parallelization overhead. Given the substantial architectural changes required in this case study that were not for parallelizing per se, we discuss the possible implications of the prevalence of multi-core processors on software design practices, at least in a scientific computing context.