Enhancing Performance in Scientific Applications with Energy and Resilience Constraints from Modern Architectures
Brandon Nesterenko · Digital Collections of Colorado (Colorado State University) · 2020
8.3 Domain-level properties of the stencils our model supports in ocean cp 141 x List of Figures 1.1 Time distribution of functions within ocean˙cp application from the Splash 2 benchmark suite . . . . . . . . . . . . . . . . . . . . . . . .1.2 Workflow to optimize a scientific application . . . . . . . . . . . . . .2.1 Three-dimensional model of a tornado using VAPOR[122] . . . . . . .2.2 Typical workflow for scientific applications . . . . . . . . . . . . . . .4.1 Process execution comparison of scheduling policies . . . . . . . . . .4.2 Workflow when starting a progress period in our scheduling environment 4.3 Time lapse for a process with a single progress period being scheduled in our environment. . . . . . . . . . . . . . . . . . . . . . . . . . . . .4.4 C code sample consisting of one progress period, the call to DGEMM.4.5 Workflow that shows the progress monitor's sub components when an application begins a progress period . . . . . . . . . . . . . . . . . . .4.6 Workflow that shows the progress monitor's sub components when an application ends a progress period . . . . . . . . . . . . . . . . . . . .4.7 Comparing the resource demand aware scheduling policies against the Linux default scheduling policy in terms of energy, in Joules, consumed by the system (CPU + cache + DRAM) when scheduling the eight workloads under different policies. . . . . . . . . . . . . . . . . . . . .4.8 Energy (Joules) consumed only by DRAM when scheduling the eight workloads under different policies . . . . . . . . . . . . . . . . . . . .xi 4.9 Performance measured in GFLOPS for each workload run under the different scheduling policies . . . . . . . . . . . . . . . . . . . . . . .44 4.10 GFLOPS per Watt of system (CPU + caches + DRAM) comparison of scheduling policies on the eight workloads . . . . . . . . . . . . . .45 4.11 Execution time taken by dgemm with no progress periods, and progress periods placed containing the outer, middle, and inner loop. . . . . .47 4.12Here we show how the actual working set size of two progress period's from the water nsquared and ocean cp applications increase with respect to the input size, and our prediction function results.The input size scales from an original input size, 1x, to be 2x, 4x, and 8x the size.49 4.13 The amount of slowdown for the largest progress period in water nsquared due to LLC interference from increased data size and concurrent processes running. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .50 5.1 The comparison of saving an application state using (a) a traditional hard disk drive approach, and (b) a non-volatile RAM based approach.58 5.2 Original workflow of Fluidanimate . . . . . . . . . . . . . . . . . . . .62 5.3 Restructured workflow of Fluidanimate . . . . . . . . . . . . . . . . .62 5.4 Our heartbeat monitor extension to the fault-tolerant version of Fluidanimate in Figure 5.3. . . . . . . . . . . . . . . . . . . . . . . . . .70 5.5 Cache miss rate of Fluidanimate when run under each fault tolerant implementation compared to the original version with no fault tolerance.74 5.6 The bandwidth observed when saving/recovering data for Fluidanimate compared to the maximum bandwidth attainable for each storage system. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .74 5.7 The bandwidth observed when checkpointing the modified version of Fluidanimate that uses arrays when compared against the maximum bandwidth attainable for a filesystem. . . . . . . . . . . . . . . . . . .75 xii 5.8 The duration of time spent in saving/recovering data in Fluidanimate when using a linked list vs an array. . . . . . . . . . . . . . . . . . . .75 5.9 Extending Figure 5.8 to show the duration of time taken to checkpoint Fluidanimate with a blocked array. . . . . . . . . . . . . . . . . . . .76 5.10 Performance scalability of the application when using each NVRAM API to support resiliency. . . . . . . . . . . . . . . . . . . . . . . . .77 5.11 Whether the four NVRAM APIs allows instant restart within the heartbeat timeout (goal, 1.7s). . . . . . . . . . . . . . . . . . . . . . .79 5.12 Whether the four NVRAM APIs allow instant restart with the modified heartbeat timeout (goal, 2.4s). . . . . . . . . . . . . . . . . . . . . . .80 6.1 Impact of dimensionality on the effectiveness of rectangular (R), partial diamond (Pd), and full diamond (Fd) tiling with thread-level parallelism with 16 threads. . . . . . . . . . . . . . . . . . . . . . . . . .87 6.2 Workflow showing the process to iteratively navigate the optimization space for a stencil . . . . . . . . .