How to measure perfectness of parallelization in hardware/software systems
János Végh, Péter Molnár · 2017
After that single-thread processing performance has stalled more than a decade ago, research has intensified in direction of using more intensive parallelism. Unfortunately, the resulting computer performance is not simply the sum of single processor performances of individual processors working in parallel. Communication, housekeeping and some inherently sequential code parts degrade efficiency of parallelization, in a strongly nonlinear way. Amdahl's law provides an upper bound for achievable parallelism, but it requires the exact knowledge of the structure of the program, and in addition, it can consider realistic systems in a rather limited way. Modifying Amdahl's law for considering modern general purpose systems working in parallel, and reverting the formula derived, a quantitative merit was delivered, which shows how perfect the result of parallelization is. The derived formulas are applied to problems on different fields, like qualifying efficacy of a load balancing compiler, comparing communication methods usable in an SoC system, or finding out which law governed the supercomputer technology in the past quarter of century.