Communication methods for hierarchical global address models in HPC
Huan Wei Zhou · OPUS Publication Server of the University of Stuttgart (University of Stuttgart) · 2016
Simulation is becoming increasingly important in the field of engineering science and viewed as an extension of the traditional theoretical and experimental investigation.However, the common obstacles we are facing in the large-scale simulations are the large volumes of data sets to be processed and the computational complexity of the methods.To simulate the large-scale problems within reasonable time, High Performance Computing (HPC) techniques are needed.One method to speeding up computational execution time is parallel processing.I.e., a single simulation program is distributed over multiple processes.Various parallel programming models are developed for writing efficient parallel applications to perform the distributed simulation on parallel computer hardware.The Message Passing Interface (MPI) remains as the dominant communication model as a result of its high performance, portability and standardization on cuttingedge computing systems.However, writing an efficient MPI application is challenging.The emergence of multi-and many-core processors fuels an explosive growth of processing capability in current-generation high-end systems.In this sense, communication speed becomes a limiting factor in the distributed simulation performance.To alleviate the communication bottleneck, an adequate communication scheme stemming from the existing high performance parallel programming models is needed.In this dissertation, I explore the approach of designing an efficient communication runtime system on top of MPI-3 Remote Memory Access (RMA) model by hiding the complexity of MPI from the user.To take the hierarchical memory layout of HPC systems into account, I adopt a hybrid strategy which combines the MPI RMA model (for internode communications) and the shared-memory programming model (for intra-node communications).Moreover, to achieve higher communication-computation overlap, I design a progress engine by using a process-based approach.The progress engine can start the asynchronous progression of communication independent of MPI synchronization calls on the origin side.I have evaluated the performance improvement offered by the proposed communication scheme using application benchmarks as well as communication microbenchmarks.The micro-benchmark evaluations demonstrate that the proposed communication runtime system can consistently yield better performance than MPI RMA operations in the intra-node case.Also, it can retain the performance of MPI RMA operations in the inter-node case without adding noticeable overheads.Most importantly, it is demonstrated that the proposed communication scheme can lead to substantial performance increase for real engineering applications.Besides that, the proposed communication runtime system shows a clear performance advantage compared with other low-level RMA libraries in the intra-node case according to the micro-benchmarks and i ii it can also match those low-level RMA libraries in terms of the application benchmarks involving inter-node data transfers.Zusammenfassung Simulation wird immer wichtiger im Bereich der Softwareentwicklung und wird als Erweiterung der traditionellen theoretischen und experimentellen Untersuchung angesehen.Allerdings, wie wir heute wissen, begegnen wir Hürden, wie große Mengen an zu verarbeitenden Datensätzen und die hohe Rechenkomplexität von Methoden.Um diese Simulationen in angemessener Zeit berechnen zu können, werden High Performance Computing (HPC) Techniken benötigt.Eine Methode, um die Rechenzeit zu beschleunigen, ist die parallele Verarbeitung.D. h., ein Simulationsprogramm wird über mehrere Prozesse hinweg verteilt.Verschiedene parallele Programmiermodelle sind für die effiziente parallele Implementierung von Anwendungen entwickelt worden, um die verteilte Anwendung auf paralleler Rechnerhardware auszuführen.Wegen der hohen Leistung, Portabilität und der Standardisierung für modernste Computersysteme wird Message Passing Interface (MPI) als das Kommunikationsmodel angesehen.Auf der anderen Seite ist die Implementierung einer effizienten MPI Applikation herausfordernd.Die Entstehung von Multiund Many-Core Prozessoren beschleunigt den explosiven Wachstum der Rechenleistung in der aktuellen High-End System Generation.In diesem Fall wird die Kommunikationsgeschwindigkeit ein begrenzender Faktor der Performance in verteilten Simulation.Um diesen Engpass zu mildern, ist ein angemessenes Kommunikationssystem aus der bestehenden Hochleistung Parallele Programmierung vonnöten.In dieser Dissertation untersuche ich die Vorgehensweise eines effizienten PGAS Laufzeit-Kommunikationssystem auf Basis des MPI-3 RMA Modells.Hierbei wird die MPI Komplexität vor dem Benutzer verborgen.Um das hierarchische Speicherlayout von HPC System mit in Betracht zu ziehen, verwende ich eine hybride Strategie, welche das traditionelle MPI RMA Modell ( für Inter-Node Kommunikation) mit der, in MPI integrierte, Shared Memory Parallelisierung (für Intra-Node Kommunication) kombiniert.Um eine höhere Überlappung zwischen Kommunikation und Berechnung zu erreichen, verwende ich den Progress-prozessbasierten Ansatz, dadurch findet der asynchrone Progress der Kommunikation unabhängig von den MPI-Synchronisation statt.Mit Applikation Benchmarks sowie Kommunikation-Mikro-Benchmarks wurden die Leistungssteigerungen des im Rahmen dieser Dissertation entwickelte Kommunikationsschemas ausgewertet.Die Auswertung von Mikro Benchmark zeigt, im Fall von Intra-Node Kommunikation, eine bessere Leistung als MPI RMA.Auch kann gezeigt werden, dass die Leistung für Inter-Node Kommunikation, mit vernachlässigbaren Overhead, im Vergleich zu MPI RMA gehalten werden kann.Der wichtigste Aspekt ist jedoch die erhebliche Leistungssteigerung des vorgeschlagenen Kommunikationsschemas für echte Engineering Anwendungen.Darüber hinaus zeigt im Intra-Node Mirkro-Benchmark Vergleich zu anderen RMA-Bibliotheken das vorgeschlagene Kommunikationsschema einen klaren iii iv Vorteil und kann im Bezug auf Anwendungen auch bei Inter-Node Kommunikation mit den RMA-Bibliotheken mithalten.Stuttgart (HLRS) for all his support and instructive comments that are necessary for accomplishing my dissertation.I would like to thank my advisor Dr. Jose Gracia for his patient guidance during the course of my PhD study.I appreciate the considerable number of time and energy that he has devoted to my dissertation.Not only has he provided me invaluable advices and encouragements that led to the completion of my dissertation, but he also