Using Variable Conflict Granularity to Improve the Performance of Transactional Memory Support for Games
Mihai Burcea · TSpace (University of Toronto) · 2015
Transactional memory (TM) is a promising parallel programming paradigm for generic applications on large scale parallel architectures. In spite of significant research work in this area in the past decade, adoption by the parallel programming community has been slow due to two reasons. First, performance of most TM support for applications has been disappointing. Second, efforts towards using TM to parallelize realistic applications or highly popular benchmarks for TM have been rare. It is widely known that two factors can have a drastic impact on the performance of parallel applications with Transactional Memory (TM) support: i) contention among execution threads, particularly contention due to false sharing, and ii) the software TM runtime conflict tracking overhead. We propose techniques to address application contention and reduce TM conflict tracking overheads by varying the memory access tracking granularity in the TM. To further optimize performance of TM support, we leverage a software transactional memory platform called HyTM based on Intel’s TM hardware support, Transactional Synchronization Extensions (TSX) in Haswell. We build on existing efforts towards developing a gaming benchmark for TM, by characterizing and optimizing the contention patterns in a realistic application, SynQuake, developed based on the open source Quake 3 code. We port SynQuake to HyTM and design, prototype and evaluate adaptive techniques for varying the conflict tracking granularity of the TM, on the fly, for SynQuake. For regions of the application space causing high degrees of contention, we reduce the tracking granularity while, conversely, we allow a coarse tracking granularity to be used for application regions with low degrees of contention, in order to decrease tracking overhead and transaction size. Our prototype implements our variable granularity adaptations either by using application-specific knowledge, or, completely transparently, by relying primarily on the support provided by a TM library. Our evaluation shows that our techniques improve performance by 16.7%, on average, compared to a range of configurations using static parameter settings. More importantly, we show that our techniques are lightweight and can provide statistics that allow both application and user to understand application contention patterns at the cost of negligible runtime overhead.