Hyper-Threading Technology: Impact on Compute-Intensive Workloads

William R. Magro, Sanjiv Shah · 2002

Intel’s recently introduced Hyper-Threading Technology promises to increase applicationand system-level performance through increased utilization of processor resources. It achieves this goal by allowing the processor to simultaneously maintain the context of multiple instruction streams and execute multiple instruction streams or threads. These multiple streams afford the processor added flexibility in internal scheduling, lowering the impact of external data latency, raising utilization of internal resources, and increasing overall performance. We compare the performance of an Intel Xeon processor enabled with Hyper-Threading Technology to that of a dual Xeon processor that does not have HyperThreading Technology on a range of compute-intensive, data-parallel applications threaded with OpenMP. The applications include both real-world codes and handcoded “kernels” that illustrate performance characteristics of Hyper-Threading Technology. The results demonstrate that, in addition to functionally decomposed applications, the technology is effective for  Intel is a registered trademark of Intel Corporation or its subsidiaries in the United States and other countries.  Xeon is a trademark of Intel Corporation or its subsidiaries in the United States and other countries. 1 OpenMP is an industry-standard specification for multithreading data-intensive and other highly structured applications in C, C++, and Fortran. See www.openmp.org for more information. many data-parallel applications. Using hardware performance counters, we identify some characteristics of applications that make them especially promising candidates for high performance on threaded processors. Finally, we explore some of the issues involved in threading codes to exploit Hyper-Threading Technology, including a brief survey of both existing and still-needed tools to support multi-threaded software development. INTRODUCTION While the most visible indicator of computer performance is its clock rate, overall system performance is also proportional to the number of instructions retired per clock cycle. Ever-increasing demand for processing speed has driven an impressive array of architectural innovations in processors, resulting in substantial improvements in clock rates and instructions per cycle. One important innovation, super-scalar execution, exploits multiple execution units to allow more than one operation to be in flight simultaneously. While the performance potential of this design is enormous, keeping these units busy requires super-scalar processors to extract independent work, or instructionlevel parallelism (ILP), directly from a single instruction stream. Modern compilers are very sophisticated and do an admirable job of exposing parallelism to the processor; nonetheless, ILP is often limited, leaving some internal processor resources unused. This can occur for a number of reasons, including long latency to main memory, branch mis -prediction, or data dependences in the instruction stream itself. Achieving additional performance often requires tedious performance Intel Technology Journal Q1, 2002. Vol. 6 Issue 1. Hyper-Threading Technology: Impact on Compute-Intensive Workloads 2 analysis, experimentation with advanced compiler optimization settings, or even algorithmic changes. Feature sets, rather than performance, drive software economics. This results in most applications never undergoing performance tuning beyond default comp iler optimization. An Intel processor with Hyper-Threading Technology offers a different approach to increasing performance. By presenting itself to the operating system as two logical processors, it is afforded the benefit of simultaneously scheduling two potentially independent instruction streams [1]. This explicit parallelism complements ILP to increase instructions retired per cycle and increase overall system utilization. This approach is known as simultaneous multi-threading, or SMT. Because the operating system treats an SMT processor as two separate processors, Hyper-Threading Technology is able to leverage the existing base of multithreaded applications and deliver immediate performance gains. To assess the effectiveness of this technology, we first measure the performance of existing multi-threaded applications on systems containing the Intel Xeon processor with Hyper-Threading Technology. We then examine the system’s performance characteristics more closely using a selection of hand-coded application kernels. Finally, we consider the issues and challenges application developers face in creating new threaded applications, including existing and needed tools for efficient multi-threaded development. APPLICATION SCOPE While many existing applications can benefit from Hyper-Threading Technology, we focus our attention on single-process, numerically intensive applications. By numerically intensive, we mean applications that rarely wait on external inputs, such as remote data sources or network requests, and instead work out of main system memory. Typical examples include mechanical design analysis, multi-variate optimization, electronic design automation, genomics, photo-realistic rendering, weather forecasting, and computational chemistry. A fast turnaround of results normally provides significant value to the users of these applications  Intel is a registered trademark of Intel Corporation or its subsidiaries in the United States and other countries.  Xeon is a trademark of Intel Corporation or its subsidiaries in the United States and other countries. through better quality products delivered more quickly to market. The data-intensive nature of these codes, paired with the demand for better performance, makes them ideal candidates for multi-threaded speed-up on shared memory multi-processor (SMP) systems. We considered a range of applications, threaded with OpenMP, that show good speed-up on SMP systems. The applications and their problem domains are listed in Table 1. Each of these applications achieves 100% processor utilization from the operating system’s point of view. Despite external appearances, however, internal processor resources often remain underutilized. For this reason, these applications appeared to be good candidates for additional speed-up via Hyper-Threading Technology. Table 1: Applications type

Read the paper · More papers on PaperTik