HPC Colony: Linux at Large Node Counts
Terry Jones, USDOE, Andrew Tauferner, Todd A. Inglett, Albert Sidelnik · 2007
As part of the HPC-Colony project sponsored by the Department of Energy’s FastOS program (http://www.hpc-colony.org), we are conducting research on OS issues on extremely large systems. In particular, we are actively investigating scaling issues associated with system software including certain coordination aspects of our smart runtime system and operating system, operating system jitter, and administrative activities. During the recent BGW-Day, we had a chance to test the scalability of some of our approaches in those areas. We were able to start our tests on configurations with a few thousand BG/L processors, and increase the test configuration to the full machine size, so that the scaling behavior of our techniques can be assessed. This document describes the tests we conducted, the results we obtained, and the conclusions that we could draw from such experiments. The remainder of this report is organized as follows. In Section 2 we briefly describe our motivation. Section 3 describes the experiments that we conducted and provides the scalability results on large BGW configurations. Finally, Section 4 contains our conclusions and a summary of our achievements during BGW-Day.