Timings of an Unstructured-Grid CFD Code on Common Hardware Platforms and Compilers
Fernando E. Camelli, Rainald Löhner, Eric L. Mestreau · 45th AIAA Aerospace Sciences Meeting and Exhibit · 2007
Timings have been conducted for several benchmark testcases using the same code on a variety of common hardware platforms and compilers. The results indicate surprisingly little variation in the performance given the considerable number of vendors, architectures and compilers. The largest differences amounted to less than 30% of run-times. For the single processor runs, an increase in cache size reduced run-times, though not dramatically (approximately 10%). Going from 32 bits to 64 bits (with the same clockspeed, cache size, memory and compiler) in most cases produced a gain of 10%, although in some cases no gain was recorded. Overall, the chip with the slowest clocktime (Intel It II at 1.50 GHz) achieved the best performance. In some cases, it beat the 64-bit, 3.40 GHz P4/Xeon machines by a considerable margin. For shared memory parallelism, the best scaling was achieved by the SGI Altix. This should not come as a surprise, as SGI’s CCNUMA technology has matured over the last decade. For the AMD Opteron, the SUN compiler exhibited the best scaling.