COTS Cluster-based Sort-last Rendering: Performance Evaluation and Pipelined Implementation
Xavier Cavin, Christophe Mion, A. Filbois · 2006
Figure 1: Views of the head section (512x512x209) of the visible female CT data with 16 nodes (a space has been left between the subvolumes to highlight their boundaries). Using a 3 years old 32-node COTS cluster, a volume dataset can be rendered at constant 13 frames per second on a 1024×768 rendering area using 5 nodes. On a 1.5 years old, fully optimized, 5-node COTS cluster, the frame rate obtained for the same rendering area reaches constant 31 frames per second. We truly expect our future work, including further algorithm optimizations and hardware tuning on a modern PC cluster, to provide higher frame rates for bigger datasets (using more nodes) on larger rendering areas. Sort-last parallel rendering is an efficient technique to visualize huge datasets on COTS clusters. The dataset is subdivided and dis-tributed across the cluster nodes. For every frame, each node ren-ders a full resolution image of its data using its local GPU, and the images are composited together using a parallel image composit-ing algorithm. In this paper, we present a performance evaluation of standard sort-last parallel rendering methods and of the different improvements proposed in the literature. This evaluation is based on a detailed analysis of the different hardware and software com-