Sort-First Parallel Rendering with a Cluster of PCs
Rudrajit Samanta, Thomas A. Funkhouser, Kai Li · 2000
The objective of our research is to investigate whether it is possible to construct a fast and in-expensive parallel rendering system leveraging the aggregate performance of multiple commodity graphics accelerators in PCs connected by a system area network. We investigate a sort-first approach [1, 2, 3]. As shown in Figure 1, our system comprises of a client PC, servers PCs, and a display PC arranged in a three-stage pipeline. During the first stage, the client partitions the screen into tiles and assigns each of them to a different server. During the next stage, every server renders all the graphics primitives at least partially overlapping its assigned tile from a replicated scene database to form a subimage and sends the resulting pixels to the display PC. Finally, the display PC composites all the subimages in its frame buffer for display. The main challenge is to develop a dynamic screen partitioning algorithm that balances the rendering load among the servers, minimizes overheads by reducing primitive-tile overlaps, and executes at interactive rates. Our method is based on the mesh-based adaptive decomposition (MAHD) algorithm described by Mueller [2]. Since it is not possible for the client PC to transform and sort every graphics