A high-performance VLSI graphics processor

Yin-Kuan Lin · 1995

The increasing complexity of real-time graphics processing demands a high performance graphics processor which is capable of processing a large number of objects to generate high resolution graphics. The use of very large scale integrated circuits (VLSI) to pack a large number of functional units in a chip has a large potential to increase performance and offer a cost-effective solution. This dissertation presents techniques for increasing performance and the efficient use of resources in the design and implementation of a high-speed CMOS VLSI graphics processor for drawing high resolution real-time dynamic graphics at three levels: algorithm, architecture, and circuit design. At the algorithmic level, computational algorithms for different system mechanisms are analyzed to select algorithms which minimize the computation time for drawing real-time high-resolution graphics. Data flow paths are investigated to discover potential parallelism. At the architectural level, an architecture which exploits the potential parallelism by maximizing the number of data operations and I/O functions, and a set of vector instructions which minimizes the number of instructions needed to draw a given object are defined. Then, we show how instruction fetching, drawing operations, and output data can be arranged in pipeline and parallel on a VLSI chip. This design demonstrates that vector instruction set architectures with interleaved registers and functional units can offer an effective solution to increase the performance of a single chip graphics processor. At the circuit design level, logic, circuit, and layout techniques are examined to increase the clock speed of the devices. To manage the increasing complexity in chip design, designers must know what tools can increase their productivity and how to convert a design into a regular or semi-regular structure. Since there is a need for a ROM in the graphics processor, techniques for designing a ROM layout generator are investigated. Our work shows how to achieve highly compacted layout in a very small amount of execution time by using a simple program. This approach can easily be expanded to include some other features for regular or semi-regular structures such as decoders, encoders, RAMs, adders, and multipliers.

Read the paper · More papers on PaperTik