A RISC-based architecture for real-time computation
Ronald Mraz · 1992
This thesis develops the design of a pipelined RISC processor that supports dynamic scheduling of real-time applications with little or no noticeable overhead. The processor design subsumes scheduling overhead through the concurrent execution of an Embedded Kernel within the idle cycles of applications. Essential resource management functions of event handling, context swapping and task scheduling are programmed in the Embedded Kernel, and execute in a way that does not hinder the application's pipelined execution. Traditional approaches require suspension of the application's execution in order to perform resource management functions. This in-line execution of the resource manager can degrade the schedulable utilization of real-time computation an average of 7 or 16 percent for interrupt and timer driven implementations, respectively. To realize concurrent execution of the resource manager, the new design maintains two independent register states to support dual stream execution within a single CPU pipeline. Given that a 5 stage RISC pipeline typically stalls 12 percent of the time due to predictable hazard conditions, this instruction injection approach can completely subsume the resource manager for many real-time applications without impacting performance. This is because the dual streams are complementary in nature, and coordinate rather than compete for CPU resources. Furthermore, since the scheduler is always present, the cumulative overhead of traditional systems to invoke and terminate the scheduler is no longer a factor. To demonstrate the viability of this dual stream approach, a design is proposed that is based on the MIPS (tm) R2000 RISC organization. This design is termed DART, for the Deluxe Architecture for Real-Time. Simulation case studies demonstrate that DART can improve real-time resource utilization as much as 28% while completely subsuming the kernel in the application stall cycles. This offers a potential performance gain of greater than 40% for traditional timer driven polling implementations. Furthermore, detailed organizational studies show that hardware support for the dual stream features can be integrated into a single chip implementation without degrading the CPU cycle time.