A Software-Hardware Integration Framework for Multiprocessor System-On-Chip Based on Tightly-Coupled Thread Model
Ullah Khan, Arif Ullah Khan · Institutional Repositories DataBase (IRDB) · 2014
The improvements in process technology is still keeping Moore law valid and increasing number of gate count at lower cost, enabling systems with increased functionality and complexity.With the growing complexity of consumer embedded products and the improvements in process technology, multiprocessor system-on-chip (MPSoC) architectures have become widespread.Future MPSoCs will be truly heterogeneous which will not only include multiple processors but also multiple dedicated hardware accelerators too.Designing and Implementation of such systems is a challenge as developers attempt to scale conventional methods of using register-transfer-level hardware description languages to build their circuits.This problem is commonly referred to as the design gap -the dierence between the number of gates available and developer productivity.As Moore's law outpaces advancements in design tools and methodologies, this gap widens.With increased gate count and performance gains of new devices, the movement to a higher level of abstraction is permitted.The goal of our research is to ll the design gap by developing a truly heterogeneous MPSoC framework to address various challenges in the MPSoC design e.g.software programming model, HW accelerator design & integration, fast performance estimation, system verication and design space exploration (DSE).In this thesis we will focus on the following design challenges in MPSoC design :• HW accelerator design & integration• Fast performance estimation • Ecient design space exploration for HW/SW based systemsThe EDA industry has identied the need for tools and methodologies for synthesis of logic from high-level languages so that designers no longer need to concern themselves with low-level circuit implementation details.Several tools known as high level synthesis tools (HLS) have emerged in the past decade that attempt to address this need.By using these HLS tools the dedicated hardware accelerators can be generated from software programs, written in high-level languages like `C'.One of the problems in integrating these HW blocks is the communication and synchronization between the HW/SW which is usually handled by the system designer.Conventional methodologies for generating HW accelerators using HLS were based on uni-processor models.Uni-processor systems have a single process and HW & SW interact in a sequential manner.In such systems the issue of communication and synchronization is not so complex.Now in the era of multiprocessor systems and parallel programming environment, with several concurrent processes communicating with each other concurrently, communication & synchronization become more complex.In our research we addressed this issue and developed a HW/SW design methodology based on the TCT model by integrating a HLS tool in our MPSoC design framework which enables HW or SW implementation for each of the concurrent process.Our approach eliminates the manual synchronization between HW & SW and relieves MPSoC designer from manual coding of the HW blocks.We have tested our HW integration method for JPEG encoder and achieved 17 % performance improvement over the SW only MPSoC solution.Another challenge in the future MPSoC design is to get fast & accurate system performance estimation in order to select best HW/SW partitioning.Traditional techniques of getting performance estimation of a system e.g.HW/SW co-simulation are very slow and time consuming for DSE.There is a strong need for methodologies that quickly and accurately estimate the performance of complex systems.We have developed a novel system level fast and accurate performance estimation method for exploring the trade-o between hardware and software implementations in MPSoCs.Our method is based on workload simulation driven by program execution traces encoded in the form of branch bitstreams.The key feature of our performance estimation is the unied timing model, in the form of a program trace graph (PTG) i for both software executions on processors as well as the hardware blocks (nite state machines) synthesized by a HLS tool.Our methodology allows highly accurate performance estimation under the existence of data dependent behavior of software and hardware components.Due to the increase complexity of the MPSoC, fast and accurate DSE for best system performance at early stage of the design process is desired.Any DSE solution is desired to provide best system partitioning scheme for best performance with ecient area utilization.In this research we propose a DSE framework for heterogeneous MPSoC based on tightly-coupled thread (TCT) parallel programing model which can handles system partition exploration and HW synthesis exploration.The proposed framework drastically reduces the exponential size design space into near-linear size by utilizing the accurate HW timing models as the indicator for system bottleneck and guiding the enumeration process of HW version combinations.Experimental results shows the accuracy of the proposed method with an average estimation error of 1.38% for HW timing of each thread, and 2.80% estimation error for the systemlevel simulation, where the simulation speedup factor was in the order of 5,000 times. Currently the proposed framework partially depends on a high level synthesis (HLS)tool eXCite, but other HLS tools can be easily integrated into the proposed framework.