A Small Multi-Threaded Microkernel for Symmetrical Multiprocessing Hardware Architectures

Dominik Moser · 1999

A symmetric multiprocessor (SMP) is the most commonly used type of multiprocessor system. All CPUs and I/O devices are tightly coupled over a common bus, sharing a global main memory to which they have symmetric and equal access. Compared to a uniprocessor system, an SMP system imposes a high demand for memory bus bandwidth. The maximum number of CPUs that can be used for practical work is limited by this bandwidth, being proportional to the number of processors connected to the memory bus. To reduce memory bus bandwidth limitations, an SMP implementation should use a secondary cache and a cache-consistency protocol. Since the need to synchronize access to shared memory locations is so common on SMP systems, most implementations provide the basic hardware support for this through atomic read-modify-write operations. However, since the sequential memory model does not guarantee a deterministic ordering of simultaneous reads and writes to the same memory location from more than one CPU, any shared data structure can cause a race condition to occur. Such nondeterministic behavior can be catastrophic to the integrity of the OS kernel’s data structures and must be prevented. The operating system for a multiprocessor must be designed to coordinate simultaneous activity by all CPUs and maintain system integrity. To simplify their design, many uniprocessor kernel systems have relied on the fact that there is never more than one process executing in the kernel at the same time. However, this policy fails on SMP systems when kernel code can be executed simultaneously on more than one processor. Therefore, a uniprocessor kernel cannot be run on an SMP system without modification. The easiest way to maintain system integrity within an SMP kernel is with spin locks. Spin locks work correctly for any number of processors but are only efficient if the associated critical section is short. Overall system performance will be lowered if the processors spend too much time waiting to acquire locks or if too many processors frequently contend for the same lock. To reduce lock contention, the kernel has to use different spin locks for different critical sections. Overall system performance can be significantly improved by allowing parallel kernel activity on multiple processors. The amount of concurrent kernel activity that is possible across all the processors is partially determined by the degree of multithreading. A coarse-grained implementation uses few locks, whereas a fine-grained implementation protects different data structures with different locks. The goal is to ensure that unrelated processes can perform their kernel activities without contention. This thesis presents TopsySMP, a small multithreaded microkernel for an SMP architecture. It is based on Topsy which has been designed for teaching purpose at the Department of Electrical Engineering at ETH Zurich. It consists of at least three independent kernel modules each represented by a control thread: the thread manager, the memory manager, and the I/O manager. Besides, there are a number of device driver threads, and a simple command-line shell. TopsySMP provides parallel thread scheduling, and allows threads to be pinned to specific processors. The number of available CPUs is determined upon startup which leads to a dynamic configuration of the kernel. The uniprocessor system call API is preserved, therefore all applications written for the original Topsy can run on the SMP port. This thesis shows that the implementation of an SMP kernel based on a multithreaded uniprocessor kernel is straightforward and results in a well structured and clear design. The amount of concurrent kernel activity is determined by the degree of multithreading and leads to a coarse-grained implementation using only a few spin locks. The spin lock times for a small-scale system with up to eight processors are reasonably short compared to a context switch; no other synchronization primitives are used within the kernel. The overhead caused by the integration of SMP functionality was kept to a minimum, resulting in a small and efficient kernel implementation. Furthermore, a suggestion is made on how to improve system performance by multiplicating the control thread of a kernel module, allowing the throughput of module specific kernel activity to be ideally multiplied.

Read the paper · More papers on PaperTik