Design and analysis of reliable corba-based software systems
Tong Luo, Kishor Shridharbhai Trivedi · 2000
The Common Object Request Broker Architecture (CORBA) has become a standard for distributed computing systems. With the success of many commercial software systems based on CORBA, more and more mission and performance critical applications are also adopting CORBA as the system architecture. But the reliability and performance issues of CORBA-based software are yet to be addressed. This dissertation studies the reliability and performance problems involved in constructing and analyzing CORBA-based software systems. We first develop new approaches to building reliable CORBA naming servers, event channels, and other critical business objects, and compare our approaches with the ones available in the research literature. We analyze the dependability of two different systems. For a system with independent repair facilities, we present a fault-tree model and develop closed-form solutions for the reliability function, MTTF, and steady-state availability. For a system with shared repair facility per subsystem, we present a hierarchical fault-tree and Markov model, and solve it using the SHARPE package. We show that the fault-tree can help to overcome the state explosion problem of Markov models when analyzing complex systems. We then analyze a fault-tolerant event channel design without the event backup queue, assuming Poisson arrivals. We derive closed-form formulae for the probability of incoming event rejection, the probability of event crash when the event channel fails, the mean number of events crashed when the event channel fails, and so on. Transient results are obtained using stochastic reward net (SRN) models. We also describe how to analyze a complex failure model using a hierarchical aggregation approach. Furthermore, we analyze the effect of the event backup queue on our fault-tolerant event channel design with more complex arrivals. We define the event rejection loss and the event crash loss, design a mixed FIFO queue, and present two propositions that capture the conditions when an event rejection loss or an event crash loss occurs. Both the transient and steady-state measures of the probability of event rejection loss, the probability of event crash loss, and the mean number of event crash losses, etc., are obtained using SRN models. The results of our reliable event channel are useful during system design in order to meet certain quality of service requirements. But a major drawback of the mixed FIFO queue SRN model is that its state space grows exponentially with its size. This makes it difficult to solve larger systems using the SRN models. Using the knowledge that a fault-tree approach can help overcome the state space explosion problem of Markov models, we try to develop more efficient combinatorial. algorithms for reliability analysis. For coherent systems, we develop an improved algorithm (named LVT) based on the well known VT algorithm, using multiple variable inversion (MVI) techniques, and show it is the best in the class. For non-coherent systems, which are usually represented as fault-trees, we develop the first MVI algorithm (named LT), and show it is faster and generates fewer terms than a single variable inversion (SVI) algorithm. LVT and LT apply to general network reliability problems. Specifically INT can be used to solve individual fault-trees or hierarchical models involving fault-trees and Markov chains as in our reliable CORBA-based software systems.