Reliability And Timing Analysis of Distributed systems
Hans Arne Hansson, Christer Norström · 2000
Summary Modelling and analysis are important tools in the development of safety critical real-time systems. The introduction of state-of-the-art analysis techniques in industry is however rather slow. One reason for this is the pessimism in models and analysis, e.g., schedulability analysis for realistic systems are typically based on simplifying assumptions which leads to pessimism that forces designers to make costly overdesigns, dimensioning the system for worst-case situations that may never occur. At the same time, the over all system requirement is to satisfy a reliability measure of, say, at most 10 -9 faults per hour. This project proposes a reliability analysis method that considers the effects of faults and timing parameter distributions (including execution time distributions, jitter distributions, and sporadic task inter-arrival time distributions) on schedulability analysis. The goal is to provide designers with well founded support that allow them to make trade-offs between timing guarantees and reliability, i.e. by allowing occasional deadline misses a less costly implementation may be used, while still satisfying the over all reliability requirement. Furthermore, it is well-known in industry that a missed deadline in most cases will not lead to a failure. This and other properties of executions will be considered. The work will have both theoretical and practical impact, since it will evaluate a new approach for integrating reliability modelling and schedulability analysis, and enable industry to use modern analysis techniques both to model components and systems in early design phases and to validate these models by measurements on the developed components/system. We apply for funding of one graduate student to develop a method for integrating schedulability analysis in reliability modelling. The result will be a unified framework for holistic analysis of systems' timing and reliability behaviour.