Planning for the Un-plannable: Redundancy, Fault Protection, Contingency Planning and Anomaly Response for the Mars Reconnaissance Orbiter Mission
Todd J. Bayer · 2007
For interplanetary spacecraft the round trip travel time for electromagnetic waves ranges from several tens of minutes to many hours depending on their distance to Earth. The round trip light time for communications with Mars Reconnaissance Orbiter (MRO) can be up to 40 minutes. With this latency, a variety of failures onboard the spacecraft could result in loss of the spacecraft before ground controllers could respond. These spacecraft must therefore be able to autonomously diagnose and fix time-critical failures. For those failures that onboard fault protection cannot diagnose or fix, ground controllers must be prepared to intervene. Using the actual MRO in-flight anomalies experienced to date, the complementary roles of redundancy, on-board fault protection software and ground-based anomaly response are examined to show how they provide a robust and multi-layered safety net for the mission. I. Introduction ars Reconnaissance Orbiter’s (MRO) crucial role in the long term strategy for Mars exploration requires a high level of reliability during its 5.4 year mission. This requires an architecture which incorporates extensive redundancy and cross-strapping. The overall MRO architecture is discussed in this context. M Because of the latency due to round trip light time (RTLT), many possible failures could cause spacecraft and hence mission failure before ground controllers could react. Therefore, in order to make use of MRO’s hardware redundancy, it must itself be able to respond to these mission-critical failures. The architecture of MRO’s semiautonomous fault protection software, known as the Spacecraft Imbedded Distributed Error Response (SPIDER), is described in the context of each phase of the mission, with emphasis on its key role as ‘first responder’ in detecting and responding to potentially threatening situations on board the spacecraft. A key aspect of this ‘first response’ is establishing a stable, power positive, thermally safe and commandable configuration termed “safe mode”. On-board fault protection software is still only semi-autonomous. Assuming it has successfully established safe mode, the ground must then take over to complete the recovery back to nominal operations. The ground would also need to intervene when fault protection is unable to recognize a potentially threatening condition, either due to known limitations or software flaws. In any of these cases it is crucial to have well thought-out plans for how the ground should proceed. Many of the commands the ground might need to send are known beforehand, and these contingency plans incorporate all the commands which might be required, fully tested and ready for uplink. The set of MRO contingency plans is discussed in terms of the relationship of each plan to the on-board fault protection responses. When anomalies actually happen, all of this redundancy, software and planning is put to the test. MRO has experienced—and survived—several significant anomalies in its mission so far, including command errors, flight software bugs and hardware failures. Each of these significant anomalies is examined in terms of the parts of the overall safety net that came into play, the root cause and lessons learned.