An engineering study of onboard checkout techniques. Task 5: Subsystem level failure modes and effects
Failure effects analysis of guidance, navigation, and control, data management, and communications subsystems
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Failure effects analysis of guidance, navigation, and control, data management, and communications subsystems
The triply redundant intercomputer network for the Advanced Information Processing System (AIPS), an architecture developed to serve as the core avionics system for a broad range of aerospace vehicles, is discussed. The AIPS intercomputer network provides a high-speed, Byzantine-fault-resilient communication service between processing sites, even in the presence of arbitrary failures of simplex and duplex processing sites on the IC network. The IC network contention poll has evolved from the Laning Poll. An analysis of the failure modes and effects and a simulation of the AIPS contention poll, demonstrate the robustness of the system.
The survivable, adaptable fiber optic embeddable network (SAFENET) is a draft standard for local area networking (LAN) developed by the Navy which, when adopted, will become a military standard. The standard is being developed for procurement specifications of computer resources to be used on ships and aircraft and has some of the real-time concerns that network standards for space vehicles have. Architecture and survivability are considered. It is noted that the token-ring LAN must implement the IEEE 802.5 recommended practice for dual ring reconfiguration, which is currently being reviewed for inclusion into the IEEE standard. A trunk coupling unit is used at each station to isolate a station from the ring in case of failure. Up to five stations can be bypassed in this fashion. Communication architecture has an OSI profile but differs from the standard concept of the seven layers by allowing alternate suits and breaking the layers into three groupings of services to allow for physical interfacing. It also provides several paths, even if only one profile is used. Management and synchronization protocols are discussed and security issues are addressed. Implications for aerospace applications are considered and it is projected that interoperability with the Navy and other U.S. Government systems may require SAFENET specifications for NASA systems.
A connected hypercube with faulty links and/or nodes is called an injured hypercube. To enable any non-faulty node to communicate with any other non-faulty node, information on component failures has to be made available to non-faulty nodes so as to route messages around the faulty components. A distributed adaptive fault tolerant routing scheme is proposed in which each node is required to know only the condition of its own links. This scheme is shown to be capable of routing messages successfully as long as the number of faulty components is less than n (the dimension of the hypercube), and to route messages via shortest paths with a rather high probability. A second routing scheme based on depth-first search is proposed which works in the presence of an arbitrary number of faulty components; however, the paths chosen by this may not always be the shortest. To guarantee shortest paths, every mode must be given information beyond that on its own links; the additional information to be kept at each node for shortest-path routing is determined. Several examples are given to illustrate the results.
Fault tolerance in future processing and switching communication satellites is addressed by showing new methods for detecting hardware failures in the first major subsystem, the multichannel demultiplexer. An efficient method for demultiplexing frequency slotted channels uses multirate filter banks which contain fast Fourier transform processing. All numerical processing is performed at a lower rate commensurate with the small bandwidth of each bandbase channel. The integrity of the demultiplexing operations is protected by using real number convolutional codes to compute comparable parity values which detect errors at the data sample level. High rate, systematic convolutional codes produce parity values at a much reduced rate, and protection is achieved by generating parity values in two ways and comparing them. Parity values corresponding to each output channel are generated in parallel by a subsystem, operating even slower and in parallel with the demultiplexer that is virtually identical to the original structure. These parity calculations may be time shared with the same processing resources because they are so similar.
System diagnosis is an integral part of any Integrated System Health Management application. Diagnostic applications make use of system information from the design phase, such as safety and mission assurance analysis, failure modes and effects analysis, hazards analysis, functional models, fault propagation models, and testability analysis. In modern process control and equipment monitoring systems, topological and analytic , models of the nominal system, derived from design documents, are also employed for fault isolation and identification. Depending on the complexity of the monitored signals from the physical system, diagnostic applications may involve straightforward trending and feature extraction techniques to retrieve the parameters of importance from the sensor streams. They also may involve very complex analysis routines, such as signal processing, learning or classification methods to derive the parameters of importance to diagnosis. The process that is used to diagnose anomalous conditions from monitored system signals varies widely across the different approaches to system diagnosis. Rule-based expert systems, case-based reasoning systems, model-based reasoning systems, learning systems, and probabilistic reasoning systems are examples of the many diverse approaches ta diagnostic reasoning. Many engineering disciplines have specific approaches to modeling, monitoring and diagnosing anomalous conditions. Therefore, there is no "one-size-fits-all" approach to building diagnostic and health monitoring capabilities for a system. For instance, the conventional approaches to diagnosing failures in rotorcraft applications are very different from those used in communications systems. Further, online and offline automated diagnostic applications are integrated into an operations framework with flight crews, flight controllers and maintenance teams. While the emphasis of this paper is automation of health management functions, striking the correct balance between automated and human-performed tasks is a vital concern.
An estimated 45 communications satellites will be in geostationary orbit in the 1980s. Based on past history, half of the subsystem failures will be due to design failures. Methods now used to achieve present availabilities are summarized in this paper. To increase the availability of a single satellite to 99.99%, new techniques are needed; several possibilities are suggested herein. New developments, including nickel-hydrogen batteries, magnetic bearings, and solid-state amplifiers, are possible solutions to long standing problem areas in communications satellites. Launching a satellite two or three years before the system is needed is one way to reduce design failures in subsequent satellites. Diversity of manufacture can be used to obtain the maximum advantage from redundant elements. Finally, unmanned module exchange at geostationary orbit has been shown to be technically feasible.
Consideration is given to some of the negative aspects of the trend toward increased automation of aircraft flight decks. The history of automated devices for navigation, communications and detection on board aircraft is reviewed. Instances of automatic system failure are identified which have led to accidents, and the events surrounding the downing of Korean Airlines Flight 747 are reexamined within the context of a computer-based system failure. Finally, new software and interactive systems to reduce navigational error due to inadequate computer-assisted flight instruction (CAI) are described, with emphasis given to speech processing and intelligent CAI systems.
Rovers will play a critical role in the exploration of Mars. Near-term mission plans call for long traverses over unknown terrain, robust navigation and instrument placement, and reliable operations for extended periods of time. Longer-term missions may visit multiple science sites in a single day and perform opportunistic science data collection, as well as complex scouting, construction, and maintenance tasks in preparation for an eventual human presence. The Pathfinder mission demonstrated the potential for robotic Mars exploration but at the same time indicated the need for more rover autonomy. The highly ground-intensive control with infrequent communication and high latency limited the effectiveness of the Sojourner rover. When failures occurred, Sojourner often sat idle for extended periods of time, awaiting further commands from earth. In future missions, the tasks will be more complex and extended; hence there will be even more situations where things do not go exactly as planned. Significant advances in rover autonomy are needed to cope with increasing task complexity and greater execution uncertainty. Towards this end, we have designed an on-board executive architecture that incorporates robust operation, resource utilization, and failure recovery. In addition, we have designed ground tools to produce and refine contingent schedules that take advantage of the on-board architecture's flexible execution characteristics. Together, the on-board executive and the ground tools constitute an integrated rover autonomy architecture. This work draws from our experience with the Deep Space One autonomy experiment, with enhancements to ensure robust operation in the face of the unpredictable, complex environment that the rover will encounter on Mars. The rover autonomy architecture is currently being developed and deployed on the Marsokhod rover platform at NASA Ames Research Center. The capabilities of the rover autonomy architecture to support autonomous operations will be demonstrated concretely in upcoming field tests.
Good requirements are the first step for good communications, and good communications are central to insure an understanding between the customer and contractor. Failure to generate good requirements is unfortunately commonplace and repeated. Waivers to requirements are discussed from a risk based point of view. The assumption that every requirement will eventually be waived is used to establish a critical review of a draft safety requirement. Validation methods of requirements are addressed. Value added that safety requirements contribute to the Project is estimated to further our critical review of draft requirements.
NASA is currently building the Space Launch System (SLS) Block-1 launch vehicle for the Exploration Mission 1 (EM-1) test flight. The next evolution of SLS, the Block-1B Exploration Mission 2 (EM-2), is currently being designed. The Block-1 and Block-1B vehicles will use the Powered Explicit Guidance (PEG) algorithm. Due to the relatively low thrust-to-weight ratio of the Exploration Upper Stage (EUS), certain enhancements to the Block-1 PEG algorithm are needed to perform Block-1B missions. In order to accommodate mission design for EM-2 and beyond, PEG has been significantly improved since its use on the Space Shuttle program. The current version of PEG has the ability to switch to different targets during Core Stage (CS) or EUS flight, and can automatically reconfigure for a single Engine Out (EO) scenario, loss of communication with the Launch Abort System (LAS), and Inertial Navigation System (INS) failure. The Thrust Factor (TF) algorithm uses measured state information in addition to a priori parameters, providing PEG with an improved estimate of propulsion information. This provides robustness against unknown or undetected engine failures. A loft parameter input allows LAS jettison while maximizing payload mass. The current PEG algorithm is now able to handle various classes of missions with burn arcs much longer than were seen in the shuttle program. These missions include targeting a circular LEO orbit with a low-thrust, long-burn-duration upper stage, targeting a highly eccentric Trans-Lunar Injection (TLI) orbit, targeting a disposal orbit using the low-thrust Reaction Control System (RCS), and targeting a hyperbolic orbit. This paper will describe the design and implementation of the TF algorithm, the strategy to handle EO in various flight regimes, algorithms to cover off-nominal conditions, and other enhancements to the Block-1 PEG algorithm. This paper illustrates challenges posed by the Block-1B vehicle, and results show that the improved PEG algorithm is capable for use on the SLS Block 1-B vehicle as part of the Guidance, Navigation, and Control System.
Aerospace electrical systems are required to withstand and adequately operate in extremely harsh environments that include, for example, high radiation exposure, temperature extremes, intense vibrational stress and drastic temperature cycling. The nature of aerospace electronics also demands high reliability since, with very few exceptions, there is no chance for hardware servicing or repairs. Common risk mitigation techniques for this type of situation are to perform a Reliability Analysis of the system throughout the development cycle, and to use electrical components that are regarded as “high reliability” because of additional controls and requirements applied in their design, manufacturing and testing. Unfortunately, studies have shown that even though these techniques are used, many systems fail to meet mission requirements well before the predicted lifetimes. This paper presents the analysis of failures of electrical parts, experienced during various stages of system development, at NASA Goddard Space Flight Center, Greenbelt MD, between the years 2001 and 2013. These components were subjected to qualification, screening and testing in which the goal was to ensure that the components would survive the stresses of the mission. The analysis categorizes failures by part type and failure mechanisms. One of the results of the analysis was the realization that a surprising proportion of failures experienced during system integration and testing were caused by human error (i.e. human induced defect). Further analysis included the determination of root failure mechanisms and any influencing factors contributing to these failures. The major causes of these defects were attributed to electrostatic damage (ESD), electrical overstress (EOS), mechanical overstress (MOS), and thermal overstress (TOS). Finally, the study proposes a risk analysis tool which incorporates these major causes for the failures, termed error-producing conditions (EPCs), and a proportionality factor representing the number of each type of failure that has occurred at the facility under study. These factors are quantified and used to communicate the risk of human induced defects for the assembly, integration and testing of space hardware based on the system’s electrical parts list. The new risk identification can trigger risk-mitigating actions more effectively, based on the presence of component categories or other hazardous conditions that have a history of failure due to human error.
National Aeronautics and Space Administration’s (NASA) Goddard Space Flight Center (GSFC) operates a constellation of ten geosynchronous Tracking and Data Relay Satellites (TDRS). The TDRS constellation consists of multiple geosynchronous communication relay satellites located around the equator so they can provide continual coverage of any mission in low earth orbit. The TDRS are located primarily in three oceanic regions around the earth. NASA’s White Sands Complex provides the ground communication support for TDRS located over the Atlantic and Pacific Oceans. Another TDRS ground station in Guam supports the TDRS over the Indian Ocean. With these satellites the TDRS network can provide continuous coverage of satellites in low-earth orbit. The NASA Space Network (SN) project office at GSFC manages the constellation of spacecraft. Major customers of the TDRS constellation include, but are not limited to, the International Space Station and the Hubble Space Telescope. The TDRS constellation has three generations of satellites and has been active for over 30 years providing reliable communication links between customer satellites and corresponding ground stations. However, one of the major concerns for TDRS, and in any space mission, is to ensure the health and safety of the spacecraft. Generally, engineers use telemetry data to monitor and analyze the performance and state of health of the spacecraft. Telemetry data contains hundreds of parameters that monitor each important component in the spacecraft, which can be utilized to recognize and characterize the behavior of the spacecraft. Each parameter contains considerable information to represent time-dependent properties of each spacecraft subsystem and component. During the entire life of a TDRS spacecraft, thousands of gigabytes of telemetry data are transmitted in real-time from the spacecraft to the ground station at the White Sands Complex in Las Cruces, New Mexico, and recorded as historical data sets for engineers to process and analyze the events that occurred on-orbit. These parameters contain the function of multiple spacecraft subsystems, such as the attitude control system (ACS), Thermal, Electrical Power Subsystem (EPS), etc. . The first and second generations have exceeded their required lifetime and NASA is keen to manage these spacecrafts carefully in order to maximize the remaining life using the spacecraft telemetry. The challenge is to know when the risk of losing a spacecraft in geosynchronous orbit exceeds the benefit of continued operations for customer support. In the TDRS fleet, the EPS is the most critical subsystem related to spacecraft operations. Failure of the EPS would strand a spacecraft in geosynchronous orbit. Since EPS provides power to the spacecraft, component failures ultimately lead to the inability to support the spacecraft loads and the communications payload. For instance, TDRS-8 has several anomalies in EPS including the Bus Voltage Limiter (BVL) shunt current, solar array loss of circuits, and failed battery cells. Any of these anomalies can cause critical issues to the spacecraft. Therefore, developing a system to analyze and perform early detection of a potential anomaly is an important issue in telemetry data analysis. In recent years, Telemetry Mining (TM) has been proposed to process telemetry data by using Data Mining (DM) techniques such as classification, clustering, regression and anomaly detection. Anomaly detection, also known as outlier detection, has been widely used in many data mining areas such as remote sensing, medical data processing and digital image processing. The goal of anomaly detection is to detect abnormal data, which contains a relatively low probability of occurrence among the entire data set. Early detection of anomalies is one of the most significant issues in managing the spacecraft configuration. If anomalies can be detected early enough, then the redundant resources can be used to extend the life of the operational spacecraft. We present an unsupervised anomaly detection method to process the EPS data extracted from TDRS-8. This is different from traditional analytical methods, which use telemetry data to illustrate behavior and physical meaning of each spacecraft component. TM connects multiple parameters as a vector and then conducts data analysis on this high dimension telemetry vector. This method is looking at the properties of a high dimensional vector that is able to consider the relationship between different parameters in the anomaly detection problem. This kind of method performs much better than the traditional limit checking method. In addition, we propose a new approach of real-time anomaly detection to process telemetry data in real-time, which can then be applied to spacecraft monitoring with high reliability, low cost and high accuracy.
Auger spectroscopy has been used to evaluate the properties of 'good' and 'poor' impregnated tungsten cathodes used in high-power microwave wave tubes. The results were interpreted to analyze failure modes in cathodes removed from TWT's because of poor emission characteristics. Most of the poor cathodes evaluated in this program were obtained from fabricated electron guns that had been employed and discarded from the 200-W TWT tubes developed for the Communication Technology Satellite program. The results of these measurements have shown there are at least two types of failure modes that one observes with poor cathodes. They are (1) chemical contamination of the cathode surface and (2) low partial layer barium coverage of the cathode surface.
Today's launch vehicles complex electronic and avionic systems heavily utilize the Field Programmable Gate Array (FPGA) integrated circuit (IC). FPGAs are prevalent ICs in communication protocols such as MIL-STD-1553B, and in control signal commands such as in solenoid/servo valves actuations. This paper will demonstrate guidelines to estimate FPGA failure rates for a launch vehicle, the guidelines will account for hardware, firmware, and radiation induced failures. The hardware contribution of the approach accounts for physical failures of the IC, FPGA memory and clock. The firmware portion will provide guidelines on the high level FPGA programming language and ways to account for software/code reliability growth. The radiation portion will provide guidelines on environment susceptibility as well as guidelines on tailoring other launch vehicle programs historical data to a specific launch vehicle.
The design is described of the Venus probe windows, which are required to measure solar flux, infrared flux, aureole, and cloud particles. Window heating and structural materials for the probe window assemblies are discussed along with the magnetometer. The command lists for science, power and communication requirements, telemetry sign characteristics, mission profile summary, mass properties of payloads, and failure modes are presented.
To enable communication between spacecraft operating in a formation or small constellation, a mesh network architecture was developed and tested using a time division multiple access (TDMA) communication scheme. The network is designed to allow for the exchange of telemetry and other data between spacecraft to enable collaboration between small spacecraft. The system uses a peer-to-peer topology with no central router, so that it does not have a single point of failure. The mesh network is dynamically configurable to allow for addition and subtraction of new spacecraft into the communication network. Flight testing was performed using an unmanned aerial system (UAS) formation acting as a spacecraft analogue and providing a stressing environment to prove mesh network performance. The mesh network was primarily devised to provide low latency, high frequency communication but is flexible and can also be configured to provide higher bandwidth for applications desiring high data throughput. The network includes a relay functionality that extends the maximum range between spacecraft in the network by relaying data from node to node. The mesh network control is implemented completely in software making it hardware agnostic, thereby allowing it to function with a wide variety of existing radios and computing platforms..
This paper addresses NASA's requirement on the 2007 Phoenix Mars Lander to provide spacecraft communications during entry, descent, and landing on Mars to allow the identification of probable root cause should any mission failure occur. The Phoenix mission launched on 4 August 2007 and will land on 25 May 2008 on the northern plains of Mars to conduct a three-month study of the Martian environment. The paper discusses the architectural trades in designing a communications link and surveys the entry, descent, and landing communications approaches taken by previous missions. It then discusses the Phoenix-specific constraints and degrees of freedoms and presents a novel and robust implementation approach to entry, descent, and landing communications. The overall methodology and conclusions described herein can serve as a pathfinder for the entry, descent, and landing communications architecture and implementation of future Mars landed missions.