Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

ISS Internal Active Thermal Control System (IATCS) Coolant Remediation Project -2006 Update

The IATCS coolant has experienced a number of anomalies in the time since the US Lab was first activated on Flight 5A in February 2001. These have included: 1) a decrease in coolant pH, 2) increases in inorganic carbon, 3) a reduction in phosphate concentration, 4) an increase in dissolved nickel and precipitation of nickel salts, and 5) increases in microbial concentration. These anomalies represent some risk to the system, have been implicated in some hardware failures and are suspect in others. The ISS program has conducted extensive investigations of the causes and effects of these anomalies and has developed a comprehensive program to remediate the coolant chemistry of the on-orbit system as well as provide a robust and compatible coolant solution for the hardware yet to be delivered. This paper presents a status of the coolant stability over the past year as well as results from destructive analyses of hardware removed from the on-orbit system and the current approach to coolant remediation.

Morrison, Russell H.↗

Localization of Ad-Hoc Lunar Constellations in Communication Failure Modes for Distributed Spacecraft Autonomy

As Lunar missions increase in complexity, inspired by NASA’s Artemis Program, they will require reliable and sufficient Position, Navigation, and Timing (PNT) capability to support the upcoming Lunar users. The navigation service should also be compatible with the smaller platforms, like CubeSats, being sent by the public and private sectors. A non-dedicated, ad-hoc Lunar navigation constellation can provide PNT services on-demand using the non-dedicated swarm assets. Swarm members cooperatively and autonomously localize themselves with minimal interaction from Earth, freeing up valuable bandwidth and ground segment resources. The autonomous localization of Lunar constellations utilizes neighbor two-way intersatellite link (ISL) measurements in a distributed extended Kalman filter (DEKF) system to minimize operating costs. Because the decentralized Lunar PNT system relies on relay communication amongst the agents, network failures or loss of assets among ad-hoc Lunar constellations may impact localization performance. This study presents an evaluation of localization performance under increasing levels of network degradation. A simulation of an ad-hoc Lunar PNT swarm is augmented to include system faults and the impacts of intermittent and permanent failures on localization performance are evaluated. We investigate three potential causes of network degradation: single spacecraft loss, multiple spacecraft loss, and antenna failure. The numerical assessments from the simulation show that the LPNT system under study, based on an autonomous decentralized concept of operation, is highly robust and resilient to communication failures. Minor faults, such as single spacecraft loss, solar interference, technical malfunctions, message delays, and antenna outages, have minimal impact on state estimation, with only a 4.47% and 3.75% degradation in median position error for assets and a representative ground user, respectively, compared to an ideal communication scenario. However, major faults, such as hardware failures or meteor strikes leading to the loss of multiple spacecrafts, are more concerning. The permanent loss of three spacecraft results in a more severe performance degradation, with median position error increasing by 23.3% for assets and 11.7% for a representative ground user, despite the Lunar PNT system remaining functional.

Yeji Kim↗

Flight Performance of Skylab Attitude and Pointing Control System

In 1967 a paper at the AIAA Guidance, Control and Flight Dynamics Conference in Huntsville, Ala. presented for the first time the prot)osed SKYLAB Attitude and Pointing Control System (APCS) The system requirements, Apollo Telescope Mount (ATM) configuration, control philosophy, and operational modes were presented and the APCS described. The Initial mission and system design requirements changed during the period of time before the SKYLAB was launched. This paper will review the Initial and final APCS requirements and goals and their relationship. The actual flight mission (and Its alterations during the flight) and known achieved APCS performance will then be presented. SKYLAB was a tremendous success in furthering man's scientific knowledge; but perhaps SKYLAB will be remembered more for the anomalies and the efforts undertaken to solve them. On May 14, 1973, the unmanned SKYLAB Orbital Workshop (OWS) was launched from Cape Kennedy. Serious hardware failures began to occur during ascent through the atmosphere and their spectre continued to haunt both the astronauts and their ground based support team. Nor were these the only surprises affecting the design and operation of the APCS. Mission requirements for pointing to various stellar targets and to nadir for earth resources experiments were added after the hardware was designed. The chance appearance of comet Kohoutek during the SKYLAB operational life-time caused NASA to add comet observation to the mission requirements and to adjust the time when the third crew would man the SKYLAB. The development of new procedures and software for the opportunity to observe this visitor to our solar system is described.

Chubb, W. B.↗

Use of a Slick-Plate as a Contingency Exercise Surface for the Treadmill With Vibration Isolation System

The treadmill with vibration isolation system (TVIS) was developed to counteract cardiovascular, musculoskeletal, and neurovestibular deconditioning during long-duration missions to the International Space Station (ISS). However, recent hardware failures have necessitated the development of a short-term, temporary contingency exercise countermeasure for TVIS until nominal operations could be restored. The purpose of our evaluation was twofold: 1) to examine whether a slick-plate/contingency exercise surface (CES) could be used as a walking/running surface and could elicit a heart rate (HR) greater than or equal to 70% HR maximum and 2) to determine the optimal hardware configuration, in microgravity, to simulate running/walking in a 1-g environment. One subject (male) participated in the slick surface evaluation and two subjects (one male, one female) participated in the microgravity evaluation of the slick surface configuration. During the slick surface evaluation, the subject was suspended in a parachute harness and bungee cord configuration to offset the subject's body weight. Using another bungee cord configuration, we added a vertical load back to the subject, who was then asked to run for 20 minutes on the slick surface. The microgravity evaluation simulated the ISS TVIS, and we evaluated two different slick surfaces (Teflon surface and an aluminum surface coated with Tufram) for use as a CES. We evaluated each surface with the subject walking and running, with and without a handrail, and while wearing either socks or nylon booties over shoes. In the slick surface evaluation, the subject ran for 20 minutes and reached a maximum HR of 170 bpm. In the microgravity evaluation, the subjects chose the aluminum plate coated with Tufram as the CES, while wearing a pair of nylon booties over running shoes and using a handrail, as the optimal hardware configuration. The results indicate that the CES may provide an interim capability to counteract aerobic deconditioning until TVIS can be returned to an operational status. No indices of musculoskeletal or neurovestibular deconditioning were evaluated. Future studies are needed to validate the efficacy of CES in countering aerobic deconditioning and the effects, if any, on musculoskeletal and neurovestibular deconditioning.

James A. Loehr↗

Fault tolerant distributed systems using Ada

This paper discusses the use of Ada on distributed systems in which failure of processors has to be tolerated. It is assumed that communication between tasks on separate processors will take place using the facilities of the Ada language, primarily the rendezvous. It is shown that there are numerous aspects of the language which make its use on a distributed system very difficult. The issues are raised from the desire to be able to recover, reconfigure, and provide continued service in the presence of hardware failure. For example, if a rendezvous takes place between two tasks on different processors, failure of the processor executing the serving task will cause the calling task to be permanently suspended because the rendezvous will never end. Extensive modifications to the execution support required for Ada are proposed which provide all the necessary facilities for programs written in Ada to withstand arbitrary processor failure. Mechanisms are suggested to allow processor failure to be detected and for tasks which would be permanently suspended to be released. Provided the required program structures are used, continued processing can be provided.

Knight, J. C.↗

Automated Mixed Traffic Vehicle (AMTV) technology and safety study

Technology and safety related to the implementation of an Automated Mixed Traffic Vehicle (AMTV) system are discussed. System concepts and technology status were reviewed and areas where further development is needed are identified. Failure and hazard modes were also analyzed and methods for prevention were suggested. The results presented are intended as a guide for further efforts in AMTV system design and technology development for both near term and long term applications. The AMTV systems discussed include a low speed system, and a hybrid system consisting of low speed sections and high speed sections operating in a semi-guideway. The safety analysis identified hazards that may arise in a properly functioning AMTV system, as well as hardware failure modes. Safety related failure modes were emphasized. A risk assessment was performed in order to create a priority order and significant hazards and failure modes were summarized. Corrective measures were proposed for each hazard.

Johnston, A. R.↗

Evaluating the Performance of the NASA LaRC CMF Motion Base Safety Devices

This paper describes the initial measured performance results of the previously documented NASA Langley Research Center (LaRC) Cockpit Motion Facility (CMF) motion base hardware safety devices. These safety systems are required to prevent excessive accelerations that could injure personnel and damage simulator cockpits or the motion base structure. Excessive accelerations may be caused by erroneous commands or hardware failures driving an actuator to the end of its travel at high velocity, stepping a servo valve, or instantly reversing servo direction. Such commands may result from single order failures of electrical or hydraulic components within the control system itself, or from aggressive or improper cueing commands from the host simulation computer. The safety systems must mitigate these high acceleration events while minimizing the negative performance impacts. The system accomplishes this by controlling the rate of change of valve signals to limit excessive commanded accelerations. It also aids hydraulic cushion performance by limiting valve command authority as the actuator approaches its end of travel. The design takes advantage of inherent motion base hydraulic characteristics to implement all safety features using hardware only solutions.

Gupton, Lawrence E.↗

Evolution of safety-critical requirements post-launch

This paper reports the results of a small study of requirements changes to the onboard software of three spacecraft subsequent to launch. Only those requirement changes that resulted from post-launch anoma-lies (i.e., during operations) were of interest here, since the goal was to better understand the relation-ship between critical anomalies during operations and how safety-critical requirements evolve. The results of the study were surprising in that anomaly-driven, post-launch requirements changes were rarely due to previous requirements having been incorrect. Instead, changes involved new requirements (1) for the software to handle rare events or (2) for the software to compensate for hardware failures or limitations. The prevalence of new requirements as a result of post-launch anomalies suggests a need for increased requirements-engineering support of maintenance activities in these systems. The results also confirm both the difficulty and the benefits of pursuing requirements completeness, especially in terms of fault tolerance, during development of critical systems.

Software requirements↗

Dynamic Modeling of Off-Nominal Operation in Advanced Life Support Systems

System failures, off-nominal operation, or unexpected interruptions in processing capability can cause unanticipated instabilities in Advanced Life Support (ALS) systems, even long after they are repaired. Much current modeling assumes ALS systems are static and linear, but ALS systems are actually dynamic and nonlinear, especially when failures and off nominal operation are considered. Modeling and simulation provide a way to study the stability and time behavior of nonlinear dynamic ALS systems under changed system configurations or operational scenarios. The dynamic behavior of a nonlinear system can be fully explored only by computer simulation over the full range of inputs and initial conditions. Previous simulations of BIO-Plex in SIMULINK, a toolbox of Matlab, were extended to model the off-nominal operation and long-term dynamics of partially closed physical/chemical and bioregenerative life support systems. System nonlinearity has many interesting potential consequences. Different equilibrium points may be reached for different initial conditions. The system stability can depend on the exact system inputs and initial conditions. The system may oscillate or even in rare cases behave chaotically. Temporary internal hardware failures or external perturbations in ALS systems can lead to dynamic instability and total ALS system failure. Appropriate control techniques can restore reliable operation and minimize the effects of dynamic instabilities due to anomalies or perturbations in a life support system.

Jones, Harry↗

Optimal Aircraft Control Upset Recovery With and Without Component Failures

This paper treats the problem of recovering sustainable nondescending (safe) flight in a transport aircraft after one or more of its control effectors fail. Such recovery can be a challenging goal for many transport aircraft currently in the operational fleet for two reasons. First, they have very little redundancy in their means of generating control forces and moments. These aircraft have, as primary control surfaces, a single rudder and pairwise elevators and aileron/spoiler units that provide yaw, pitch, and roll moments with sufficient bandwidth to be used in stabilizing and maneuvering the airframe. Beyond this, throttling the engines can provide additional moments, but on a much slower time scale. Other aerodynamic surfaces, such as leading and trailing edge flaps, are only intended to be placed in a position and left, and are, hence, very slow-moving. Because of this, loss of a primary control surface strongly degrades the controllability of the vehicle, particularly when the failed effector becomes stuck in a non-neutral position where it exerts a disturbance moment that must be countered by the remaining operating effectors. The second challenge in recovering safe flight is that these vehicles are not agile, nor can they tolerate large accelerations. This is of special importance when, at the outset of the recovery maneuver, the aircraft is flying toward the ground, as is frequently the case when there are major control hardware failures. Recovery of safe flight is examined in this paper in the context of trajectory optimization. For a particular transport aircraft, and a failure scenario inspired by an historical air disaster, recovery scenarios are calculated with and without control surface failures, to bring the aircraft to safe flight from the adverse flight condition that it had assumed, apparently as a result of contact with a vortex from a larger aircraft's wake. An effort has been made to represent relevant airframe dynamics, acceleration limits, and actuator limits faithfully, since these contribute to the lack of agility and control power that plays an important role in defining what can be achieved with the vehicle when it is in extremis.

Sparks, Dean W.↗

Evaluation of methods for determining hardware projected life

An investigation of existing methods of predicting hardware life is summarized by reviewing programs having long life requirements, current research efforts on long life problems, and technical papers reporting work on life predicting techniques. The results indicate that there are no accurate quantitative means to predict hardware life for system level hardware. The effectiveness of test programs and the cause of hardware failures is considered.

Source record↗

Saving Skylab

Initial difficulties encountered with the launch and operational deployment of the Skylab vehicle are described in terms of early indications of hardware failures, ground-based efforts at diagnosing the nature and significance of damage, and organization of improvized means and procedures for repair operations. Attention is given to heating and power-supply problems stemming from the loss of a micrometeoroid shield and a solar cell array. Systems engineering problems arising from the necessity of achieving a compromise among thermal, electrical, and stabilization requirements are explained, and repair solutions and operations in space and on the ground are described in detail.-

Schneider, W. C.↗

Flight performance of Skylab attitude and pointing control system

The Skylab attitude and pointing control system (APCS) requirements are briefly reviewed and the way in which they became altered during the prelaunch phase of development is noted. The actual flight mission (including mission alterations during flight) is described. The serious hardware failures that occurred, beginning during ascent through the atmosphere, also are described. The APCS's ability to overcome these failures and meet mission changes are presented. The large around-the-clock support effort on the ground is discussed. Salient design points and software flexibility that should afford pertinent experience for future spacecraft attitude and pointing control system designs are included.

Chubb, W. B.↗

Impact of coverage on the reliability of a fault tolerant computer

A mathematical reliability model is established for a reconfigurable fault tolerant avionic computer system utilizing state-of-the-art computers. System reliability is studied in light of the coverage probabilities associated with the first and second independent hardware failures. Coverage models are presented as a function of detection, isolation, and recovery probabilities. Upper and lower bonds are established for the coverage probabilities and the method for computing values for the coverage probabilities is investigated. Further, an architectural variation is proposed which is shown to enhance coverage.

Bavuso, S. J.↗

D-1A equipment module structure test

The Centaur Equipment Module (E/M) structural test program was performed in two parts due to an unscheduled hardware failure in the first test series. The objectives of the initial test program were to define the flexibility characteristics of the E/M, verify the design load capability, and determine its ultimate strength capability by loading to structural failure. However, during the first failure test attempt, the Intelsat IV MPA failed instead of the E/M. Therefore a new adapter was fabricated to simulate the HEAO mission adapter and the second series of tests were then performed. They concluded with the failure of the E/M forward interface ring resulting in about 3.5 degrees of permanent set on the high compression side. Nevertheless, the linear or useable strength capability of the E/M is greater or equal to that which is required for the HEAO missions. The E/M is deemed structurally qualified for the HEAO missions.

Niezgoda, T. F.↗

Instrumentation complex for Langley Research Center's National Transonic Facility

The instrumentation discussed in the present paper was developed to ensure reliable operation for a 2.5-meter cryogenic high-Reynolds-number fan-driven transonic wind tunnel. It will incorporate four CPU's and associated analog and digital input/output equipment, necessary for acquiring research data, controlling the tunnel parameters, and monitoring the process conditions. Connected in a multipoint distributed network, the CPU's will support data base management and processing; research measurement data acquisition and display; process monitoring; and communication control. The design will allow essential processes to continue, in the case of major hardware failures, by switching input/output equipment to alternate CPU's and by eliminating nonessential functions. It will also permit software modularization by CPU activity and thereby reduce complexity and development time.

Russell, C. H.↗

Orbiter subsystem hardware/software interaction analysis. Volume 8: Forward reaction control system

The results of the orbiter hardware/software interaction analysis for the AFT reaction control system are presented. The interaction between hardware failure modes and software are examined in order to identify associated issues and risks. All orbiter subsystems and interfacing program elements which interact with the orbiter computer flight software are analyzed. The failure modes identified in the subsystem/element failure mode and effects analysis are discussed.

Becker, D. D.↗

RAMP - A fault tolerant distributed microcomputer structure for aircraft navigation and control

Design methodologies for realizing future high authority autoflight control systems are being investigated, taking into account also the study of distributed microcomputer architectures. Attention is given to the redundant asynchronous microprocessor (RAMP) structure. RAMP comprises a connected network of microcomputers which has as input command and sensor information, and which generates servo information to drive actuators, and thrust linkages. Tolerance to hardware failures is achieved by static redundancy. Results of a failed microcomputer are simply rejected. This is done in lieu of dynamic redundancy wherein the distributed computer system performs real time fault detection and reconfiguration of the system. Attention is given to the RAMP network structure and operation, flight control with parallel asynchronous computers, and intermittent fault tolerance.

Dunn, W. R.↗