Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Method of Testing and Predicting Failures of Electronic Mechanical Systems

A method employing a knowledge base of human expertise comprising a reliability model analysis implemented for diagnostic routines is disclosed. The reliability analysis comprises digraph models that determine target events created by hardware failures human actions, and other factors affecting the system operation. The reliability analysis contains a wealth of human expertise information that is used to build automatic diagnostic routines and which provides a knowledge base that can be used to solve other artificial intelligence problems.

Iverson, David L.↗

Independent Orbiter Assessment (IOA): FMEA/CIL assessment

The results of the Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis (FMEA) and Critical Items List (CIL) are presented. Direction was given by the Orbiter and GFE Projects Office to perform the hardware analysis and assessment using the instructions and ground rules defined in NSTS 22206. The IOA analysis features a top-down approach to determine hardware failure modes, criticality, and potential critical items. To preserve independence, the anlaysis was accomplished without reliance upon the results contained within the NASA and prime contractor FMEA/CIL documentation. The assessment process compares the independently derived failure modes and criticality assignments to the proposed NASA Post 51-L FMEA/CIL documentation. When possible, assessment issues are discussed and resolved with the NASA subsystem managers. The assessment results for each subsystem are summarized. The most important Orbiter assessment finding was the previously unknown stuck autopilot push-button criticality 1/1 failure mode, having a worst case effect of loss of crew/vehicle when a microwave landing system is not active.

Saiidi, Mo J.↗

Test and evaluation of the generalized gate logic system simulator

The results of the initial testing of the Generalized Gate Level Logic Simulator (GGLOSS) are discussed. The simulator is a special purpose fault simulator designed to assist in the analysis of the effects of random hardware failures on fault tolerant digital computer systems. The testing of the simulator covers two main areas. First, the simulation results are compared with data obtained by monitoring the behavior of hardware. The circuit used for these comparisons is an incomplete microprocessor design based upon the MIL-STD-1750A Instruction Set Architecture. In the second area of testing, current simulation results are compared with experimental data obtained using precursors of the current tool. In each case, a portion of the earlier experiment is confirmed. The new results are then viewed from a different perspective in order to evaluate the usefulness of this simulation strategy.

Miner, Paul S.↗

Earth to Moon Transfer: Direct vs Via Libration Points (L1, L2)

For some three decades, the Apollo-style mission has served as a proven baseline technique for transporting flight crews to the Moon and back with expendable hardware. This approach provides an optimal design for expeditionary missions, emphasizing operational flexibility in terms of safely returning the crew in the event of a hardware failure. However, its application is limited essentially to low-latitude lunar sites, and it leaves much to be desired as a model for exploratory and evolutionary programs that employ reusable space-based hardware. This study compares the performance requirements for a lunar orbit rendezvous mission type with one using the cislunar libration point (L1) as a stopover and staging point for access to arbitrary sites on the lunar surface. For selected constraints and mission objectives, it contrasts the relative uniformity of performance cost when the L1 staging point is used with the wide variation of cost for the Apollo-style lunar orbit rendezvous.

Condon, Gerald L.↗

ISS Internal Active Thermal Control System (IATCS) Coolant Remediation Project

The IATCS coolant has experienced a number of anomalies in the time since the US Lab was first activated on Flight 5A in February 2001. These have included: 1) a decrease in coolant pH, 2) increases in inorganic carbon, 3) a reduction in phosphate buffer concentration, 4) an increase in dissolved nickel and precipitation of nickel salts, and 5) increases in microbial concentration. These anomalies represent some risk to the system, have been implicated in some hardware failures and are suspect in others. The ISS program has conducted extensive investigations of the causes and effects of these anomalies and has developed a comprehensive program to remediate the coolant chemistry of the on-orbit system as well as provide a robust and compatible coolant solution for the hardware yet to be delivered. The remediation steps include changes in the coolant chemistry specification, development of a suite of new antimicrobial additives, and development of devices for the removal of nickel and phosphate ions from the coolant. This paper presents an overview of the anomalies, their known and suspected system effects, their causes, and the actions being taken to remediate the coolant.

Morrison, Russell H.↗

ISS Internal Active Thermal Control System (IATCS) Coolant Remediation Project -2006 Update

The IATCS coolant has experienced a number of anomalies in the time since the US Lab was first activated on Flight 5A in February 2001. These have included: 1) a decrease in coolant pH, 2) increases in inorganic carbon, 3) a reduction in phosphate concentration, 4) an increase in dissolved nickel and precipitation of nickel salts, and 5) increases in microbial concentration. These anomalies represent some risk to the system, have been implicated in some hardware failures and are suspect in others. The ISS program has conducted extensive investigations of the causes and effects of these anomalies and has developed a comprehensive program to remediate the coolant chemistry of the on-orbit system as well as provide a robust and compatible coolant solution for the hardware yet to be delivered. This paper presents a status of the coolant stability over the past year as well as results from destructive analyses of hardware removed from the on-orbit system and the current approach to coolant remediation.

Morrison, Russell H.↗

Localization of Ad-Hoc Lunar Constellations in Communication Failure Modes for Distributed Spacecraft Autonomy

As Lunar missions increase in complexity, inspired by NASA’s Artemis Program, they will require reliable and sufficient Position, Navigation, and Timing (PNT) capability to support the upcoming Lunar users. The navigation service should also be compatible with the smaller platforms, like CubeSats, being sent by the public and private sectors. A non-dedicated, ad-hoc Lunar navigation constellation can provide PNT services on-demand using the non-dedicated swarm assets. Swarm members cooperatively and autonomously localize themselves with minimal interaction from Earth, freeing up valuable bandwidth and ground segment resources. The autonomous localization of Lunar constellations utilizes neighbor two-way intersatellite link (ISL) measurements in a distributed extended Kalman filter (DEKF) system to minimize operating costs. Because the decentralized Lunar PNT system relies on relay communication amongst the agents, network failures or loss of assets among ad-hoc Lunar constellations may impact localization performance. This study presents an evaluation of localization performance under increasing levels of network degradation. A simulation of an ad-hoc Lunar PNT swarm is augmented to include system faults and the impacts of intermittent and permanent failures on localization performance are evaluated. We investigate three potential causes of network degradation: single spacecraft loss, multiple spacecraft loss, and antenna failure. The numerical assessments from the simulation show that the LPNT system under study, based on an autonomous decentralized concept of operation, is highly robust and resilient to communication failures. Minor faults, such as single spacecraft loss, solar interference, technical malfunctions, message delays, and antenna outages, have minimal impact on state estimation, with only a 4.47% and 3.75% degradation in median position error for assets and a representative ground user, respectively, compared to an ideal communication scenario. However, major faults, such as hardware failures or meteor strikes leading to the loss of multiple spacecrafts, are more concerning. The permanent loss of three spacecraft results in a more severe performance degradation, with median position error increasing by 23.3% for assets and 11.7% for a representative ground user, despite the Lunar PNT system remaining functional.

Yeji Kim↗

Flight Performance of Skylab Attitude and Pointing Control System

In 1967 a paper at the AIAA Guidance, Control and Flight Dynamics Conference in Huntsville, Ala. presented for the first time the prot)osed SKYLAB Attitude and Pointing Control System (APCS) The system requirements, Apollo Telescope Mount (ATM) configuration, control philosophy, and operational modes were presented and the APCS described. The Initial mission and system design requirements changed during the period of time before the SKYLAB was launched. This paper will review the Initial and final APCS requirements and goals and their relationship. The actual flight mission (and Its alterations during the flight) and known achieved APCS performance will then be presented. SKYLAB was a tremendous success in furthering man's scientific knowledge; but perhaps SKYLAB will be remembered more for the anomalies and the efforts undertaken to solve them. On May 14, 1973, the unmanned SKYLAB Orbital Workshop (OWS) was launched from Cape Kennedy. Serious hardware failures began to occur during ascent through the atmosphere and their spectre continued to haunt both the astronauts and their ground based support team. Nor were these the only surprises affecting the design and operation of the APCS. Mission requirements for pointing to various stellar targets and to nadir for earth resources experiments were added after the hardware was designed. The chance appearance of comet Kohoutek during the SKYLAB operational life-time caused NASA to add comet observation to the mission requirements and to adjust the time when the third crew would man the SKYLAB. The development of new procedures and software for the opportunity to observe this visitor to our solar system is described.

Chubb, W. B.↗

Use of a Slick-Plate as a Contingency Exercise Surface for the Treadmill With Vibration Isolation System

The treadmill with vibration isolation system (TVIS) was developed to counteract cardiovascular, musculoskeletal, and neurovestibular deconditioning during long-duration missions to the International Space Station (ISS). However, recent hardware failures have necessitated the development of a short-term, temporary contingency exercise countermeasure for TVIS until nominal operations could be restored. The purpose of our evaluation was twofold: 1) to examine whether a slick-plate/contingency exercise surface (CES) could be used as a walking/running surface and could elicit a heart rate (HR) greater than or equal to 70% HR maximum and 2) to determine the optimal hardware configuration, in microgravity, to simulate running/walking in a 1-g environment. One subject (male) participated in the slick surface evaluation and two subjects (one male, one female) participated in the microgravity evaluation of the slick surface configuration. During the slick surface evaluation, the subject was suspended in a parachute harness and bungee cord configuration to offset the subject's body weight. Using another bungee cord configuration, we added a vertical load back to the subject, who was then asked to run for 20 minutes on the slick surface. The microgravity evaluation simulated the ISS TVIS, and we evaluated two different slick surfaces (Teflon surface and an aluminum surface coated with Tufram) for use as a CES. We evaluated each surface with the subject walking and running, with and without a handrail, and while wearing either socks or nylon booties over shoes. In the slick surface evaluation, the subject ran for 20 minutes and reached a maximum HR of 170 bpm. In the microgravity evaluation, the subjects chose the aluminum plate coated with Tufram as the CES, while wearing a pair of nylon booties over running shoes and using a handrail, as the optimal hardware configuration. The results indicate that the CES may provide an interim capability to counteract aerobic deconditioning until TVIS can be returned to an operational status. No indices of musculoskeletal or neurovestibular deconditioning were evaluated. Future studies are needed to validate the efficacy of CES in countering aerobic deconditioning and the effects, if any, on musculoskeletal and neurovestibular deconditioning.

James A. Loehr↗

Fault tolerant distributed systems using Ada

This paper discusses the use of Ada on distributed systems in which failure of processors has to be tolerated. It is assumed that communication between tasks on separate processors will take place using the facilities of the Ada language, primarily the rendezvous. It is shown that there are numerous aspects of the language which make its use on a distributed system very difficult. The issues are raised from the desire to be able to recover, reconfigure, and provide continued service in the presence of hardware failure. For example, if a rendezvous takes place between two tasks on different processors, failure of the processor executing the serving task will cause the calling task to be permanently suspended because the rendezvous will never end. Extensive modifications to the execution support required for Ada are proposed which provide all the necessary facilities for programs written in Ada to withstand arbitrary processor failure. Mechanisms are suggested to allow processor failure to be detected and for tasks which would be permanently suspended to be released. Provided the required program structures are used, continued processing can be provided.

Knight, J. C.↗

Automated Mixed Traffic Vehicle (AMTV) technology and safety study

Technology and safety related to the implementation of an Automated Mixed Traffic Vehicle (AMTV) system are discussed. System concepts and technology status were reviewed and areas where further development is needed are identified. Failure and hazard modes were also analyzed and methods for prevention were suggested. The results presented are intended as a guide for further efforts in AMTV system design and technology development for both near term and long term applications. The AMTV systems discussed include a low speed system, and a hybrid system consisting of low speed sections and high speed sections operating in a semi-guideway. The safety analysis identified hazards that may arise in a properly functioning AMTV system, as well as hardware failure modes. Safety related failure modes were emphasized. A risk assessment was performed in order to create a priority order and significant hazards and failure modes were summarized. Corrective measures were proposed for each hazard.

Johnston, A. R.↗

Evaluating the Performance of the NASA LaRC CMF Motion Base Safety Devices

This paper describes the initial measured performance results of the previously documented NASA Langley Research Center (LaRC) Cockpit Motion Facility (CMF) motion base hardware safety devices. These safety systems are required to prevent excessive accelerations that could injure personnel and damage simulator cockpits or the motion base structure. Excessive accelerations may be caused by erroneous commands or hardware failures driving an actuator to the end of its travel at high velocity, stepping a servo valve, or instantly reversing servo direction. Such commands may result from single order failures of electrical or hydraulic components within the control system itself, or from aggressive or improper cueing commands from the host simulation computer. The safety systems must mitigate these high acceleration events while minimizing the negative performance impacts. The system accomplishes this by controlling the rate of change of valve signals to limit excessive commanded accelerations. It also aids hydraulic cushion performance by limiting valve command authority as the actuator approaches its end of travel. The design takes advantage of inherent motion base hydraulic characteristics to implement all safety features using hardware only solutions.

Gupton, Lawrence E.↗

Evolution of safety-critical requirements post-launch

This paper reports the results of a small study of requirements changes to the onboard software of three spacecraft subsequent to launch. Only those requirement changes that resulted from post-launch anoma-lies (i.e., during operations) were of interest here, since the goal was to better understand the relation-ship between critical anomalies during operations and how safety-critical requirements evolve. The results of the study were surprising in that anomaly-driven, post-launch requirements changes were rarely due to previous requirements having been incorrect. Instead, changes involved new requirements (1) for the software to handle rare events or (2) for the software to compensate for hardware failures or limitations. The prevalence of new requirements as a result of post-launch anomalies suggests a need for increased requirements-engineering support of maintenance activities in these systems. The results also confirm both the difficulty and the benefits of pursuing requirements completeness, especially in terms of fault tolerance, during development of critical systems.

Software requirements↗

Dynamic Modeling of Off-Nominal Operation in Advanced Life Support Systems

System failures, off-nominal operation, or unexpected interruptions in processing capability can cause unanticipated instabilities in Advanced Life Support (ALS) systems, even long after they are repaired. Much current modeling assumes ALS systems are static and linear, but ALS systems are actually dynamic and nonlinear, especially when failures and off nominal operation are considered. Modeling and simulation provide a way to study the stability and time behavior of nonlinear dynamic ALS systems under changed system configurations or operational scenarios. The dynamic behavior of a nonlinear system can be fully explored only by computer simulation over the full range of inputs and initial conditions. Previous simulations of BIO-Plex in SIMULINK, a toolbox of Matlab, were extended to model the off-nominal operation and long-term dynamics of partially closed physical/chemical and bioregenerative life support systems. System nonlinearity has many interesting potential consequences. Different equilibrium points may be reached for different initial conditions. The system stability can depend on the exact system inputs and initial conditions. The system may oscillate or even in rare cases behave chaotically. Temporary internal hardware failures or external perturbations in ALS systems can lead to dynamic instability and total ALS system failure. Appropriate control techniques can restore reliable operation and minimize the effects of dynamic instabilities due to anomalies or perturbations in a life support system.

Jones, Harry↗

Optimal Aircraft Control Upset Recovery With and Without Component Failures

This paper treats the problem of recovering sustainable nondescending (safe) flight in a transport aircraft after one or more of its control effectors fail. Such recovery can be a challenging goal for many transport aircraft currently in the operational fleet for two reasons. First, they have very little redundancy in their means of generating control forces and moments. These aircraft have, as primary control surfaces, a single rudder and pairwise elevators and aileron/spoiler units that provide yaw, pitch, and roll moments with sufficient bandwidth to be used in stabilizing and maneuvering the airframe. Beyond this, throttling the engines can provide additional moments, but on a much slower time scale. Other aerodynamic surfaces, such as leading and trailing edge flaps, are only intended to be placed in a position and left, and are, hence, very slow-moving. Because of this, loss of a primary control surface strongly degrades the controllability of the vehicle, particularly when the failed effector becomes stuck in a non-neutral position where it exerts a disturbance moment that must be countered by the remaining operating effectors. The second challenge in recovering safe flight is that these vehicles are not agile, nor can they tolerate large accelerations. This is of special importance when, at the outset of the recovery maneuver, the aircraft is flying toward the ground, as is frequently the case when there are major control hardware failures. Recovery of safe flight is examined in this paper in the context of trajectory optimization. For a particular transport aircraft, and a failure scenario inspired by an historical air disaster, recovery scenarios are calculated with and without control surface failures, to bring the aircraft to safe flight from the adverse flight condition that it had assumed, apparently as a result of contact with a vortex from a larger aircraft's wake. An effort has been made to represent relevant airframe dynamics, acceleration limits, and actuator limits faithfully, since these contribute to the lack of agility and control power that plays an important role in defining what can be achieved with the vehicle when it is in extremis.

Sparks, Dean W.↗

Asynchronous distributed-memory task-parallel algorithm for compressible flows on unstructured 3D Eulerian grids

Here, we discuss the implementation of a finite element method, used to numerically solve the Euler equations of compressible flows, using an asynchronous runtime system (RTS). The algorithm is implemented for distributed-memory machines, using stationary unstructured 3D meshes, combining data-, and task-parallelism on top of the Charm++ RTS. Charm++’s execution model is asynchronous by default, allowing arbitrary overlap of computation and communication. Task-parallelism allows scheduling parts of an algorithm independently of, or dependent on, each other. Built-in automatic load balancing enables continuous redistribution of computational load by migration of work units based on real-time CPU load measurement. The RTS also features automatic checkpointing, fault tolerance, resilience against hardware failure, and supports power-, and energy-aware computation. We demonstrate scalability up to 25 x 10 9 cells at $\mathscr{O}$10 4 compute cores and the benefits of automatic load balancing for irregular workloads. The full source code with documentation is available at https://quinoacomputing.org.

42 ENGINEERING↗

Anomaly Detection in Accelerator Facilities Using Machine Learning

Synchrotron light sources are user facilities and usually run about 5000 hours per year to support many beamlines operations in parallel. Reliability is a key parameter to evaluate machine performance. Even many facilities have achieved >95% beam reliability, there are still many hours of unscheduled downtime and every hour lost is a waste of operation costs along with a big impact on individual scheduled user experiments. Preventive maintenance on subsystems and quick recovery from machine trips are the basic strategies to achieve high reliability, which heavily depends on experts’ dedication. Recently, SLAC, APS, and NSLS-II collaborated to develop machine-learning-based approaches aiming to solve both situations, hardware failure prediction and machine failure diagnosis to find the root sources. In this paper, we report our facility operation status, development progress, and plans.

Accelerator Physics↗

ECE 4396 (Final Report)

In the summer of 2025, I was fortunate enough intern at Sandia National Laboratories in Albuquerque, New Mexico. I was hired into the Southwest Analysis Laboratories for Semiconductor Advancement (SALSA) intern program. In this internship, I applied my knowledge and skills in electrical engineering to conduct hardware failure analysis. I utilized various failure analyze techniques involving the use of Infrared Thermography (IRT) and Laser Scanning Microscopy (LSM) to test different Application-Specific Integrated Circuits (ASIC) chips that are available in the public market.

42 ENGINEERING↗