Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Having a Come-Apart: Lessons Learned from Additively Manufactured Hardware Failures

NASA has been engaged with additively manufactured (AM) process and component development since the 2000’s. AM offers various technical advantages, such as enhanced hardware design complexity, part consolidation, and processing of novel alloys in addition to programmatic advantages for reduction in processing time and cost. The focus of much of the AM development at NASA has been to mature the various processes, characterize material properties, develop standards, produce demonstrator parts, and integrate AM hardware in liquid rocket engines. These aspects have been demonstrated through process and design iterations using a methodical characterization, test-fail-fix cycles, as well as application and dissemination of lessons learned. In addition to these fundamental demonstrations of the AM process and hardware development, alloys that provide performance advantages in the high temperature and high-pressure environments have been matured for use in rocket engines. These environments are challenging for any alloy and any design, and the AM process is required to fully meet the intended design requirements. The importance of proper AM process was made evident in the failure of a Laser Powder Bed Fusion (L-PBF) copper-alloy combustion chamber during a hot-fire test due to a degraded material quality resulted from an AM process issue. The hot-fire test aimed to demonstrate high duty cycle under a risk-tolerant development project, where consequences of component failure would be minimal. However, the unintentional component failure emphasized the necessity of robust material characterization and rigorous process control procedures for the safe use of AM components in critical applications. In part, such concerns motivate the AM certification approach that NASA has recently adopted in NASA-STD-6030 “Additive Manufacturing Requirements for Spaceflight Systems”. This presentation provides an overview of the previously mentioned failure, a discussion on the evaluation of the failed chamber and supplemental chambers produced at the same time, a representative material samples that included intentional build witness lines, and a summary of the key results and recommendations from the evaluations. NASA continues to approach AM processes and designs with a level of risk and acceptance of failures that is appropriate for the project objectives, with the overall goal of safe implementation of AM technology and transferring AM technology into commercial space applications. The objective of this presentation is to provide awareness to the community working critical and non-critical AM components and the lessons learned on proper implementation of AM.

Additive Manufacturing↗

Study of intermittent field hardware failure data in digital electronics

The collection and analysis of data concerning intermittent dailures in digital devices was performed using data from a computer design for shipboard usage. The failure data consisted of actual field failures classified by failure mechanisms and their likelihood of having been intermittent, potentially intermittent, or hard. Each class was studies with respect to computer operation in the ranges of 0 to 2,000 hours, 0 to 5, hours, and 0 to 10,000 hours. The study was done at the computer level as well as the microcircuit level. Results indicate that as age increases, the quasi-intermittent failure rate increases and the mean time to failure descreases.

Oneill, E. J.↗

Diagnosing faults in autonomous robot plan execution

A major requirement for an autonomous robot is the capability to diagnose faults during plan execution in an uncertain environment. Many diagnostic researches concentrate only on hardware failures within an autonomous robot. Taking a different approach, the implementation of a Telerobot Diagnostic System that addresses, in addition to the hardware failures, failures caused by unexpected event changes in the environment or failures due to plan errors, is described. One feature of the system is the utilization of task-plan knowledge and context information to deduce fault symptoms. This forward deduction provides valuable information on past activities and the current expectations of a robotic event, both of which can guide the plan-execution inference process. The inference process adopts a model-based technique to recreate the plan-execution process and to confirm fault-source hypotheses. This technique allows the system to diagnose multiple faults due to either unexpected plan failures or hardware errors. This research initiates a major effort to investigate relationships between hardware faults and plan errors, relationships which were not addressed in the past. The results of this research will provide a clear understanding of how to generate a better task planner for an autonomous robot and how to recover the robot from faults in a critical environment.

Lam, Raymond K.↗

Diagnosing faults in autonomous robot plan execution

A major requirement for an autonomous robot is the capability to diagnose faults during plan execution in an uncertain environment. Many diagnostic researches concentrate only on hardware failures within an autonomous robot. Taking a different approach, the implementation of a Telerobot Diagnostic System that addresses, in addition to the hardware failures, failures caused by unexpected event changes in the environment or failures due to plan errors, is described. One feature of the system is the utilization of task-plan knowledge and context information to deduce fault symptoms. This forward deduction provides valuable information on past activities and the current expectations of a robotic event, both of which can guide the plan-execution inference process. The inference process adopts a model-based technique to recreate the plan-execution process and to confirm fault-source hypotheses. This technique allows the system to diagnose multiple faults due to either unexpected plan failures or hardware errors. This research initiates a major effort to investigate relationships between hardware faults and plan errors, relationships which were not addressed in the past. The results of this research will provide a clear understanding of how to generate a better task planner for an autonomous robot and how to recover the robot from faults in a critical environment.

Lam, Raymond K.↗

An experimental evaluation of software redundancy as a strategy for improving reliability

The strategy of using multiple versions of independently developed software as a means to tolerate residual software design faults is suggested by the success of hardware redundancy for tolerating hardware failures. Although, as generally accepted, the independence of hardware failures resulting from physical wearout can lead to substantial increases in reliability for redundant hardware structures, a similar conclusion is not immediate for software. The degree to which design faults are manifested as independent failures determines the effectiveness of redundancy as a method for improving software reliability. Interest in multi-version software centers on whether it provides an adequate measure of increased reliability to warrant its use in critical applications. The effectiveness of multi-version software is studied by comparing estimates of the failure probabilities of these systems with the failure probabilities of single versions. The estimates are obtained under a model of dependent failures and compared with estimates obtained when failures are assumed to be independent. The experimental results are based on twenty versions of an aerospace application developed and certified by sixty programmers from four universities. Descriptions of the application, development and certification processes, and operational evaluation are given together with an analysis of the twenty versions.

Eckhardt, Dave E., Jr.↗

Flying U.S. science on the U.S.S.R. Cosmos biosatellites

The USSR Cosmos Biosatellites are unmanned missions with durations of approximately 14 days. They are capable of carrying a wide variety of biological specimens such as cells, tissues, plants, and animals, including rodents and rhesus monkeys. The absence of a crew is an advantage with respect to the use of radioisotopes or other toxic materials and contaminants, but a disadvantage with respect to the performance of inflight procedures or repair of hardware failures. Thus, experiments hardware and procedures must be either completely automated or remotely controlled from the ground. A serious limiting factor for experiments is the amount of electrical powers available, so when possible experiments should be self-contained with their own batteries and data recording devices. Late loading is restricted to approximately 48 hours before launch and access time upon recovery is not precise since there is a ballistic reentry and the capsule must first be located and recovery vehicles dispatched to the site. Launches are quite reliable and there is a proven track record of nine previous Biosatellite flights. This paper will present data and experience from the seven previous Cosmos flights in which the US has participated as well as the key areas of consideration in planning a flight investigation aboard this Biosatellite platform.

Flight Experiment↗

Digital avionics design and reliability analyzer

The description and specifications for a digital avionics design and reliability analyzer are given. Its basic function is to provide for the simulation and emulation of the various fault-tolerant digital avionic computer designs that are developed. It has been established that hardware emulation at the gate-level will be utilized. The primary benefit of emulation to reliability analysis is the fact that it provides the capability to model a system at a very detailed level. Emulation allows the direct insertion of faults into the system, rather than waiting for actual hardware failures to occur. This allows for controlled and accelerated testing of system reaction to hardware failures. There is a trade study which leads to the decision to specify a two-machine system, including an emulation computer connected to a general-purpose computer. There is also an evaluation of potential computers to serve as the emulation computer.

Source record↗

Detecting servo failures with software

Program detects hardware failure in servosystems by comparing actual servo valve position with predictions of software model. In addition, system will also pick up most computer input/output failures. Process presents faster and more reliable results than previous failure detection methods.

Lew, D.↗

Catastrophic Fault Recovery with Self-Reconfigurable Chips

Mission critical systems typically employ multi-string redundancy to cope with possible hardware failure. Such systems are only as fault tolerant as there are many redundant strings. Once a particular critical component exhausts its redundant spares, the multi-string architecture cannot tolerate any further hardware failure. This paper aims at addressing such catastrophic faults through the use of 'Self-Reconfigurable Chips' as a last resort effort to 'repair' a faulty critical component.

fault tolerance↗

Catastrophic fault recovery with self-reconfigurable chips

Mission critical systems typically employ multi-string redundancy to cope with possible hardware failure. Such systems are only as fault tolerant as there are many redundant strings. Once a particular critical component exhausts its redundant spares, the multi-string architecture cannot tolerate any further hardware failure. This paper aims at addressing such catastrophic faults through the use of "Self-Reconfigurable Chips" as a last resort effort to "repair" a faulty critical component.

Chau, Savio N.↗

CHESS 2025: Waveform LiDAR data from NEON AOP surveys

This dataset provides Level 1 (L1) full-waveform light detection and ranging (LiDAR) data collected for the 2025 Colorado Headwaters Ecological Spectroscopy Study (CHESS). These data were acquired to enable characterization of vegetation structure and other three-dimensional features of the land surface, and to evaluate structural changes that may have occurred between a prior LiDAR acquisition in 2018 and the 2025 overflight. Waveform LiDAR data can provide more detailed information about objects on the ground than discrete point clouds typically do, and they are often used for granular target segmentation and characterization of subcanopy vegetation. The data were acquired over three study domains in the Upper Gunnison river basin: the upper East River watershed (CRBU); Almont Triangle and Taylor Canyon (ALMO); and Upper Taylor River watershed (UPTA) between 2025-06-13 and 2025-07-15. LiDAR data were acquired using the Optech Galaxy Prime Airborne LiDAR Terrain Mapper onboard the National Ecological Observatory Network (NEON) Airborne Observation Platform (AOP). These are the primary waveform LiDAR data delivered by NEON and are provided per flightline in compressed Pulsewaves format, an open-source binary file standard. A Pulsewaves object comprises a two files: a pulse (.pls) file, which stores the geographic origin, outgoing vector, and metadata for every laser pulse emitted by the scanner, and a wave file (.wvs), which stores the sequential amplitude samples of the outgoing pulse and the returning signals. The files are published here in their compressed forms (.plz, .wvz). All waveform data were processed following the theoretical workflow described in the NEON L0-to-L1 Waveform LiDAR Algorithm Theoretical Basis Document (Krause and Goulden 2022a); however, the Pulsewaves output format differs from a legacy format described in that document. Waveform amplitude samples are recorded at 1 nanosecond intervals. All coordinates are provided in meters. Horizontal coordinates are referenced in Universal Transverse Mercator (UTM) zone 13N and the World Geodetic System (WGS) 1984 ensemble datum. Elevations are referenced to Geoid12A. Waveform data for the UPTA survey area were collected without incident and the published records are complete. However, both the ALMO and CRBU collections experienced issues that resulted in incomplete data for those areas. On collection day 2018-06-16 a hardware failure caused the waveform digitizer to lose data from the eastern edge of the ALMO site (Figure 22). The waveform data for flightlines 2–20 could not be extracted from the digitizer, and the data proved unrecoverable. As a result, a portion of the site does not have coverage with waveform data. Although no hardware failure was observed during collection over the CRBU area, final waveform files generated by vendor software contained only ~25% of the expected number of return pulses. After discovery, NEON initiated troubleshooting with the vendor. The root cause of the data ablation had not been identified at the time of publication. Additional data will be published in an update to this package if further recovery proves successful. CHESS Project Description: The Colorado Headwaters Ecological Spectroscopy Study (CHESS) comprised a multi-week airborne remote sensing and field observation campaign in the Upper Gunnison Basin, Colorado, conducted in June and July of 2025. Airborne remote sensing was conducted by the National Ecological Observatory Network Airborne Observation Platform (NEON AOP), concurrent with a field campaign run by the Rocky Mountain Biological Laboratory (RMBL), the Lawrence Berkeley National Laboratory (LBNL) and SLAC National Accelerator Laboratory Watershed Function Science Focus Area (SFA), and NASA-JPL (Jet Propulsion Laboratory) Earth Surface Mineral Dust Source Investigation (EMIT) program. Between June 10 and July 18, 2025, the NEON AOP flight team collected high-resolution aerial imaging spectroscopy and Light Detection and Ranging (LiDAR) data over three domains: the Upper East River (CRBU), Almont Triangle (ALMO), and the Upper Taylor Basin (UPTA). In coordination with the flights, a field campaign acquired ground-truth observations, including observations of vegetation composition, foliar traits, forest demography, and subsurface properties in 18 core sampling areas within the domains. Additional surface water observations were taken at over 380 point locations. All CHESS campaign datasets can be found within the CHESS ESS-DIVE data portal: https://data.ess-dive.lbl.gov/portals/chess. Funding Acknowledgement: Field and remote-sensing data acquisition was performed under a grant from the National Aeronautics and Space Administration (80NSSC24K1005). This work was also supported by the Watershed Function Science Focus Area at Lawrence Berkeley National Laboratory funded by the US Department of Energy, Office of Science, Biological and Environmental Research under Contract No. DE-AC02-05CH11231.

2018 NEON and 2025 CHESS Campaigns↗

Predicting service life margins

Margins are developed for equipment susceptible to malfunction due to excessive time or operation cycles, and for identifying limited life equipment so monitoring and replacing is accomplished before hardware failure. Method applies to hardware where design service is established and where reasonable expected usage prediction is made.

Egan, G. F.↗

Lunar Landing Operational Risk Model

Characterizing the risk of spacecraft goes beyond simply modeling equipment reliability. Some portions of the mission require complex interactions between system elements that can lead to failure without an actual hardware fault. Landing risk is currently the least characterized aspect of the Altair lunar lander and appears to result from complex temporal interactions between pilot, sensors, surface characteristics and vehicle capabilities rather than hardware failures. The Lunar Landing Operational Risk Model (LLORM) seeks to provide rapid and flexible quantitative insight into the risks driving the landing event and to gauge sensitivities of the vehicle to changes in system configuration and mission operations. The LLORM takes a Monte Carlo based approach to estimate the operational risk of the Lunar Landing Event and calculates estimates of the risk of Loss of Mission (LOM) - Abort Required and is Successful, Loss of Crew (LOC) - Vehicle Crashes or Cannot Reach Orbit, and Success. The LLORM is meant to be used during the conceptual design phase to inform decision makers transparently of the reliability impacts of design decisions, to identify areas of the design which may require additional robustness, and to aid in the development and flow-down of requirements.

Mattenberger, Chris↗

Effects of hydrogen on metals

Several rules to guide choice of materials, and methods of welding, electroplating, and heat treatment will provide a method for minimizing failures in storage tanks and related hardware. Failures are caused by high-pressure hydrogen effects, the formation of hydrides in titanium, and hydrogen absorption through various metals processing techniques.

Cataldo, C. E.↗

Tapered Roller Bearing Damage Detection Using Decision Fusion Analysis

A diagnostic tool was developed for detecting fatigue damage of tapered roller bearings. Tapered roller bearings are used in helicopter transmissions and have potential for use in high bypass advanced gas turbine aircraft engines. A diagnostic tool was developed and evaluated experimentally by collecting oil debris data from failure progression tests conducted using health monitoring hardware. Failure progression tests were performed with tapered roller bearings under simulated engine load conditions. Tests were performed on one healthy bearing and three pre-damaged bearings. During each test, data from an on-line, in-line, inductance type oil debris sensor and three accelerometers were monitored and recorded for the occurrence of bearing failure. The bearing was removed and inspected periodically for damage progression throughout testing. Using data fusion techniques, two different monitoring technologies, oil debris analysis and vibration, were integrated into a health monitoring system for detecting bearing surface fatigue pitting damage. The data fusion diagnostic tool was evaluated during bearing failure progression tests under simulated engine load conditions. This integrated system showed improved detection of fatigue damage and health assessment of the tapered roller bearings as compared to using individual health monitoring technologies.

Dempsey, Paula J.↗

Deep learning model to detect various synchrophasor data anomalies

High-density synchrophasors provide valuable information for power grid situational awareness, operation and control. Unfortunately, due to factors including communication instability and hardware failure, their data quality can be greatly deteriorated by anomalies. Since the anomalies can impact the performance of the synchrophasor applications, it is of paramount significance to propose a model to detect anomalies in synchrophasor. In this study, a convolutional neural network model is established to detect and classify the anomalies in the synchrophasor measurements. Additionally, four types of anomalies observed in actual synchrophasors including erroneous patterns, random spikes, missing points and high-frequency interferences are considered in this study. The proposed model is extensively evaluated via field-collected measurements from the synchrophasor network in Jiangsu grid, China. The superior performance of the proposed model indicates the great potential of using deep learning for the detection of abnormal synchrophasor measurements.

42 ENGINEERING↗