Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Improving Logistics and Waste Management for Deep Space Human Exploration

NASA's Advanced Exploration Systems Logistics Reduction Project is developing technologies that reduce mission mass and volume for exploration. Recently there has been increasing interest in determining the quantity of consumable logistics and system spares necessary to ensure a certain level of reliability. This is influenced by a technology's criticality and degree of impact to the overall mission. Technologies that directly reduce mass (e.g. longer wear crew clothing) are relatively straightforward for calculating the savings and understanding the mission impacts. Waste management technologies that process waste can reduce mass, but spares and contingency modes are more interwoven with other vehicle systems, so assessment is more complex. This paper considers mission benefits while also considering impacts from hardware failures for technologies including: crew clothing, reusable cargo bags for habitat outfitting, automated RFID cargo tracking, trash processing/storage/repurposing, and high reliability toilets.

Broyan, James Lee, Jr.↗

"Making Safety Happen" Through Probabilistic Risk Assessment at NASA

NASA is using Probabilistic Risk Assessment (PRA) as one of the tools in its Safety & Mission Assurance (S&MA) tool belt to identify and quantify risks associated with human spaceflight. This paper discusses some of the challenges and benefits associated with developing and using PRA for NASA human space programs. Some programs have entered operation prior to developing a PRA, while some have implemented PRA from the start of the program. It has been observed that the earlier a design change is made in the concept or design phase, the less impact it has on cost and schedule. Not finding risks until the operation phase yields much costlier design changes and major delays, which can result in discussions of just accepting the risk. Risk contributors identified by PRA are not just associated with hardware failures. They include but are not limited to crew fatality due to medical causes, the environment the vehicle and crew are exposed to, the software being used, and the reliability of the crew performing required actions. Some programs have entered operation prior to developing a PRA, and while PRA can still provide a benefit for operations and future design trades, the benefit of implementing PRA from the start of the program provides the added benefit of informing design and reducing risk early in program development. Currently, NASA’s International Space Station (ISS) program is in its 20th year of on-orbit operations around the Earth and has several new programs in the design phase preparing to enter the operation phase all of which have active (or living) PRAs. These programs incorporate PRA as part of their Risk-Informed, Decision-Making (RIDM) process. For new NASA human spaceflight programs discussion begins with mission concept, establishing requirements, forming the PRA team, and continues through the design cycles into the operational phase. Several examples of PRA related applications and observed lessons are included.

Applications↗

NASA Engineering and Safety Center Technical Bulletin No. 21-01: Experimental and Computational Study of Cavitation in Hydrogen Peroxide

Cavitation in liquid propulsion systems can lead to performance degradation and hardware failures. The NESC sponsored an investigation to measure and model cavitation in pressurized hydrogen peroxide fl ow. The experimentally measured and computationally predicted cavitation lengths were compared as a function of cavitation number. The measured and predicted data exhibited close agreement over the range of pressures and temperatures studied, and no calibration of the cavitation model coefficients was needed.

NESC Technical Bulletin No. 21-01↗

Fault Propagation, EMI Propagation, and Fault Containment in Aerospace Systems

The occurrence of faults in aerospace system hardware and software have consequences ranging from minor effects to catastrophic effects, and such faults can directly affect the safety of hardware and personnel. There are many origins to fault conditions, and the hardware that is capable of still meeting its performance requirements after experiencing itself a fault is said to be fault tolerant. A fault tolerant hardware is capable of detecting, isolating, and recovering from a fault condition; and this is a subfield of control engineering. An aerospace system that has been shown to have electromagnetic compatibility (EMC) in all its subsystems and systems cannot induced faults caused by electromagnetic interference (EMI). It can be proposed that the presence of EMI (or lack of EMC) is analogous to a potential fault initiator and the effects can likewise range from minor to severe. This paper starts by addressing the consequences of hardware failure in aerospace systems from a fault perspective, because the design of fault tolerant system is a major endeavor in aerospace. To arrive to this goal the paper starts with the concepts of fault, fault propagation, and a new concept called fault containment region. The paper then proceeds to provide two very recent examples in the aircraft industry of fault propagation with catastrophic effects. The paper proceeds to introduce the concept of EMI fault containment and a brief introduction to another new concept called the EMI containment region. The paper proceeds with an example of EMI fault containment region. The paper ends with a lesson learned conclusions.

Perez, Reinaldo↗

Qualification and Performance of a High-Efficiency Laser Transmitter for Deep-Space Optical Communications

A high-power Laser Transmitter Assembly (LTA) was developed to support the Deep Space Optical Communications (DSOC) technology demonstration being developed by the Jet Propulsion Laboratory. NASA’s Psyche Mission plans to host the DSOC flight subsystem for testing space-to-ground high-bandwidth laser communications en route to the 16 Psyche asteroid. We review the design, performance, and qualification of the LTA Engineering Model and Flight Model (EM and FM) delivered to JPL. The LTA uses a master-oscillator power amplifier (MOPA) design and delivers up to 4.5 W at 1550 nm, with a highly efficient, cladding-pumped, polarization-maintaining erbium-ytterbium fiber amplifier. The master oscillator generates a range of pulse widths and repetition rates to support modulation formats from 16- to 128-PPM for optical data transmission at >100 Mbps. The LTA was designed for high reliability and radiation hardness, and includes redundant signal and pumping paths to reduce single points of failure, hardware interlocks to ensure safe operation and protection against damage, closed-loop control of optical power, and detailed health and status via telemetry. The LTA EM and FM were subjected to unit-appropriate space qualification testing. We describe the performance testing of the EM and FM, for the characterization of key metrics such as wavelength stability, signal linewidth, optical pulse width, jitter, and extinction ratio, and polarization extinction ratio. The management of optical nonlinearities (self-phase modulation, Brillouin scattering, or pulse-to-pulse energy variation), which could result in an optical link penalty or damage to the LTA, is also detailed, and factors affecting the power efficiency are discussed.

Jaques, J.↗

NASA Johnson Space Center Astromaterials Research Exploration Science Image Science and Analysis Group (ISAG) ISAG Support to the ISS Program

- Primary focus is on maintaining the safety of the crew and vehicle. - ISAG personnel request and screen imagery to: - Monitor changes in the ISS external condition - Anomalous indications - Hardware out of configuration - Detect Micro-Meteoroid or Orbital Debris (MMOD) impacts leading to: - Hardware failure - EVA sharp edges - ISAG personnel derive engineering data from imagery for: - Anomaly investigations - Structural dynamics measurements - Clearance assessments - Verifying ISS configuration against models and requirements - Jettison trajectory calculation - Provide Real-Time Support to the ISS and Anomalies

Dwight Osborne↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This paper presents the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Fault Management Algorithm Risk Assessment for the NASA Space Launch System

This presentation describes the false positive (FP) and false negative (FN) risk assessment process currently being conducted for the Space Launch System (SLS) Artemis II Fault Management (FM) detection functions. The analysis scope, general assumptions and guide rules, and key modeling concepts were discussed to establish the basis of the risk assessments conducted. Initial analyses indicated a dominance in the total risk by software and firmware failures. This paper presents efforts applied to refine the software risks and the overall impact of implementing those modifications. Current analyses conducted on the detection functions implemented for the SLS Artemis II mission indicate primary risk drivers for the individual FM detection functions are flight software failures, firmware design failures, and hardware Common Cause Failures (CCFs). There still remains issues of how to account for time and redundancy in the software risk estimations.

probability risk analysis↗

Hybrid Modeling for Scenario-Based Evaluation of Failure Effects in Advanced Hardware-Software Designs

This paper describes an incremental scenario-based simulation approach to evaluation of intelligent software for control and management of hardware systems. A hybrid continuous/discrete event simulation of the hardware dynamically interacts with the intelligent software in operations scenarios. Embedded anomalous conditions and failures in simulated hardware can lead to emergent software behavior and identification of missing or faulty software or hardware requirements. An approach is described for extending simulation-based automated incremental failure modes and effects analysis, to support concurrent evaluation of intelligent software and the hardware controlled by the software

Malin, Jane T.↗

Laser Peening Effects on Friction Stir Welding

The laser peening process can result in considerable improvement to crack initiation, propagation, and mechanical properties in FSW which equates to longer hardware service life Processed hardware safety is improved by producing higher failure tolerant hardware, and reducing risk. Lowering hardware maintenance cost produces longer hardware service life, and lower hardware down time. Application of this proposed technology will result in substantial benefits and savings throughout the life of the treated components

Hatameleh, Omar↗

Reliability Growth in Space Life Support Systems

A hardware system's failure rate often increases over time due to wear and aging, but not always. Some systems instead show reliability growth, a decreasing failure rate with time, due to effective failure analysis and remedial hardware upgrades. Reliability grows when failure causes are removed by improved design. A mathematical reliability growth model allows the reliability growth rate to be computed from the failure data. The space shuttle was extensively maintained, refurbished, and upgraded after each flight and it experienced significant reliability growth during its operational life. In contrast, the International Space Station (ISS) is much more difficult to maintain and upgrade and its failure rate has been constant over time. The ISS Carbon Dioxide Removal Assembly (CDRA) reliability has slightly decreased. Failures on ISS and with the ISS CDRA continue to be a challenge.

life support↗

A Comprehensive Reliability Methodology for Assessing Risk of Reusing Failed Hardware Without Corrective Actions with and Without Redundancy

This paper deals with the development of a reliability methodology to assess the consequences of using hardware, without failure analysis or corrective action, that has previously demonstrated that it did not perform per specification. The subject of this paper arose from the need to provide a detailed probabilistic analysis to calculate the change in probability of failures with respect to the base or non-failed hardware. The methodology used for the analysis is primarily based on principles of Monte Carlo simulation. The random variables in the analysis are: Maximum Time of Operation (MTO) and operation Time of each Unit (OTU) The failure of a unit is considered to happen if (OTU) is less than MTO for the Normal Operational Period (NOP) in which this unit is used. NOP as a whole uses a total of 4 units. Two cases are considered. in the first specialized scenario, the failure of any operation or system failure is considered to happen if any of the units used during the NOP fail. in the second specialized scenario, the failure of any operation or system failure is considered to happen only if any two of the units used during the MOP fail together. The probability of failure of the units and the system as a whole is determined for 3 kinds of systems - Perfect System, Imperfect System 1 and Imperfect System 2. in a Perfect System, the operation time of the failed unit is the same as that of the MTO. In an Imperfect System 1, the operation time of the failed unit is assumed as 1 percent of the MTO. In an Imperfect System 2, the operation time of the failed unit is assumed as zero. in addition, simulated operation time of failed units is assumed as 10 percent of the corresponding units before zero value. Monte Carlo simulation analysis is used for this study. Necessary software has been developed as part of this study to perform the reliability calculations. The results of the analysis showed that the predicted change in failure probability (P(sub F)) for the previously failed units is as high as 49 percent above the baseline (perfect system) for the worst case. The predicted change in system P(sub F) for the previously failed units is as high as 36% for single unit failure without any redundancy. For redundant systems, with dual unit failure, the predicted change in P(sub F) for the previously failed units is as high as 16%. These results will help management to make decisions regarding the consequences of using previously failed units without adequate failure analysis or corrective action.

Putcha, Chandra S.↗

ISS Fiber Optic Failure Investigation Root Cause Report

In August of 1999, Boeing Corporation (Boeing) engineers began investigating failures of optical fiber being used on International Space Station flight hardware. Catastrophic failures of the fiber were linked to a defect in the glass fiber. Following several meetings of Boeing and NASA engineers and managers, Boeing created and led an investigation team, which examined the reliability of the cable installed in the U.S. Lab. NASA Goddard Space Flight Center's Components Technologies and Radiation Effects Branch (GSFC) led a team investigating the root cause of the failures. Information was gathered from: regular telecons and other communications with the investigation team, investigative trips to the cable distributor's plant, the cable manufacturing plant and the fiber manufacturing plant (including a review of build records), destructive and non-destructive testing, and expertise supplied by scientists from Dupont, and Lucent-Bell Laboratories. Several theories were established early on which were not able to completely address the destructive physical analysis and experiential evidence. Lucent suggested hydrofluoric acid (HF) etching of the glass and successfully duplicated the "rocket engine" defect. Strength testing coupled with examination of the low strength break sites linked features in the polyimide coating with latent defect sites. The information provided below explains what was learned about the susceptibility of the pre-cabled fiber to failure when cabled as it was for Space Station and the nature of the latent defects.

Leidecker, Henning↗

Pre-Installation Acceptance (PIA) Functional Performance of the Design Verification Test (DVT) Exploration Extra-vehicular Mobility Unit (xEMU)

In an effort that began with technology investment by NASA in a few key components during the Constellation Program and then evolved to demonstrate a packaged Portable Life Support System (PLSS) as part of the Advanced Exploration Systems (AES) Program, the next evolution of the PLSS is now a key component of the Exploration Extra-vehicular Mobility Unit (xEMU) and is assembled as a Design Verification Test (DVT) unit. The xEMU has been detailed with respect to completing a demonstration on the International Space Station (ISS) with support of units for initial lunar capability. The xEMU completed the Preliminary Design Review (PDR) with subsequent Safety Review Panel (SRP) Phase I reviews in 2019-2020. The objectives for DVT are to validate requirements, train the team, learn how to fabricate the hardware with appropriate process controls, assemble the hardware, test the hardware, determine the failure mechanisms/limits of the hardware design and buy down the most risk possible for the qualification and flight phases of the development. With completion of the assembly and initial functional testing of the DVT PLSS in laboratory ambient conditions and vacuum conditions the xEMU has progressed significantly into the DVT objectives. A key part of the test sequences for DVT and all future phases is the Pre-Installation Acceptance (PIA) functional testing which validates the performance of integrated systems including: primary thermal control, auxiliary thermal control, suit ventilation, primary oxygen, secondary oxygen, power distribution, as well as caution and warning all with respect to the applied requirements. This discussion will include an overview of the assembly, summary of the PIA functional testing, lessons learned, and corrective actions implemented moving forward into the remainder of the DVT phase for xEMU.

Colin Campbell↗

Electrostatic Discharge (ESD) Failures in Thin Film Resistors

Field failures of nichrome thin film resistors have been investigated recently for several pieces of space flight hardware. These failures have involved resistance shifts ranging from a few percent to complete open circuits. Failure analysis and duplication of these failures have revealed that the failures were caused by electrostatic discharge. The failure characteristics and the circuit conditions necessary for failure have been studied for several types of thin film resistors including nichrome and tantalum nitride resistive elements. The effects of latent damage and resistive pattern design will also be discussed.

Hull, Scott M.↗

Contamination Examples and Lessons from Low Earth Orbit Experiments and Operational Hardware

Flight experiments flown on the Space Shuttle, the International Space Station, Mir, Skylab, and free flyers such as the Long Duration Exposure Facility, the European Retrievable Carrier, and the EFFU, provide multiple opportunities for the investigation of molecular contamination effects. Retrieved hardware from the Solar Maximum Mission satellite, Mir, and the Hubble Space Telescope has also provided the means gaining insight into contamination processes. Images from the above mentioned hardware show contamination effects due to materials processing, hardware storage, pre-flight cleaning, as well as on-orbit events such as outgassing, mechanical failure of hardware in close proximity, impacts from man-made debris, and changes due to natural environment factors.. Contamination effects include significant changes to thermal and electrical properties of thermal control surfaces, optics, and power systems. Data from several flights has been used to develop a rudimentary estimate of asymptotic values for absorptance changes due to long-term solar exposure (4000-6000 Equivalent Sun Hours) of silicone-based molecular contamination deposits of varying thickness. Recommendations and suggestions for processing changes and constraints based on the on-orbit observed results will be presented.

Pippin, Gary↗

R-Hope: Development Approach to Extreme Non-volatile Memory Reuse Onboard the Curiosity Rover

The MSL Curiosity rover landed on Mars on August~5, 2012. Over time, one of its two computers experienced critical hardware memory failure. This non-volatile NAND flash memory held file system partitions and tunable parameters needed for running rover flight software. The project assembled a design and development team to re-purpose a NOR flash memory hardware chip, only 1.5\% of the size of the NAND, to hold the file systems and parameters. The usable NOR memory required major software changes to accommodate the new limitations of slower access speeds, vastly different physical layout, and smaller size. This presentation discusses the approach, challenges, and outcomes of restoring function to the computer so it can act as a ``lifeboat'' in event of problems with the primary computer.

Peper, Nick↗

Should we attempt global (inlet engine airframe) control design?

The feasibility of multivariable design of the entire airplane control system is briefly addressed. An intermediate step in that direction is to design a control for an inlet engine augmentor system by using multivariable techniques. The supersonic cruise large scale inlet research program is described which will provide an opportunity to develop, integrate, and wind tunnel test a control for a mixed compression inlet and variable cycle engine. The integrated propulsion airframe control program is also discussed which will introduce the problem of implementing MVC within a distributed processing avionics architecture, requiring real time decomposition of the global design into independent modules in response to hardware communication failures.

Carlin, C. M.↗