Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “hardware failure”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Pyrotechnic system failures: Causes and prevention

Although pyrotechnics have successfully accomplished many critical mechanical spacecraft functions, such as ignition, severance, jettisoning and valving (excluding propulsion), failures continue to occur. Provided is a listing of 84 failures of pyrotechnic hardware with completed design over a 23-year period, compiled informally by experts from every NASA Center, as well as the Air Force Space Division and the Naval Surface Warfare Center. Analyses are presented as to when and where these failures occurred, their technical source or cause, followed by the reasons why and how these kinds of failures persist. The major contributor is a fundamental lack of understanding of the functional mechanisms of pyrotechnic devices and systems, followed by not recognizing pyrotechnics as an engineering technology, insufficient manpower with hands-on experience, too few test facilities, and inadequate guidelines and specifications for design, development, qualification and acceptance. Recommendations are made on both a managerial and technical basis to prevent failures, increase reliability, improve existing and future designs, and develop the technology to meet future requirements.

Bement, Laurence J.↗

The Certification of Environmental Chambers for Testing Flight Hardware

The JPL chamber certification process for ensuring that test chambers used to test flight hardware meet a minimum standard is critical to the safety of the hardware and personnel. Past history has demonstrated that this process is important due to the catastrophic incidents that could occur if the chamber is not set up correctly. Environmental testing is one of the last phases in the development of a subsystem, and it typically occurs just before integration of flight hardware into the fully assembled flight system. A seemingly insignificant -miscalculation or missed step can necessitate rebuilding or replacing a subsystem due to over-testing or damage from the test chamber. Conversely, under-testing might fail to detect weaknesses that might cause failure when the hardware is in service. This paper describes the process that identifies the many variables that comprise the testing scenario and screening of as built chambers, the training of qualified operators, and a general "what-to-look-for" in minimum standards.

Vacuum Chamber↗

The Certification of Environmental Chambers for Testing Flight Hardware

The JPL chamber certification process for ensuring that test chambers used to test flight hardware meet a minimum standard is critical to the safety of the hardware and personnel. Past history as demonstrated that this process is important due to the catastrophic incidents that could occur if the chamber is not set up correctly. Environmental testing is one of the last phases in the development of a subsystem, and it typically occurs just before integration of flight hardware into the fully assembled flight system. A seemingly insignificant -miscalculation or missed step can necessitate rebuilding or replacing a subsystem due to over-testing or damage from the test chamber. Conversely, under-testing might fail to detect weaknesses that might cause failure when the hardware is in service. This paper describes the process that identifies the many variables that comprise the testing scenario and screening of as built chambers, the training of qualified operators, and a general "what-to-look-for" in minimum standards.

Thermal↗

Prototype part task trainer: A remote manipulator system simulator

The Part Task Trainer program (PTT) is a kinematic simulation of the Remote Manipulator System (RMS) for the orbiter. The purpose of the PTT is to supply a low cost man-in-the-loop simulator, allowing the student to learn operational procedures which then can be used in the more expensive full scale simulators. PTT will allow the crew members to work on their arm operation skills without the need for other people running the simulation. The controlling algorithms for the arm were coded out of the Functional Subsystem Requirements Document to ensure realistic operation of the simulation. Relying on the hardware of the workstation to provide fast refresh rates for full shaded images allows the simulation to be run on small low cost stand alone work stations, removing the need to be tied into a multi-million dollar computer for the simulation. PTT will allow the student to make errors which in full scale mock up simulators might cause failures or damage hardware. On the screen the user is shown a graphical representation of the RMS control panel in the aft cockpit of the orbiter, along with a main view window and up to six trunion and guide windows. The dials drawn on the panel may be turned to select the desired mode of operation. The inputs controlling the arm are read from a chair with a Translational Hand Controller (THC) and a Rotational Hand Controller (RHC) attached to it.

Shores, David↗

Elimination of Potential Electrical Stress During EMC (CS01) Testing

This viewgraph presentation reviews possible ways to eliminate electrical stress during Electromagneticic Compatibility (EMC) testing. The presentation reviews tests that have had problems due to electrical stress. On December 5, 1995 Cassini Radar instrument failed a functional test in preparation for EMC conducted susceptibility (CSO 1 ) testing. The instrument power supply did not turn on as required, and failure occurred prior to injection of CS test stimulus. A investigation of the failure was conducted. A PSPICE simulation of Cassini Radar 30V line using the EMC test setup was performed; the result of the simulation was an oscillation on the 30V input of the power supply. In another case: on December 28, 1999 an oscillation occurred on the input power line of the SlRTF Infrared Array Camera (IRAC) while preparing to perform CSOI testing, Resulted in damage to flight hardware. Subsequent to failure, JPL provided GSFC history and corrective action from Cassini Radar CSOI test failure GSFC implemented the same corrective action as JPL, except that the value of the resistor connected across the isolation transformer primary winding is 2.5 ohms instead of 50 ohms. Three recommendations are made: (1) Make EMC test community aware of the problem and potential solutions by presenting papers at major environmental test conferences (2) Include warnings and safeguards in EMC test requirements and procedures (3) Try to convince EMC test equipment suppliers to design a CSOl test fixture similar to fixture shown in the diagram

Cassini radar↗

Independent Orbiter Assessment (IOA): Analysis of the DPS subsystem

The results of the Independent Orbiter Assessment (IOA) of the Failure Modes and Effects Analysis/Critical Items List (FMEA/CIL) is presented. The IOA approach features a top-down analysis of the hardware to independently determine failure modes, criticality, and potential critical items. The independent analysis results corresponding to the Orbiter Data Processing System (DPS) hardware are documented. The DPS hardware is required for performing critical functions of data acquisition, data manipulation, data display, and data transfer throughout the Orbiter. Specifically, the DPS hardware consists of the following components: Multiplexer/Demultiplexer (MDM); General Purpose Computer (GPC); Multifunction CRT Display System (MCDS); Data Buses and Data Bus Couplers (DBC); Data Bus Isolation Amplifiers (DBIA); Mass Memory Unit (MMU); and Engine Interface Unit (EIU). The IOA analysis process utilized available DPS hardware drawings and schematics for defining hardware assemblies, components, and hardware items. Each level of hardware was evaluated and analyzed for possible failure modes and effects. Criticality was assigned based upon the severity of the effect for each failure mode. Due to the extensive redundancy built into the DPS the number of critical items are few. Those identified resulted from premature operation and erroneous output of the GPCs.

Lowery, H. J.↗

PRELIMINARY JPSS-3 VIIRS POLARIZATION SENSITIVITY AND COMPARISON WITH S-NPP, JPSS-1 AND -2

The Visible-Infrared Imaging Radiometer Suite (VIIRS) was first launched on-board the Suomi National Polar-orbiting Partnership (S-NPP) spacecraft in October of 2011. There have been three subsequent builds of the VIIRS sensor for the Joint Polar Satellite System (JPSS) program with JPSS-1, -2 and -3 having launch dates of November 2017, March 2022 and 2026 respectively. There is also a JPSS-4 VIIRS, that is in hardware integration during 2020, with a launch date of 2031. VIIRS has 22 bands: 7 thermal emissive bands (TEBs), 14 reflective solar bands (RSBs) and a Day Night Band (DNB). Ocean Color/Chlorophyll (OCC) products use calibrated Science Data Records (SDRs) for bands M1-M7(0.412-0.865μm) to compute their ocean chemistry products. These bands require accurate polarization sensitivity characterization to compensate for polarized upwelling Rayleigh scatter and produce accurate OCC Environment Data Products (EDRs). VIIRS polarization sensitivity requirement failures have driven hardware modifications to the bandpass filters and dichroic beam splitter over the program. This paper will discuss the preliminary JPSS-3 polarization results and how these hardware modifications, as the JPSS program progresses, have affected the sensor performance. Comparisons of the polarization sensitivities between sensor builds will be discussed along with the hardware modifications that contributed to their differences.

VIIRS↗

Preliminary JPSS-3 VIIRS Polarization Sensitivity and Comparison with S-NPP, JPSS-1 and -2

The Visible-Infrared Imaging Radiometer Suite (VIIRS) was first launched on-board the Suomi National Polar-orbiting Partnership (S-NPP) spacecraft in October of 2011. There have been three subsequent builds of the VIIRS sensor for the Joint Polar Satellite System (JPSS) program with JPSS-1, -2 and -3 having launch dates of November 2017, March 2022 and 2026 respectively. There is also a JPSS-4 VIIRS, that is in hardware integration during 2020, with a launch date of 2031. VIIRS has 22 bands: 7 thermal emissive bands (TEBs), 14 reflective solar bands (RSBs) and a Day Night Band (DNB). Ocean Color/Chlorophyll (OCC) products use calibrated Science Data Records (SDRs) for bands M1-M7 (0.412-0.865μm) to compute their ocean chemistry products. These bands require accurate polarization sensitivity characterization to compensate for polarized upwelling Rayleigh scatter and produce accurate OCC Environment Data Products (EDRs). VIIRS polarization sensitivity requirement failures have driven hardware modifications to the bandpass filters and dichroic beam splitter over the program. This paper will discuss the preliminary JPSS-3 polarization results and how these hardware modifications, as the JPSS program progresses, have affected the sensor performance. Comparisons of the polarization sensitivities between sensor builds will be discussed along with the hardware modifications that contributed to their differences.

VIIRS↗

Lox/Gox related failures during Space Shuttle Main Engine development

Specific rocket engine hardware and test facility system failures are described which were caused by high pressure liquid and/or gaseous oxygen reactions. The failures were encountered during the development and testing of the space shuttle main engine. Failure mechanisms are discussed as well as corrective actions taken to prevent or reduce the potential of future failures.

Cataldo, C. E.↗

Lunar Reconnaissance Orbiter (LRO) Sun Safe Mode

The Lunar Reconnaissance Orbiter (LRO), a spacecraft designed and built at the National Aeronautics and Space Administration s (NASA) Goddard Space Flight Center (GSFC) in Greenbelt, MD, was launched on June 18, 2009 from Cape Canaveral. It is currently in orbit about the Moon taking detailed science measurements and providing a highly accurate mapping of the suface in preparation for the future return of astronauts to a permanent moon base. Onboard the spacecraft is a complex set of algorithms designed by the attitude control engineers at GSFC to control the pointig for all operational events, including anomalies that require the spacecraft to be put into a well known attitude configuration for a sufficiently long duration to allow for the investigation and correction of the anomaly. GSFC level requirements state that each spacecraft s control system design must include a configuration for this pointing and lso be able to maintain a thermally safe and power positive attitude. This stable control algorithm for anomalous events is commonly referred to as the safe mode and consists of control logic thatwill put the spacecraft in this safe configuration defined by the spacecraft s hardware, power and environment capabilities and limitations. The LRO Sun Safe mode consists of a coarse sun-pointing set of algorithms that puts the spacecraft into this thermally safe and power positive attitude and can be achieved wihin a required amount of time from any initial attitude, provided that the system momentum is within the momentum capability of the reaction wheels. On LRO the Sun Safe mode makes use of coarse sun sensors (CSS), an inertial reference unit (IRU) and reaction wheels (RW) to slew the spacecraft to a solar inertial pointing. The CSS and reaction wheels have some level of redundancy because of their numbers. However, the IRU is a single-point-failure piece of hardware. Without the rate information provided by the IRU, the Sun Safe control algorithms could not maintain the required pointing, so a sub-mode of the Sun Safe mode that does not use the IRU was designed. This submode, referred to as the Sun Safe Gyroless control mode, consists of an algorithm that estimates rate information from the CSS and the RW measurements. RW momentum information is used to estimate the body rate parallel to the target sunline, which CSS alone would not be able to observe. Sun Safe can be autonomously, or via ground command, entered from any other control mode and in the event the IRU is not providing rate information, the control mode is switched to the gyroless submode. This paper looks at the design of the Sun Safe modes and discusses the constraints placed on the algorithm and how the mode wored around these constraints. Items of particular interest include CSS placement on the Solar Array (SA) and its implications to design, estimation of body rate information for the Sun Safe Gyroless control mode, and the effect of solar eclipse on each of the Sun Safe modes. Placing CSS on the SA was necessary for the means to put the Sun along the targeted sun-line, nominally normal to the SA panels, for all operational considerations. This had design implications for determining a sun vector during normal SA operations, if one or both gimbals become inoperable and when the SA is in a stowed configuration. The ability of body rate estimation in Sun Safe Gyroless not only uses CSS sun vector data but requires RW momentum measuremens to estimate rates parallel to the sun-line. LRO encounters solar eclipses of some length for most of its orbits about the Moon. With the lack of CSS measurement data a design was implemented in both Sun Safe and Sun Safe Gyroless, they differ because of having or not having IRU measurement data, to carry the spacecraft through these eclipse periods. This paper also includes some discussion of sun avoidance and how it affected design decisions during nominal and eclipse perids for each of the Sun Safe modes.

Garrick, Joseph↗

Modeling interconnections of safety and financial performance of nuclear power plants, part 3: Spatiotemporal probabilistic physics-of-failure analysis and its connection to safety and financial performance

Here, this paper is a byproduct of a line of research by the authors to analyze interrelationships of safety and financial performance of nuclear power plants (NPPs). The result of this line of research is summarized in three parts: Part 1 covers a categorical review of relevant literature and the theoretical bases that support the methodological developments in Part 2. Part 2 introduces an Integrated Enterprise Risk Management (I-ERM) methodological framework to quantify the interconnections of safety and financial performance with a focus on operation and maintenance (O&M) of NPPs. Part 2 has also demonstrated the applicability and values of the I-ERM methodology through an NPP case study. This paper is Part 3, where detailed development and implementation of one of the I-ERM modules, i.e., probabilistic physics-of-failure (PPoF) analysis, and its connection with safety and financial performance is reported. In this article, the physical failure modeling for hardware components is advanced by incorporating finite element analysis (FEA) into PPoF analysis and coupling the FEA-based PPoF with the maintenance performance through a renewal process model. This article covers two scientific contributions: (i) first-of-its-kind incorporation of FEA into the PPoF model of thermal fatigue for NPP components; and (ii) advancing the interface between the PPoF analysis and the renewal process model in order to deal with spatiotemporal FEA outputs and to efficiently estimate the physical transition rates even when the PPoF outputs are dominated by success data. Through the incorporation of FEA, the resolution of the PPoF analysis is enhanced as spatiotemporal conditions such as stress and temperature can be considered explicitly instead of relying on simplified assumptions or analytical models with reduced spatiotemporal dimensions. To demonstrate an application of the FEA-based PPoF analysis and its coupling with maintenance through the renewal process model, a case study is conducted using excess letdown elbow piping in the chemical and volume control system of a Pressurized Water Reactor.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Report on local data recovery approaches suitable for weather and climate prediction (Deliverable 1.3) (V.1.0)

Numerical weather and climate prediction rates as one of the scientific applications whose accuracy improvements greatly depend on the growth of the available computing power. As the number of cores in top computing facilities pushes into the millions, increasing average frequency of hardware and software failures forces users to review their algorithms and systems in order to protect simulations from breakdown. This report surveys approaches for fault-tolerance in numerical algorithms and system resilience in parallel simulations from the perspective of numerical weather and climate prediction systems. A selection of existing strategies is analyzed, featuring interpolation-restart and compressed checkpointing for the numerics, in-memory checkpointing, user-level failure mitigation-based and backup-based methods for the systems. Numerical examples showcase the performance of the techniques in addressing faults, with particular emphasis on iterative solvers for linear systems, a staple of atmospheric fluid flow solvers. The potential impact of these strategies is discussed in relation to current development of numerical weather prediction algorithms and systems towards the exascale. Trade-offs between performance, efficiency and effectiveness of resiliency strategies are analyzed and some recommendations outlined for future developments.

97 MATHEMATICS AND COMPUTING↗

A specification-based approach to concurrent structure verification in multiprocessor systems

A recently initiated research project concerned with the concurrent detection of software errors and errors due to physical failures in the hardware of multiprocessor systems is described in this paper. An approach to error detection is described, which is specification based and relies on the structural verification of program control flow and data structure integrity. The techniques discussed utilize the hardware redundancy inherent in parallel processing systems to provide verification of both program structure and data concurrently with program execution.

Fuchs, W. Kent↗

The Management and Security Expert (MASE)

The Management and Security Expert (MASE) is a distributed expert system that monitors the operating systems and applications of a network. It is capable of gleaning the information provided by the different operating systems in order to optimize hardware and software performance; recognize potential hardware and/or software failure, and either repair the problem before it becomes an emergency, or notify the systems manager of the problem; and monitor applications and known security holes for indications of an intruder or virus. MASE can eradicate much of the guess work of system management.

Miller, Mark D.↗

A fault-tolerant intelligent robotic control system

This paper describes the concept, design, and features of a fault-tolerant intelligent robotic control system being developed for space and commercial applications that require high dependability. The comprehensive strategy integrates system level hardware/software fault tolerance with task level handling of uncertainties and unexpected events for robotic control. The underlying architecture for system level fault tolerance is the distributed recovery block which protects against application software, system software, hardware, and network failures. Task level fault tolerance provisions are implemented in a knowledge-based system which utilizes advanced automation techniques such as rule-based and model-based reasoning to monitor, diagnose, and recover from unexpected events. The two level design provides tolerance of two or more faults occurring serially at any level of command, control, sensing, or actuation. The potential benefits of such a fault tolerant robotic control system include: (1) a minimized potential for damage to humans, the work site, and the robot itself; (2) continuous operation with a minimum of uncommanded motion in the presence of failures; and (3) more reliable autonomous operation providing increased efficiency in the execution of robotic tasks and decreased demand on human operators for controlling and monitoring the robotic servicing routines.

Marzwell, Neville I.↗

Expert System for UNIX System Reliability and Availability Enhancement

Highly reliable and available systems are critical to the airline industry. However, most off-the-shelf computer operating systems and hardware do not have built-in fault tolerant mechanisms, the UNIX workstation is one example. In this research effort, we have developed a rule-based Expert System (ES) to monitor, command, and control a UNIX workstation system with hot-standby redundancy. The ES on each workstation acts as an on-line system administrator to diagnose, report, correct, and prevent certain types of hardware and software failures. If a primary station is approaching failure, the ES coordinates the switch-over to a hot-standby secondary workstation. The goal is to discover and solve certain fatal problems early enough to prevent complete system failure from occurring and therefore to enhance system reliability and availability. Test results show that the ES can diagnose all targeted faulty scenarios and take desired actions in a consistent manner regardless of the sequence of the faults. The ES can perform designated system administration tasks about ten times faster than an experienced human operator. Compared with a single workstation system, our hot-standby redundancy system downtime is predicted to be reduced by more than 50 percent by using the ES to command and control the system.

Xu, Catherine Q.↗