Engineering PapersSearch

SEARCH · Engineering Papers

Results for “failure cause”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Metrics for Anomalous Charging currents in Polymer Tantalum Capacitors

Anomalous charging currents (ACC) in polymer tantalum capacitors may appear as a temporary short circuit that can last for dozens of milliseconds, cause failures to the parts, or cause malfunctions to fast operating electronic systems. Currently, there is no standard technique or set of metrics to evaluate the level of ACC which compares results obtained by different users and manufacturers. In this work, ACC in different types of capacitors were characterized by analysis of current transients during the constant voltage ramp or power surge testing (PST) techniques. It is shown that although the shape of transients may vary at different test conditions, the transfer charge and energy remain practically the same. The level of ACC is characterized by the energy dissipated in the part during PST. Variations of the dissipated energy with moisture content, test temperature, and voltage are evaluated. Effects of different reflow soldering conditions and long-term (up to 3000 hours) storage at 125 °C are discussed.

polymer tantalum capacitor

Achieving Improved Reliability with Failure Analysis

Reliability is the ability of a product to properly function, within specified performance limits, for a specified period of time, under the life cycle application conditions. Failure analysis is a vital tool in the effort to ensure reliability of electronic products and systems throughout their product lifecycle. Today, organizations involved in activities within the electronics supply chain are facing new challenges, not just from complex assembly styles, harsher lifecycle environments, and sophisticated supply chains, but also from customers who are demanding a quicker turn-around. Unfortunately, root cause failure analysis is often performed incompletely, leading to a poor understanding of failure mechanisms and causes and, customer dissatisfaction due to recurring failures. The PDC starts with an introduction to reliability concepts, physics of failure and an overview of failure mechanisms that affect PCBs, PCBAs and components. The PDC then dives into root cause hypothesizing techniques (Pareto, FMEA, fishbone, FTA), non-destructive and destructive analysis and, materials characterization will be discussed. Numerous failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies. What Will You Learn: Topics include: Overview of Reliability Concepts Failure mechanisms of electronic products Root cause analysis Failure analysis techniques -Non-destructive techniques (optical, CSAM etc.) -Destructive analysis (DPA, Decap, FIB etc.) -Materials characterization (XRF, EDS, TMA/DSC etc.) Who Will Benefit: Reliability engineers, failure analysis engineers, engineering managers, design engineers, component engineers, quality assurance functions and, personnel involved with reliability activities within their company.

non-destructive techniques

Independent Review Support for Phoenix Mars Mission Robotic Arm Brush Motor Failure

The Phoenix Project requested the NASA Engineering and Safety Center (NESC) perform an independent peer review of the Robotic Arm (RA) Direct Current (DC) motor brush anomalies that originated during the Mars Exploration Rover (MER) Project and recurred during the Phoenix Project. The request was to evaluate the Phoenix Project investigation efforts and provide an independent risk assessment. This includes a recommendation for additional work and assessment of the flight worthiness of the RA DC motors. Based on the investigation and findings contained within this report, the IRT concurs with the risk assessment Failure Cause / Corrective Action (FC/CA) by the project, "Failure Effect Rating "3"; Major Degradation or Total Loss of Function, Failure Cause/Corrective Action Rating Currently "4"; Unknown Cause, Uncertainty in Corrective Action."

McManamen, John P.

Flight experience with Apollo spacecraft propulsion systems

Apollo 17 ended the most successful application of rocket propulsion systems in man's history. A total of 23 developmental and manned operational flights were made. Seven hundred and sixty-three spacecraft rocket engines were flown in the program. Over 6 h of manned rocket flights were logged by the spacecraft propulsion systems and approximately one million rocket engine firings were made. One engine failure was encountered on an early unmanned flight as a result of a failure in the guidance programmer which caused the engine to operate in a manner known to cause failures. Numerous operational problems and malfunctions were observed; however, system and component redundancy prevented loss of mission objectives and never jeopardized crew safety. Performance of all systems was usually nominal and most problems were merely nuisances. This paper will present some highlights of Apollo propulsion performance and will provide a bibliography of all flight results.

Thibodaux, J. G., Jr.

Investigation of heat transfer in zirconium potassium perchlorate at low temperature: A study of the failure mechanism of the NASA standard initiator

The objective of this work was to study the reasons for the failure of pyrotechnic initiators at very low temperatures (10 to 100 K). A two-dimensional model of the NASA standard initiator was constructed to model heat transfer from the electrically heated stainless steel bridgewire to the zirconium potassium perchlorate explosive charge and the alumina charge cup. Temperature dependent properties were used in the model to simulate initiator performance over a wide range of initial temperatures (10 to 500 K). A search of the thermophysical property data base showed that pure alumina has a very high thermal conductivity at low temperatures. It had been assumed to act as a thermal insulator in all previous analyses. Rapid heat transfer from the bridgewire to the alumina at low initial temperatures was shown to cause failure of the initiators if the wire did not also make good contact with the zirconium potassium perchlorate charge. The mode is able to reproduce the results of the tests that had been conducted to investigate the cause for failure. It also provides an explanation for previously puzzling results and suggests simple design changes that will increase reliability at very low initial temperatures.

Varghese, Philip L.

Effects of Near Field Pyroshock on the Performance of a Nitramine Nitrocellulose Propellant

The overall purpose of this study is to investigate the effects of a pyroshock environment on the performance characteristics of a propellant used in pyrotechnic devices such as guillotine cutters. Near field pyroshock which is defined by acceleration amplitudes in excess of 10,000g at a frequency of greater than 10,000 Hz is a highly transient environment that has a known potential to cause failure in both structural and electronic components. A heritage pressure cartridge assembly which uses a nitramine nitrocellulose propellant with a known performance baseline will be exposed to a near field pyroshock event. The pressure cartridge will then be fired in an ambient closed bomb firing to collect pressure time history. The two performance characteristics that will be evaluated are the pressure amplitude and time to peak pressure. This data will be compared to the base-lined ambient closed bomb data to evaluate the effects of the shock on the performance of the propellant. It is expected that the pyroshock environment will cause brittle failures of the propellant increasing the surface area of said propellant. This increase of surface area should result in increased combustion rate which should show as an increased pressure peak and decreased time to peak pressure in the pressure time data.

Baca, Arcenio B.

Operating Experience Data Analysis for Digital Instrumentation and Control System Reliability and Risk Assessment in Nuclear Power Plants

The implementation of advanced digital instrumentation and control (DI&C) systems in U.S. nuclear power plants (NPPs) can bring significant advancements in reliability, monitoring, and control capabilities. However, these systems also introduce new challenges, particularly in assessing risks such as common-cause failures (CCFs) and establishing robust reliability estimates for DI&C components. Addressing these challenges is critical for ensuring the safe and efficient operation of NPPs. Recently, Idaho National Laboratory was tasked by the U.S. Nuclear Regulatory Commission (NRC) to conduct a DI&C reliability study using operating experience data from the nuclear industry. The two operating experience data sources for the study are the Institute of Nuclear Power Operations’ Industry Reporting and Information System (IRIS) and the NRC’s Licensee Event Report database which is hosted at Idaho National Laboratory at https://lersearch.inl.gov/LERSearchCriteria.aspx. This report provides a comprehensive examination of DI&C systems, including their architecture, operational advantages, and associated challenges. It reviews existing industry DI&C studies and failure mode taxonomies, along with reliability data from various industries. Through a detailed analysis of these databases, the study provides insights into DI&C system performance. Considerations should be given to incorporate DI&C failure data into the NRC's Integrated Data Collection and Coding System and updating the Reliability and Availability Data System to support ongoing DI&C reliability studies. Recommendations are also provided for modeling DI&C reliability and CCF in probabilistic risk assessment, thereby supporting risk-informed decision-making and enhancing the reliability and safety of NPPs.

22 GENERAL STUDIES OF NUCLEAR REACTORS

Diagnosing faults in autonomous robot plan execution

A major requirement for an autonomous robot is the capability to diagnose faults during plan execution in an uncertain environment. Many diagnostic researches concentrate only on hardware failures within an autonomous robot. Taking a different approach, the implementation of a Telerobot Diagnostic System that addresses, in addition to the hardware failures, failures caused by unexpected event changes in the environment or failures due to plan errors, is described. One feature of the system is the utilization of task-plan knowledge and context information to deduce fault symptoms. This forward deduction provides valuable information on past activities and the current expectations of a robotic event, both of which can guide the plan-execution inference process. The inference process adopts a model-based technique to recreate the plan-execution process and to confirm fault-source hypotheses. This technique allows the system to diagnose multiple faults due to either unexpected plan failures or hardware errors. This research initiates a major effort to investigate relationships between hardware faults and plan errors, relationships which were not addressed in the past. The results of this research will provide a clear understanding of how to generate a better task planner for an autonomous robot and how to recover the robot from faults in a critical environment.

Lam, Raymond K.

Diagnosing faults in autonomous robot plan execution

A major requirement for an autonomous robot is the capability to diagnose faults during plan execution in an uncertain environment. Many diagnostic researches concentrate only on hardware failures within an autonomous robot. Taking a different approach, the implementation of a Telerobot Diagnostic System that addresses, in addition to the hardware failures, failures caused by unexpected event changes in the environment or failures due to plan errors, is described. One feature of the system is the utilization of task-plan knowledge and context information to deduce fault symptoms. This forward deduction provides valuable information on past activities and the current expectations of a robotic event, both of which can guide the plan-execution inference process. The inference process adopts a model-based technique to recreate the plan-execution process and to confirm fault-source hypotheses. This technique allows the system to diagnose multiple faults due to either unexpected plan failures or hardware errors. This research initiates a major effort to investigate relationships between hardware faults and plan errors, relationships which were not addressed in the past. The results of this research will provide a clear understanding of how to generate a better task planner for an autonomous robot and how to recover the robot from faults in a critical environment.

Lam, Raymond K.

Predicting Time Series Outputs and Time-to-Failure for an Aircraft Controller Using Bayesian Modeling

Safety of unmanned aerial systems (UAS) is paramount, but the large number of dynamically changing controller parameters makes it hard to determine if the system is currently stable, and the time before loss of control if not. We propose a hierarchical statistical model using Treed Gaussian Processes to predict (i) whether a flight will be stable (success) or become unstable (failure), (ii) the time-to-failure if unstable, and (iii) time series outputs for flight variables. We first classify the current flight input into success or failure types, and then use separate models for each class to predict the time-to-failure and time series outputs. As different inputs may cause failures at different times, we have to model variable length output curves. We use a basis representation for curves and learn the mappings from input to basis coefficients. We demonstrate the effectiveness of our prediction methods on a NASA neuro-adaptive flight control system.

Statistics

Component damage analysis

Semiconductor breakdown in aircraft was investigated since lightning strikes induce large current and voltage pulses which may cause failure. Work was done to determine whether or not these voltages and currents cause upset or damage to active or passive components. Failure thresholds were studied extensively and an assessment was made of the vulnerability of a system to a transient environment.

Franklin A Fisher

Autonomous power management and distribution

The goal of the Autonomous Power System program is to develop and apply intelligent problem solving and control to the Space Station Freedom's electric power testbed being developed at NASA's Lewis Research Center. Objectives are to establish artificial intelligence technology paths, craft knowledge-based tools and products for power systems, and integrate knowledge-based and conventional controllers. This program represents a joint effort between the Space Station and Office of Aeronautics and Space Technology to develop and demonstrate space electric power automation technology capable of: (1) detection and classification of system operating status, (2) diagnosis of failure causes, and (3) cooperative problem solving for power scheduling and failure recovery. Program details, status, and plans will be presented.

Dolce, Jim

High Reliability Requires More than Providing Spares

It is sometimes optimistically hoped that a space life support system can be kept working throughout a long duration mission by repairing failed components, as long as sufficient spares are flown. It is usually assumed that the components have constant known failure rates. Then the needed numbers of spares can be computed to have any particular probability that all failed components can be replaced by available spares. This approach can provide high reliability if its favorable assumptions, including constant known failure rates, are satisfied. Other favorable assumptions are that the failures are statistically independent, repair will be successful without causing further failures, and all failures are due to internal component failures. These assumptions are not usually justified. The failure rates may be estimates that are inadequately verified because of insufficient testing. Failure rates may change due to materials substitutions, manufacturing changes, redesigns to fix failures, and new failures caused by redesigns. Failures that are not statistically independent may result from one common cause, such as a design or manufacturing error or a cascade of cause and effect, possibly caused by an external event such as a power outage. Repair may be unsuccessful or cause damage. Many failures occur at component interfaces or at the overall systems level, not within isolated components. Other failures causes are completely external to the system, due to assembly, maintenance, and operational errors or to unexpected environmental challenges. Replacement with sufficient spares can compensate for expected internal component failures but may not be able to cope with unpredictable design and manufacturing flaws, human errors, and environmental impacts. Reliability estimates based on providing sufficient spares to compensate for expected failures may be far too high. They are essentially upper bounds on reliability that might be approached if many frequent but often unconsidered failure causes can be eliminated.

spares

Derivation and application of hard deadlines for real-time control systems

The computation-time delay in the feedback controller of a real-time control system may cause failure to update the control input during one or more sampling periods. If this delay exceeds a certain limit called a hard deadline, either the necessary conditions for system stability are violated or the system leaves the allowed state-space. In such a case a dynamic failure is said to occur to the system. A method for calculating the hard deadlines in linear time-invariant control systems by considering system stability and the allowed state-space is presented. To derive necessary conditions for (asymptotic) system stability, the state difference equation is modified based on an assumed maximum delay and the probability distribution of delays whose magnitudes are less than, or equal to, the assumed maximum delay. Moreover, the allowed state-space - which is derived from given input and state constraints - is used to calculate the hard deadline as a function of time and the system state. A one-shot delay model in which a single event causes a dynamic failure is also considered. The knowledge of hard deadline is then applied to the design of error recovery in a triple modular redundant (TMR) controller computer.

Shin, Kang G.

Derivation of hard deadlines for real-time control systems

The computation-time delay in the feedback controller of a real-time control system may cause failure to update the control input during one or more sampling periods. A dynamic failure is said to occur if this delay exceeds a certain limit called a hard deadline. The authors present a method for calculating the hard deadlines in linear time-invariant control systems. To derive necessary conditions for (asymptotic) system stability, the state difference equation is modified based on an assumed maximum delay and the probability distribution of delays whose magnitudes are less than, or equal to, the assumed maximum delay. Moreover, the allowed state-space-which is derived from given input and state constraints-is used to calculate the hard deadline as a function of time and the system state. The authors consider a one-shot delay model in which a single event causes a dynamic failure.

Shin, Kang G.

Reliability and Maintainability Analysis of a High Air Pressure Compressor Facility

This paper discusses a Reliability, Availability, and Maintainability (RAM) independent assessment conducted to support the refurbishment of the Compressor Station at the NASA Langley Research Center (LaRC). The paper discusses the methodologies used by the assessment team to derive the repair by replacement (RR) strategies to improve the reliability and availability of the Compressor Station (Ref.1). This includes a RAPTOR simulation model that was used to generate the statistical data analysis needed to derive a 15-year investment plan to support the refurbishment of the facility. To summarize, study results clearly indicate that the air compressors are well past their design life. The major failures of Compressors indicate that significant latent failure causes are present. Given the occurrence of these high-cost failures following compressor overhauls, future major failures should be anticipated if compressors are not replaced. Given the results from the RR analysis, the study team recommended a compressor replacement strategy. Based on the data analysis, the RR strategy will lead to sustainable operations through significant improvements in reliability, availability, and the probability of meeting the air demand with acceptable investment cost that should translate, in the long run, into major cost savings. For example, the probability of meeting air demand improved from 79.7 percent for the Base Case to 97.3 percent. Expressed in terms of a reduction in the probability of failing to meet demand (1 in 5 days to 1 in 37 days), the improvement is about 700 percent. Similarly, compressor replacement improved the operational availability of the facility from 97.5 percent to 99.8 percent. Expressed in terms of a reduction in system unavailability (1 in 40 to 1 in 500), the improvement is better than 1000 percent (an order of magnitude improvement). It is worthy to note that the methodologies, tools, and techniques used in the LaRC study can be used to evaluate similar high value equipment components and facilities. Also, lessons learned in data collection and maintenance practices derived from the observations, findings, and recommendations of the study are extremely important in the evaluation and sustainment of new compressor facilities.

Safie, Fayssal M.

Oversimplification of Systems Engineering Goals, Processes, and Criteria in NASA Space Life Support

This paper investigates the oversimplification of the inherently complex systems engineering process in space life support. The standard systems engineering process steps are described. The International Space Station (ISS) life support system is explained with its goals and performance criteria. Although it is not usually emphasized, the essential function of developing a hierarchy of systems and subsystems is to simplify the design process. The System Complexity Metric (SCM) shows how this di-vide-and-conquer approach also reduces the system complexity. The complete systems engineering process has many detailed steps. It is often simplified because of the effort required and the human limitations on working memory and decision span. Systems analysis demands slow, logical, and fo-cused thinking but is often bypassed in favor of quick, intuitive, subconscious “gut feel.” A study of 100 system designs found examples of 12 specific mental mistakes, such as ignoring stakeholder needs, and these mistakes are essentially oversimplifications of the systems engineering process. An analysis of space life support goals, options, criteria, and processes found 11 examples of oversimplifications in systems engineering, such as neglecting safety and cost. All these 11 oversimplifications could be traced to one or more of the 12 previously identified mental mistakes or other well-known ones, such as ig-noring sunk costs. Oversimplification of the systems engineering process is rarely noticed but is a common and harmful problem. A study of failures in 50 different space systems found that problems in systems engineering caused failures and often led to errors in design, development, and test that further contributed to failure. It seems that more diligent systems engineering could prevent many project problems and failures, but projects seem to be more guided by “gut feel” based on tradition, authority, and consensus than on the logical, rational systems engineering approach.

Simplified systems engineering

A scheme for fault tolerance in earth sensors

A system is presented that uses dual-redundant earth sensors to measure pitch and roll errors of a three-axis stabilized spacecraft, with provision for (1) autonomously detecting and identifying a faulty earth sensor, and (2) automatically selecting the outputs of the fault-free sensor for closed-loop attitude control, before failures cause major problems. A brief description is given of the system, and various failure modes of earth sensors and their effects are discussed. Novel techniques and algorithms for automatic fault detection, identification, and reconfiguration (FDIR) of dual-redundant earth sensors are developed. The algorithms are validated through computer simulations, and the results are presented. The proposed scheme can easily be implemented without much penalty on hardware, power consumption, and processing time.

Murugesan, S.