Engineering PapersSearch

SEARCH · Engineering Papers

Results for “root cause analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Root cause analysis of a molten salt pump in FLUSTFA

The primary salt pump installed in the high-temperature FLUoride Salt Test Facility (FLUSTFA) was successfully operated for some time, but later ceased operation. To understand what occurred, a Root Cause Analysis (RCA) was performed. Steps taken to try to get the pump operational include adjusting the shaft position, increasing the heating power of the tape heaters on the pump volute, and manually rotating the pump shaft. While removing the insulation, corrosion was noted on the outside of the pump volute, and decolorization of the insulation and tape heaters was observed. Significant corrosion products were also observed in the pump itself and the piping connected to the pump. The nitrogen cover gas was maintained from before salt was introduced into the loop until the pump was dismounted and continues to be maintained even after the pump was removed. After considering probable scenarios, causes were assigned and corrective actions were developed to prevent those causes. Then, the RCA was presented to an advisory committee for review, the “Review Committee,” consisting of experts in large molten salt systems: Brandon Haugh, David Holcomb, Kevin Robb, and Vicente Rojas. As a result, the advisory committee provided comprehensive feedback, which have been incorporated into a revised RCA. Findings have then been summarized and reported in this publication.

22 - GENERAL STUDIES OF NUCLEAR REACTORS

Root cause analysis

Explore the source record for details and available documents.

root cause analysis

Root Cause Analysis of the Data Refinement Process – Medical Conditions Capability Resource Tables

The medical system for spaceflight thus far has been designed to support missions in low earth orbit (LEO). Crew capabilities are limited and heavily dependent on the team of medical support staff at Mission Control Center (MCC) to guide diagnosis and management. However, missions to the Moon and Mars will suffer from several constraints that will make this ground support focused approach to care ineffective. In order to update and modify medical system design, NASA has relied on Probabilistic Risk Assessment (PRA) modeling to mitigate medical risk through trade space analysis. Specifically, capability resource tables (CRT’s) were developed to create a dataset of resources required to manage a list of accepted medical conditions significant in exploration spaceflight. With 120 conditions, this dataset contained hundreds of capabilities and thousands of resources with tens of thousands of cells of data. Initially these tables were built in excel for high throughput during development, but ultimately had to be transferred, managed, and modified into the Evidence Library database for modeling purposes. The process of collating and reviewing the Evidence Library revealed numerous errors in the dataset that had to be corrected through iterative changes. Several error types emerged during this process and can be broken into specific classifications defined as “input”, “transcription”, “structural”, “branching”, and “information”. In reviewing these error types through the root cause analysis (RCA) approach, we were able to identify the contributors to these errors which included single data review points, changing product end goals, limited software selection, time constraints and several others. By reviewing and evaluating the underlying causes we can provide possible system improvements that can be implemented for current and future data management in PRA model inputs.

A. Anderson

NASA's Evolutionary Xenon Thruster (NEXT) Power Processing Unit (PPU) Capacitor Failure Root Cause Analysis

The NASA's Evolutionary Xenon Thruster (NEXT) project is developing an advanced ion propulsion system for future NASA missions for solar system exploration. A critical element of the propulsion system is the Power Processing Unit (PPU) which supplies regulated power to the key components of the thruster. The PPU contains six different power supplies including the beam, discharge, discharge heater, neutralizer, neutralizer heater, and accelerator supplies. The beam supply is the largest and processes up to 93+% of the power. The NEXT PPU had been operated for approximately 200+ hours and has experienced a series of three capacitor failures in the beam supply. The capacitors are in the same, nominally non-critical location the input filter capacitor to a full wave switching inverter. The three failures occurred after about 20, 30, and 135 hours of operation. This paper provides background on the NEXT PPU and the capacitor failures. It discusses the failure investigation approach, the beam supply power switching topology and its operating modes, capacitor characteristics and circuit testing. Finally, it identifies root cause of the failures to be the unusual confluence of circuit switching frequency, the physical layout of the power circuits, and the characteristics of the capacitor.

Soeder, James F.

NASA's Evolutionary Xenon Thruster (NEXT) Power Processing Unit (PPU) Capacitor Failure Root Cause Analysis

The NASA s Evolutionary Xenon Thruster (NEXT) project is developing an advanced ion propulsion system for future NASA missions for solar system exploration. A critical element of the propulsion system is the Power Processing Unit (PPU) which supplies regulated power to the key components of the thruster. The PPU contains six different power supplies including the beam, discharge, discharge heater, neutralizer, neutralizer heater, and accelerator supplies. The beam supply is the largest and processes up to 93+% of the power. The NEXT PPU had been operated for approximately 200+ hr and has experienced a series of three capacitor failures in the beam supply. The capacitors are in the same, nominally non-critical location-the input filter capacitor to a full wave switching inverter. The three failures occurred after about 20, 30, and 135 hr of operation. This paper provides background on the NEXT PPU and the capacitor failures. It discusses the failure investigation approach, the beam supply power switching topology and its operating modes, capacitor characteristics and circuit testing. Finally, it identifies root cause of the failures to be the unusual confluence of circuit switching frequency, the physical layout of the power circuits, and the characteristics of the capacitor.

Soeder, James F.

Feedback and Oscillations: Constructing Feedback Systems for Root Cause Analysis of Oscillations in Power Grids

Dynamic phenomena linked to inverter-based resources (IBRs) have gained global attention. Several IBR-induced dynamics have caused bulk power system-connected wind or solar power plants to trip, and some have even led to widespread outages. In addition, many oscillations have been observed involving IBR power plants. In 2023, the IEEE Power & Energy Society (PES) IBR Subsynchronous Oscillations (SSO) task force published a journal article, “Real-World Subsynchronous Oscillation Events in Power Grids With High Penetrations of Inverter-Based Resources,” in which 19 IBR oscillation events were examined for their causation. Earlier in 2020, another PES task force article, “Definition and Classification of Power System Stability-Revisited & Extended,” authored by prominent academics, introduced converter-driven stability as a new category of stability. The international power grid industry community also took action by publishing the CIGRE Green Book, Power System Dynamic Modelling and Analysis in Evolving Networks (led by Babak Badrzadeh and Zia Emin) in 2024. In August 2024, the Energy Systems Integration Group (ESIG) released a practical guide led by Nick Miller, “Diagnosis and Mitigation of Observed Oscillations in IBR-Dominant Power System: A Practical Guide.” The goal of the guide is to assist practicing engineers in making initial judgments and conducting detailed analyses about oscillations. Finally, when addressing the classification of stability and oscillations, the guide emphasizes a causality-based taxonomy for grouping, such as voltage control-induced oscillations, synchronization-induced oscillations, and frequency or active power control-induced oscillations.

Fan, Lingling [Univ. of South Florida, Tampa, FL (

VIIRS On-Orbit Optical Anomaly - Investigation, Analysis, Root Cause Determination and Lessons Learned

A gradual, but persistent, decrease in the optical throughput was detected during the early commissioning phase for the Suomi National Polar-Orbiting Partnership (SNPP) Visible Infrared Imager Radiometer Suite (VIIRS) Near Infrared (NIR) bands. Its initial rate and unknown cause were coincidently coupled with a decrease in sensitivity in the same spectral wavelength of the Solar Diffuser Stability Monitor (SDSM) raising concerns about contamination or the possibility of a system-level satellite problem. An anomaly team was formed to investigate and provide recommendations before commissioning could resume. With few hard facts in hand, there was much speculation about possible causes and consequences of the degradation. Two different causes were determined as will be explained in this paper. This paper will describe the build and test history of VIIRS, why there were no indicators, even with hindsight, of an on-orbit problem, the appearance of the on-orbit anomaly, the initial work attempting to understand and determine the cause, the discovery of the root cause and what Test-As-You-Fly (TAYF) activities, can be done in the future to greatly reduce the likelihood of similar optical anomalies. These TAYF activities are captured in the lessons learned section of this paper.

Iona, Glenn

A Perspective on DSN System Performance Analysis

This paper discusses the performance analysis effort being carried out in the NASA Deep Space Network. The activity involves root cause analysis of failures and assessment of key performance metrics. The root cause analysis helps pinpoint the true cause of observed problems so that proper correction can be effected. The assessment currently focuses on three aspects: (1) data delivery metrics such as Quantity, Quality, Continuity, and Latency; (2) link-performance metrics such as antenna pointing, system noise temperature, Doppler noise, frequency and time synchronization, wide-area-network loading, link-configuration setup time; and (3) reliability, maintainability, availability metrics. The analysis establishes whether the current system is meeting its specifications and if so, how much margin is available. The findings help identify the weak points in the system and direct attention of programmatic investment for performance improvement.

Deep Space Network (DSN)

Printed Circuit Board Inspection and Quality Control - PCB Failure Causes and Cures

This two day workshop will discuss a range of topics, including root cause analysis, physics-of-failure principles and failure mechanisms in printed circuit boards. Printed circuit boards (PCBs) are the baseline for electronics manufacturing upon which electronic components are mounted and formed into electronic systems. PCBs are used in a variety of electronic circuits from simple one-transistor amplifiers to large super computers. A PCB serves three main functions: 1) it provides the necessary mechanical support for the components in the circuit 2) it provides the necessary electrical interconnections, and 3) it bears some form of legend which identifies the components it carries. The failure modes on the PCBs can be categorized in a hierarchical structure, in which the mechanisms and causes are site or location dependant. Specimen preparation techniques, non-destructive and destructive analysis, and materials characterization will also be discussed. The first day of the workshop will present methodologies for identifying potential failure mechanisms in electronics based on the failure history and, systematic approaches to root cause analysis. The second day will cover failure analysis techniques geared towards various failure mechanisms, along with numerous component and PCB assembly failure analysis case studies that illustrate the techniques and analysis. Failure analysis case studies will be used to illustrate the techniques and analysis principles to arrive at the root cause(s) of field failures on printed circuit boards, active components, and assemblies.

Printed circuit boards

Quality Interaction Between Mission Assurance and Project Team Members

Mission Assurance independent assessments started during the development cycle and continued through post launch operations. In operations, Health and Safety of the Observatory is of utmost importance. Therefore, Mission Assurance must ensure requirements compliance and focus on process improvements required across the operational systems including new/modified products, tools, and procedures. The deployment of the interactive model involves three objectives: Team member Interaction, Good Root Cause Analysis Practices, and Risk Assessment to avoid reoccurrences. In applying this model, we use a metric based measurement process and was found to have the most significant effect, which points to the importance of focuses on a combination of root cause analysis and risk approaches allowing the engineers the ability to prioritize and quantify their corrective actions based on a well-defined set of root cause definitions (i.e. closure criteria for problem reports), success criteria and risk rating definitions.

quality

Integrated System Health Management (ISHM) for Test Stand and J-2X Engine: Core Implementation

ISHM capability enables a system to detect anomalies, determine causes and effects, predict future anomalies, and provides an integrated awareness of the health of the system to users (operators, customers, management, etc.). NASA Stennis Space Center, NASA Ames Research Center, and Pratt & Whitney Rocketdyne have implemented a core ISHM capability that encompasses the A1 Test Stand and the J-2X Engine. The implementation incorporates all aspects of ISHM; from anomaly detection (e.g. leaks) to root-cause-analysis based on failure mode and effects analysis (FMEA), to a user interface for an integrated visualization of the health of the system (Test Stand and Engine). The implementation provides a low functional capability level (FCL) in that it is populated with few algorithms and approaches for anomaly detection, and root-cause trees from a limited FMEA effort. However, it is a demonstration of a credible ISHM capability, and it is inherently designed for continuous and systematic augmentation of the capability. The ISHM capability is grounded on an integrating software environment used to create an ISHM model of the system. The ISHM model follows an object-oriented approach: includes all elements of the system (from schematics) and provides for compartmentalized storage of information associated with each element. For instance, a sensor object contains a transducer electronic data sheet (TEDS) with information that might be used by algorithms and approaches for anomaly detection, diagnostics, etc. Similarly, a component, such as a tank, contains a Component Electronic Data Sheet (CEDS). Each element also includes a Health Electronic Data Sheet (HEDS) that contains health-related information such as anomalies and health state. Some practical aspects of the implementation include: (1) near real-time data flow from the test stand data acquisition system through the ISHM model, for near real-time detection of anomalies and diagnostics, (2) insertion of the J-2X predictive model providing predicted sensor values for comparison with measured values and use in anomaly detection and diagnostics, and (3) insertion of third-party anomaly detection algorithms into the integrated ISHM model.

Figueroa, Jorge F.

Structural Dynamics Observations in Space Launch System Green Run Hot Fire Testing

The Space Launch System (SLS) Core Stage (CS) Thrust Vector Control (TVC) system is comprised of eight mechanical feedback Shuttle heritage Type III TVC actuators and four RS-25 engines, each attached to a Shuttle heritage gimbal block/bearing. Two actuators are used to move each engine in two planes perpendicular to one another (i.e., pitch and yaw). The TVC system design leverages hardware from the Space Shuttle program as well as new hardware designed specifically for the Core Stage. The Green Run Hot Fire (GRHF) of the SLS Core Stage provided a flight-like ground test environment for verification of integrated vehicle TVC performance. A TVC model coupled to a vehicle structural dynamic model has been developed previously and incrementally validated in subsystem tests and simulations. Still, some aspects of TVC performance in GRHF were not anticipated. The ensuing investigation demonstrated the need for well-instrumented test environments, various levels of modeling fidelity, test-representative structural models, and caution in reuse of legacy components. This paper is the sixth installment in a seven-paper series surveying the design, engineering, test validation, and flight performance of the Core Stage Thrust Vector Control system. It introduces the salient structural dynamic phenomena uncovered in ambient and hot fire testing. During the Green Run test campaign, a comparison of ambient and hot fire step responses showed a significant change in apparent damping due to the presence of friction, challenging long standing assumptions that friction could be neglected. Additionally, the characteristic response of the engine and thrust structure during GRHF proved to be more complex than anticipated, as evidenced by the available actuator, thrust structure, and engine measurements. While the string-potentiometer based test instrumentation was intended to allow for reconstruction of the engine angles along the two control axes, the geometric placement, location uncertainty, and responses in overlapping frequency spectra revealed additional phenomena requiring further analysis and post-processing. The observations from both modal and frequency response testing during the Green Run ambient and hot fire configurations led to Engine and Core Stage FEM (finite element model) updates. When evidence of unexpected engine motion was found in engine section accelerometer data, the authors pursued additional structural analysis leading to FEM updates associated with the TVC gimbal and thrust structure. Through collaboration between structures, TVC, and flight control disciplines, the test-informed models and root-cause analysis led to confident flight rationale for the first flight of the SLS launch vehicle.

Richard K. Moore

Structural Dynamics Observations in Space Launch System Green Run Hot Fire Testing

The Space Launch System (SLS) Core Stage (CS) Thrust Vector Control (TVC) system is comprised of eight mechanical feedback Shuttle heritage Type III TVC actuators and four RS-25 engines, each attached to a Shuttle heritage gimbal block/bearing. Two actuators are used to move each engine in two planes perpendicular to one another (i.e., pitch and yaw). The TVC system design leverages hardware from the Space Shuttle program as well as new hardware designed specifically for the Core Stage. The Green Run Hot Fire (GRHF) of the SLS Core Stage provided a flight-like ground test environment for verification of integrated vehicle TVC performance. A TVC model coupled to a vehicle structural dynamic model has been developed previously and incrementally validated in subsystem tests and simulations. Still, some aspects of TVC performance in GRHF were not anticipated. The ensuing investigation demonstrated the need for well-instrumented test environments, various levels of modeling fidelity, test-representative structural models, and caution in reuse of legacy components. This paper is the sixth installment in a seven-paper series surveying the design, engineering, test validation, and flight performance of the Core Stage Thrust Vector Control system. It introduces the salient structural dynamic phenomena uncovered in ambient and hot fire testing. During the Green Run test campaign, a comparison of ambient and hot fire step responses showed a significant change in apparent damping due to the presence of friction, challenging long standing assumptions that friction could be neglected. Additionally, the characteristic response of the engine and thrust structure during GRHF proved to be more complex than anticipated, as evidenced by the available actuator, thrust structure, and engine measurements. While the string-potentiometer based test instrumentation was intended to allow for reconstruction of the engine angles along the two control axes, the geometric placement, location uncertainty, and responses in overlapping frequency spectra revealed additional phenomena requiring further analysis and post-processing. The observations from both modal and frequency response testing during the Green Run ambient and hot fire configurations led to Engine and Core Stage FEM (finite element model) updates. When evidence of unexpected engine motion was found in engine section accelerometer data, the authors pursued additional structural analysis leading to FEM updates associated with the TVC gimbal and thrust structure. Through collaboration between structures, TVC, and flight control disciplines, the test-informed models and root-cause analysis led to confident flight rationale for the first flight of the SLS launch vehicle.

Richard Moore

Fault localization in a microfabricated surface ion trap using diamond nitrogen-vacancy center magnetometry

Here, as quantum computing hardware becomes more complex with ongoing design innovations and growing capabilities, the quantum computing community needs increasingly powerful techniques for fabrication failure root-cause analysis. This is especially true for trapped-ion quantum computing. As trapped-ion quantum computing aims to scale to thousands of ions, the electrode numbers are growing to several hundred, with likely integrated photonic components also adding to the electrical and fabrication complexity, making faults even harder to locate. In this work, we used a high-resolution quantum magnetic imaging technique, based on nitrogen-vacancy centers in diamond, to investigate short-circuit faults in an ion trap chip. We imaged currents from these short-circuit faults to ground and compared them to intentionally created faults, finding that the root cause of the faults was failures in the on-chip trench capacitors. This work, where we exploited the performance advantages of a quantum magnetic sensing technique to troubleshoot a piece of quantum computing hardware, is a unique example of the evolving synergy between emerging quantum technologies to achieve capabilities that were previously inaccessible.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC

Systems Modeling to Implement Integrated System Health Management Capability

ISHM capability includes: detection of anomalies, diagnosis of causes of anomalies, prediction of future anomalies, and user interfaces that enable integrated awareness (past, present, and future) by users. This is achieved by focused management of data, information and knowledge (DIaK) that will likely be distributed across networks. Management of DIaK implies storage, sharing (timely availability), maintaining, evolving, and processing. Processing of DIaK encapsulates strategies, methodologies, algorithms, etc. focused on achieving high ISHM Functional Capability Level (FCL). High FCL means a high degree of success in detecting anomalies, diagnosing causes, predicting future anomalies, and enabling health integrated awareness by the user. A model that enables ISHM capability, and hence, DIaK management, is denominated the ISHM Model of the System (IMS). We describe aspects of the IMS that focus on processing of DIaK. Strategies, methodologies, and algorithms require proper context. We describe an approach to define and use contexts, implementation in an object-oriented software environment (G2), and validation using actual test data from a methane thruster test program at NASA SSC. Context is linked to existence of relationships among elements of a system. For example, the context to use a strategy to detect leak is to identify closed subsystems (e.g. bounded by closed valves and by tanks) that include pressure sensors, and check if the pressure is changing. We call these subsystems Pressurizable Subsystems. If pressure changes are detected, then all members of the closed subsystem become suspect of leakage. In this case, the context is defined by identifying a subsystem that is suitable for applying a strategy. Contexts are defined in many ways. Often, a context is defined by relationships of function (e.g. liquid flow, maintaining pressure, etc.), form (e.g. part of the same component, connected to other components, etc.), or space (e.g. physically close, touching the same common element, etc.). The context might be defined dynamically (if conditions for the context appear and disappear dynamically) or statically. Although this approach is akin to case-based reasoning, we are implementing it using a software environment that embodies tools to define and manage relationships (of any nature) among objects in a very intuitive manner. Context for higher level inferences (that use detected anomalies or events), primarily for diagnosis and prognosis, are related to causal relationships. This is useful to develop root-cause analysis trees showing an event linked to its possible causes and effects. The innovation pertaining to RCA trees encompasses use of previously defined subsystems as well as individual elements in the tree. This approach allows more powerful implementations of RCA capability in object-oriented environments. For example, if a pressurizable subsystem is leaking, its root-cause representation within an RCA tree will show that the cause is that all elements of that subsystem are suspect of leak. Such a tree would apply to all instances of leak-events detected and all elements in all pressurizable subsystems in the system. Example subsystems in our environment to build IMS include: Pressurizable Subsystem, Fluid-Fill Subsystem, Flow-Thru-Valve Subsystem, and Fluid Supply Subsystem. The software environment for IMS is designed to potentially allow definition of any relationship suitable to create a context to achieve ISHM capability.

Figueroa, Jorge F.

Root Source Analysis/ValuStream[Trade Mark] - A Methodology for Identifying and Managing Risks

Root Source Analysis (RoSA) is a systems engineering methodology that has been developed at NASA over the past five years. It is designed to reduce costs, schedule, and technical risks by systematically examining critical assumptions and the state of the knowledge needed to bring to fruition the products that satisfy mission-driven requirements, as defined for each element of the Work (or Product) Breakdown Structure (WBS or PBS). This methodology is sometimes referred to as the ValuStream method, as inherent in the process is the linking and prioritizing of uncertainties arising from knowledge shortfalls directly to the customer's mission driven requirements. RoSA and ValuStream are synonymous terms. RoSA is not simply an alternate or improved method for identifying risks. It represents a paradigm shift. The emphasis is placed on identifying very specific knowledge shortfalls and assumptions that are the root sources of the risk (the why), rather than on assessing the WBS product(s) themselves (the what). In so doing RoSA looks forward to anticipate, identify, and prioritize knowledge shortfalls and assumptions that are likely to create significant uncertainties/ risks (as compared to Root Cause Analysis, which is most often used to look back to discover what was not known, or was assumed, that caused the failure). Experience indicates that RoSA, with its primary focus on assumptions and the state of the underlying knowledge needed to define, design, build, verify, and operate the products, can identify critical risks that historically have been missed by the usual approaches (i.e., design review process and classical risk identification methods). Further, the methodology answers four critical questions for decision makers and risk managers: 1. What s been included? 2. What's been left out? 3. How has it been validated? 4. Has the real source of the uncertainty/ risk been identified, i.e., is the perceived problem the real problem? Users of the RoSA methodology have characterized it as a true bottoms up risk assessment.

Brown, Richard Lee

Usability/Sentiment for the Enterprise and ENTERPRISE

The purpose of the Sentiment of Search Study for NASA Johnson Space Center (JSC) is to gain insight into the intranet search environment. With an initial usability survey, the authors were able to determine a usability score based on the Systems Usability Scale (SUS). Created in 1986, the freely available, well cited, SUS is commonly used to determine user perceptions of a system (in this case the intranet search environment). As with any improvement initiative, one must first examine and document the current reality of the situation. In this scenario, a method was needed to determine the usability of a search interface in addition to the user's perception on how well the search system was providing results. The use of the SUS provided a mechanism to quickly ascertain information in both areas, by adding one additional open-ended question at the end. The first ten questions allowed us to examine the usability of the system, while the last questions informed us on how the users rated the performance of the search results. The final analysis provides us with a better understanding of the current situation and areas to focus on for improvement. The power of search applications to enhance knowledge transfer is indisputable. The performance impact for any user unable to find needed information undermines project lifecycle, resource and scheduling requirements. Ever-increasing complexity of content and the user interface make usability considerations for the intranet, especially for search, a necessity instead of a 'nice-to-have'. Despite these arguments, intranet usability is largely disregarded due to lack of attention beyond the functionality of the infrastructure (White, 2013). The data collected from users of the JSC search system revealed their overall sentiment by means of the widely-known System Usability Scale. Results of the scores suggest 75%, +/-0.04, of the population rank the search system below average. In terms of a grading scaled, this equated to D or lower. It is obvious JSC users are not satisfied with the current situation, however they are eager to provide information and assistance in improving the search system. A majority of the respondents provided feedback on the issues most troubling them. This information will be used to enrich the next phase, root cause analysis and solution creation.

Meza, David