Engineering PapersSearch

SEARCH · Engineering Papers

Results for “root cause analysis”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Systems Modeling to Implement Integrated System Health Management Capability

ISHM capability includes: detection of anomalies, diagnosis of causes of anomalies, prediction of future anomalies, and user interfaces that enable integrated awareness (past, present, and future) by users. This is achieved by focused management of data, information and knowledge (DIaK) that will likely be distributed across networks. Management of DIaK implies storage, sharing (timely availability), maintaining, evolving, and processing. Processing of DIaK encapsulates strategies, methodologies, algorithms, etc. focused on achieving high ISHM Functional Capability Level (FCL). High FCL means a high degree of success in detecting anomalies, diagnosing causes, predicting future anomalies, and enabling health integrated awareness by the user. A model that enables ISHM capability, and hence, DIaK management, is denominated the ISHM Model of the System (IMS). We describe aspects of the IMS that focus on processing of DIaK. Strategies, methodologies, and algorithms require proper context. We describe an approach to define and use contexts, implementation in an object-oriented software environment (G2), and validation using actual test data from a methane thruster test program at NASA SSC. Context is linked to existence of relationships among elements of a system. For example, the context to use a strategy to detect leak is to identify closed subsystems (e.g. bounded by closed valves and by tanks) that include pressure sensors, and check if the pressure is changing. We call these subsystems Pressurizable Subsystems. If pressure changes are detected, then all members of the closed subsystem become suspect of leakage. In this case, the context is defined by identifying a subsystem that is suitable for applying a strategy. Contexts are defined in many ways. Often, a context is defined by relationships of function (e.g. liquid flow, maintaining pressure, etc.), form (e.g. part of the same component, connected to other components, etc.), or space (e.g. physically close, touching the same common element, etc.). The context might be defined dynamically (if conditions for the context appear and disappear dynamically) or statically. Although this approach is akin to case-based reasoning, we are implementing it using a software environment that embodies tools to define and manage relationships (of any nature) among objects in a very intuitive manner. Context for higher level inferences (that use detected anomalies or events), primarily for diagnosis and prognosis, are related to causal relationships. This is useful to develop root-cause analysis trees showing an event linked to its possible causes and effects. The innovation pertaining to RCA trees encompasses use of previously defined subsystems as well as individual elements in the tree. This approach allows more powerful implementations of RCA capability in object-oriented environments. For example, if a pressurizable subsystem is leaking, its root-cause representation within an RCA tree will show that the cause is that all elements of that subsystem are suspect of leak. Such a tree would apply to all instances of leak-events detected and all elements in all pressurizable subsystems in the system. Example subsystems in our environment to build IMS include: Pressurizable Subsystem, Fluid-Fill Subsystem, Flow-Thru-Valve Subsystem, and Fluid Supply Subsystem. The software environment for IMS is designed to potentially allow definition of any relationship suitable to create a context to achieve ISHM capability.

Figueroa, Jorge F.

Root Source Analysis/ValuStream[Trade Mark] - A Methodology for Identifying and Managing Risks

Root Source Analysis (RoSA) is a systems engineering methodology that has been developed at NASA over the past five years. It is designed to reduce costs, schedule, and technical risks by systematically examining critical assumptions and the state of the knowledge needed to bring to fruition the products that satisfy mission-driven requirements, as defined for each element of the Work (or Product) Breakdown Structure (WBS or PBS). This methodology is sometimes referred to as the ValuStream method, as inherent in the process is the linking and prioritizing of uncertainties arising from knowledge shortfalls directly to the customer's mission driven requirements. RoSA and ValuStream are synonymous terms. RoSA is not simply an alternate or improved method for identifying risks. It represents a paradigm shift. The emphasis is placed on identifying very specific knowledge shortfalls and assumptions that are the root sources of the risk (the why), rather than on assessing the WBS product(s) themselves (the what). In so doing RoSA looks forward to anticipate, identify, and prioritize knowledge shortfalls and assumptions that are likely to create significant uncertainties/ risks (as compared to Root Cause Analysis, which is most often used to look back to discover what was not known, or was assumed, that caused the failure). Experience indicates that RoSA, with its primary focus on assumptions and the state of the underlying knowledge needed to define, design, build, verify, and operate the products, can identify critical risks that historically have been missed by the usual approaches (i.e., design review process and classical risk identification methods). Further, the methodology answers four critical questions for decision makers and risk managers: 1. What s been included? 2. What's been left out? 3. How has it been validated? 4. Has the real source of the uncertainty/ risk been identified, i.e., is the perceived problem the real problem? Users of the RoSA methodology have characterized it as a true bottoms up risk assessment.

Brown, Richard Lee

Usability/Sentiment for the Enterprise and ENTERPRISE

The purpose of the Sentiment of Search Study for NASA Johnson Space Center (JSC) is to gain insight into the intranet search environment. With an initial usability survey, the authors were able to determine a usability score based on the Systems Usability Scale (SUS). Created in 1986, the freely available, well cited, SUS is commonly used to determine user perceptions of a system (in this case the intranet search environment). As with any improvement initiative, one must first examine and document the current reality of the situation. In this scenario, a method was needed to determine the usability of a search interface in addition to the user's perception on how well the search system was providing results. The use of the SUS provided a mechanism to quickly ascertain information in both areas, by adding one additional open-ended question at the end. The first ten questions allowed us to examine the usability of the system, while the last questions informed us on how the users rated the performance of the search results. The final analysis provides us with a better understanding of the current situation and areas to focus on for improvement. The power of search applications to enhance knowledge transfer is indisputable. The performance impact for any user unable to find needed information undermines project lifecycle, resource and scheduling requirements. Ever-increasing complexity of content and the user interface make usability considerations for the intranet, especially for search, a necessity instead of a 'nice-to-have'. Despite these arguments, intranet usability is largely disregarded due to lack of attention beyond the functionality of the infrastructure (White, 2013). The data collected from users of the JSC search system revealed their overall sentiment by means of the widely-known System Usability Scale. Results of the scores suggest 75%, +/-0.04, of the population rank the search system below average. In terms of a grading scaled, this equated to D or lower. It is obvious JSC users are not satisfied with the current situation, however they are eager to provide information and assistance in improving the search system. A majority of the respondents provided feedback on the issues most troubling them. This information will be used to enrich the next phase, root cause analysis and solution creation.

Meza, David

Coincident learning for beam-based rf station fault identification using phase information at the SLAC linac coherent light source

Anomalies in radio-frequency (rf) stations can result in unplanned downtime and performance degradation in linear accelerators such as SLAC’s Linac Coherent Light Source (LCLS). Detecting these anomalies is challenging due to the complexity of accelerator systems, high data volume, and scarcity of labeled fault data. Prior work identified faults using beam-based detection, combining rf amplitude and beam position monitor data. Due to the simplicity of the rf amplitude data, classical methods are sufficient to identify faults, but the recall is constrained by the low-frequency and asynchronous characteristics of the data. In this work, we leverage high-frequency, time-synchronous rf phase data to enhance anomaly detection in the LCLS accelerator. Due to the complexity of phase data, classical methods fail, and we instead train deep neural networks within the Coincident Anomaly Detection (CoAD) framework. We find that applying CoAD to phase data detects nearly 3 times as many anomalies as when applied to amplitude data, while achieving broader coverage across rf stations. Furthermore, the rich structure of phase data enables us to cluster anomalies into distinct physical categories. Through the integration of auxiliary system status bits, we link clusters to specific fault signatures, providing additional granularity for uncovering the root cause of faults. We also investigate interpretability via Shapley values, confirming that the learned models focus on the most informative regions of the data and providing insight for cases where the model makes mistakes. This work demonstrates that phase-based anomaly detection for rf stations improves both diagnostic coverage and root cause analysis in accelerator systems and that deep neural networks are essential for effective analysis.

Accelerator Physics (physics.acc-ph)

Mitigating field emission in SSR2 cryomodules for PIP-II

The SSR2 cavities for PIP-II have consistently been affected by field emission since the first cold tests with the unity coupler. A comprehensive root cause analysis was conducted to investigate the origin of this issue and to identify the fabrication, processing, and handling factors that have the greatest impact on field emission onset. New techniques were developed and effectively implemented to achieve field emission-free SSR2 cavities. In addition, effort was dedicated both to relating the radiation level measured at the test stand with the expected levels in the LINAC tunnel and to understand the evolution of field emission through the assembly steps. The challenge of overcoming field emission also led to a reassessment of design choices, enhancing our understanding of their effects on cavity performance.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS

Preparation and Qualification of Preproduction SSR2 Jacketed Cavities for PIP-II

The qualification of 325 MHz Single Spoke Resonators type 2 (SSR2) jacketed cavities to meet technical requirements represents a significant milestone in the development of the SSR2 cryomodules for the PIP-II Project at Fermilab. This poster reports the procedures and lessons learned in processing and preparing these cavities for horizontal cold testing prior to integration into a cavity string assembly, with a focus on addressing the field emission issues observed during the cold testing. A comprehensive root cause analysis identified critical fabrication, processing, and handling factors impacting field emission onset. New techniques were successfully developed and implemented to achieve field emission-free SSR2 cavities, and efforts were made to correlate radiation levels measured at the test stand with expected levels in the LINAC tunnel. Additionally, the evolution of field emission through assembly steps was thoroughly investigated, leading to a reassessment of design choices and enhancing our understanding of their effects on cavity performance.

Grassellino, L. [Fermilab]

Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS)

Opportunities exist for realizing transformative advances in productivity and reductions in energy footprint through ubiquitous sensing in manufacturing environments. Enhanced Preparation for Intelligent Cybermanufacturing Systems (EPICS) is a 21-month (4 academic semesters, plus one summer) experience for graduate students that focuses on scaling the knowledge, understanding and leadership skills in the cyber manufacturing area. Masters students (8/year, 32 total) complete 2-year projects on industrially-driven project topics, rotating to internships in summer semester to work on scoping and implementation at project partners. Students complete academic training in embedded systems, process modeling, data science, and cloud-based systems design. Their projects are targeted toward sensor retrofit, process monitoring, root cause analysis, and sensor fusion.

Advanced Manufacturing

Software Sleuth

NASA's need to trace mistakes to their source to try and eliminate them in the future has resulted in software known as Root Cause Analysis (RoCA). Fair, Isaac & Co., Inc. has applied RoCA software, originally developed under an SBIR contract with Kennedy, to its predictive software technology. RoCA can generate graphic reports to make analysis of problems easier and more efficient.

Source record

Applying Model-Based Reasoning to the FDIR of the Command and Data Handling Subsystem of the International Space Station

All of the International Space Station (ISS) systems which require computer control depend upon the hardware and software of the Command and Data Handling System (C&DH) system, currently a network of over 30 386-class computers called Multiplexor/Dimultiplexors (MDMs)[18]. The Caution and Warning System (C&W)[7], a set of software tasks that runs on the MDMs, is responsible for detecting, classifying, and reporting errors in all ISS subsystems including the C&DH. Fault Detection, Isolation and Recovery (FDIR) of these errors is typically handled with a combination of automatic and human effort. We are developing an Advanced Diagnostic System (ADS) to augment the C&W system with decision support tools to aid in root cause analysis as well as resolve differing human and machine C&DH state estimates. These tools which draw from sources in model-based reasoning[ 16,291, will improve the speed and accuracy of flight controllers by reducing the uncertainty in C&DH state estimation, allowing for a more complete assessment of risk. We have run tests with ISS telemetry and focus on those C&W events which relate to the C&DH system itself. This paper describes our initial results and subsequent plans.

Robinson, Peter

Intelligent Elements for ISHM

There are a number of architecture models for implementing Integrated Systems Health Management (ISHM) capabilities. For example, approaches based on the OSA-CBM and OSA-EAI models, or specific architectures developed in response to local needs. NASA s John C. Stennis Space Center (SSC) has developed one such version of an extensible architecture in support of rocket engine testing that integrates a palette of functions in order to achieve an ISHM capability. Among the functional capabilities that are supported by the framework are: prognostic models, anomaly detection, a data base of supporting health information, root cause analysis, intelligent elements, and integrated awareness. This paper focuses on the role that intelligent elements can play in ISHM architectures. We define an intelligent element as a smart element with sufficient computing capacity to support anomaly detection or other algorithms in support of ISHM functions. A smart element has the capabilities of supporting networked implementations of IEEE 1451.x smart sensor and actuator protocols. The ISHM group at SSC has been actively developing intelligent elements in conjunction with several partners at other Centers, universities, and companies as part of our ISHM approach for better supporting rocket engine testing. We have developed several implementations. Among the key features for these intelligent sensors is support for IEEE 1451.1 and incorporation of a suite of algorithms for determination of sensor health. Regardless of the potential advantages that can be achieved using intelligent sensors, existing large-scale systems are still based on conventional sensors and data acquisition systems. In order to bring the benefits of intelligent sensors to these environments, we have also developed virtual implementations of intelligent sensors.

Schmalzel, John L.

Streamlining Payload Integration

Payload integration onto space transport vehicles and the International Space Station (ISS) is a complex process. Yet, cargo transport is the sole reason for any space mission, be it for ferrying humans, science, or hardware. As the largest such effort in history, the ISS offers a wide variety of payload experience. However, for any payload to reach the Space Station under the current process, Payload Developers face a list of daunting tasks that go well beyond just designing the payload to the constraints of the transport vehicle and its stowage topology. Payload customers are required to prove their payload s functionality, structural integrity, and safe integration - including under less than nominal situations. They must also plan for or provide training, procedures, hardware labeling, ground support, and communications. In addition, they must deal with negotiating shared consumables, integrating software, obtaining video, and coordinating the return of data and hardware. All the while, they must meet export laws, launch schedules, budget limits, and the consensus of more than 12 panel and board reviews. Despite the cost and infrastructure overhead, payload proposals have increased. Just in the span from FY08 to FY09, the NASA Payload Space Station Support Office budget rose from $78M to $96M in attempt to manage the growing manifest, but the potential number of payloads still exceeds available Payload Integration Management manpower. The growth has also increased management difficulties due to the fact that payloads are more frequently added to a flight schedule late in the flow. The current standard ISS template for payload integration from concept to payload turn-over is 36 months, or 18 months if the payload already has a preliminary design. Customers are increasingly requiring a turn-around of 3 to 6-months to meet market needs. The following paper suggests options for streamlining the current payload integration process in order to meet customer schedule needs and reduce costs for both the integration support teams and the developers, without reducing quality or compromising safety. Issues for the key integration areas of planning, training, verification, and safety are presented in a Root-Cause Analysis study, with plausible solutions provided that involve technology and tools already available to the ISS community. Although based upon the ISS process, the payload integration techniques outlined herein also offer an integration template for any space transport endeavor.

Lufkin, Susan N.

Smart Sensor Demonstration Payload

Sensors are a critical element to any monitoring, control, and evaluation processes such as those needed to support ground based testing for rocket engine test. Sensor applications involve tens to thousands of sensors; their reliable performance is critical to achieving overall system goals. Many figures of merit are used to describe and evaluate sensor characteristics; for example, sensitivity and linearity. In addition, sensor selection must satisfy many trade-offs among system engineering (SE) requirements to best integrate sensors into complex systems [1]. These SE trades include the familiar constraints of power, signal conditioning, cabling, reliability, and mass, and now include considerations such as spectrum allocation and interference for wireless sensors. Our group at NASA s John C. Stennis Space Center (SSC) works in the broad area of integrated systems health management (ISHM). Core ISHM technologies include smart and intelligent sensors, anomaly detection, root cause analysis, prognosis, and interfaces to operators and other system elements [2]. Sensor technologies are the base fabric that feed data and health information to higher layers. Cost-effective operation of the complement of test stands benefits from technologies and methodologies that contribute to reductions in labor costs, improvements in efficiency, reductions in turn-around times, improved reliability, and other measures. ISHM is an active area of development at SSC because it offers the potential to achieve many of those operational goals [3-5].

Schmalzel, John

Model-Based Fault Diagnosis: Performing Root Cause and Impact Analyses in Real Time

Generic, object-oriented fault models, built according to causal-directed graph theory, have been integrated into an overall software architecture dedicated to monitoring and predicting the health of mission- critical systems. Processing over the generic fault models is triggered by event detection logic that is defined according to the specific functional requirements of the system and its components. Once triggered, the fault models provide an automated way for performing both upstream root cause analysis (RCA), and for predicting downstream effects or impact analysis. The methodology has been applied to integrated system health management (ISHM) implementations at NASA SSC's Rocket Engine Test Stands (RETS).

Figueroa, Jorge F.

Photogrammetry Measurements During a Tanking Test on the Space Shuttle External Tank, ET-137

On November 5, 2010, a significant foam liberation threat was observed as the Space Shuttle STS-133 launch effort was scrubbed because of a hydrogen leak at the ground umbilical carrier plate. Further investigation revealed the presence of multiple cracks at the tops of stringers in the intertank region of the Space Shuttle External Tank. As part of an instrumented tanking test conducted on December 17, 2010, a three dimensional digital image correlation photogrammetry system was used to measure radial deflections and overall deformations of a section of the intertank region. This paper will describe the experimental challenges that were overcome in order to implement the photogrammetry measurements for the tanking test in support of STS-133. The technique consisted of configuring and installing two pairs of custom stereo camera bars containing calibrated cameras on the 215-ft level of the fixed service structure of Launch Pad 39-A. The cameras were remotely operated from the Launch Control Center 3.5 miles away during the 8 hour duration test, which began before sunrise and lasted through sunset. The complete deformation time history was successfully computed from the acquired images and would prove to play a crucial role in the computer modeling validation efforts supporting the successful completion of the root cause analysis of the cracked stringer problem by the Space Shuttle Program. The resulting data generated included full field fringe plots, data extraction time history analysis, section line spatial analyses and differential stringer peak ]valley motion. Some of the sample results are included with discussion. The resulting data showed that new stringer crack formation did not occur for the panel examined, and that large amounts of displacement in the external tank occurred because of the loads derived from its filling. The measurements acquired were also used to validate computer modeling efforts completed by NASA Marshall Space Flight Center (MSFC).

Littell, Justin D.

Arecibo Observatory Auxiliary M4N Socket Termination Failure Investigation

The NASA Engineering and Safety Center (NESC) was requested to support the Arecibo Observatory failure investigation in determining the root cause of the Auxiliary M4N cable failure. The NESC and Kennedy Space Center led an integrated NASA investigation with support from Marshall Space Flight Center and The Aerospace Corporation that included forensic investigation of failed hardware, finite element modeling, materials characterization, and root cause analysis. This document contains the outcome of the NESC Assessment.

Zinc Spelter Sockets

A Visual Analytics Approach to Debugging Cooperative, Autonomous Multi-Robot Systems’ Worldviews

Autonomous multi-robot systems, where a team of robots shares information to perform tasks that are beyond an individual robot’s abilities, hold great promise for a number of applications, such as planetary exploration missions. Each robot in a multi- robot system autonomously schedules which robots should perform a given task and when, using its worldview–the robot’s internal representation of its belief about the environment and other robots’ states. A key problem for operators is that robots’ worldviews can fall out of sync (often due to weak communication links), leading to desynchronization of the robots’ scheduling decisions and inconsistent emergent behavior (e.g., tasks not performed, or performed by multiple robots). Operators face the time-consuming and difficult task of making sense of the robots’ scheduling decisions, detecting de-synchronizations, and pinpointing their cause by comparing every robot’s worldview. To address these challenges, we introduce MOSAIC Viewer, a visual analytics system that helps operators (i) make sense of the robots’ schedules and (ii) detect and conduct a root cause analysis of the robots’ desynchronized worldviews. Over a year-long partnership with roboticists at the NASA Jet Propulsion Laboratory, a formative study was performed to identify the necessary system design requirements, which supports the design of the system. A qualitative study with 12 roboticists reveals that MOSAIC Viewer is faster- and easier-to-use than the users’ current approaches, and it allows them to stitch low-level details to formulate a high-level understanding of the robots’ schedules and detect and pin-point the cause of desynchronized worldviews.

Ma, Kwan-Liu

BERT-Based Topic Modeling and Information Retrieval to Support Fishbone Diagramming for Safe Integration of Unmanned Aircraft Systems in Wildfire Response

Recent concepts for emerging wildfire response operations have included unmanned aircraft systems (UAS) due to their increasing accessibility and capabilities. To integrate UAS into wildfire response safely, researchers have studied the use of large repositories of historic incident reports to improve the scope of root cause analysis. Recent work has emphasized applying state-of-the-art natural language processing techniques to extract useful information from these repositories. However, it has not yet been studied how these results can be interpreted and integrated into the systems engineering process. In this work, we propose a process in which Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling and information retrieval are applied to a relevant set of documents in order to support the development of a fishbone diagram in a semiautomated process. High-level themes in the document set are identified using topic modeling, which are then refined and interpreted by a human analyst. Then, the themes are used to guide a finer search using information retrieval, which returns specific incident reports of relevance. This provides traceability to specific incidents as well as broader categorizations that comprise the fishbone branches. We apply the proposed process to relevant documents from NASA’s Aviation Safety Reporting System (ASRS). The proposed process is widely applicable when relevant documents are available, and the results from this study will be useful to identifying potential causes of wildfire response UAS incidents.

Hazard analysis

BERT-based Topic Modeling and Information Retrieval to Support Fishbone Diagramming for Safe Integration of Unmanned Aircraft Systems in Wildfire Response

Recent concepts for emerging wildfire response operations have included unmanned aircraft systems (UAS) due to their increasing accessibility and capabilities. To integrate UAS into wildfire response safely, researchers have studied the use of large repositories of historic incident reports to improve the scope of root cause analysis. Recent work has emphasized applying state-of-the-art natural language processing techniques to extract useful information from these repositories. However, it has not yet been studied how these results can be interpreted and integrated into the systems engineering process. In this work, we propose a process in which Bidirectional Encoder Representations from Transformers (BERT)-based topic modeling and information retrieval are applied to a relevant set of documents in order to support the development of a fishbone diagram in a semiautomated process. High-level themes in the document set are identified using topic modeling, which are then refined and interpreted by a human analyst. Then, the themes are used to guide a finer search using information retrieval, which returns specific incident reports of relevance. This provides traceability to specific incidents as well as broader categorizations that comprise the fishbone branches. We apply the proposed process to relevant documents from NASA’s Aviation Safety Reporting System (ASRS). The proposed process is widely applicable when relevant documents are available, and the results from this study will be useful to identifying potential causes of wildfire response UAS incidents.

hazard analysis