Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Operational Resiliency”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Real-time Unimpeded Taxi Out Machine Learning Service

This presentation describes a study on the estimation of the unimpeded taxi out time using Machine Learning (ML) tools and proposes an implementation that can be used to make real-time predictions at any airport in the National Airspace System. Kedro, an open-source pipeline framework, is used to develop the model definition and training. Models are stored in scikit-learn containers on a MLFlow server where they can be retrieved and served to make predictions in the live system. These open source frameworks provide common structures between ML services, allow for easier maintenance and updates, and overall deliver an easier CI/CD (Continuous Integration/Continuous Deployment) process. The current models were trained on data acquired at KCLT and KDFW from June 1st to December 31st, 2019 and compute taxi time in the ramp, airport movement area (AMA) and total (from gates to runways). The current versions of the models achieve relatively low uncertainties of about 10 to 15% for the total and AMA taxi times and about 20% for the ramp taxi time at both KCLT and KDFW. Initial tests on offline data from 2020 and 2021 show a small degradation (10 to 15%) in accuracy performance indicating the model’s resilience to operational changes over time.

Machine Learning↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in mission or safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes a set of 55 historic aerospace software error incidents from 1962 to 2023 to determine trends of how and where automation is most likely to fail, behaving unexpectedly. A distinction between software producing unexpected (erroneous) output versus no output (failsilent) is introduced. Of the historical incidents analyzed, 85% were from software producing wrong output rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and not reliable to recover from silent failures. Error origin was within the code/logic itself in 58% of cases, 16% from configurable data, 15% from unexpected sensor input, and 11% from command/operator input. A substantial forty percent (40%) of unexpected software behavior was indicated by the absence of code, arising from unanticipated situations and missing requirements, and 16% of incidents were subjectively deemed “unknown-unknowns”. No incidents were found to be the result of programming language, compiler, tool, or operating system; and only sixteen percent (16%) of all incidents were considered errors traditional computer science/programming in nature. These findings indicate that for fault tolerance, erroneous automation behavior must be a primary consideration especially at critical moments, and reboot recoverability may not be viable. Special care should be taken to validate configurable data and commands prior to use. “Test-like-you-fly”, including hardware-in-the-loop combined with robust off-nominal testing should be used to uncover missing logic arising from unanticipated situations not covered by requirements alone. This study uniquely focuses on manifestations of unexpected flight software behavior, independent of ultimate root cause. We characterize software error behavior and origin to improve software design, test, and operations for resilience to the most common manifestations, and provide a rich dataset for further study.

Aerospace↗

Assembly and Integration Status of a High Fidelity Ground Test Bed for the Water Processor Assembly

The Water Recovery System (WRS) is a critical component of life support aboard the International Space Station (ISS) and will play an essential role in future missions beyond Low Earth Orbit (LEO). Its primary functional units – the Urine Processor Assembly (UPA), Brine Processor Assembly (BPA), and Water Processor Assembly (WPA) – must be evaluated for extended operation, dormancy resilience, material obsolescence, and reliability under exploration-driven constraints. Ground testing is vital for developing these technologies and generating statistically relevant reliability assessments, which requires extended runtime under integrated, Flight-like conditions. Currently, no high-fidelity, fully integrated WPA ground test bed exists to support these objectives. To address this gap, NASA is developing a WPA test bed at Marshall Space Flight Center (MSFC) that combines downgraded ISS flight hardware with functionally flight-like components in a cost-effective configuration while maintaining priority hardware investigations. This paper describes the current status of hardware assembly and integration, outlines key challenges such as simulating microgravity effects and mitigating obsolescence, and presents future test objectives including software development, reliability assessments, dormancy studies, and exploration-oriented upgrades.

Mary-Elizabeth Davis↗

Assembly and Integration Status of a High Fidelity Ground Test Bed for the Water Processor Assembly

The Water Recovery System (WRS) is a critical component of life support aboard the International Space Station (ISS) and will play an essential role in future missions beyond Low Earth Orbit (LEO). Its primary functional units – the Urine Processor Assembly (UPA), Brine Processor Assembly (BPA), and Water Processor Assembly (WPA) – must be evaluated for extended operation, dormancy resilience, material obsolescence, and reliability under exploration-driven constraints. Ground testing is vital for developing these technologies and generating statistically relevant reliability assessments, which requires extended runtime under integrated, Flight-like conditions. Currently, no high-fidelity, fully integrated WPA ground test bed exists to support these objectives. To address this gap, NASA is developing a WPA test bed at Marshall Space Flight Center (MSFC) that combines downgraded ISS flight hardware with functionally flight-like components in a cost-effective configuration while maintaining priority hardware investigations. This paper describes the current status of hardware assembly and integration, outlines key challenges such as simulating microgravity effects and mitigating obsolescence, and presents future test objectives including software development, reliability assessments, dormancy studies, and exploration-oriented upgrades.

Water Processor Assembly↗

Using Degradation Modeling to Identify Fragile Operational Conditions in Human- and Component-driven Resilience Assessment

Studying failure events shows that many high-impact events result from the complex interactions between precipitating failure events and degraded operational conditions. Often, when a system is put in operations, unforeseen practical realities (e.g., maintenance and/or workforce availability) lead the system to be operated in configurations outside its envisioned nominal range. However, design-time failure models often assume that the failure events are initiated in an idealized, nominal state of system operation, resulting in an incomplete assessment of future risk. To solve this, this paper develops a framework to consider degraded operational performance in scenario-based resilience models which uses a corresponding model of performance degradation to determine the values of deteriorated model parameters in the resilience model. This framework is demonstrated on a remotely-piloted rover to determine the (individual and combined) effect of drive-train wear and operator fatigue on the resilience of the rover to drive-train faults. This demonstration showed the substantial impact that degradation has on resilience, highlighting the need to account for degradation in resilience models–specifically, unconsidered degradation can lead to overestimates of resilience (and thus underestimates of safety margin) and because resilience can degrade prior to visible unreliability, which can lead to an operational environment with a high propensity for high-impact unforeseen failure events.

resilience↗

Using Degradation Modeling to Identify Fragile Operational Conditions in Human- and Component-driven Resilience Assessment

Studying failure events shows that many high-impact events result from the complex interactions between precipitating failure events and degraded operational conditions. Often, when a system is put in operations, unforeseen practical realities (e.g., maintenance and/or workforce availability) lead the system to be operated in configurations outside its envisioned nominal range. However, design-time failure models often assume that the failure events are initiated in an idealized, nominal state of system operation, resulting in an incomplete assessment of future risk. To solve this, this paper develops a framework to consider degraded operational performance in scenario-based resilience models which uses a corresponding model of performance degradation to determine the values of deteriorated model parameters in the resilience model. This framework is demonstrated on a remotely-piloted rover to determine the (individual and combined) effect of drive-train wear and operator fatigue on the resilience of the rover to drive-train faults. This demonstration showed the substantial impact that degradation has on resilience, highlighting the need to account for degradation in resilience models--specifically, unconsidered degradation can lead to overestimates of resilience (and thus underestimates of safety margin) and because resilience can degrade prior to visible unreliability, which can lead to an operational environment with a high propensity for high-impact unforeseen failure events.

Daniel Hulse↗

Resilience Modeling in Complex Engineered Systems with Human-Machine Interactions

In recent times, there has been a growing interest in resilience-based design. Resilience-based design operates on the concept that failures and unexpected events will happen, and when they occur, complex engineered systems should be able to operate within acceptable bounds and recover reasonably. Humans can contribute to the resilience of a system by quickly detecting unforeseen events and taking corrective measures. To this effect, researchers have proposed guidelines and design approaches that can help promote human-system resilience. However, there is no early design stage tool to validate if a system is indeed resilient after applying these guidelines and design methods. In this research, we integrate the Human Error and Functional Failure Reasoning (HEFFR) framework into the fmdtools toolkit to enable designers to model the combined (machine, human, and joint) failures, including their propagation and dynamic effects, during early design stages. This integrated tool also allows designers to model the effects of performance shaping factors, team dynamics, and human-machine interactions in systems of systems. A demonstrative example of a remotely operated rover is explored to demonstrate how this approach can be applied to understand resilience in complex engineered systems with human interactions.

Lukman Irshad↗

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad↗

Achieving Resilient In-Flight Performance for Advanced Air Mobility through Simplified Vehicle Operations

A research and development (R&D) approach is proposed for developing and validating concepts and technologies to achieve vehicle autonomy goals of Advanced Air Mobility (AAM) through Simplified Vehicle Operations (SVO). The approach applies resilience-engineering and human-automation teaming (HAT) principles to a framework for defining vehicle-based functions for the management of missions and flight trajectories, focusing initially on the en route flight domain. To achieve the SVO goal of reducing pilot training requirements and thereby increasing the pilot pool for AAM, while at the same time promoting ever-safer operations, a framework for identifying essential functions is proposed. In this framework, functions are first categorized by high-level functional purpose (mission management, flightpath management, tactical operations, and vehicle control) and then subcategorized by attributes of resilient-performing systems (abilities to monitor, respond, learn, and anticipate). The categorization by functional purpose provides structure within which HAT designs can be holistically explored and total levels of human vs. automation responsibility can be varied. The subcategorization by resilient-system attributes provides a mechanism for capturing safety-critical functions that may not be codified in current operational procedures and training curricula, particularly those where humans proactively enhance safety in currently undocumented ways. An R&D approach consisting of seven strategies is proposed in which automation engineering and human-factors communities can collaborate in the research, development, and design of an SVO roadmap to enable the ambitious objectives of AAM.

AAM↗

Design and Implementation of a Scenario Development Process for a 2040 Trajectory-Based Operations Simulation

The National Aeronautics and Space Administration (NASA) is supporting research to transition from a legacy, air traffic management (ATM) system to Trajectory-Based Operations (TBO) environment targeting the 2035-2045 timeframe by creating simulation scenarios for current and future studies. TBO in the National Airspace System (NAS) focuses on modernizing the current operating paradigm to increase efficiency, predictability, resilience, and flexibility while migrating toward greater operational autonomy across the airspace. These simulation scenarios offer a wide dissemination and use of system-level constraints and aircraft state/intent data, routine use of data communications to request complex trajectory modifications and to receive clearances, and aircraft that are flying on 4-D trajectories with flexibility when desired, but structure where required. This report describes the end-to-end development of scenarios for the NASA Langley Research Center simulation environment to represent a 2040 Trajectory-Based Operations airspace for air transport operations research.

Trajectory-Based Operations↗

Operating characteristics of a cantilever-mounted resilient-pad gas-lubricated thrust bearing

A resilient-pad gas thrust bearing consisting of pads mounted on cantilever beams was tested to determine its operating characteristic. The bearing was run at a thrust load of 74 newtons to a speed of 17000 rpm. The pad film thickness and bearing friction torque were measured and compared with theory. The measured film thickness was less than that predicted by theory. The bearing friction torque was greater than that predicted by theory.

Nemeth, Z. N.↗

Implementing Artificial Thinking Autonomy with Model-Based System Engineering

Complex autonomous systems capable of successfully operating independently under ‘known unknowns’ and harsh conditions require paradigm innovation in modern development strategies. In the field of autonomy, developing a system-of-systems which can ostensibly think for itself in the face of ‘unknown unknowns’ is still a field of ongoing research. Maturing the systems architecting and modeling methodologies for developing henceforth named Thinking Autonomous Systems, which are verified with digital mission simulation, can potentially usher in the next generation of artificial intelligence for space exploration. The concept presented in this paper incorporates multiple Model-Based Systems Engineering and simulation methodologies combined as a new paradigm to design a novel, biomimetic thinking autonomy strategy. Anachronistic concepts from classical Kantian philosophy will be leveraged to inspire architectural designs that could be used for complex distributed systems in deep space. To accomplish this, digital transformation of a document-based implementation plan for Thinking Autonomous Systems, generated by experienced NASA software engineers, is implemented for NASA’s Platform for Autonomous Systems by creating descriptive and executable software models in SysML to prototype real-time operating capabilities. This conceptual implementation has been developed by incorporating model-based digital simulations to theorize how a cyberphysical thinking system would achieve specific strategies without crew reliance, while simultaneously being resilient to all operating conditions and remaining functional when devoid of ground communication. Additionally, ensuring that an autonomous system framework is an ethical Artificial Intelligence requires careful consideration of system behavior and accountability, human factors for teaming with a thinking autonomous system, and comparison to other modern approaches used for implementing true autonomy. This paper presents the first steps in formalizing the metacognition required for instantiating a truly Thinking Autonomous System; the approach described symphonizes autonomy characteristics from classical philosophical into a unified software architecture describing human thought. In the future, the foundational models described in this paper can be further leveraged to help advance research into thinking autonomy requirements for future deep space missions as well as for current near-term applications, i.e., living aboard crewed spacecraft like a NASA Gateway cislunar habitat.

Artificial Thought↗

Design of Space Systems to Enable In-space Assembly and Servicing

For several decades, NASA has employed in-space systems to enhance the performance and extend the useful life of operational orbital assets. In at least one case, an operational mission was not only enhanced, but enabled – the International Space Station was made possible by crewed and robotic in-space assembly, and continues to support installation and operation of new science and technology payloads. In several cases (Hubble Space Telescope, Intelsat 401, Westar and Palapa), major operational assets were rescued or repaired soon after launch when otherwise mission-ending anomalies occurred or were detected. In addition to the original rescue, Hubble was upgraded four times, enabling high-demand, world class science over four decades. More recently, two Northrop Grumman Mission Extension Vehicles have captured two Intelsat spacecraft near the end of their life and fuel capacity, to take over maneuvering duties. In spite of these recent operational achievements, and with the exception of large human exploration vehicles and large space telescopes, space architects rarely consider in-orbit servicing and assembly capabilities in their future planning. Technologies such as multi-launch mission architectures (and rendezvous and proximity operations systems), docking systems, external robotics, advanced tools, modular systems and structures, and fluid transfer systems are available today to support these missions. In-space manufacturing will soon be operational to enable resilient missions that recover from on-orbit failures, and expand the utilization of space. We envision a future that includes these capabilities, and discuss the cultural, engineering, and technological challenges to achieving this vision. We discuss the vision, the proverbial chicken and the egg (which came first, the serviceable spacecraft or the servicer?), the cost, risk, and perceptions thereof of in-space operations, a “spectrum” of cooperative servicing design considerations, and the current status of the space industry’s slow but steady march to widespread operational use of on-orbit servicing, assembly, and manufacturing.

Bo Naasz↗

Automated Target Planning for FUSE Using the SOVA Algorithm

The SOVA algorithm was originally developed under the Resilient Systems and Operations Project of the Engineering for Complex Systems Program from NASA s Aerospace Technology Enterprise as a conceptual framework to support real-time autonomous system mission and contingency management. The algorithm and its software implementation were formulated for generic application to autonomous flight vehicle systems, and its efficacy was demonstrated by simulation within the problem domain of Unmanned Aerial Vehicle autonomous flight management. The approach itself is based upon the precept that autonomous decision making for a very complex system can be made tractable by distillation of the system state to a manageable set of strategic objectives (e.g. maintain power margin, maintain mission timeline, and et cetera), which if attended to, will result in a favorable outcome. From any given starting point, the attainability of the end-states resulting from a set of candidate decisions is assessed by propagating a system model forward in time while qualitatively mapping simulated states into margins on strategic objectives using fuzzy inference systems. The expected return value of each candidate decision is evaluated as the product of the assigned value of the end-state with the assessed attainability of the end-state. The candidate decision yielding the highest expected return value is selected for implementation; thus, the approach provides a software framework for intelligent autonomous risk management. The name adopted for the technique incorporates its essential elements: Strategic Objective Valuation and Attainability (SOVA). Maximum value of the approach is realized for systems where human intervention is unavailable in the timeframe within which critical control decisions must be made. The Far Ultraviolet Spectroscopic Explorer (FUSE) satellite, launched in 1999, has been collecting science data for eight years.[1] At its beginning of life, FUSE had six gyros in two IRUs and four reaction wheels. Over time through various failures, the satellite has been left with one reaction wheel on the vehicle skew axis and two gyros. To remain operational, a control scheme has been implemented using the magnetic torque rods and the remaining momentum wheel.[2] As a consequence, there are attitude regions where there is insufficient torque authority to overcome environmental disturbances (e.g. gravity gradient torques). The situation is further complicated by the fact that these attitude regions shift inertially with time as the spacecraft moves through earth s magnetic field during the course of its orbit. Under these conditions, the burden of planning targets and target-to-target slew maneuvers has increased significantly since the beginning of the mission.[3] Individual targets must be selected so that the magnetic field remains roughly aligned with the skew wheel axis to provide enough control authority to the other two orthogonal axes. If the field moves too far away from the skew axis, the lack of control authority allows environmental torques to pull the satellite away from the target and can potentially cause it to tumble. Slew maneuver planning must factor the stability of targets at the beginning and end, and the torque authority at all points along the slew. Due to the time varying magnetic field geometry relative to any two inertial targets, small modifications in slew maneuver timing can make large differences in the achievability of a maneuver.

Heatwole, Scott↗

Evaluating the Use of High-Fidelity Simulator Research Methods to Study Airline Flight Crew Resilience

As it evolves, aviation will continue to require integration of a wide range of safety systems and practices, some of which are already in place and others that are yet to be developed. New concepts in system safety thinking have emerged to consider not only what may go wrong, but also what can be learned when things go right during commercial flight operations. Taken together, these complementary perspectives form a more comprehensive approach to systemsafety thinking that can help to recognize and preserve the resilient performance capabilities currently provided by humans. A need exists, however, for research methods to enable better understanding of the human contributions to aviation safety. NASA’s System-Wide Safety Project supports research on using flight simulation methods to study operator resilience and safety-producing behaviors. Building on prior NASA efforts investigating procedural non-adherences during area navigation standard terminal route arrivals, a high-fidelity commercial aviation line operational simulation (LOS) experiment has been designed to study how flight crews anticipate, monitor for, respond to, and learn from expected and unexpected disturbances during these operations. A diverse set of LOS scenarios were developed to simulate highly realistic, complex, but routinely encountered operational situations. Each scenario provided multiple opportunities to collect data on how flight crews manage threats and errors, as well as novel opportunities to observe resilient and safety-producing behaviors. The experimental design, implications for the study of safety-producing behaviors using simulation, and considerations for airline pilot training will be discussed.

Chad L Stephens↗

People are the Weak Link in the System.

People are often considered the "weak link" in every system. It is assumed that there is a causal "chain" of events, and each link in the chain provides some protection against failure. In this notion, the chain is only as strong as its weakest link and people are often blamed for failures. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are only the source of error and failure.

operations↗

Simulation of a Representative Future Trajectory-Based Operations Environment

Trajectory-Based Operations in the National Airspace System is a key aspect of advanced air traffic management research. Trajectory-Based Operations focuses on modernizing the current operating paradigm to increase efficiency, predictability, resilience, and flexibility while migrating toward greater operational autonomy across the airspace. Research conducted at the National Aeronautics and Space Administration supports the transition from current airspace operations to Trajectory-Based Operations targeting a 2035-2045 implementation timeframe. Simulation scenarios that demonstrate a representative Trajectory-Based Operations environment in that timeframe, and enable the evaluation of advanced airborne tools such as strategic airborne trajectory management services, are necessary to support this research effort. This report describes characteristics and assumptions made about the future operating environment that were applied to a scenario development methodology to study a representative 2040 Trajectory-Based Operations environment in a simulation use case. This report also includes descriptions of the study design and analysis approach, discussion of the simulation results, and application of these scenarios to future research activities.

Trajectory Based Operations↗