Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “System Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Human Factors of Remotely Piloted Aircraft Systems: Lessons from Incident Reports

It has been estimated that aviation accidents are typically preceded by numerous minor incidents arising from the same causal factors that ultimately produced the accident. Accident databases provide in-depth information on a relatively small number of occurrences, however incident databases have the potential to provide insights into the human factors of Remotely Piloted Aircraft System (RPAS) operations based on a larger volume of less-detailed reports. Currently, there is a lack of incident data dealing with the human factors of unmanned aircraft systems. An exploratory study is being conducted to examine the feasibility of collecting voluntary critical incident reports from RPAS pilots. Twenty-three experienced RPAS pilots volunteered to participate in focus groups in which they described critical incidents from their own experience. Participants were asked to recall (1) incidents that revealed a system flaw, or (2) highlighted a case where the human operator contributed to system resilience or mission success. A total of 90 incidents were reported. This presentation begins with an overview of RPAS human factors and then considers the emerging issues identified by pilot reports.

incidents↗

The Collection of Human Factors Incident Reports via UAS Pilot Focus Groups

There is a need for incident data relevant to the operation of civilian unmanned aircraft systems (UAS) in the National Air Space (NAS). Currently, the tightly-restricted civilian UAS industry has produced relatively few incident reports that shed light on the human factors of UAS operations. An exploratory study is being conducted to examine the feasibility of using the critical incident technique (Flanagan, 1954) to collect reports from UAS pilots. The information is being used to identify areas where human factors guidelines will be of assistance. Experienced UAS pilots are participating in small focus groups in which they are prompted to describe critical incidents that either reveal a system flaw, or highlight a case where the human operator contributed to system resilience or mission success. To date, 90 incidents have been reported by UAS pilots. The de-identified incidents are being analyzed to identify contributing factors, with a focus on design issues that either hindered or assisted the pilot in dealing with the incident..

Hobbs, Alan↗

NASA’s Optical Communications Programs

The way we communicate in space and how we stay connected on Earth is experiencing a change: while we traditionally use radio frequency communication methods from satellites to ground stations, there is a growing interest in laser communication connections back to Earth, but also in the space environment. Meanwhile, the need for higher data rates also push for a shift in the use of frequency bands, with new bands such as high-power KA, KU, Q and V being developed.These changes require new technologies and new issues need to be addressed. Bringing together satellite operators, satellite manufacturers, and component suppliers, this panel will discuss the following questions and topics: (1)New technology requirements for different communication bands: from antennas to optimized ground systems; (2) New technology challenges when implementing new communication systems: from smaller spaces, to 'noise on the line' and radiation issues; (4) How can we leverage safely what is already out there?; (5) How to ensure communication system resiliency and security?

Edwards, Bernard L.↗

NASA's Implementation of Cloud Services for Human Space Flight

Cloud is a tried-and-true technology used throughout United States government agencies, including the National Aeronautics and Space Administration (NASA). With reliable results and infrequent downtimes, cloud allows for secure remote access, customizability, and streamlined monitoring options, creating an environment for better data integrity and availability. As NASA increasingly migrates functions to the cloud, the Space Communications and Navigation Program (SCaN) program has been investigating how this capability can be leveraged to provide communication services to its users and customers. Currently, missions such as NASA-ISRO Synthetic Aperture Radar (NISAR), Plankton, Aerosol, Cloud, ocean Ecosystem (PACE), and Roman Space Telescope (RST) are planned to incorporate cloud into their data delivery architecture. However, SCaN is looking to expand further. This conversion to using cloud services allows for greater availability of mission data for both robotic and human space flight (HSF)missions. The SCaN program and the Near Space Network (NSN) are working to consolidate resources and create a cloud environment suitable for the entirety of the SCaN program network architecture. SCaN is in the process of finalizing its cloud architecture and soon will be implementing cloud services. The new services used will adhere to federal regulations including Federal Risk and Authorization Management Program (FedRAMP), which is built upon National Institute of Standards and Technology (NIST)documentation. While keeping in mind these security requirements, an auxiliary objective of the cloud integration is to ensure the most cost-efficient solution; providing a scalable, robust and resilient system. Using cloud services, NASA will gain access to better centralized monitoring and management features, along with customizable services on a pay-per-use plan. With the ever-growing NASA mission data volume needs, maintaining ample storage space is another major constraint. Processing and storing such large amounts of data, on the order of terabytes a day, requires dynamic processing capability which is inherently a strength of cloud computing. By routing this data from ground stations through the cloud, there will be greater ease of access for both SCaN and the user community. Artificial intelligence and other built-in cloud functions can also enhance efficiency, improving data processing time. Thereby also allowing for better data availability. As we look to the future of cloud services, NASA will continue to leverage capabilities that will benefit NASA’s ability to provide cost-effective communication services. This paper further outlines the evolution of cloud use by SCaN in the context of Human Space Flight.

cloud storage↗

Earth Independent Medical Operations [EIMO] Definition Workshop

As the vanguard of human spaceflight increasingly transitions from one of low Earth orbit (LEO) to Lunar and subsequently Martian missions, there is a commensurate imperative for Earth-based medical authority to transition to space-based assets for the continued assurance of optimal astronaut health and performance. The transition to Earth Independent Medical Operations (EIMO) will be a process that enables progressively resilient systems and crews to reduce risk and enhance wellness and overall mission success for deep space exploration. Terrestrial assets will continue to be paramount in pre-mission screening, planning, maintenance, and prevention. Yet on-board care, response to unexpected medical events, and management of communication delays and dropouts will increasingly become the purview of the crew for primary management.

John Lemery↗

Space Crop Considerations for Human Exploration

NASA has been actively working to both determine how many crops will be needed for early exploration missions as well as updating the “Crop Readiness Level” (CRL) for a library of crops that can be selected for supporting a long-term mission. The Crop Readiness Level (CRL) is modelled after NASA’s Technology Readiness Level (TRL) approach for developing and advancing new technologies for space, first suggested by Barry Finger and published by Wheeler and Strayer [2]. The CRL model has nine levels from “crop identification” to “consumed in space.” The number and variety of crops needed is impacted by both primary factors (nutrition, menu fatigue, behavioral health system resiliency) as well as secondary factors such as ECLSS considerations, crop robustness, and hardware considerations.

Gioia D Massa↗

Hot Water, Cold Reality: Feasibility Assessment of Iodine Removal in Heated Spacecraft Potable Water Systems

The current eXploration Potable Water Dispenser (xPWD) design removes iodine upstream of the heated leg due to concerns with the Activated Carbon and Ion Exchange (ACTEX) functionality in hot water, leaving the downstream volume without residual biocide. The NESC determined that this non-iodinated volume is a concern for microbial growth during exploration missions and proposed 33 biocide architecture options for future missions that could address this concern. The top-ranked architecture out of the report was Option 1: moving the iodine removal media as close to the dispensing needle as possible to minimize the wetted components without biocide in the xPWD. Three main challenges were identified with this proposed configuration. First, the hot water at 175 ± 25 °F is a concern for the potential physical degradation of the ion exchange resin and lowered adsorption capacity in activated carbon. Second, moving the ACTEX or alternative sorption media closer to the dispense needle increases the unheated volume downstream of the heater, challenging the ability for dispensed water to meet temperature requirements. Finally, bubbles evolved from dissolved gas coming out of solution in the heater could clog or reduce the efficiency of the sorption media. To address the first challenge, more thermally robust ion exchange resins were identified and adsorption capacity tests were planned and will be discussed in a companion ICES paper (ICES-2026-5). To address the dispense temperature concerns, allowable bed size and architectural configuration changes are proposed. The value of adding phase separators to remove bubbles and potential implementation schemes are discussed. These findings support the development of potable water systems resilient to microbial risks during long-duration space missions.

Biocide↗

Hot Water, Cold Reality: Feasibility Assessment of Iodine Removal in Heated Spacecraft Potable Water Systems

The current eXploration Potable Water Dispenser (xPWD) design removes iodine upstream of the heated leg due to concerns with the Activated Carbon and Ion Exchange (ACTEX) functionality in hot water, leaving the downstream volume without residual biocide. The NESC determined that this non-iodinated volume is a concern for microbial growth during exploration missions and proposed 33 biocide architecture options for future missions that could address this concern. The top-ranked architecture out of the report was Option 1: moving the iodine removal media as close to the dispensing needle as possible to minimize the wetted components without biocide in the xPWD. Three main challenges were identified with this proposed configuration. First, the hot water at 175 ± 25 °F is a concern for the potential physical degradation of the ion exchange resin and lowered adsorption capacity in activated carbon. Second, moving the ACTEX or alternative sorption media closer to the dispense needle increases the unheated volume downstream of the heater, challenging the ability for dispensed water to meet temperature requirements. Finally, bubbles evolved from dissolved gas coming out of solution in the heater could clog or reduce the efficiency of the sorption media. To address the first challenge, more thermally robust ion exchange resins were identified and adsorption capacity tests were planned and will be discussed in a companion ICES paper (ICES-2026-5). To address the dispense temperature concerns, allowable bed size and architectural configuration changes are proposed. The value of adding phase separators to remove bubbles and potential implementation schemes are discussed. These findings support the development of potable water systems resilient to microbial risks during long-duration space missions.

PWD↗

Measuring the Resilience of Advanced Life Support Systems

Despite the central importance of crew safety in designing and operating a life support system, the metric commonly used to evaluate alternative Advanced Life Support (ALS) technologies does not currently provide explicit techniques for measuring safety. The resilience of a system, or the system s ability to meet performance requirements and recover from component-level faults, is fundamentally a dynamic property. This paper motivates the use of computer models as a tool to understand and improve system resilience throughout the design process. Extensive simulation of a hybrid computational model of a water revitalization subsystem (WRS) with probabilistic, component-level faults provides data about off-nominal behavior of the system. The data can then be used to test alternative measures of resilience as predictors of the system s ability to recover from component-level faults. A novel approach to measuring system resilience using a Markov chain model of performance data is also developed. Results emphasize that resilience depends on the complex interaction of faults, controls, and system dynamics, rather than on simple fault probabilities.

Bell, Ann Maria↗

The System Modeling and Analysis of Resiliency in STEReO (SMARt-STEReO)

Wildfire emergency response has remained rooted in relatively low-tech solutions for coordination between ground and aerial assets. These low-tech solutions are robust for the remote environments in which wildfires are usually fought, but limit strategic cross-organizational support and the ability to deploy and effectively utilize aerial assets. As aircraft become more advanced and new technology, including drones, become available to firefighters, a new, more modern method of asset coordination is needed. NASA is working on a project called ‘Scalable Traffic Management for Emergency Response Operations’ (STEReO) to integrate unmanned aerial systems (UAS)and UAS traffic management (UTM)into wildfire response. STEReO’s goals include simplifying the coordination of aerial assets, improving the existing UAS framework, and increasing the role of additional autonomous systems to reduce human risk and to increase system resilience. This paper describes the development of the ‘System Modeling and Analysis of Resiliency in STEReO’ (SMARt-STEReO) project, which aims to model wildfire response and to quantify the additional system resilience that STEReO technology provides firefighters. This paper verifies SMARt-STEReO and defines its scope; it includes experimental and statistical analysis of the impact that the addition of UAS has on both performance metrics and also on performance resiliency response to a given fault. SMARt-STEReO is a grid-based model of fire propagation that incorporates varying crew responses. Through the use of a Python package called ‘fmdtools’, the model easily allows for the addition of faults to the system. These faults allow analysts to investigate various response parameters. Factors including terrain, fuel type and wind speed can be modified to affect the fire propagation; additionally, the number of ground crews, engines, fixed wing aircraft, helicopters, and UAS can be changed to affect the crew response. The communication lines between actors mimic those used in real life situations. This paper explains the development of SMARt-STEReO including background research, verification and validation, and preliminary experimental analysis of system resilience to both a minor and major fault in systems with and without UAS.

Resiliency↗

Indicators and Monitoring Systems for Urban Climate Resiliency

Cities in the USA and around the world have begun to take an active role in responding to climate change. A central requirement for effective urban climate strategies is the capacity to understand and measure how the climate is changing, the physical, environmental, and social impacts of the changes, and whether adaptation and resiliency policies and programs put in place in response are working. The objective of this paper is to review and assess how urban climate change and resiliency efforts can be measured and to define what might serve as meaningful indicator and monitoring protocols. The New York City Panel on Climate Change (NPCC) is used as a case study along with a reviews of the emerging literature of urban climate change indicators to analyze the requirements and processes needed for a successful urban climate resiliency indicator and monitoring (I and M) system. In the paper, the basic requirements of a proposed Urban Climate Resilience Indicators and Monitoring System are presented. A specific illustration of an I and M system for tracking the urban heat island highlights challenges as well as potential solutions embedded within such systems. Discussions how these protocols can be translated to other locales and settings, as well as the relationship to the US National Climate Assessment indicator process, are presented.

New York City Panel on Climate Change↗

Going Beyond Reliability to Robustness and Resilience in Space Life Support Systems

The words reliability, robustness, and resilience are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environment is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to unanticipated damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes can invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging conditions. These three qualities are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from unanticipated disruptions. Increasing the resilience of space systems would greatly increase space crew safety. Improving reliability and robustness requires dealing with known problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. The need for robustness and resilience has been stated for decades but little has been done. Systems designers often assume that they understand everything they need to know. The potential failures caused by changes, failures, accidents, unknown environments, and unknown unknowns can be ignored. Such overconfidence can lead to neglect of reliability, robustness, and resilience.

Harry W. Jones↗

Going Beyond Reliability to Robustness and Resilience in Space Life Support Systems

The words reliability, robustness, and resilience are often used interchangeably to describe tough and dependable systems but the distinctions between them suggest how to design more serviceable space systems. Reliability is simply the quality of consistently performing well. A system that dependably meets its design requirements in the specified environment is reliable. The designers may not consider themselves responsible for failures under unanticipated conditions. Robustness is the capability of performing without failure under a wide range of conditions, which can go beyond the expected range to include possible off-nominal conditions. Resilience is the ability to recover from or adapt to unanticipated damaging events, such as failures, accidents, external disruptions, and repurposing. Such changes can invalidate the usual operating assumptions and cause system failure. Reliability, robustness, and resilience describe dependable performance under increasingly difficult conditions, first the specified environment, then a wider possible environment, and finally unanticipated damaging conditions. These three qualities are increasingly desirable and increasingly difficult to achieve. Engineering for resilience would design systems that can ignore or repair failures, survive accidents, and recover from unanticipated disruptions. Increasing the resilience of space systems would greatly increase space crew safety. Improving reliability and robustness requires dealing with known problems, but improving resilience requires implementing a general approach to reducing the impact of unknown future events. The need for robustness and resilience has been stated for decades but little has been done. Systems designers often assume that they understand everything they need to know. The potential failures caused by changes, failures, accidents, unknown environments, and unknown unknowns can be ignored. Such overconfidence can lead to neglect of reliability, robustness, and resilience.

Harry W. Jones↗

Understanding Resilience Optimization Architectures With an Optimization Problem Repository

Optimizing a system’s resilience can be challenging, especially when it involves considering both the inherent resilience of a robust design and the active resilience of a health management system to a set of computationally-expensive hazard simulations. While prior work has developed specialized architectures to effectively and efficiently solve combined design and resilience optimization problems, the comparison of these architectures has been limited to a single case study. To further study resilience optimization formulations, this work develops a problem repository which includes previously-developed resilience optimization problems and additional problems presented in this work: a notional system resilience model, a pandemic response model, and a cooling tank hazard prevention model. This work then uses models in the repository at large to understand the characteristics of resilience optimization problems and study the applicability of optimization architectures and decomposition strategies. Based on the comparisons in the repository, applying an optimization architecture effectively requires understanding the alignment and coupling relationships between the design and resilience models, as well as the efficiency characteristics of the algorithms. While alignment determines the necessity of a surrogate of resilience cost in the upper-level design problem, coupling determines the overall applicability of a sequential, alternating, or bilevel structure. Additionally, the application of decomposition strategies is dependent on there being limited interactions between variable sets, which often does not hold when a resilience policy is parameterized in terms of actions to take in hazardous model states rather than specific given scenarios.

Resilience↗

Resilience Modeling in Complex Engineered Systems with Human-Machine Interactions

In recent times, there has been a growing interest in resilience-based design. Resilience-based design operates on the concept that failures and unexpected events will happen, and when they occur, complex engineered systems should be able to operate within acceptable bounds and recover reasonably. Humans can contribute to the resilience of a system by quickly detecting unforeseen events and taking corrective measures. To this effect, researchers have proposed guidelines and design approaches that can help promote human-system resilience. However, there is no early design stage tool to validate if a system is indeed resilient after applying these guidelines and design methods. In this research, we integrate the Human Error and Functional Failure Reasoning (HEFFR) framework into the fmdtools toolkit to enable designers to model the combined (machine, human, and joint) failures, including their propagation and dynamic effects, during early design stages. This integrated tool also allows designers to model the effects of performance shaping factors, team dynamics, and human-machine interactions in systems of systems. A demonstrative example of a remotely operated rover is explored to demonstrate how this approach can be applied to understand resilience in complex engineered systems with human interactions.

Lukman Irshad↗