Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “System Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Achieving Resilient In-Flight Performance for Advanced Air Mobility through Simplified Vehicle Operations

A research and development (R&D) approach is proposed for developing and validating concepts and technologies to achieve vehicle autonomy goals of Advanced Air Mobility (AAM) through Simplified Vehicle Operations (SVO). The approach applies resilience-engineering and human-automation teaming (HAT) principles to a framework for defining vehicle-based functions for the management of missions and flight trajectories, focusing initially on the en route flight domain. To achieve the SVO goal of reducing pilot training requirements and thereby increasing the pilot pool for AAM, while at the same time promoting ever-safer operations, a framework for identifying essential functions is proposed. In this framework, functions are first categorized by high-level functional purpose (mission management, flightpath management, tactical operations, and vehicle control) and then subcategorized by attributes of resilient-performing systems (abilities to monitor, respond, learn, and anticipate). The categorization by functional purpose provides structure within which HAT designs can be holistically explored and total levels of human vs. automation responsibility can be varied. The subcategorization by resilient-system attributes provides a mechanism for capturing safety-critical functions that may not be codified in current operational procedures and training curricula, particularly those where humans proactively enhance safety in currently undocumented ways. An R&D approach consisting of seven strategies is proposed in which automation engineering and human-factors communities can collaborate in the research, development, and design of an SVO roadmap to enable the ambitious objectives of AAM.

AAM↗

Modeling Distributed Situation Awareness in Resilience-based Design of Complex Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling↗

New York City Panel on Climate Change 2015 Report Introduction

The climate of the New York City metropolitan region is changing annual temperatures are hotter, heavy downpours are increasingly frequent, and the sea is rising. These trends, which are also occurring in many parts of the world, are projected to continue and even worsen in the coming decades because of higher concentrations of greenhouse gases in the atmosphere caused by burning of fossil fuels and clearing of forests for agriculture. These changing climate hazards increase the risks for the people, economy, and infrastructure of New York City. As was demonstrated by Hurricane Sandy, coastal and low-lying areas, the elderly and very young, and lower-income neighborhoods are highly vulnerable. In response to these climate challenges, New York City is developing a broad range of climate resiliency policies and programs, as well as the knowledge base to support them. The knowledge base includes up-to-date climate, sea level rise, and coastal flooding projections; a Climate Resiliency Indicators and Monitoring System; and resiliency studies. A special attribute of the New York City response to these challenges is the recognition that both the knowledge base and the programs and policies it supports need to evolve through time as climate risks unfold in the coming decades. In early September 2012, just weeks before Hurricane Sandy hit, the New York City Council passed Local Law 42 that established the New York City Panel on Climate Change (NPCC) as an ongoing body serving the City of New York. The NPCC is required to meet at least twice each calendar year to review recent scientific data on climate change and its potential impacts, and to make recommendations on climate projections for the coming decades to the end of the century. These projections are due within one year of the publication of the Intergovernmental Panel on Climate Change Assessment Reports (http:www.ipcc.ch), or at least every three years. The NPCC also advises the Mayor's Office of Sustainability and the Mayor's Office of Recovery and Resiliency (ORR) on the development of a community- or borough-level communications strategy intended to ensure that the public is informed about the findings of the panel, including the creation of a summary of the climate change projections for dissemination to city residents.

hurricanes↗

The System Modeling and Analysis of Resiliency in STEReO (SMARt-STEReO)

NASA's Scalable Traffic Management for Emergency Response Operations (STEReO) project aims to leverage Unmanned Aerial Systems (UAS) and UAS Traffic Management (UTM) to improve asset coordination and overall emergency response. One application of STEReO is wildfire response, which is the focus of this research. In order to implement the operations described in the STEReO project, these additions must have tangible benefits and proven safety. To this end, the System Modeling and Analysis of Resiliency in STEReO (SMARt-STEReO) project constructs a simulation model, developed through the Python modeling and resiliency analysis package fmdtools. The model describes wildfire response operations, including current operational concepts and emerging concepts utilizing UAS as described in STEReO. While previous simulation models focus primarily on fire propagation with some models including emergency response intervention, SMART-STEReO evaluates the system performance and resilience benefits gained by the addition of UAS and UTM. Due to the novelty and complexity of the model, initial model verification and validation efforts are conducted and a detailed description of the model is provided. Preliminary results from experimental analysis on the SMARt-STEReO model indicate that when compared to current operations, the addition of UAS in wildfire operations results in improved response efforts, in terms of fewer acres burned, as well as improved system resilience in response to a given fault.

Sequoia Andrade↗

Modeling Distributed Situation Awareness in Resilience-Based Design of Complex Engineered Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling↗

Experiments for Securing Air Traffic Against Cyber-Physical System Attacks

This presentation describes experiments conducted with single board computers to investigate methods for creating trust for enabling the development of cyber-resilient air transportation systems. Methods included secure communication to prevent unauthorized access to data, consistency of data obtained via sensors and by processing, and built-in safeguards to prevent mission failure. The motivation for this work are the following. The future air transportation system needs to ensure availability, integrity, confidentiality and safety of operations. Safety of vehicles and operations is paramount for successful integration of Urban Air Mobility (UAM), Unmanned Aerial Systems (UAS), supersonic aircraft and launch vehicles with conventional aviation operations in the National Airspace System. Security is becoming critical because the sensors, networks and computers are far more vulnerable to bad actors than their mechanical or human predecessors. The goal therefore is to design and develop cyber-resilient systems that continue to function even in degraded states. The main findings are (1) off-the-shelf hardware can support development of cyber-resilient onboard flight computers and (2) trust in system design and implementation can be accomplished by integrating layers in depth (detail) and in breadth (scope).

cyber-resilient autonomy, trust, secure communicat↗

Can Resilience Assessments Inform Early Design Human Factors Decision-making?

There is a growing call among researchers for tighter coupling between human factors and human reliability assessments. In this research, we explore if early design stage resilience assessments can help bridge some of the gaps between human factors and human reliability assessments. Resilience in systems is their ability to recover reasonably and operate within acceptable bounds during failures and unexpected events. The fmdtools toolkit allows designers to assess the resilience of a system by modeling the human error and machine-related failure propagation in both nominal and faulty scenarios during the early design stages. As a result, the fmdtools toolkit has a low-fidelity dynamic human reliability assessment component built into it. In this paper, we study if the results from the fmdtools simulations can help inform and prioritize human factors design decision-making, resulting in tighter coupling between human factors and human reliability assessments during the design process. Specifically, we explore the results from a rover design example to understand the types of information that can help guide human factor-related decision-making.

Resilience-based Design↗

Verification and Validation of Safety-Critical Aircraft Systems Operating under Off-Nominal, Contingency, and Emergency Conditions

Verification and validation (V&V) of safety-critical technologies developed for loss of control (LOC) prevention and recovery and other aviation safety concerns pose significant challenges. Aircraft LOC can result from a wide spectrum of hazards, often occurring in combination, which cannot be fully replicated during evaluation. Technologies developed for LOC prevention and recovery must therefore be effective under a wide variety of hazardous and uncertain conditions, and the verification and validation of these technologies must provide some measure of assurance that the new vehicle safety technologies do no harm (i.e., that they themselves do not introduce new safety risks). V&V technologies must also enable the identification of system limitations and constraints, as well as enable the identification of safe and unsafe operating conditions (and their boundaries). Additionally, the V&V of complex, increasingly autonomous systems is a fundamental concern. Scalable, reproducible and cost-effective techniques for the assurance of safety critical systems during their design and operation is a key barrier to fielding new systems or updating current systems. Moreover, these techniques need to provide artifacts that enable a comprehensive evidence-based approach to certification. This briefing summarizes research performed under NASA’s Aviation Safety Program and follow-on research for the V&V of safety-critical aircraft system technologies developed for LOC prevention and recovery and increasingly autonomous systems, and for a broad assurance capability in both current and emerging aviation applications. Note that, in this briefing, the term “validation” refers to a confirmation that the system implementation (e.g., algorithms etc.) is performing the intended function(s), as well as an affirmation of effectiveness in these functions. “Verification” refers to a confirmation that the system implementation in the software and hardware meets its (hopefully validated) specifications (e.g., correctly executes algorithms as designed).

Validation↗

Strategies for the Design and Operation of Resilient Extraterrestrial Habitats

An Earth-independent permanent extraterrestrial habitat system must function as intended under continuous disruptive conditions, and with significantly limited Earth support and extended uncrewed periods. Designing for the demands that extreme environments such as wild temperature fluctuations, galactic cosmic rays, destructive dust, meteoroid impacts (direct or indirect), vibrations, and solar particle events, will place on long-term deep space habitats represents one of the greatest challenges in this endeavor. This context necessitates that we establish the know-how and technologies to build habitat systems that are resilient. Resilience is not simply robustness or redundancy: it is a system property that accounts for both anticipated and unanticipated disruptions via the design choices and maintenance processes, and adapts to them in operation. We currently lack the frameworks and technologies needed to achieve a high level of resilience in a habitat system. The Resilient Extra Terrestrial Habitats Institute (RETHi) has the mission of leveraging existing novel technologies to provide situational awareness and autonomy to enable the design of habitats that are able to adapt, absorb and rapidly recover from expected and unexpected disruptions. We are establishing both fully virtual and coupled physical-virtual simulation capabilities that will enable us to explore a wide range of potential deep space Smart Hab configurations and operating modes.

Space habitats↗

Strategies for the Design and Operation of Resilient Extraterrestrial Habitats

An Earth-independent permanent extraterrestrial habitat system must function as intended under continuous disruptive conditions, and with significantly limited Earth support and extended uncrewed periods. Designing for the demands that extreme environments such as wild temperature fluctuations, galactic cosmic rays, destructive dust, meteoroid impacts (direct or indirect), vibrations, and solar particle events, will place on long-term deep space habitats represents one of the greatest challenges in this endeavor. This context necessitates that we establish the know-how and technologies to build habitat systems that are resilient. Resilience is not simply robustness or redundancy: it is a system property that accounts for both anticipated and unanticipated disruptions via the design choices and maintenance processes and adapts to them in operation. We currently lack the frameworks and technologies needed to achieve a high level of resilience in a habitat system. The Resilient ExtraTerrestrial Habitats Institute (RETHi) has the mission of leveraging existing novel technologies to provide situational awareness and autonomy to enable the design of habitats that are able to adapt, absorb and rapidly recover from expected and unexpected disruptions. We are establishing both fully virtual and coupled physical-virtual simulation capabilities that will enable us to explore a wide range of potential deep space SmartHab configurations and operating modes.

Space habitats↗

Foundational Concepts in Simulation-Based Resilience Analysis and Design

Resilience is a topic of increasing interest–with ever-present calls from policymakers to increase the resilience of complex systems and infrastructure. However, resilience as a concept can be confusing, because of a lack of a common unified definition and frame of reference. Sometimes it can appear as if resilience analysis is merely duplicating other, more mature fields such as safety, reliability, or risk, while other times it seems as if resilience is providing an “alternative” view with limited rigor. To better understand the resilience concept (and its relation to the broader field of risk management), this paper will present the perspective of simulation-based resilience analysis and design, including some of the foundational precepts and resultant concepts defining the resilience concept. It will further present the motivation for using simulation to understand resilience and highlight some ongoing work and research challenges in this area. From this frame of reference, one can better understand the field of resilience, including how different aspects and definitions of resilience relate to each other, and how resilience relates to broader design considerations and practices.

resilience↗

Foundational Concepts in Simulation-Based Resilience Analysis and Design

Resilience is a topic of increasing interest–with ever-present calls from policymakers to increase the resilience of complex systems and infrastructure. However, resilience as a concept can be confusing, because of a lack of a common unified definition and frame of reference. Sometimes it can appear as if resilience analysis is merely duplicating other, more mature fields such as safety, reliability, or risk, while other times it seems as if resilience is providing an “alternative” view with limited rigor. To better understand the resilience concept (and its relation to the broader field of risk management), this paper will present the perspective of simulation-based resilience analysis and design, including some of the foundational precepts and resultant concepts defining the resilience concept. It will further present the motivation for using simulation to understand resilience and highlight some ongoing work and research challenges in this area. From this frame of reference, one can better understand the field of resilience, including how different aspects and definitions of resilience relate to each other, and how resilience relates to broader design considerations and practices.

resilience↗

MAX - An advanced parallel computer for space applications

MAX is a fault-tolerant multicomputer hardware and software architecture designed to meet the needs of NASA spacecraft systems. It consists of conventional computing modules (computers) connected via a dual network topology. One network is used to transfer data among the computers and between computers and I/O devices. This network's topology is arbitrary. The second network operates as a broadcast medium for operating system synchronization messages and supports the operating system's Byzantine resilience. A fully distributed operating system supports multitasking in an asynchronous event and data driven environment. A large grain dataflow paradigm is used to coordinate the multitasking and provide easy control of concurrency. It is the basis of the system's fault tolerance and allows both static and dynamical location of tasks. Redundant execution of tasks with software voting of results may be specified for critical tasks. The dataflow paradigm also supports simplified software design, test and maintenance. A unique feature is a method for reliably patching code in an executing dataflow application.

Lewis, Blair F.↗

The Influence of Mental Workload in Causes of System Degradation in Air Traffic Control

System safety and resilience is a critical concern in the air traffic domain. An important element of maintaining system safety and resilience is the ability of systems to ‘degrade gracefully’. However, previous research on the causes of system degradation in the air traffic domain are sporadic, and the potential interaction between the causes of degradation, and the resulting possible compound effect on the entire system, has been under-researched. An interview study was conducted with 12 retired controllers as participants. The results of a thematic analysis revealed the key causes of system degradation, and the associated impact on the ability of the controllers to prevent system degradation or recover the system. Findings have direct implications for identifying and mitigating potential risks of increasingly automated air traffic control systems.

Edwards, Tamsyn E.↗

The Influence of Mental Workload in Causes of System Degradation in Air Traffic Control

System safety and resilience is a critical concern in the air traffic domain. An important element of maintaining system safety and resilience is the ability of systems to ‘degrade gracefully'. However, previous research on the causes of system degradation in the air traffic domain are sporadic, and the potential interaction between the causes of degradation, and the resulting possible compound effect on the entire system, has been under-researched. An interview study was conducted with 12 retired controllers as participants. The results of a thematic analysis revealed the key causes of system degradation, and the associated impact on the ability of the controllers to prevent system degradation or recover the system. Findings have direct implications for identifying and mitigating potential risks of increasingly automated air traffic control systems.

Edwards, Tamsyn↗

Cyber-Threat Assessment for the Air Traffic Management System: A Network Controls Approach

Air transportation networks are being disrupted with increasing frequency by failures in their cyber- (computing, communication, control) systems. Whether these cyber- failures arise due to deliberate attacks or incidental errors, they can have far-reaching impact on the performance of the air traffic control and management systems. For instance, a computer failure in the Washington DC Air Route Traffic Control Center (ZDC) on August 15, 2015, caused nearly complete closure of the Centers airspace for several hours. This closure had a propagative impact across the United States National Airspace System, causing changed congestion patterns and requiring placement of a suite of traffic management initiatives to address the capacity reduction and congestion. A snapshot of traffic on that day clearly shows the closure of the ZDC airspace and the resulting congestion at its boundary, which required augmented traffic management at multiple locations. Cyber- events also have important ramifications for private stakeholders, particularly the airlines. During the last few months, computer-system issues have caused several airlines fleets to be grounded for significant periods of time: these include United Airlines (twice), LOT Polish Airlines, and American Airlines. Delays and regional stoppages due to cyber- events are even more common, and may have myriad causes (e.g., failure of the Department of Homeland Security systems needed for security check of passengers, see [3]). The growing frequency of cyber- disruptions in the air transportation system reflects a much broader trend in the modern society: cyber- failures and threats are becoming increasingly pervasive, varied, and impactful. In consequence, an intense effort is underway to develop secure and resilient cyber- systems that can protect against, detect, and remove threats, see e.g. and its many citations. The outcomes of this wide effort on cyber- security are applicable to the air transportation infrastructure, and indeed security solutions are being implemented in the current system. While these security solutions are important, they only provide a piecemeal solution. Particular computers or communication channels are protected from particular attacks, without a holistic view of the air transportation infrastructure. On the other hand, the above-listed incidents highlight that a holistic approach is needed, for several reasons. First, the air transportation infrastructure is a large scale cyber-physical system with multiple stakeholders and diverse legacy assets. It is impractical to protect every cyber- asset from known and unknown disruptions, and instead a strategic view of security is needed. Second, disruptions to the cyber- system can incur complex propagative impacts across the air transportation network, including its physical and human assets. Also, these implications of cyber- events are exacerbated or modulated by other disruptions and operational specifics, e.g. severe weather, operator fatigue or error, etc. These characteristics motivate a holistic and strategic perspective on protecting the air transportation infrastructure from cyber- events. The analysis of cyber- threats to the air traffic system is also inextricably tied to the integration of new autonomy into the airspace. The replacement of human operators with cyber functions leaves the network open to new cyber threats, which must be modeled and managed. Paradoxically, the mitigation of cyber events in the airspace will also likely require additional autonomy, given the fast time scale and myriad pathways of cyber-attacks which must be managed. The assessment of new vulnerabilities upon integration of new autonomy is also a key motivation for a holistic perspective on cyber threats.

Complex Networks↗

Contingency Software in Autonomous Systems: Technical Level Briefing

Contingency management is essential to the robust operation of complex systems such as spacecraft and Unpiloted Aerial Vehicles (UAVs). Automatic contingency handling allows a faster response to unsafe scenarios with reduced human intervention on low-cost and extended missions. Results, applied to the Autonomous Rotorcraft Project and Mars Science Lab, pave the way to more resilient autonomous systems.

autonomous systems↗

Development of a Weather Capability for the Urban Air Mobility Airspace Research Roadmap

Traditionally, the transportation system’s resiliency to the impacts of weather is an area where neglected or incorrect assumptions can lead to difficulties later in the research and development lifecycle. To mitigate this, NASA has ongoing efforts to develop a set of research roadmaps for organizing, integrating, and communicating research into new aviation infrastructure and transportation modalities, within which weather is being addressed early on. An effort has been undertaken to add weather assumptions and requirements to an already-existing roadmap for the Urban Air Mobility (UAM) airspace, seeking to integrate weather requirements early in the system design. This effort addresses the way in which state-of-the art and evolving weather science and technology can enable safe and efficient travel with increasing tempo of UAM operations over time. This paper describes the addition of weather as one of 10 capabilities into the UAM Airspace research roadmap, laying out the anticipated weather technology and information requirements needed to facilitate operations at various UAM Maturity Levels. The process developed and exercised by MIT Lincoln Laboratory researchers produced 41 unique requirements to be satisfied by a Weather capability for the UAM ecosystem, with more than 300 dependencies identified across the system. These requirements cover measurement, analysis, modeling, forecasting, decision support, dissemination, and overarching policy, and are provided with an overview of weather challenges for UAM. The requirements were mainly defined based on subject matter expert review of existing UAM Airspace system requirements, and refined based on iterative feedback with various stakeholders including regulators, academia, and industry. Going forward, this roadmap will help researchers and developers align to a common vision in ensuring that weather is appropriately considered in the UAM ecosystem.

Timothy Bonin↗