Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

To Create Safety Is Human

It is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations, are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

operations↗

Changing the Narrative About the Human Role in Accidents

It is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

operations↗

What Do People Do?

It is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

operations↗

So What Do People Actually Do?

I is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

safety↗

Localization of Ad-Hoc Lunar Constellations in Communication Failure Modes for Distributed Spacecraft Autonomy

As Lunar missions increase in complexity, inspired by NASA’s Artemis Program, they will require reliable and sufficient Position, Navigation, and Timing (PNT) capability to support the upcoming Lunar users. The navigation service should also be compatible with the smaller platforms, like CubeSats, being sent by the public and private sectors. A non-dedicated, ad-hoc Lunar navigation constellation can provide PNT services on-demand using the non-dedicated swarm assets. Swarm members cooperatively and autonomously localize themselves with minimal interaction from Earth, freeing up valuable bandwidth and ground segment resources. The autonomous localization of Lunar constellations utilizes neighbor two-way intersatellite link (ISL) measurements in a distributed extended Kalman filter (DEKF) system to minimize operating costs. Because the decentralized Lunar PNT system relies on relay communication amongst the agents, network failures or loss of assets among ad-hoc Lunar constellations may impact localization performance. This study presents an evaluation of localization performance under increasing levels of network degradation. A simulation of an ad-hoc Lunar PNT swarm is augmented to include system faults and the impacts of intermittent and permanent failures on localization performance are evaluated. We investigate three potential causes of network degradation: single spacecraft loss, multiple spacecraft loss, and antenna failure. The numerical assessments from the simulation show that the LPNT system under study, based on an autonomous decentralized concept of operation, is highly robust and resilient to communication failures. Minor faults, such as single spacecraft loss, solar interference, technical malfunctions, message delays, and antenna outages, have minimal impact on state estimation, with only a 4.47% and 3.75% degradation in median position error for assets and a representative ground user, respectively, compared to an ideal communication scenario. However, major faults, such as hardware failures or meteor strikes leading to the loss of multiple spacecrafts, are more concerning. The permanent loss of three spacecraft results in a more severe performance degradation, with median position error increasing by 23.3% for assets and 11.7% for a representative ground user, despite the Lunar PNT system remaining functional.

Yeji Kim↗

Design of Solar Sailing Trajectories Resilient to Safe Mode Events

Solar sails are an enabling technology for stand-alone small-satellite deep-space exploration. However, their always-on nature, combined with the time of flight required for deep-space missions, makes them particularly susceptible to safe mode events. Unlike missions using electric propulsion, exclusively solar sail- ing missions cannot carry extra propellant to make up for a safe mode event, and more powerful methods like expected availability cannot be directly applied. This work extends the expected availability and duty cycle approaches to solar sailing. Through an application to NASA’s Solar Cruiser mission, a 46 % increase in tra- jectory resilience is obtained at negligible change in the mission’s time of flight. Additionally, it is shown that when a safe mode event stops the spacecraft from reaching its target orbit, the developed methods reduce the expected final error by approximately an order of magnitude.

Johnson, Les↗

The Contribution of Pilots to Resilience in Normal Operations: A Survey Approach

Much of our knowledge about human performance in flight safety has come from the analysis of undesired events, whether accidents, incidents, or crew behaviors identified via flight exceedance monitoring or observational techniques. In recent years, there has been an acknowledgement that operational personnel are not merely sources of “human error”, but also make a unique human contribution to safe outcomes. In a few celebrated cases, this takes the form of “heroic saves”, but on many more occasions, operational personnel contribute to safety through everyday, often-unnoticed actions that turn potentially hazardous situations into non-events. An emerging approach to safety, frequently referred to as “Safety II,” proposes that the positive human contribution is an important and largely untapped source of safety information. Some airlines have successfully trained observers to identify and record the positive behaviors exhibited by the crew over the course of a flight. In other cases, flight crew are interviewed about good practices or positive behaviors. However, each of these methods are relatively limited in scale and resource intensive. A survey could provide a relatively low-cost approach to systematically gather this information on a larger scale. The primary purpose of the research was to develop and assess a surveys methodology for assessing crews' activities in normal flights and the operational perturbations encountered during normal operations. We hope that such a survey could be both a research tool as well as a safety management aid for the aviation industry. We collected responses concerning revenue flights from two groups of airline pilots (N = 25 & N= 65). The results indicated that relatively few flights proceeded exactly as in the original flight plan. Pilots routinely anticipated and adapted to changing circumstances. We will review the challenges encountered in developing the survey and summarize preliminary findings from two administrations of the survey to airline pilots.

human contribution safety↗

The Contribution of Pilots to Resilience in Normal Operations. Part II: A closer look at briefings: Anticipation and Monitoring Also Known as Planning and Coordination

Much of our knowledge about human performance in flight safety has come from the analysis of undesired events, whether accidents, incidents, or crew behaviors identified via flight exceedance monitoring or observational techniques. In recent years, there has been an acknowledgement that operational personnel are not merely sources of “human error”, but also make a unique human contribution to safe outcomes. In a few celebrated cases, this takes the form of “heroic saves”, but on many more occasions, operational personnel contribute to safety through everyday, often-unnoticed actions that turn potentially hazardous situations into non-events. An emerging approach to safety, frequently referred to as “Safety II,” proposes that the positive human contribution is an important and largely untapped source of safety information. Some airlines have successfully trained observers to identify and record the positive behaviors exhibited by the crew over the course of a flight. In other cases, flight crew are interviewed about good practices or positive behaviors. However, each of these methods are relatively limited in scale and resource intensive. A survey could provide a relatively low-cost approach to systematically gather this information on a larger scale. The primary purpose of the research was to develop and assess a surveys methodology for assessing crews' activities in normal flights and the operational perturbations encountered during normal operations. We hope that such a survey could be both a research tool as well as a safety management aid for the aviation industry. We collected responses concerning revenue flights from two groups of airline pilots (N = 25 & N= 65). The results indicated that relatively few flights proceeded exactly as in the original flight plan. Pilots routinely anticipated and adapted to changing circumstances. We will review the challenges encountered in developing the survey and summarize preliminary findings from two administrations of the survey to airline pilots.

human contribution safety↗

Cloud Influence on ERA5 and AMPS Surface Downwelling Longwave Radiation Biases in West Antarctica

The surface downwelling longwave radiation component (LW[down arrow]) is crucial for the determination of the surface energy budget and has significant implications for the resilience of ice surfaces in the polar regions. Accurate model evaluation of this radiation component requires knowledge about the phase, vertical distribution, and associated temperature of water in the atmosphere, all of which control the LW[down arrow] signal measured at the surface. In this study, we examine the LW[down arrow] model errors found in the Antarctic Mesoscale Prediction System (AMPS) operational forecast model and the ERA5 reanalysis model relative to observations from the AWARE campaign at McMurdo Station and the West Antarctic Ice Sheet (WAIS) Divide. The errors are calculated separately for observed clear-sky conditions, ice-cloud occurrences, and liquid-bearing cloud layer (LBCL) occurrences. The analysis results show a tendency in both models at each site to underestimate the LW[down arrow] during clear sky conditions, high error variability (standard deviations > 20 W/m[exp2]) during any type of cloud occurrence, and negative LW biases when LBCLs are observed (bias magnitudes > 15 W/m[exp2] in tenuous LBCL cases; > 43 W/m[exp2] in optically thick/opaque LBCLs instances). We suggest that a generally dry and liquid-deficient atmosphere responsible for the identified LW[down arrow] biases in both models is the result of excessive ice formation and growth, which could stem from model initial and lateral boundary conditions, microphysics scheme, aerosol representation, and/or limited vertical resolution.

Israel Silber↗

Assessing the Viability of Using GEOS-Forecast Product for Landslides Forecasting: A Step Toward Early Warning System

Landslides across the globe are mostly triggered by extreme rainfall events affecting infrastructure, transportation and livelihoods. The risks are rarely quantified due to lack of data, analytical skills and limited modeling techniques. Knowledge of local to global scale landslide risks provides communities and national agencies the ability to adapt disaster management practices to mitigate and recover from these hazards. In order to minimize the risks and improve characterization of community resilience to landslides, it is vital to have reliable information about the factors triggering landslides such as rainfall, well ahead in time. Forecasting potential landslide activity and impacts can be achieved through reliable precipitation forecast models. However, it is challenging because of the temporal and spatial variability of precipitation, an important factor in triggering landslides. Evaluation of the precipitation field, associated errors, and sampling uncertainties is integral for development of efficient and reliable landslide forecasting and early warning system. This study develops a methodology to assess the viability of using a precipitation field provided by a global model and its potential integration in the landslide forecasting system. The study focuses on the comparison between the IMERG (Integrated Multi-satellitE Retrievals for Global Precipitation Mission) and GEOS (NASA Goddard Earth Observing System)-Forecast product over contiguous United States (CONUS). GEOS model assimilates new observations every 6 hours, at 00, 06, 12, and 18 UTC. The framework is tested on the GEOS-Forecast Model initialized at 00 UTC using daily IMERG early product as reference using both categorical and continuous statistics. The categorical statistics includes the probability of detection (POD), success ratio (SR), critical success index (CSI), and the hit bias. Continuous statistics such as correlation, normalized standard deviation, and root-mean-square error are also evaluated. Overall, GEOS-Forecast precipitation field over the analysis period (~1 year) show underestimation with respect to IMERG early for the daily accumulated rainfall. However, the probability distribution function and cumulative distribution function of both show similar patterns. In terms of correlations, POD, SR, CSI, hit bias, the performance varies with respect to the rainfall threshold used.

Sana Khan↗

A Scheduling Algorithm for Replicated Real-Time Tasks

We present an algorithm for scheduling real-time periodic tasks on a multiprocessor system under fault-tolerant requirement. Our approach incorporates both the redundancy and masking technique and the imprecise computation model. Since the tasks in hard real-time systems have stringent timing constraints, the redundancy and masking technique are more appropriate than the rollback techniques which usually require extra time for error recovery. The imprecise computation model provides flexible functionality by trading off the quality of the result produced by a task with the amount of processing time required to produce it. It therefore permits the performance of a real-time system to degrade gracefully. We evaluate the algorithm by stochastic analysis and Monte Carlo simulations. The results show that the algorithm is resilient under hardware failures.

Yu, Albert C.↗

A Byzantine resilient processor with an encoded fault-tolerant shared memory

The memory requirements for ultra-reliable computers are expected to increase due to future increases in mission functionality and operating-system requirements. This increase will have a negative effect on the reliability and cost of the system. Increased memory size will also reduce the ability to reintegrate a channel after a transient fault, since the time required to reintegrate a channel in a conventional fault-tolerant processor is dominated by memory realignment time. A Byzantine Resilient Fault-Tolerant Processor with Fault-Tolerant Shared Memory (FTP/FTSM) is presented as a solution to these problems. The FTSM uses an encoded memory system, which reduces the memory requirement by one-half compared to a conventional quad-FTP design. This increases the reliability and decreases the cost of the system. The realignment problem is also addressed by the FTSM. Because any single error is corrected upon a read from the FTSM, a faulty channel's corrupted memory does not need realignment before reintegration of the faulty channel. A combination of correct-on-access and background scrubbing is proposed to prevent the accumulation of transient errors in the memory. With a hardware-implemented scrubber, the scrubbing cycle time, and therefore the memory fault latency, can be upper-bounded at a small value. This technique increases the reliability of the memory system and facilitates validation of its reliability model.

Butler, Bryan↗

Psychophysiological Methods to Assess Pilot Productive Safety Behaviors

The NASA System-Wide Safety (SWS) Project is focused on developing new technologies and operational concepts for the aviation industry to meet the increasing global demand while maintaining the current ultra-safe system safety levels. To achieve this, the SWS Project is developing research priorities, including In-time System-wide Safety Assurance (ISSA) and In-time Aviation Safety Management System (IASMS; Ellis et al., 2019). A critical component of the IASMS is the human as pilot and in other roles in aviation operations as demonstrated by SWS human factors research on rare occurrences of human error and the far more prevalent human safety producing behaviors (e.g., Hollnagel, 2016). The talk presented by Chad Stephens of NASA Langley Research Center and NASA SWS Project will describe the history of human factors research involving psychophysiological and biocybernetics methods supporting aviation safety conducted at NASA. Specific examples of recent NASA crew state monitoring research focused on a psychophysiological assessment method and system to enable Training for Attention Management will be demonstrated. Current SWS research including the SWS Operations and Technologies for Enabling Resilient In-Time Assurance (SOTERIA) flight simulation study and a data testbed created to enable study of Human Contributions to Safety (HC2S) will be presented. Ongoing collaborative research efforts with Boeing researchers will be highlighted and opportunities for further collaboration will be discussed.

psychophysiology↗

A GPS-Based Pitot-Static Calibration Method Using Global Output-Error Optimization

Pressure-based airspeed and altitude measurements for aircraft typically require calibration of the installed system to account for pressure sensing errors such as those due to local flow field effects. In some cases, calibration is used to meet requirements such as those specified in Federal Aviation Regulation Part 25. Several methods are used for in-flight pitot-static calibration including tower fly-by, pacer aircraft, and trailing cone methods. In the 1990 s, the introduction of satellite-based positioning systems to the civilian market enabled new inflight calibration methods based on accurate ground speed measurements provided by Global Positioning Systems (GPS). Use of GPS for airspeed calibration has many advantages such as accuracy, ease of portability (e.g. hand-held) and the flexibility of operating in airspace without the limitations of test range boundaries or ground telemetry support. The current research was motivated by the need for a rapid and statistically accurate method for in-flight calibration of pitot-static systems for remotely piloted, dynamically-scaled research aircraft. Current calibration methods were deemed not practical for this application because of confined test range size and limited flight time available for each sortie. A method was developed that uses high data rate measurements of static and total pressure, and GPSbased ground speed measurements to compute the pressure errors over a range of airspeed. The novel application of this approach is the use of system identification methods that rapidly compute optimal pressure error models with defined confidence intervals in nearreal time. This method has been demonstrated in flight tests and has shown 2- bounds of approximately 0.2 kts with an order of magnitude reduction in test time over other methods. As part of this experiment, a unique database of wind measurements was acquired concurrently with the flight experiments, for the purpose of experimental validation of the optimization method. This paper describes the GPS-based pitot-static calibration method developed for the AirSTAR research test-bed operated as part of the Integrated Resilient Aircraft Controls (IRAC) project in the NASA Aviation Safety Program (AvSP). A description of the method will be provided and results from recent flight tests will be shown to illustrate the performance and advantages of this approach. Discussion of maneuver requirements and data reduction will be included as well as potential applications.

Foster, John V.↗

The Human Challenges of Remotely Piloted Aircraft Systems

Remotely Piloted Aircraft (RPA) range from small electric quadcopters that are flown close to the ground within visual range of the operator, to larger systems capable of extended flight in airspace shared with conventional aircraft. Before RPA can operate routinely and safely in civilian airspace, we need to understand the unique human factors associated with these aircraft. RPA pilots participated in focus groups where they were asked to recall critical incidents that either presented a threat to safety, or highlighted a case where the pilot contributed to system resilience or mission success. Ninety incidents were gathered from focus-groups. Human factor issues included the impact of reduced sensory cues, traffic separation in the absence of an out-the-window view, control latencies, vigilance during monotonous and ultra-long endurance flights, control station design considerations, transfer of control between control stations, the management of lost link procedures, and decision-making during emergencies. Some of these concerns have received significant attention in the literature, or are analogous to human factors of manned aircraft. Although many of the reported incidents were related to pilot error, the participants also provided examples of the positive contribution that humans make to the operation of highly-automated systems.

Hobbs, Alan↗

Adaptive Controller Effects on Pilot Behavior

Adaptive control provides robustness and resilience for highly uncertain, and potentially unpredictable, flight dynamics characteristic. Some of the recent flight experiences of pilot-in-the-loop with an adaptive controller have exhibited unpredicted interactions. In retrospect, this is not surprising once it is realized that there are now two adaptive controllers interacting, the software adaptive control system and the pilot. An experiment was conducted to categorize these interactions on the pilot with an adaptive controller during control surface failures. One of the objectives of this experiment was to determine how the adaptation time of the controller affects pilots. The pitch and roll errors, and stick input increased for increasing adaptation time and during the segment when the adaptive controller was adapting. Not surprisingly, altitude, cross track and angle deviations, and vertical velocity also increase during the failure and then slowly return to pre-failure levels. Subjects may change their behavior even as an adaptive controller is adapting with additional stick inputs. Therefore, the adaptive controller should adapt as fast as possible to minimize flight track errors. This will minimize undesirable interactions between the pilot and the adaptive controller and maintain maneuvering precision.

Trujillo, Anna C.↗

Cyber-Threat Assessment for the Air Traffic Management System: A Network Controls Approach

Air transportation networks are being disrupted with increasing frequency by failures in their cyber- (computing, communication, control) systems. Whether these cyber- failures arise due to deliberate attacks or incidental errors, they can have far-reaching impact on the performance of the air traffic control and management systems. For instance, a computer failure in the Washington DC Air Route Traffic Control Center (ZDC) on August 15, 2015, caused nearly complete closure of the Centers airspace for several hours. This closure had a propagative impact across the United States National Airspace System, causing changed congestion patterns and requiring placement of a suite of traffic management initiatives to address the capacity reduction and congestion. A snapshot of traffic on that day clearly shows the closure of the ZDC airspace and the resulting congestion at its boundary, which required augmented traffic management at multiple locations. Cyber- events also have important ramifications for private stakeholders, particularly the airlines. During the last few months, computer-system issues have caused several airlines fleets to be grounded for significant periods of time: these include United Airlines (twice), LOT Polish Airlines, and American Airlines. Delays and regional stoppages due to cyber- events are even more common, and may have myriad causes (e.g., failure of the Department of Homeland Security systems needed for security check of passengers, see [3]). The growing frequency of cyber- disruptions in the air transportation system reflects a much broader trend in the modern society: cyber- failures and threats are becoming increasingly pervasive, varied, and impactful. In consequence, an intense effort is underway to develop secure and resilient cyber- systems that can protect against, detect, and remove threats, see e.g. and its many citations. The outcomes of this wide effort on cyber- security are applicable to the air transportation infrastructure, and indeed security solutions are being implemented in the current system. While these security solutions are important, they only provide a piecemeal solution. Particular computers or communication channels are protected from particular attacks, without a holistic view of the air transportation infrastructure. On the other hand, the above-listed incidents highlight that a holistic approach is needed, for several reasons. First, the air transportation infrastructure is a large scale cyber-physical system with multiple stakeholders and diverse legacy assets. It is impractical to protect every cyber- asset from known and unknown disruptions, and instead a strategic view of security is needed. Second, disruptions to the cyber- system can incur complex propagative impacts across the air transportation network, including its physical and human assets. Also, these implications of cyber- events are exacerbated or modulated by other disruptions and operational specifics, e.g. severe weather, operator fatigue or error, etc. These characteristics motivate a holistic and strategic perspective on protecting the air transportation infrastructure from cyber- events. The analysis of cyber- threats to the air traffic system is also inextricably tied to the integration of new autonomy into the airspace. The replacement of human operators with cyber functions leaves the network open to new cyber threats, which must be modeled and managed. Paradoxically, the mitigation of cyber events in the airspace will also likely require additional autonomy, given the fast time scale and myriad pathways of cyber-attacks which must be managed. The assessment of new vulnerabilities upon integration of new autonomy is also a key motivation for a holistic perspective on cyber threats.

Complex Networks↗

A Prognostics Framework Development for Swarm Satellite Formations

Prognostics is the science of predicting the failure(s) of a component or a system and understanding how the performance will change in the event of a failure or degradation mechanism. With accurate predictions of possible failures, autonomous mitigative actions can be taken to correct/repair any issues or alert human operators of a failure threshold exceedance requiring condition-based maintenance. Although there is extensive research on failure predictions for a component or a system, there are significantly more opportunities to foray into failure predictions and prognostics for a system of systems such as an airspace consisting of multiple aircraft, a fleet of unmanned aerial vehicles, and a swarm of intelligent satellite systems. Failure prediction and mitigation are particularly important in autonomous systems such as satellite swarm systems that need effective resource management and minimal human interactions. Based on NASA's decadal survey, there is a clear need to prioritize the development of satellite swarm technology for studies of space physics and Earth science. The science community will propose future missions that return in-situ measurements from a 3-D (three-dimensional) volume of space, with relative spacecraft motion and inter-satellite baselines controlled according to the mission objectives. For such multi-spacecraft missions, it is required that ground operations resources do not scale with the number of satellites, thus compromising the swarm or leading to inefficiencies in resource allocation. Swarms of tens or hundreds of small satellites will require autonomy in attitude control, navigation and failure. Although significant research has been conducted in the areas of autonomous formation flying algorithms, less attention has been given to the development of resilient systems robust to failures.The focus of this research paper is the integration of model-based prognostics into the swarm dynamics control and decision-making algorithms. We simulate swarm management strategies for a subsystem failure to demonstrate the importance of failure predictions by comparing two cases: (i) no health information is provided to the system and utilized in the decision-making process and (2) system health information is obtained using prognostics and employed by the control system. One example scenario presented is for the GPS (Global Positioning System) system of an individual satellite to perform off-nominally due to increasing estimated error. In this scenario, the keep-out zone for that satellite would become more conservative, thereby decreasing the risk of collision. This is achieved via tuning the individual artificial repulsive functions assigned to each satellite.This paper is structured as follows. First we provide an overview of current swarm technology development, where we specifically use the term swarm to define multiple satellites flying in formation in similar orbits, with cross-link communication and station-keeping capabilities. Second, we give an introduction to the Swarm Orbital Dynamics Advisor (SODA), a tool that accepts high-level configuration commands and provides the orbital maneuvers required to achieve the prescribed formation configuration. Third, we provide the details of the model-based prognostics algorithm implementation in SODA. Finally, we present different case studies for potential component/subsystem failures and the swarm responses based with and without failure prediction information.

prognostics↗