Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

A Real-Time High Performance Data Compression Technique For Space Applications

A high performance lossy data compression technique is currently being developed for space science applications under the requirement of high-speed push-broom scanning. The technique is also error-resilient in that error propagation is contained within a few scan lines. The algorithm is based on block-transform combined with bit-plane encoding; this combination results in an embedded bit string with exactly the desirable compression rate. The lossy coder is described. The compression scheme performs well on a suite of test images typical of images from spacecraft instruments. Hardware implementations are in development; a functional chip set is expected by the end of 2001.

Yeh, Pen-Shu↗

Visually Lossless Data Compression for Real-Time Frame/Pushbroom Space Science Imagers

A visually lossless data compression technique is currently being developed for space science applications under the requirement of high-speed push-broom scanning. The technique is also applicable to frame based imaging and is error-resilient in that error propagation is contained within a few scan lines. The algorithm is based on a block transform of a hybrid of modulated lapped transform (MLT) and discrete cosine transform (DCT), or a 2-dimensional lapped transform, followed by bit-plane encoding; this combination results in an embedded bit string with exactly the desirable compression rate as desired by the user. The approach requires no unique table to maximize its performance. The compression scheme performs well on a suite of test images typical of images from spacecraft instruments. Flight qualified hardware implementations are in development; a functional chip set is expected by the end of 2001. The chip set is being designed to compress data in excess of 20 Msamples/sec and support quantizations from 2 to 16 bits.

Yeh, Pen-Shu↗

Optical storage media data integrity studies

Optical disk-based information systems are being used in private industry and many Federal Government agencies for on-line and long-term storage of large quantities of data. The storage devices that are part of these systems are designed with powerful, but not unlimited, media error correction capacities. The integrity of data stored on optical disks does not only depend on the life expectancy specifications for the medium. Different factors, including handling and storage conditions, may result in an increase of medium errors in size and frequency. Monitoring the potential data degradation is crucial, especially for long term applications. Efforts are being made by the Association for Information and Image Management Technical Committee C21, Storage Devices and Applications, to specify methods for monitoring and reporting to the user medium errors detected by the storage device while writing, reading or verifying the data stored in that medium. The Computer Systems Laboratory (CSL) of the National Institute of Standard and Technology (NIST) has a leadership role in the development of these standard techniques. In addition, CSL is researching other data integrity issues, including the investigation of error-resilient compression algorithms. NIST has conducted care and handling experiments on optical disk media with the objective of identifying possible causes of degradation. NIST work in data integrity and related standards activities is described.

Podio, Fernando L.↗

Research in computer science

Several short summaries of the work performed during this reporting period are presented. Topics discussed in this document include: (1) resilient seeded errors via simple techniques; (2) knowledge representation for engineering design; (3) analysis of faults in a multiversion software experiment; (4) implementation of parallel programming environment; (5) symbolic execution of concurrent programs; (6) two computer graphics systems for visualization of pressure distribution and convective density particles; (7) design of a source code management system; (8) vectorizing incomplete conjugate gradient on the Cyber 203/205; (9) extensions of domain testing theory and; (10) performance analyzer for the pisces system.

Ortega, J. M.↗

Resilience, ASRS, and the Narrative about Human Error

We present a study of weather-related incident reports submitted to NASA’s Aviation Safety Reporting System (ASRS) by air carrier pilots in the US. Using specific examples, we examine the relevant aspects of human performance and resilience management exhibited during these incidents. We describe the common narrative about human error and how ASRS data can be used to change it.

resilience↗

Modeling Distributed Situation Awareness in Resilience-Based Design of Complex Engineered Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling↗

Modeling Distributed Situation Awareness in Resilience-based Design of Complex Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling↗

Can Resilience Assessments Inform Early Design Human Factors Decision-making?

There is a growing call among researchers for tighter coupling between human factors and human reliability assessments. In this research, we explore if early design stage resilience assessments can help bridge some of the gaps between human factors and human reliability assessments. Resilience in systems is their ability to recover reasonably and operate within acceptable bounds during failures and unexpected events. The fmdtools toolkit allows designers to assess the resilience of a system by modeling the human error and machine-related failure propagation in both nominal and faulty scenarios during the early design stages. As a result, the fmdtools toolkit has a low-fidelity dynamic human reliability assessment component built into it. In this paper, we study if the results from the fmdtools simulations can help inform and prioritize human factors design decision-making, resulting in tighter coupling between human factors and human reliability assessments during the design process. Specifically, we explore the results from a rover design example to understand the types of information that can help guide human factor-related decision-making.

Resilience-based Design↗

An Approach for the Assessment of System Upset Resilience

This report describes an approach for the assessment of upset resilience that is applicable to systems in general, including safety-critical, real-time systems. For this work, resilience is defined as the ability to preserve and restore service availability and integrity under stated conditions of configuration, functional inputs and environmental conditions. To enable a quantitative approach, we define novel system service degradation metrics and propose a new mathematical definition of resilience. These behavioral-level metrics are based on the fundamental service classification criteria of correctness, detectability, symmetry and persistence. This approach consists of a Monte-Carlo-based stimulus injection experiment, on a physical implementation or an error-propagation model of a system, to generate a system response set that can be characterized in terms of dimensional error metrics and integrated to form an overall measure of resilience. We expect this approach to be helpful in gaining insight into the error containment and repair capabilities of systems for a wide range of conditions.

Torres-Pomales, Wilfredo↗

Evaluating the Use of High-Fidelity Simulator Research Methods to Study Airline Flight Crew Resilience

As it evolves, aviation will continue to require integration of a wide range of safety systems and practices, some of which are already in place and others that are yet to be developed. New concepts in system safety thinking have emerged to consider not only what may go wrong, but also what can be learned when things go right during commercial flight operations. Taken together, these complementary perspectives form a more comprehensive approach to systemsafety thinking that can help to recognize and preserve the resilient performance capabilities currently provided by humans. A need exists, however, for research methods to enable better understanding of the human contributions to aviation safety. NASA’s System-Wide Safety Project supports research on using flight simulation methods to study operator resilience and safety-producing behaviors. Building on prior NASA efforts investigating procedural non-adherences during area navigation standard terminal route arrivals, a high-fidelity commercial aviation line operational simulation (LOS) experiment has been designed to study how flight crews anticipate, monitor for, respond to, and learn from expected and unexpected disturbances during these operations. A diverse set of LOS scenarios were developed to simulate highly realistic, complex, but routinely encountered operational situations. Each scenario provided multiple opportunities to collect data on how flight crews manage threats and errors, as well as novel opportunities to observe resilient and safety-producing behaviors. The experimental design, implications for the study of safety-producing behaviors using simulation, and considerations for airline pilot training will be discussed.

Chad L Stephens↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in mission or safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes a set of 55 historic aerospace software error incidents from 1962 to 2023 to determine trends of how and where automation is most likely to fail, behaving unexpectedly. A distinction between software producing unexpected (erroneous) output versus no output (failsilent) is introduced. Of the historical incidents analyzed, 85% were from software producing wrong output rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and not reliable to recover from silent failures. Error origin was within the code/logic itself in 58% of cases, 16% from configurable data, 15% from unexpected sensor input, and 11% from command/operator input. A substantial forty percent (40%) of unexpected software behavior was indicated by the absence of code, arising from unanticipated situations and missing requirements, and 16% of incidents were subjectively deemed “unknown-unknowns”. No incidents were found to be the result of programming language, compiler, tool, or operating system; and only sixteen percent (16%) of all incidents were considered errors traditional computer science/programming in nature. These findings indicate that for fault tolerance, erroneous automation behavior must be a primary consideration especially at critical moments, and reboot recoverability may not be viable. Special care should be taken to validate configurable data and commands prior to use. “Test-like-you-fly”, including hardware-in-the-loop combined with robust off-nominal testing should be used to uncover missing logic arising from unanticipated situations not covered by requirements alone. This study uniquely focuses on manifestations of unexpected flight software behavior, independent of ultimate root cause. We characterize software error behavior and origin to improve software design, test, and operations for resilience to the most common manifestations, and provide a rich dataset for further study.

Aerospace↗

People are the Weak Link in the System.

People are often considered the "weak link" in every system. It is assumed that there is a causal "chain" of events, and each link in the chain provides some protection against failure. In this notion, the chain is only as strong as its weakest link and people are often blamed for failures. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are only the source of error and failure.

operations↗

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad↗

Why Didn't They Just Follow the Procedure?

It is often said that to err is human and that procedures are in place to prevent people from making mistakes. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. And while it's true that good procedures can help avoid error, it's also the case that procedures have their own limitations. This talk highlights the limits of procedures and the resilience people bring to operations. It discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure, and it proposes an approach to the design of procedures that supports the human operator.

operations↗

Resilience Modeling in Complex Engineered Systems with Human-Machine Interactions

In recent times, there has been a growing interest in resilience-based design. Resilience-based design operates on the concept that failures and unexpected events will happen, and when they occur, complex engineered systems should be able to operate within acceptable bounds and recover reasonably. Humans can contribute to the resilience of a system by quickly detecting unforeseen events and taking corrective measures. To this effect, researchers have proposed guidelines and design approaches that can help promote human-system resilience. However, there is no early design stage tool to validate if a system is indeed resilient after applying these guidelines and design methods. In this research, we integrate the Human Error and Functional Failure Reasoning (HEFFR) framework into the fmdtools toolkit to enable designers to model the combined (machine, human, and joint) failures, including their propagation and dynamic effects, during early design stages. This integrated tool also allows designers to model the effects of performance shaping factors, team dynamics, and human-machine interactions in systems of systems. A demonstrative example of a remotely operated rover is explored to demonstrate how this approach can be applied to understand resilience in complex engineered systems with human interactions.

Lukman Irshad↗

What Can We Learn From Resilient Pilot Behaviors? The Case of Energy Management While Flying a Star

Recently, there has been increased interest in documenting flightcrew behaviors that contribute to safe operations. Instead of only capturing errors, new efforts are attempting to understand how pilots manage complexity and variability in the operational environment to ensure a safe mission. This approach highlights pilot responses to events and conditions that fall outside typical TEM threats; e.g., revised ATC clearances. This approach presents a two-sided coin: characterize flightcrew resilience /or/ generate insights regarding complexity in the operational environment that is not adequately managed by current flight deck interface designs, procedures, and training. To capture operational complexity, we have been analyzing flight path management tied to flying an RNAV STAR. Because ATC often requests revisions—e.g., descend late—and because RNAV STARs may not align with airplane performance limits, flightcrews need to monitor, anticipate threats to RNAV STAR compliance, and devise ways to accommodate unexpected challenges. In this paper, we identify general strategies that can support response adaptation and explore methods to facilitate training these strategies.

flight operations↗

Thinking outside the box: The human role in increasingly automated aviation systems

Rapid advances in artificial intelligence are enabling automated systems to operate in an increasingly autonomous manner in domains that previously required the involvement of human operators. Examples are rail transport systems, self-driving cars, and warehouse delivery systems. From time to time, such automation encounters operational conditions that fall outside a “competency box” within which the system has been designed to operate. Human operators add resilience because they can see and act outside the competency box of scenarios and environments for which the system was designed. The system’s competencies can be expanded over time with modifications to software, sensors, etc.; however, it is unclear at what point the competency box becomes large enough to safely eliminate the role of the human operator. One area where advanced automation may be applied is Urban Air Mobility (UAM). Current UAM concepts envision fleets of highly automated air vehicles providing on-demand transport for people and goods. A phased development of UAM has been proposed, beginning with on-board pilots and transitioning to a future state where automated vehicles operate with minimal human involvement. Proponents of UAM note that this final state reduces cost as well as eliminating pilot error, identified as a contributing factor in many aircraft accidents. However, eliminating human involvement also risks eliminating their positive contributions to system resilience. Here we examine Concepts of Operation proposed for future UAM systems and explore how humans can best be incorporated to maintain resilience while minimizing cost and risk. A human-autonomy teaming approach is suggested.

Advanced Air Mobility↗

How Humans Contribute to Safety

We have all heard, and much too often, how human error is the leading cause of accidents. What we haven’t been hearing is how humans produce safety far more often than reduce safety. Before we embark on developing technologies to replace the error-prone human, it behooves us to understand how humans produce safety lest we lose that primary source of resilience in our aviation system.

safety↗