Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Resilient State Recovery Using Prior Measurement Support Information

Resilient state recovery of cyber-physical systems has attracted much research attention due to the unique challenges posed by the tight coupling between communication, computation, and the underlying physics of such systems. By modeling attacks as additive adversary signals to a sparse subset of measurements, this resilient recovery problem can be formulated as an error correction problem. To achieve exact state recovery, most existing results require less than 50% of the measurement nodes to be compromised, which limits the resiliency of the estimators. In this paper, we show that observer resiliency can be further improved by incorporating data-driven prior information. Here, we provide an analytical bridge between the precision of prior information and the resiliency of the estimator. By quantifying the relationship between the estimation error of the weighted ℓ 1 observer and the precision of the support prior, this quantified relationship provides guidance for the estimator’s weight design to achieve optimal resiliency. Several numerical simulations and an application case study are presented to validate the theoretical claims.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Can Resilience Assessments Inform Early Design Human Factors Decision-making?

There is a growing call among researchers for tighter coupling between human factors and human reliability assessments. In this research, we explore if early design stage resilience assessments can help bridge some of the gaps between human factors and human reliability assessments. Resilience in systems is their ability to recover reasonably and operate within acceptable bounds during failures and unexpected events. The fmdtools toolkit allows designers to assess the resilience of a system by modeling the human error and machine-related failure propagation in both nominal and faulty scenarios during the early design stages. As a result, the fmdtools toolkit has a low-fidelity dynamic human reliability assessment component built into it. In this paper, we study if the results from the fmdtools simulations can help inform and prioritize human factors design decision-making, resulting in tighter coupling between human factors and human reliability assessments during the design process. Specifically, we explore the results from a rover design example to understand the types of information that can help guide human factor-related decision-making.

Resilience-based Design↗

Quantum Zeno Monte Carlo for computing observables

The recent development of logical quantum processors marks a pivotal transition from the noisy intermediate-scale quantum (NISQ) era to the fault-tolerant quantum computing (FTQC) era. These devices have the potential to address classically challenging problems with polynomial computational time using quantum properties. However, they remain susceptible to noise, necessitating noise resilient algorithms. We introduce Quantum Zeno Monte Carlo (QZMC), a classical-quantum hybrid algorithm that demonstrates resilience to device noise and Trotter errors while showing polynomial computational cost for a gapped system. QZMC computes static and dynamic properties without requiring initial state overlap or variational parameters, offering reduced quantum circuit depth.

Han, Mancheon [Korea Institute for Advanced Study ↗

Control optimization for parametric Hamiltonians by pulse reconstruction

Optimal control techniques provide a means to tailor the control pulses required to generate customized quantum gates, which helps to improve the resilience of quantum simulations to gate errors and device noise. However, the significant amount of (classical) computation required to generate customized gates can quickly undermine the effectiveness of this approach, especially when pulse optimization needs to be iterated. We propose a method to reduce the computational time required to generate the control pulse for a Hamiltonian that is parametrically dependent on a time-varying quantity. We use simple interpolation schemes to accurately reconstruct the control pulses from a set of pulses obtained in advance for a discrete set of predetermined parameter values. We obtain a reconstruction with very high fidelity and a significant reduction in computational effort. We report the results of the application of the proposed method to device-level quantum simulations of the unitary (real) time evolution of two interacting neutrons based on superconducting qubits.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Visual Comparison of Silent Error Propagation

High-performance computing (HPC) systems play a critical role in facilitating scientific discoveries. Their scale and complexity (e.g., the number of computational units and software stack) continue to grow as new systems are expected to process increasingly more data and reduce computing time. However, with more processing elements, the probability that these systems will experience a random bit-flip error that corrupts a program's output also increases, which is often recognized as silent data corruption. Analyzing the resiliency of HPC applications in extreme-scale computing to silent data corruption is crucial but difficult. An HPC application often contains a large number of computation units that need to be tested, and error propagation caused by error corruption is complex and difficult to interpret. Here, to accommodate this challenge, we propose an interactive visualization system that helps HPC researchers understand the resiliency of HPC applications and compare their error propagation. Our system models an application's error propagation to study a program's resiliency by constructing and visualizing its fault tolerance boundary. Coordinating with multiple interactive designs, our system enables domain experts to efficiently explore the complicated spatial and temporal correlation between error propagations. At the end, the system integrated a nonmonotonic error propagation analysis with an adjustable graph propagation visualization to help domain experts examine the details of error propagation and answer such questions as why an error is mitigated or amplified by program execution.

97 MATHEMATICS AND COMPUTING↗

An Approach for the Assessment of System Upset Resilience

This report describes an approach for the assessment of upset resilience that is applicable to systems in general, including safety-critical, real-time systems. For this work, resilience is defined as the ability to preserve and restore service availability and integrity under stated conditions of configuration, functional inputs and environmental conditions. To enable a quantitative approach, we define novel system service degradation metrics and propose a new mathematical definition of resilience. These behavioral-level metrics are based on the fundamental service classification criteria of correctness, detectability, symmetry and persistence. This approach consists of a Monte-Carlo-based stimulus injection experiment, on a physical implementation or an error-propagation model of a system, to generate a system response set that can be characterized in terms of dimensional error metrics and integrated to form an overall measure of resilience. We expect this approach to be helpful in gaining insight into the error containment and repair capabilities of systems for a wide range of conditions.

Torres-Pomales, Wilfredo↗

Characterizing HVDC Transmission Flexibility under Extreme Operating Conditions

System operators rely on system flexibility, traditionally mainly from generation, to handle unexpected reliability and resilience events, ranging from excessive resource forecast errors to extreme events like heatwaves, earthquakes, and cyberattacks. Flexible transmission, such as controllable high voltage direct current (HVDC) transmission systems present an opportunity to increase overall system flexibility to accommodate operational challenges. This paper provides a methodology to study contributions to system flexibility by controllable, power electronics based transmission. The Western Electricity Coordinating Council (WECC) system is used as an example to study contributions from existing and future HVDC lines. Under extreme system conditions, it is identified that HVDC transmission flexibility can contribute with 24.8 to 28% of avoided unserved energy, and in some areas, the benefits amount to 50 to 70%.

HVDC transmission, PCM, Balancing authorities↗

Evaluating the Use of High-Fidelity Simulator Research Methods to Study Airline Flight Crew Resilience

As it evolves, aviation will continue to require integration of a wide range of safety systems and practices, some of which are already in place and others that are yet to be developed. New concepts in system safety thinking have emerged to consider not only what may go wrong, but also what can be learned when things go right during commercial flight operations. Taken together, these complementary perspectives form a more comprehensive approach to systemsafety thinking that can help to recognize and preserve the resilient performance capabilities currently provided by humans. A need exists, however, for research methods to enable better understanding of the human contributions to aviation safety. NASA’s System-Wide Safety Project supports research on using flight simulation methods to study operator resilience and safety-producing behaviors. Building on prior NASA efforts investigating procedural non-adherences during area navigation standard terminal route arrivals, a high-fidelity commercial aviation line operational simulation (LOS) experiment has been designed to study how flight crews anticipate, monitor for, respond to, and learn from expected and unexpected disturbances during these operations. A diverse set of LOS scenarios were developed to simulate highly realistic, complex, but routinely encountered operational situations. Each scenario provided multiple opportunities to collect data on how flight crews manage threats and errors, as well as novel opportunities to observe resilient and safety-producing behaviors. The experimental design, implications for the study of safety-producing behaviors using simulation, and considerations for airline pilot training will be discussed.

Chad L Stephens↗

Resilience–runtime tradeoff relations for quantum algorithms

Abstract A leading approach to algorithm design aims to minimize the number of operations in an algorithm’s compilation. One intuitively expects that reducing the number of operations may decrease the chance of errors. This paradigm is particularly prevalent in quantum computing, where gates are hard to implement and noise rapidly decreases a quantum computer’s potential to outperform classical computers. Here, we find that minimizing the number of operations in a quantum algorithm can be counterproductive, leading to a noise sensitivity that induces errors when running the algorithm in non-ideal conditions. To show this, we develop a framework to characterize the resilience of an algorithm to perturbative noises (including coherent errors, dephasing, and depolarizing noise). Some compilations of an algorithm can be resilient against certain noise sources while being unstable against other noises. We condense these results into a tradeoff relation between an algorithm’s number of operations and its noise resilience. We also show how this framework can be leveraged to identify compilations of an algorithm that are better suited to withstand certain noises.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

EFCOG Best Practice HPI for Knowledge Workers ISM-HPI-22-02

This document is a collection of these best practices as determined by team members. This best practice attempts to: Realize opportunities to break the myth where people believe that HPI does not apply to them as they perform no physical work. Recommend options to create an environment that promotes intellectual collaboration and trust, enabling candor and vulnerability. Explain how errors manifest differently from the same human fallibility. Knowledge workers (KW) have errors that take different perspectives to find and mitigate the unique manifestation of these conditions. Help KW identify the critical steps (or risk important steps) in their processes. Reduce risk/consequence from KW errors (limit latent errors as well as finding latent conditions), building resiliency into KW tasks. Mitigation strategies may be different.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Adaptive Cyber-Physical Resilience for Building Control Systems

The main goal of the project is to develop an AI-based process layer cybersecurity suite for detection, isolation and mitigation of cyber-attack effects on operation of building energy management systems (BEMS). The following constituent key technologies were developed under the program towards fulfilling the program objectives: (1) developed a high fidelity BEMS testbed for generation of training data and validation of developed technologies; (2) developed a physics informed ML based attack detection and localization module (ADL) capable of detecting high impact stealthy attacks (HISA - attacks causing 30% energy utilization but no immediate visible impact otherwise) with 98% accuracy; (3) developed a methodology to determine ’representative days’ to limit the data required for training; (4) developed a virtual sensing system that can reconstruct affected sensors with 10% error for the same HISA set; (5) developed a resilient model predictive control system that can continue operation of the BEMS without jeopardizing stability for the HISA set; and (6) integrated and deployed all the constituent modules and demonstrated the efficacy of the technology in real-time in a hardware in loop simulation.

42 ENGINEERING↗

Historical Aerospace Software Errors Categorized to Influence Fault Tolerance

Since the first use of computers in space and aircraft, software errors have occurred. These errors can manifest as loss-of-life or less catastrophically. As the demand for automation increases, software in mission or safety-critical systems should be designed to be tolerant to the most likely software faults. This paper categorizes a set of 55 historic aerospace software error incidents from 1962 to 2023 to determine trends of how and where automation is most likely to fail, behaving unexpectedly. A distinction between software producing unexpected (erroneous) output versus no output (failsilent) is introduced. Of the historical incidents analyzed, 85% were from software producing wrong output rather than simply stopping. Rebooting was found to be ineffective to clear erroneous behavior, and not reliable to recover from silent failures. Error origin was within the code/logic itself in 58% of cases, 16% from configurable data, 15% from unexpected sensor input, and 11% from command/operator input. A substantial forty percent (40%) of unexpected software behavior was indicated by the absence of code, arising from unanticipated situations and missing requirements, and 16% of incidents were subjectively deemed “unknown-unknowns”. No incidents were found to be the result of programming language, compiler, tool, or operating system; and only sixteen percent (16%) of all incidents were considered errors traditional computer science/programming in nature. These findings indicate that for fault tolerance, erroneous automation behavior must be a primary consideration especially at critical moments, and reboot recoverability may not be viable. Special care should be taken to validate configurable data and commands prior to use. “Test-like-you-fly”, including hardware-in-the-loop combined with robust off-nominal testing should be used to uncover missing logic arising from unanticipated situations not covered by requirements alone. This study uniquely focuses on manifestations of unexpected flight software behavior, independent of ultimate root cause. We characterize software error behavior and origin to improve software design, test, and operations for resilience to the most common manifestations, and provide a rich dataset for further study.

Aerospace↗

People are the Weak Link in the System.

People are often considered the "weak link" in every system. It is assumed that there is a causal "chain" of events, and each link in the chain provides some protection against failure. In this notion, the chain is only as strong as its weakest link and people are often blamed for failures. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are only the source of error and failure.

operations↗

Quantum error mitigation by hidden inverses protocol in superconducting quantum devices

We present a method to improve the convergence of variational algorithms based on hidden inverses (HIs) to mitigate coherent errors. In the context of error mitigation, this means replacing the hardware implementation of certain Hermitian gates with their inverses. Doing so results in noise cancellation and a more resilient quantum circuit. This approach improves performance in a variety of two-qubit error models where the noise operator also inverts with the gate inversion. We apply the mitigation scheme on superconducting quantum processors running the variational quantum eigensolver (VQE) algorithm to find the H 2 ground-state energy. When implemented on superconducting hardware we find that the mitigation scheme effectively reduces the energy fluctuations in the parameter learning path in VQE, reducing the number of iterations for a converged value. We also provide a detailed numerical simulation of VQE performance under different noise models and explore how HIs and randomized compiling affect the underlying loss landscape of the learning problem. These simulations help explain our experimental hardware outcomes, helping to connect lower-level gate performance to application-specific behavior in contrast to metrics like fidelity which often do not provide an intuitive insight into observed high level performance.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Noise-Resilient and Reduced Depth Approximate Adders for NISQ Quantum Computing

The "Noisy intermediate-scale quantum" NISQ machine era primarily focuses on mitigating noise, controlling errors, and executing high-fidelity operations, hence requiring shallow circuit depth and noise robustness. Approximate computing is a novel computing paradigm that produces imprecise results by relaxing the need for fully precise output for error-tolerant applications including multimedia, data mining, and image processing. We investigate how approximate computing can improve the noise resilience of quantum adder circuits in NISQ quantum computing. We propose five designs of approximate quantum adders to reduce depth while making them noise-resilient, in which three designs are with carryout, while two are without carryout. We have used novel design approaches that include approximating the Sum only from the inputs (pass-through designs) and having zero depth, as they need no quantum gates. The second design style uses a single CNOT gate to approximate the SUM with a constant depth of O(1). We performed our experimentation on IBM Qiskit on noise models including thermal, depolarizing, amplitude damping, phase damping, and bitflip: (i) Compared to exact quantum ripple carry adder without carryout the proposed approximate adders without carryout have improved fidelity ranging from 8.34% to 219.22%, and (ii) Compared to exact quantum ripple carry adder with carryout the proposed approximate adders with carryout have improved fidelity ranging from 8.23% to 371%. Further, the proposed approximate quantum adders are evaluated in terms of various error metrics.

Gaur, Bhaskar↗

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad↗