Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Optical storage media data integrity studies

Optical disk-based information systems are being used in private industry and many Federal Government agencies for on-line and long-term storage of large quantities of data. The storage devices that are part of these systems are designed with powerful, but not unlimited, media error correction capacities. The integrity of data stored on optical disks does not only depend on the life expectancy specifications for the medium. Different factors, including handling and storage conditions, may result in an increase of medium errors in size and frequency. Monitoring the potential data degradation is crucial, especially for long term applications. Efforts are being made by the Association for Information and Image Management Technical Committee C21, Storage Devices and Applications, to specify methods for monitoring and reporting to the user medium errors detected by the storage device while writing, reading or verifying the data stored in that medium. The Computer Systems Laboratory (CSL) of the National Institute of Standard and Technology (NIST) has a leadership role in the development of these standard techniques. In addition, CSL is researching other data integrity issues, including the investigation of error-resilient compression algorithms. NIST has conducted care and handling experiments on optical disk media with the objective of identifying possible causes of degradation. NIST work in data integrity and related standards activities is described.

Podio, Fernando L.↗

Resilience of the surface code to error bursts

Quantum error correction works effectively only if the error rate of gate operations is sufficiently low. However, some rare physical mechanisms can cause a temporary increase in the error rate that affects many qubits; examples include ionizing radiation in superconducting hardware and large deviations in the global control of atomic systems. We refer to such rare transient spikes in the gate error rate as error bursts. In this work, we investigate the resilience of the rotated surface code to generic error bursts. We assume that, after appropriate mitigation strategies, the spike in the error rate lasts for only a single syndrome-extraction cycle; we also assume that the enhanced error rate is uniform across the code block. Under these assumptions, and for a circuit-level depolarizing noise model, we perform Monte Carlo simulations to determine the regime in burst error rate and background error rate for which the memory time becomes arbitrarily long as the code block size grows. Our results indicate that suitable hardware mitigation methods combined with standard decoding methods may suffice to protect against transient error bursts in the rotated surface code.

Quantum error correction↗

Research in computer science

Several short summaries of the work performed during this reporting period are presented. Topics discussed in this document include: (1) resilient seeded errors via simple techniques; (2) knowledge representation for engineering design; (3) analysis of faults in a multiversion software experiment; (4) implementation of parallel programming environment; (5) symbolic execution of concurrent programs; (6) two computer graphics systems for visualization of pressure distribution and convective density particles; (7) design of a source code management system; (8) vectorizing incomplete conjugate gradient on the Cyber 203/205; (9) extensions of domain testing theory and; (10) performance analyzer for the pisces system.

Ortega, J. M.↗

Resilience, ASRS, and the Narrative about Human Error

We present a study of weather-related incident reports submitted to NASA’s Aviation Safety Reporting System (ASRS) by air carrier pilots in the US. Using specific examples, we examine the relevant aspects of human performance and resilience management exhibited during these incidents. We describe the common narrative about human error and how ASRS data can be used to change it.

resilience↗

Tough Errors are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This project builds toward a comprehensive error-mitigating toolkit that makes quantum programming more robust and adaptive to the noisy, resource-limited nature of today’s quantum hardware. To that end, it integrates established error-mitigation methods — such as zero-noise extrapolation and dynamical decoupling — directly into compiler infrastructures. These techniques will be packaged as modules that can automatically adjust and combine based on performance analysis, enabling compilers to explore large design spaces and produce optimized, low-noise quantum programs with minimal manual intervention. In parallel, this project also explores new approaches to analog quantum programming or quantum simulation, and has developed the programming language SimuQ which treats quantum Hamiltonian evolution as the central object.

97 MATHEMATICS AND COMPUTING↗

RDPM: An Extensible Tool for Resilience Design Patterns Modelling

Resilience to faults, errors, and failures in extreme-scale high-performance computing (HPC) systems is a critical challenge. Resilience design patterns offer a new, structured hardware and software design approach for improving resilience. While prior work focused on developing performance, reliability, and availability models for resilience design patterns, this paper extends it by providing a Resilience Design Patterns Modeling (RDPM) tool which allows (1) exploring performance, reliability, and availability of each resilience design pattern, (2) offering customization of parameters to optimize performance, reliability, and availability, and (3) allowing investigation of trade-off models for combining multiple patterns for practical resilience solutions.

Kumar, Mohit↗

Tough Errors Are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This report summarizes research performed under the Tough Errors Are no Match (TEAM) project. The primary focus of TEAM has been to research and develop a compilation toolbox leveraging techniques from quantum characterization and control, probabilistic programming, and approximate computing. Our goal was to develop robust protocols that can be integrated into quantum compilers to optimize and enhance the robustness of noisy computation. Here, we provide a summary of TEAM work focused on characterization and control of quantum systems.

97 MATHEMATICS AND COMPUTING↗

Ansatz-Free Hamiltonian Learning with Heisenberg-Limited Scaling

Learning the unknown interactions that govern a quantum system is crucial for quantum information processing, device benchmarking, and quantum sensing. The problem, known as Hamiltonian learning, is well understood under the assumption that interactions are local, but this assumption may not hold for arbitrary Hamiltonians. Previous methods all require high-order inverse polynomial dependency with precision, unable to surpass the standard quantum limit and reach the gold-standard Heisenberg-limited scaling. Whether Heisenberg-limited Hamiltonian learning is possible without prior assumptions about the interaction structures, a challenge we term ansatz-free Hamiltonian learning , remains an open question. In this work, we present a quantum algorithm to learn arbitrary sparse Hamiltonians without any structure constraints using only black-box queries of the system’s real-time evolution and minimal digital controls to attain Heisenberg-limited scaling in estimation error. Our method is also resilient to state-preparation-and-measurement errors, enhancing its practical feasibility. We numerically demonstrate our ansatz-free protocol for learning physical Hamiltonians and validating analog quantum simulations, benchmarking our performance against the state-of-the-art Heisenberg-limited learning approach. Moreover, we establish a fundamental trade-off between total evolution time and quantum control on learning arbitrary interactions, revealing the intrinsic interplay between controllability and total evolution-time complexity for any learning algorithm. These results pave the way for further exploration into Heisenberg-limited Hamiltonian learning in complex quantum systems under minimal assumptions, potentially enabling new benchmarking and verification protocols.

machine learning↗

Tough Errors are no Match (TEAM): Optimizing the Quantum Compiler for Noise Resilience

This report summarizes our contributions to the Department of Energy’s Tough Errors are no Match (TEAM) project (DE-SC0020377) under Thrust 2: Quantum Programming and Compilation. The central outcomes of this work included a novel efficient quantum compiling algorithm which works without requiring the quantum computer to exactly invert its operations, answering a longstanding open problem in quantum compiling. Additional results include the implementation of zero-noise extrapolation error mitigation in collaboration with the Unitary Fund, as well as novel quantum algorithms for entanglement detection and pseudorandomness.

Bouland, Adam [Stanford Univ., CA (United States)]↗

Intelligent Process Visualization through Nuclear Operation Process Modeling, Reasoning, and Object Detection from Field Videos (Final Report)

This report is a deliverable for the “Final Report” task of DOE NEET Project 19-16790, "Context-Aware Safety Information Display for Nuclear Field Workers." This project's overall goal is to test the hypothesis that integrating computer vision and process reasoning methods will enable proactive visualization of the safe operation and maintenance processes of Nuclear Power Plants (NPP) for field workers. Augmented Reality (AR) glasses adopting such proactive safety information visualization techniques can significantly increase personnel safety and reduce the NPP’s operating costs. The current practice of monitoring NPPs requires workers to switch between digital models, data, and physical workspaces in identifying relevant but potentially occluded objects and in assessing the risks of operation and maintenance processes. On the other hand, frequently changed field conditions require field workers to report to supervisors for real-time guidance. Such guidance is essential to ensure that changing conditions will not invalidate or endanger the work order and other ongoing processes that may jeopardize NPP operations. Additionally, incorrect recognition of equipment objects can result in communication errors and safety problems. AR techniques can assist engineers in viewing the physical workspaces with objects labeled with detailed operation procedures and safety reminders during field operations. The project team developed an “Intelligent Context-Aware Safety Information Display” (ICAD) for supporting Nuclear Power Plant (NPP) field workers in achieving safe and efficient execution of a series of operational tasks in uncertain and changing workspaces of an NPP. Before designing the ICAD-AR prototype, the project team synthesized NPP operational knowledge models through literature review studies, surveys, interviews with domain experts, and knowledge modeling. The project team conducted an extensive study of the operational procedures of various NPPs, and digital technologies that can support the safe and efficient execution of those procedures in different NPP operational contexts. This literature review helped the project team conduct surveys and interviews with nuclear engineers and field workers to identify three categories of information. The NPP knowledge modeling efforts reveal that the three categories of information identified have different levels of importance in a typical procedure of carrying out a series of tasks to achieve a specific NPP operation goal (e.g., shutdown, mode changes). These three categories of information include 1) Workspace dynamics – the changing spatial arrangements of workspaces, tools, protection equipment, and supporting materials, 2) Workflow prognostics – the dynamic dependencies between different parts of an NPP that functionally support and influence each other in terms of safety and efficiency, and 3) Hazards – objects and spaces that contain hazardous materials or physical conditions that can pose risks to workers or mechanical systems. The project team has profiled the importance levels of these categories of information into a knowledge model. This knowledge model specifies what types of information are more critical for a given task in a given workspace so that computers can automatically identify critical objects and sensors in a scene for delivering context-ware safety information to field workers through AR devices. Significant research development of this project results in technical research outcomes and a prototyping system that illustrates the technical feasibility of establishing an ICAD-AR system supporting the proactive safety information display for nuclear field workers. This final report summarizes the project team’s technological achievements in the past three years. Overall, the project team completed the development and integration of five techniques into a prototype ICAD Augmented Reality (ICAD-AR) system and demonstrated the developed system’s real-time execution in a mechanical room. The project team completed the analysis of using this prototype in other types of workspaces based on 3D image data and digital design models collected from two additional workspaces (a water treatment plant and a flow loop training facility). The integrated techniques include 1) Natural Language Processing (NLP) algorithms supporting the generation and updates of nuclear fieldwork process models based on text analysis of work packages and operation manuals; 2) sensor log analysis for predicting control actions in given sensor reading contexts; 3) computer vision algorithms for automatic localization and navigation of workers; 4) object detection algorithms for identifying task-related objects and correlated sensors for safety checking; 5) AR technique as a platform for supporting the integration. The testing results of these five techniques have shown that 1) the sensor log analysis model can predict the next control action with an accuracy of 0.883; 2) the trained natural language processing model can extract more than 80% of the critical information from paper-based procedures (PBPs); 3) the navigation algorithm with the integration of Visual Inertial Odometry (VIO) and Non-Recursive Bayesian Filter methods make operator’s trajectory estimation resilient to drift error; 4) the computer vision algorithm can detect task-specific and safety-critical objects with an average accuracy of 95.3%. The project team used work procedures collected from a flow loop training facility and two datasets collected from two mechanical rooms simulating the workspaces of NPPs to demonstrate the technical capabilities of the developed ICAD-AR prototype. The demonstration validated the technical feasibility of establishing the ICAD-AR system for nuclear field workers and identified the challenges in 1) automatic text analysis of work packages; 2) use of limited samples of sensor logs for predicting the proper timings of control actions; 3) reliably tracking workers and their task progress in mechanical rooms with many similar objects.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Access and limits of RMP ELM suppression with n = 1 fields in DIII-D

This work reports on DIII-D experiments aimed at extending resonant magnetic perturbation (RMP) suppression of edge localized modes (ELMs) to n = 1 fields, where n is the toroidal mode number. Modeling of the 3D ideal MHD plasma response to the RMPs using the GPEC code is used to quantify edge and core resonant fluxes, guiding experimental strategies to increase plasma resilience against core error field penetration, optimize multicoil phasing, and explore higher q 95 operation. In DIII-D, ELM mitigation is regularly observed across a wide range of n = 1 RMP scenarios. A ∼100 ms phase of complete ELM suppression was achieved at q 95 ∼ 3.9 using an odd-parity coil configuration. The suppressed phase exhibited clear signatures of RMP ELM suppression, including the elimination of Dα spikes, increased pedestal rotation, enhanced magnetic response, and elevated broadband density turbulence. An optimized coil configuration for edge-to-core resonant flux did show increased edge resonance indicated by increased density pumpout, but did not yield RMP ELM suppression. At q 95 ∼ 5.1, a bifurcation to a grassy-like ELM regime occurred, while large type-I ELMs persisted. These results demonstrate progress in experimental access to n = 1 RMP ELM suppression in DIII-D, motivating further study for robust access. This work also highlights the potential role of 3D edge stability as well as rational surface alignment in RMP ELM suppression access, which has important implications for the use of low-n RMPs in future reactor-scale devices.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Tough Errors Are no Match (TEAM): Optimizing the quantum compiler for noise resilience

This report summarizes Unitary Fund’s contributions to the Department of Energy’s TEAM project (DE-SC0020266) under Thrust 2: Quantum Programming and Compilation. The central outcomes of this work have been the development of Mitiq, an open-source Python toolkit for applying quantum error mitigation (QEM) techniques to noisy quantum programs, and the invention, benchmarking and theoretical investigation of novel QEM techniques. Additional outcomes include the development of other open source software packages for the usage, simulation and control of quantum computers.

97 MATHEMATICS AND COMPUTING↗

HAPPA: A Modular Platform for HPC Application Resilience Analysis with LLMs Embedded

High-performance computing (HPC) systems are increasingly vulnerable to soft errors, which pose significant challenges in maintaining computational accuracy and reliability. Predicting the resilience of HPC applications to these errors is crucial for robust code protection and detailed resilience analysis. In this study, we present HAppA, a modular platform designed for HPC Application Resilience Analysis. Embedding Large Language Models (LLMs), HAppA addresses understanding the context information of long code sequences typical in HPC applications. HAppA implements a novel code representation module that chunks the code into fixed-size segments and aggregates the embeddings of these segments. Three aggregation methods have been explored: MeanPooling, MaxPooling, and LSTM-based techniques. We built a DAtaset for REsilience analysis using Fault Injection (FI), named DARE. Using our DARE dataset, HAppA is trained for regression prediction tasks. Our evaluation results demonstrate the predictive accuracy of HAppA compared to other models, particularly noting that the LSTM-based aggregation method -- HAppA-LSTM -- achieves a mean squared error (MSE) of 0.078 for SDC prediction, surpassing the existing state-of-the-art PARIS model, which recorded an MSE of 0.1172. Additionally, HAppA with the KeyBERT model extracts a list of keywords representing the source code. A comprehensive importance analysis of these keywords further elucidates the code patterns contributing to the error rate. These findings highlight the effectiveness of HAppA in analyzing the resilience of HPC applications and establish a new benchmark for predictive accuracy in resilience.

Jiang, Hailong [Kent State University]↗

Modeling Distributed Situation Awareness in Resilience-Based Design of Complex Engineered Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling↗

Modeling Distributed Situation Awareness in Resilience-based Design of Complex Systems

Human operators play a major role in the resilience of complex systems–while human error is one of the biggest contributors to hazardous events, operators additionally play a critical role in mitigating hazardous events. A key factor underlying this operator resilience is situation awareness–the ability of operators to understand their environment and each other to achieve desired system functions. In contrast to situation awareness-related accident models in the literature, which are largely conceptual in nature, this work proposes the use of a dynamic simulation framework to concretely model both the effects of situation awareness-related human errors and situation awareness-related hazard-mitigating properties using the distributed situation awareness theory. This work then presents specialized model constructs to enable agents’ individual perceptions of the system state and transactions with other agents (and thus distributed situation awareness) to be represented in simulation. To demonstrate this framework, it is then adapted to an aircraft taxiway case study, where it is used to model aircraft conflicts due to lack of vision and poor communications from the air traffic controller. This demonstration shows the potential of using simulation models to rigorously understand situation awareness-related human errors and thus inform the design of resilience.

Resilience Modeling↗

Designing Resilience for Advanced Energy Systems

Advancements in energy technologies are making the grid more powerful, more efficient, and cleaner. As these promising innovations flourish across our communities, it is essential that our nation's infrastructure also becomes more resilient. Natural disasters, cyberattacks, user error, and a changing climate all present risks that could have devastating consequences. Individual emergencies are unique, but holistic planning and proactive measures can harden the grid and help anticipate and respond to any number of threats. The National Renewable Energy Laboratory (NREL) is at the forefront of this transition, establishing a vision for resilient energy systems today, and in the future.

energy disruption↗

Long-time simulations for fixed input states on quantum hardware

Publicly accessible quantum computers open the exciting possibility of experimental dynamical quantum simulations. While rapidly improving, current devices have short coherence times, restricting the viable circuit depth. Despite these limitations, we demonstrate long-time, high fidelity simulations on current hardware. Specifically, we simulate an XY-model spin chain on Rigetti and IBM quantum computers, maintaining a fidelity over 0.9 for 150 times longer than is possible using the iterated Trotter method. Our simulations use an algorithm we call fixed state Variational Fast Forwarding (fsVFF). Recent work has shown an approximate diagonalization of a short time evolution unitary allows a fixed-depth simulation. fsVFF substantially reduces the required resources by only diagonalizing the energy subspace spanned by the initial state, rather than over the total Hilbert space. We further demonstrate the viability of fsVFF through large numerical simulations, and provide an analysis of the noise resilience and scaling of simulation errors.

97 MATHEMATICS AND COMPUTING↗

Model Agnostic Bayesian Framework for Online Anomaly/Event Detection in PMU Data

Phasor measurement units (PMU) are integral to the modernization and automation plan of the electric power industry. A PMU data signature contains system-level events (e.g., faults, generation/load change, etc.) and any measurement/device-related errors. Therefore, the reliable and resilient operation of power systems is equivalent to the quality of the PMU data and the situation awareness provided by its data signature. Despite recent progress, current state-of-the-art methods are not fool-proof and have certain limitations tracing an error/abnormality to sensor sub-components and grid systems. This is because of technical challenges imposed by the scarcity of the labeled information, loss of data quality, and non-stationarity of data. In this paper, we consider the online PMU data stream as an output of a stochastic process and pose the anomaly/event detection as a changepoint detection problem dealing with detecting parameter changes in the underlying stochastic processes. The proposed model-agnostic framework relies on: (a) feature extraction utilizing the minimum volume enclosing ellipsoids (MVEE) method from raw PMU observations and (b) a Bayesian framework of changepoint detection. The validity of the proposed methodology is discussed through numerical experiments on real-world utility-scale PMU data.

Hossain, Ramij Raja↗