Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Resilience–runtime tradeoff relations for quantum algorithms

Abstract A leading approach to algorithm design aims to minimize the number of operations in an algorithm’s compilation. One intuitively expects that reducing the number of operations may decrease the chance of errors. This paradigm is particularly prevalent in quantum computing, where gates are hard to implement and noise rapidly decreases a quantum computer’s potential to outperform classical computers. Here, we find that minimizing the number of operations in a quantum algorithm can be counterproductive, leading to a noise sensitivity that induces errors when running the algorithm in non-ideal conditions. To show this, we develop a framework to characterize the resilience of an algorithm to perturbative noises (including coherent errors, dephasing, and depolarizing noise). Some compilations of an algorithm can be resilient against certain noise sources while being unstable against other noises. We condense these results into a tradeoff relation between an algorithm’s number of operations and its noise resilience. We also show how this framework can be leveraged to identify compilations of an algorithm that are better suited to withstand certain noises.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

EFCOG Best Practice HPI for Knowledge Workers ISM-HPI-22-02

This document is a collection of these best practices as determined by team members. This best practice attempts to: Realize opportunities to break the myth where people believe that HPI does not apply to them as they perform no physical work. Recommend options to create an environment that promotes intellectual collaboration and trust, enabling candor and vulnerability. Explain how errors manifest differently from the same human fallibility. Knowledge workers (KW) have errors that take different perspectives to find and mitigate the unique manifestation of these conditions. Help KW identify the critical steps (or risk important steps) in their processes. Reduce risk/consequence from KW errors (limit latent errors as well as finding latent conditions), building resiliency into KW tasks. Mitigation strategies may be different.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Adaptive Cyber-Physical Resilience for Building Control Systems

The main goal of the project is to develop an AI-based process layer cybersecurity suite for detection, isolation and mitigation of cyber-attack effects on operation of building energy management systems (BEMS). The following constituent key technologies were developed under the program towards fulfilling the program objectives: (1) developed a high fidelity BEMS testbed for generation of training data and validation of developed technologies; (2) developed a physics informed ML based attack detection and localization module (ADL) capable of detecting high impact stealthy attacks (HISA - attacks causing 30% energy utilization but no immediate visible impact otherwise) with 98% accuracy; (3) developed a methodology to determine ’representative days’ to limit the data required for training; (4) developed a virtual sensing system that can reconstruct affected sensors with 10% error for the same HISA set; (5) developed a resilient model predictive control system that can continue operation of the BEMS without jeopardizing stability for the HISA set; and (6) integrated and deployed all the constituent modules and demonstrated the efficacy of the technology in real-time in a hardware in loop simulation.

42 ENGINEERING↗

Occupant Preference-Aware Load Scheduling for Resilient Communities

The load scheduling of resilient communities in the islanded mode is subject to many uncertainties such as weather forecast errors and occupant behavior stochasticity. To date, it remains unclear how occupant preferences affect the effectiveness of the load scheduling of resilient communities. This paper proposes an occupant thermal preference-aware load scheduler for resilient communities operating in the islanded mode. First, key resilience indicators are selected to quantify its impacts on the load scheduling of a resilient community. A deterministic model predictive control-based load scheduling framework is adopted as the baseline. Then, a chance-constrained controller is proposed to address the occupant-induced uncertainty in room temperature setpoints. Finally, the chance-constrained controller is compared with the deterministic controller on a virtual community testbed based on a real-world net-zero energy community in Florida, U.S. Results have shown that the proposed chance-constrained controller performs better in terms of serving occupants’ thermal preference and the required battery sizes compared to the deterministic controller with the presence of the assumed stochastic occupant behavior. This work indicates that it is necessary to consider the stochasticity of the occupant behavior when designing optimal load schedulers for resilient communities.

Wang, Jing↗

Quantum error mitigation by hidden inverses protocol in superconducting quantum devices

We present a method to improve the convergence of variational algorithms based on hidden inverses (HIs) to mitigate coherent errors. In the context of error mitigation, this means replacing the hardware implementation of certain Hermitian gates with their inverses. Doing so results in noise cancellation and a more resilient quantum circuit. This approach improves performance in a variety of two-qubit error models where the noise operator also inverts with the gate inversion. We apply the mitigation scheme on superconducting quantum processors running the variational quantum eigensolver (VQE) algorithm to find the H 2 ground-state energy. When implemented on superconducting hardware we find that the mitigation scheme effectively reduces the energy fluctuations in the parameter learning path in VQE, reducing the number of iterations for a converged value. We also provide a detailed numerical simulation of VQE performance under different noise models and explore how HIs and randomized compiling affect the underlying loss landscape of the learning problem. These simulations help explain our experimental hardware outcomes, helping to connect lower-level gate performance to application-specific behavior in contrast to metrics like fidelity which often do not provide an intuitive insight into observed high level performance.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Noise-Resilient and Reduced Depth Approximate Adders for NISQ Quantum Computing

The "Noisy intermediate-scale quantum" NISQ machine era primarily focuses on mitigating noise, controlling errors, and executing high-fidelity operations, hence requiring shallow circuit depth and noise robustness. Approximate computing is a novel computing paradigm that produces imprecise results by relaxing the need for fully precise output for error-tolerant applications including multimedia, data mining, and image processing. We investigate how approximate computing can improve the noise resilience of quantum adder circuits in NISQ quantum computing. We propose five designs of approximate quantum adders to reduce depth while making them noise-resilient, in which three designs are with carryout, while two are without carryout. We have used novel design approaches that include approximating the Sum only from the inputs (pass-through designs) and having zero depth, as they need no quantum gates. The second design style uses a single CNOT gate to approximate the SUM with a constant depth of O(1). We performed our experimentation on IBM Qiskit on noise models including thermal, depolarizing, amplitude damping, phase damping, and bitflip: (i) Compared to exact quantum ripple carry adder without carryout the proposed approximate adders without carryout have improved fidelity ranging from 8.34% to 219.22%, and (ii) Compared to exact quantum ripple carry adder with carryout the proposed approximate adders with carryout have improved fidelity ranging from 8.23% to 371%. Further, the proposed approximate quantum adders are evaluated in terms of various error metrics.

Gaur, Bhaskar↗

A Deep Learning Approach for In-Network Synchrophasor Missing Data Recovery Using Programmable Network Switches

Phasor measurement unit (PMU) networks deliver accurate and timely measurements, which is essential for managing today’s electric power systems. To ensure data quality and enhance the cyber-resilience of PMU networks against malicious attacks and data errors, this study presents an online PMU missing data recovery scheme by leveraging P4 programmable switches. The data plane incorporates a customized PMU protocol parser that abstracts the necessary payload data for recovery. Recovery processes are executed in the control plane using a pre-trained machine learning model. Both traditional and advanced ML models, such as transformer and TimeGPT, are explicitly employed for data prediction. This approach ensures rapid and precise data recovery. Performance evaluations focus on recovery speed and accuracy, using a real dataset from a campus microgrid. With 20% missing PMU data, the mean absolute percentage error for voltage magnitude is 0.0384%, and the phase angle error discrepancy is approximately 0.4064%.

Phasor Measurement Unit, Machine Learning, Program↗

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

Recommendations for Minimum Required Error Codes for Electric Vehicle Charging Infrastructure

OCPP protocol manages the interaction between the EVSE and its respective back-end communication network. It plays a pivotal role in both error reporting and troubleshooting, carried out primarily through the CSMS. OCPP defines both standard error codes and a flexible framework for creating and communicating custom error codes. The OCPI protocol orchestrates the communication between different backhaul communication networks, incorporating the exchange of error codes. These error codes are instrumental in pinpointing and rectifying issues that can surface prior to, during, or after charging operations, fortifying the reliability and resilience of the EV charging infrastructure. The flexibility offered by the OCPP and OCPI frameworks through the introduction of custom error codes also creates its own set of challenges. While the integration of custom error codes allows for enhanced granularity, it also introduces inconsistencies and fragmentation within the overarching diagnostic reporting system. To address the challenges with custom error codes this report proposes a set of Minimum Required Error Codes (MRECs) for streamlined error reporting, interpretability, and diagnostics. Recommendations in this report are based on independent analysis of custom error codes from multiple stakeholders within the EV charging ecosystem. For better error resolution, this report also assigns one or more entities responsible for the resolution of every mentioned error code. Finally, a functional classification for each mentioned error code is also identified to describe the nature of the error. In summary, the purpose of this document is to simplify the troubleshooting process and increase charging reliability for all EV users. This report serves as a recommendation for industry stakeholders, encouraging a unified methodology to define, transmit, and interpret common error codes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A survey on checkpointing strategies: Should we always checkpoint à la Young/Daly?

The Young/Daly formula provides an approximation of the optimal checkpointing period for a parallel application executing on a supercomputing platform. It was originally designed to handle fail-stop errors for preemptible tightly-coupled applications, but has been extended to other application and resilience frameworks. Here, we provide some background and survey various scenarios to assess the usefulness and limitations of the formula, both for preemptible applications and workflow applications represented as a graph of tasks. We also discuss scenarios with uncertainties, and extend the study to silent errors. We exhibit cases where the optimal period is of a different order than that dictated by the Young/Daly formula, and finally we explain how checkpointing can be further combined with replication.

97 MATHEMATICS AND COMPUTING↗

Quantum algorithms for geologic fracture networks

Abstract Solving large systems of equations is a challenge for modeling natural phenomena, such as simulating subsurface flow. To avoid systems that are intractable on current computers, it is often necessary to neglect information at small scales, an approach known as coarse-graining. For many practical applications, such as flow in porous, homogenous materials, coarse-graining offers a sufficiently-accurate approximation of the solution. Unfortunately, fractured systems cannot be accurately coarse-grained, as critical network topology exists at the smallest scales, including topology that can push the network across a percolation threshold. Therefore, new techniques are necessary to accurately model important fracture systems. Quantum algorithms for solving linear systems offer a theoretically-exponential improvement over their classical counterparts, and in this work we introduce two quantum algorithms for fractured flow. The first algorithm, designed for future quantum computers which operate without error, has enormous potential, but we demonstrate that current hardware is too noisy for adequate performance. The second algorithm, designed to be noise resilient, already performs well for problems of small to medium size (order 10–1000 nodes), which we demonstrate experimentally and explain theoretically. We expect further improvements by leveraging quantum error mitigation and preconditioning.

58 GEOSCIENCES↗

BioSecure Digital Twin: Manufacturing Innovation and Cybersecurity Resilience

U.S. national security, prosperity, economy, and well-being require secure, flexible, and resilient Biopharmaceutical Manufacturing. The COVID-19 pandemic reaffirmed that the biomedical production value-chain is vulnerable to disruption and has been under attack from sophisticated nation-state adversaries. Current cyber defenses are inadequate, and the integrity of critical production systems and processes are inherently vulnerable to cyber-attacks, human error, and supply chain disruptions. The following chapter explores how a BioSecure Digital Twin will improve U.S. manufacturing resilience and preparedness to respond to these hazards by significantly improving monitoring, integrity, security, and agility of our manufacturing infrastructure and systems. The BioSecure Digital Twin combines a scalable manufacturing framework with a robust platform for monitoring and control to increase U.S. biopharma manufacturing resilience. Then, the chapter discusses some of the inherent vulnerabilities and challenges at the nexus of health and advanced manufacturing. Next, the chapter highlights that as the Pandemic evolves, we need agility and resilience to overcome significant obstacles. This section highlights an innovative application of Cyber Informed Engineering to developing and deploying a BioSecure Digital Twin to improve the resilience and security of the biopharma industrial supply chain and production processes. Finally, the chapter concludes with a process framework to complement the Digital Twin platform, called the Biopharma (Observe, Orient, Decide, Act) OODA Loop Framework (BOLF), a four-step approach to decision-making outputs from the Digital Twin. The BOLF will help end users leverage twin technology by distilling the available information, focusing the data on context, and rapidly making the best decision while remaining cognizant of changes that can be made as more data becomes available.

99 GENERAL AND MISCELLANEOUS↗

Assurance by Design for Cyber Physical Data-Driven Systems

Currently, Cyber Physical Data-Driven Systems (CPDDS) employ machine learning for the classification, data fusion, and control of our nation’s infrastructure, such as the power grid, transportation networks (e.g., fuel distribution, air traffic control), and DoD long-duration collaborative autonomous platforms including unmanned underwater, ground, surface, space, and aerial systems. Many CPDDSs are system-of-systems that should be designed to communicate over disadvantaged networks. It is important to assure that the CPDDSs are resilient against physical and cyber threats by design. Additionally, their design should tolerate misclassification errors resulting from natural and/or adversarial distribution shifts within their data driven components. The all-domain nature of the problem of assuring the design of CPDDSs requires a multi-disciplinary perspective as outlined in this chapter.

Chikkagoudar, Satish↗

Corrigendum to “Cool Rooms for Indoor Heat Resilience: Evaluating Affordable Cooling Strategies in Heat-Stressed California Homes” [Building and Environment 287 (2026) 113877]

The authors regret an error in the acknowledgments section regarding the U.S. Department of Energy Solar Energy Technologies Office award number. The previously listed grant number, 2597–1625, was incorrect. The corrected acknowledgment should read: “This work was supported by the Assistant Secretary for Energy Efficiency and Renewable Energy, Office of Building Technologies of the United States Department of Energy (DOE), under Contract No. DE-AC02–05CH11231. This material is based upon work supported by the U.S. Department of Energy's Office of Energy Efficiency and Renewable Energy (EERE) under the Solar Energy Technologies Office Award Number DE-EE00040384. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the Department of Energy.” The authors would like to apologise for any inconvenience caused.

Malik, Jeetika↗

SPARC 3D Field Physics and Support of the Non-Axisymmetric Coil Assessment

Commonwealth Fusion Systems (CFS) is deploying SPARC, a net-energy tokamak, by 2025. A notable challenge facing all tokamak approaches to fusion energy production is maintaining a stable plasma and thereby steady energy production. “Disruptions” to plasma operation are often encountered in devices when they operate at high normalized pressure, high normalized density, or high normalized current. SPARC is unique in that it is expected to demonstrate Q>=2 in a plasma with low normalized pressure and density, but with a more modest current limit buffer sufficient to avoid disruptions. Due to both the high normalized current and high magnetic field of the SPARC design, it is expected to decrease its relative resilience to instabilities driven by non-axisymmetry in the externally applied magnetic field, termed ‘error fields’. To prevent error field driven instabilities, a primary line of defense is strict engineering tolerances in magnet fabrication and installation. A secondary line of defense is a purpose-built set of magnetic coils to correct these errors, termed error field correction coils (EFCCs). The overarching goals of this INFUSE project were to evaluate the effectiveness of the EFCCs in the SPARC design and provide guidance as to how to improve this effectiveness.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Robust Scheduling of Networked Microgrids for Economics and Resilience Improvement

The benefits of networked microgrids in terms of economics and resilience are investigated and validated in this work. Considering the stochastic unintentional islanding conditions and conventional forecast errors of both renewable generation and loads, a two-stage adaptive robust optimization is proposed to minimize the total operating cost of networked microgrids in the worst scenario of the modeled uncertainties. By coordinating the dispatch of distributed energy resources (DERs) and responsive demand among networked microgrids, the total operating cost is minimized, which includes the start-up and shut-down cost of distributed generators (DGs), the operation and maintenance (O&M) cost of DGs, the cost of buying/selling power from/to the utility grid, the degradation cost of energy storage systems (ESSs), and the cost associated with load shedding. The proposed optimization is solved with the column and constraint generation (C&CG) algorithm. The results of case studies demonstrate the advantages of networked microgrids over independent microgrids in terms of reducing total operating cost and improving the resilience of power supply.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Robust Resilient Signal Reconstruction under Adversarial Attacks

We consider the problem of signal reconstruction for a system under sparse signal corruption by a malicious agent. The reconstruction problem follows the standard error coding problem that has been studied extensively in the literature. We include a new challenge of robust estimation of the attack support. The problem is then cast as a constrained optimization problem merging promising techniques in the area of deep learning and estimation theory. A pruning algorithm is developed to reduce the "false positive" uncertainty of data-driven attack localization results, thereby improving the probability of correct signal reconstruction. Sufficient conditions for the correct reconstruction and the associated reconstruction error bounds are obtained for both exact and inexact attack support estimation. Moreover, a simulation of a water distribution system is presented to validate the proposed techniques.

Robust, Signal reconstruction, Resilient estimator↗

Uncertainty-Guided Prediction Horizon of Phase-Resolved Ocean Wave Forecasting Under Data Sparsity: Experimental and Numerical Evaluation

Accurate short-term wave forecasting is critical for the safe and efficient operation of marine structures that rely on real-time, phase-resolved ocean wave information for control and monitoring purposes (e.g., digital twins). These systems often depend on environmental sensors (e.g., waverider buoys, wave-sensing LIDAR). Challenges arise when upstream sensor data are missing, sparse, or phase-shifted due to drift. This study investigates the performance of two machine learning models, time-series dense encoder (TiDE) and long short-term memory (LSTM), for forecasting phase-resolved ocean surface elevations under varying degrees of data degradation. We introduce the τ-trimming algorithm, which adapts the prediction horizon based on uncertainty thresholds derived from historical forecasts. Numerical wave tank (NWT) and wave basin experiments are used to benchmark model performance under short- and long-term data masking, spatially coarse sensor grids, and upstream phase shifts. Results show under a 50% probability of upstream data loss, the τ-trimmed TiDE model achieves a 46% reduction in error at the most upstream target, compared to 22% for LSTM. Furthermore, phase misalignment in upstream data introduces a near-linear increase in forecast error. Under moderate model settings, a ±3 s misalignment increases the mean absolute error by approximately 0.5 m, while the same error is accumulated at ±4 s using the more conservative approach. These findings inform the design of resilient, uncertainty-aware wave forecasting systems suited for realistic offshore sensing environments.

42 ENGINEERING↗