Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

On the Use of Resilience Models as Digital Twins for Operational Support and In time Decision Making

Human error is a major contributor to accidents and performance losses in complex engineered systems. If one examines these human error caused failures further, a specific cause, the lack of situation awareness, has dominated as a major cause of human errors that instigate latent or catastrophic failures in complex systems. Studies of aviation accidents involving major air carriers revealed that situation awareness was the root cause of around 90% of accidents involving pilot error. Another study explored offshore drilling accidents involving human error and found that 40% of accidents were directly attributed to the loss of situation awareness. Studies of human errors in other domains such as nuclear power, air traffic control, process industry, and advanced driving show that loss of SA was a root cause in a majority of the events. Situation awareness-related failures are not only common but also costly and fatal (e.g., Bhopal Gas Leak, Air France 447 Flight Crash). Thus, the concept of situation awareness has emerged as an important construct in human factors, resulting in numerous models and measurement methods to aid in promoting appropriate levels of situation awareness.

Lukman Irshad↗

Recommendations for Minimum Required Error Codes for Electric Vehicle Charging Infrastructure

OCPP protocol manages the interaction between the EVSE and its respective back-end communication network. It plays a pivotal role in both error reporting and troubleshooting, carried out primarily through the CSMS. OCPP defines both standard error codes and a flexible framework for creating and communicating custom error codes. The OCPI protocol orchestrates the communication between different backhaul communication networks, incorporating the exchange of error codes. These error codes are instrumental in pinpointing and rectifying issues that can surface prior to, during, or after charging operations, fortifying the reliability and resilience of the EV charging infrastructure. The flexibility offered by the OCPP and OCPI frameworks through the introduction of custom error codes also creates its own set of challenges. While the integration of custom error codes allows for enhanced granularity, it also introduces inconsistencies and fragmentation within the overarching diagnostic reporting system. To address the challenges with custom error codes this report proposes a set of Minimum Required Error Codes (MRECs) for streamlined error reporting, interpretability, and diagnostics. Recommendations in this report are based on independent analysis of custom error codes from multiple stakeholders within the EV charging ecosystem. For better error resolution, this report also assigns one or more entities responsible for the resolution of every mentioned error code. Finally, a functional classification for each mentioned error code is also identified to describe the nature of the error. In summary, the purpose of this document is to simplify the troubleshooting process and increase charging reliability for all EV users. This report serves as a recommendation for industry stakeholders, encouraging a unified methodology to define, transmit, and interpret common error codes.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

A survey on checkpointing strategies: Should we always checkpoint à la Young/Daly?

The Young/Daly formula provides an approximation of the optimal checkpointing period for a parallel application executing on a supercomputing platform. It was originally designed to handle fail-stop errors for preemptible tightly-coupled applications, but has been extended to other application and resilience frameworks. Here, we provide some background and survey various scenarios to assess the usefulness and limitations of the formula, both for preemptible applications and workflow applications represented as a graph of tasks. We also discuss scenarios with uncertainties, and extend the study to silent errors. We exhibit cases where the optimal period is of a different order than that dictated by the Young/Daly formula, and finally we explain how checkpointing can be further combined with replication.

97 MATHEMATICS AND COMPUTING↗

Why Didn't They Just Follow the Procedure?

It is often said that to err is human and that procedures are in place to prevent people from making mistakes. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. And while it's true that good procedures can help avoid error, it's also the case that procedures have their own limitations. This talk highlights the limits of procedures and the resilience people bring to operations. It discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure, and it proposes an approach to the design of procedures that supports the human operator.

operations↗

Resilience Modeling in Complex Engineered Systems with Human-Machine Interactions

In recent times, there has been a growing interest in resilience-based design. Resilience-based design operates on the concept that failures and unexpected events will happen, and when they occur, complex engineered systems should be able to operate within acceptable bounds and recover reasonably. Humans can contribute to the resilience of a system by quickly detecting unforeseen events and taking corrective measures. To this effect, researchers have proposed guidelines and design approaches that can help promote human-system resilience. However, there is no early design stage tool to validate if a system is indeed resilient after applying these guidelines and design methods. In this research, we integrate the Human Error and Functional Failure Reasoning (HEFFR) framework into the fmdtools toolkit to enable designers to model the combined (machine, human, and joint) failures, including their propagation and dynamic effects, during early design stages. This integrated tool also allows designers to model the effects of performance shaping factors, team dynamics, and human-machine interactions in systems of systems. A demonstrative example of a remotely operated rover is explored to demonstrate how this approach can be applied to understand resilience in complex engineered systems with human interactions.

Lukman Irshad↗

Quantum algorithms for geologic fracture networks

Abstract Solving large systems of equations is a challenge for modeling natural phenomena, such as simulating subsurface flow. To avoid systems that are intractable on current computers, it is often necessary to neglect information at small scales, an approach known as coarse-graining. For many practical applications, such as flow in porous, homogenous materials, coarse-graining offers a sufficiently-accurate approximation of the solution. Unfortunately, fractured systems cannot be accurately coarse-grained, as critical network topology exists at the smallest scales, including topology that can push the network across a percolation threshold. Therefore, new techniques are necessary to accurately model important fracture systems. Quantum algorithms for solving linear systems offer a theoretically-exponential improvement over their classical counterparts, and in this work we introduce two quantum algorithms for fractured flow. The first algorithm, designed for future quantum computers which operate without error, has enormous potential, but we demonstrate that current hardware is too noisy for adequate performance. The second algorithm, designed to be noise resilient, already performs well for problems of small to medium size (order 10–1000 nodes), which we demonstrate experimentally and explain theoretically. We expect further improvements by leveraging quantum error mitigation and preconditioning.

58 GEOSCIENCES↗

BioSecure Digital Twin: Manufacturing Innovation and Cybersecurity Resilience

U.S. national security, prosperity, economy, and well-being require secure, flexible, and resilient Biopharmaceutical Manufacturing. The COVID-19 pandemic reaffirmed that the biomedical production value-chain is vulnerable to disruption and has been under attack from sophisticated nation-state adversaries. Current cyber defenses are inadequate, and the integrity of critical production systems and processes are inherently vulnerable to cyber-attacks, human error, and supply chain disruptions. The following chapter explores how a BioSecure Digital Twin will improve U.S. manufacturing resilience and preparedness to respond to these hazards by significantly improving monitoring, integrity, security, and agility of our manufacturing infrastructure and systems. The BioSecure Digital Twin combines a scalable manufacturing framework with a robust platform for monitoring and control to increase U.S. biopharma manufacturing resilience. Then, the chapter discusses some of the inherent vulnerabilities and challenges at the nexus of health and advanced manufacturing. Next, the chapter highlights that as the Pandemic evolves, we need agility and resilience to overcome significant obstacles. This section highlights an innovative application of Cyber Informed Engineering to developing and deploying a BioSecure Digital Twin to improve the resilience and security of the biopharma industrial supply chain and production processes. Finally, the chapter concludes with a process framework to complement the Digital Twin platform, called the Biopharma (Observe, Orient, Decide, Act) OODA Loop Framework (BOLF), a four-step approach to decision-making outputs from the Digital Twin. The BOLF will help end users leverage twin technology by distilling the available information, focusing the data on context, and rapidly making the best decision while remaining cognizant of changes that can be made as more data becomes available.

99 GENERAL AND MISCELLANEOUS↗

Assurance by Design for Cyber Physical Data-Driven Systems

Currently, Cyber Physical Data-Driven Systems (CPDDS) employ machine learning for the classification, data fusion, and control of our nation’s infrastructure, such as the power grid, transportation networks (e.g., fuel distribution, air traffic control), and DoD long-duration collaborative autonomous platforms including unmanned underwater, ground, surface, space, and aerial systems. Many CPDDSs are system-of-systems that should be designed to communicate over disadvantaged networks. It is important to assure that the CPDDSs are resilient against physical and cyber threats by design. Additionally, their design should tolerate misclassification errors resulting from natural and/or adversarial distribution shifts within their data driven components. The all-domain nature of the problem of assuring the design of CPDDSs requires a multi-disciplinary perspective as outlined in this chapter.

Chikkagoudar, Satish↗

Corrigendum to “Cool Rooms for Indoor Heat Resilience: Evaluating Affordable Cooling Strategies in Heat-Stressed California Homes” [Building and Environment 287 (2026) 113877]

The authors regret an error in the acknowledgments section regarding the U.S. Department of Energy Solar Energy Technologies Office award number. The previously listed grant number, 2597–1625, was incorrect. The corrected acknowledgment should read: “This work was supported by the Assistant Secretary for Energy Efficiency and Renewable Energy, Office of Building Technologies of the United States Department of Energy (DOE), under Contract No. DE-AC02–05CH11231. This material is based upon work supported by the U.S. Department of Energy's Office of Energy Efficiency and Renewable Energy (EERE) under the Solar Energy Technologies Office Award Number DE-EE00040384. Any opinions, findings, and conclusions or recommendations expressed in this material are those of the author(s) and do not necessarily reflect the views of the Department of Energy.” The authors would like to apologise for any inconvenience caused.

Malik, Jeetika↗

SPARC 3D Field Physics and Support of the Non-Axisymmetric Coil Assessment

Commonwealth Fusion Systems (CFS) is deploying SPARC, a net-energy tokamak, by 2025. A notable challenge facing all tokamak approaches to fusion energy production is maintaining a stable plasma and thereby steady energy production. “Disruptions” to plasma operation are often encountered in devices when they operate at high normalized pressure, high normalized density, or high normalized current. SPARC is unique in that it is expected to demonstrate Q>=2 in a plasma with low normalized pressure and density, but with a more modest current limit buffer sufficient to avoid disruptions. Due to both the high normalized current and high magnetic field of the SPARC design, it is expected to decrease its relative resilience to instabilities driven by non-axisymmetry in the externally applied magnetic field, termed ‘error fields’. To prevent error field driven instabilities, a primary line of defense is strict engineering tolerances in magnet fabrication and installation. A secondary line of defense is a purpose-built set of magnetic coils to correct these errors, termed error field correction coils (EFCCs). The overarching goals of this INFUSE project were to evaluate the effectiveness of the EFCCs in the SPARC design and provide guidance as to how to improve this effectiveness.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Robust Scheduling of Networked Microgrids for Economics and Resilience Improvement

The benefits of networked microgrids in terms of economics and resilience are investigated and validated in this work. Considering the stochastic unintentional islanding conditions and conventional forecast errors of both renewable generation and loads, a two-stage adaptive robust optimization is proposed to minimize the total operating cost of networked microgrids in the worst scenario of the modeled uncertainties. By coordinating the dispatch of distributed energy resources (DERs) and responsive demand among networked microgrids, the total operating cost is minimized, which includes the start-up and shut-down cost of distributed generators (DGs), the operation and maintenance (O&M) cost of DGs, the cost of buying/selling power from/to the utility grid, the degradation cost of energy storage systems (ESSs), and the cost associated with load shedding. The proposed optimization is solved with the column and constraint generation (C&CG) algorithm. The results of case studies demonstrate the advantages of networked microgrids over independent microgrids in terms of reducing total operating cost and improving the resilience of power supply.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

What Can We Learn From Resilient Pilot Behaviors? The Case of Energy Management While Flying a Star

Recently, there has been increased interest in documenting flightcrew behaviors that contribute to safe operations. Instead of only capturing errors, new efforts are attempting to understand how pilots manage complexity and variability in the operational environment to ensure a safe mission. This approach highlights pilot responses to events and conditions that fall outside typical TEM threats; e.g., revised ATC clearances. This approach presents a two-sided coin: characterize flightcrew resilience /or/ generate insights regarding complexity in the operational environment that is not adequately managed by current flight deck interface designs, procedures, and training. To capture operational complexity, we have been analyzing flight path management tied to flying an RNAV STAR. Because ATC often requests revisions—e.g., descend late—and because RNAV STARs may not align with airplane performance limits, flightcrews need to monitor, anticipate threats to RNAV STAR compliance, and devise ways to accommodate unexpected challenges. In this paper, we identify general strategies that can support response adaptation and explore methods to facilitate training these strategies.

flight operations↗

Robust Resilient Signal Reconstruction under Adversarial Attacks

We consider the problem of signal reconstruction for a system under sparse signal corruption by a malicious agent. The reconstruction problem follows the standard error coding problem that has been studied extensively in the literature. We include a new challenge of robust estimation of the attack support. The problem is then cast as a constrained optimization problem merging promising techniques in the area of deep learning and estimation theory. A pruning algorithm is developed to reduce the "false positive" uncertainty of data-driven attack localization results, thereby improving the probability of correct signal reconstruction. Sufficient conditions for the correct reconstruction and the associated reconstruction error bounds are obtained for both exact and inexact attack support estimation. Moreover, a simulation of a water distribution system is presented to validate the proposed techniques.

Robust, Signal reconstruction, Resilient estimator↗

Uncertainty-Guided Prediction Horizon of Phase-Resolved Ocean Wave Forecasting Under Data Sparsity: Experimental and Numerical Evaluation

Accurate short-term wave forecasting is critical for the safe and efficient operation of marine structures that rely on real-time, phase-resolved ocean wave information for control and monitoring purposes (e.g., digital twins). These systems often depend on environmental sensors (e.g., waverider buoys, wave-sensing LIDAR). Challenges arise when upstream sensor data are missing, sparse, or phase-shifted due to drift. This study investigates the performance of two machine learning models, time-series dense encoder (TiDE) and long short-term memory (LSTM), for forecasting phase-resolved ocean surface elevations under varying degrees of data degradation. We introduce the τ-trimming algorithm, which adapts the prediction horizon based on uncertainty thresholds derived from historical forecasts. Numerical wave tank (NWT) and wave basin experiments are used to benchmark model performance under short- and long-term data masking, spatially coarse sensor grids, and upstream phase shifts. Results show under a 50% probability of upstream data loss, the τ-trimmed TiDE model achieves a 46% reduction in error at the most upstream target, compared to 22% for LSTM. Furthermore, phase misalignment in upstream data introduces a near-linear increase in forecast error. Under moderate model settings, a ±3 s misalignment increases the mean absolute error by approximately 0.5 m, while the same error is accumulated at ±4 s using the more conservative approach. These findings inform the design of resilient, uncertainty-aware wave forecasting systems suited for realistic offshore sensing environments.

42 ENGINEERING↗

Resilience Measurement Framework For Post-deployment Artificial Intelligence (ai) Integrated Systems

Resilience is largely defined as the ability to adapt or recover from adverse conditions, stresses, attacks, or compromises on systems that use or are enabled by digital resources. In Artificial Intelligence Management and Research for Advanced Networked Testbed Hub (AMARANTH), resilience is measured in the amount of time it took from the beginning of a testing period for the model to reach predictions outside of the original 95% confidence interval or using the Kullback-Leibler (KL) divergence theorem, the Population Stability Index (PSI), and traditional methods such as root mean squared error (RMSE) threshold. Artificial Intelligence (AI) model drift is of significant concern when deploying AI-integrated systems into critical and/or secure environments. Drift can impact resilience of the AI-integrated system post-deployment and requires consistent maintenance and upkeep to ensure the model is accurate and precise. To quantify model drift and predict the point when a model's drift becomes unacceptable, we describe using Kullback-Leibler (KL) divergence, Population Stability Index (PSI) and/or confidence interval width estimations to determine the point of failure and time to failure of a model post-deployment. Through simple code functions, the KL-divergence, PSI, confidence interval, and root mean squared (RMSE) point of failures can be used to derive when a model needs to be maintained as well as the impact of adversarial action through statistical means.

Yockey, Patience [Idaho National Laboratory (INL),↗

Thinking outside the box: The human role in increasingly automated aviation systems

Rapid advances in artificial intelligence are enabling automated systems to operate in an increasingly autonomous manner in domains that previously required the involvement of human operators. Examples are rail transport systems, self-driving cars, and warehouse delivery systems. From time to time, such automation encounters operational conditions that fall outside a “competency box” within which the system has been designed to operate. Human operators add resilience because they can see and act outside the competency box of scenarios and environments for which the system was designed. The system’s competencies can be expanded over time with modifications to software, sensors, etc.; however, it is unclear at what point the competency box becomes large enough to safely eliminate the role of the human operator. One area where advanced automation may be applied is Urban Air Mobility (UAM). Current UAM concepts envision fleets of highly automated air vehicles providing on-demand transport for people and goods. A phased development of UAM has been proposed, beginning with on-board pilots and transitioning to a future state where automated vehicles operate with minimal human involvement. Proponents of UAM note that this final state reduces cost as well as eliminating pilot error, identified as a contributing factor in many aircraft accidents. However, eliminating human involvement also risks eliminating their positive contributions to system resilience. Here we examine Concepts of Operation proposed for future UAM systems and explore how humans can best be incorporated to maintain resilience while minimizing cost and risk. A human-autonomy teaming approach is suggested.

Advanced Air Mobility↗

Towards robust laser beam propagation in atmospheric turbulence

High-fidelity optical propagation through the atmosphere is essential for free-space optical technologies, including laser-based remote sensing and optical communication. However, atmospheric turbulence severely distorts beams and compromises system performance. In this work, we employ hypergeometric-Gaussian (HyGG) vortex beams as probes to characterize and mitigate atmospheric turbulence. Using over 250,000 experimental and simulated frames, we show that refining the power spectrum density (PSD) can reduce numerical prediction errors by up to 79.8%. Concurrently, experimental observations supported by numerical simulations demonstrate that HyGG beams exhibit superior turbulence resilience across multiple metrics compared to conventional Gaussian beams, particularly in their ability to withstand over 5 times stronger turbulence while maintaining similar intensity fluctuations. These dual investigations, on both turbulence mitigation and robust beam solutions, converge to form a unified strategy for enhancing free-space optical system performance. Collectively, our findings provide new insights into light–turbulence interactions and highlight the practical utility of vortex beams under atmospheric conditions.

Zhang, Boyu↗