Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Error Resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Thinking outside the box: The human role in increasingly automated aviation systems

Rapid advances in artificial intelligence are enabling automated systems to operate in an increasingly autonomous manner in domains that previously required the involvement of human operators. Examples are rail transport systems, self-driving cars, and warehouse delivery systems. From time to time, such automation encounters operational conditions that fall outside a “competency box” within which the system has been designed to operate. Human operators add resilience because they can see and act outside the competency box of scenarios and environments for which the system was designed. The system’s competencies can be expanded over time with modifications to software, sensors, etc.; however, it is unclear at what point the competency box becomes large enough to safely eliminate the role of the human operator. One area where advanced automation may be applied is Urban Air Mobility (UAM). Current UAM concepts envision fleets of highly automated air vehicles providing on-demand transport for people and goods. A phased development of UAM has been proposed, beginning with on-board pilots and transitioning to a future state where automated vehicles operate with minimal human involvement. Proponents of UAM note that this final state reduces cost as well as eliminating pilot error, identified as a contributing factor in many aircraft accidents. However, eliminating human involvement also risks eliminating their positive contributions to system resilience. Here we examine Concepts of Operation proposed for future UAM systems and explore how humans can best be incorporated to maintain resilience while minimizing cost and risk. A human-autonomy teaming approach is suggested.

Advanced Air Mobility↗

Towards robust laser beam propagation in atmospheric turbulence

High-fidelity optical propagation through the atmosphere is essential for free-space optical technologies, including laser-based remote sensing and optical communication. However, atmospheric turbulence severely distorts beams and compromises system performance. In this work, we employ hypergeometric-Gaussian (HyGG) vortex beams as probes to characterize and mitigate atmospheric turbulence. Using over 250,000 experimental and simulated frames, we show that refining the power spectrum density (PSD) can reduce numerical prediction errors by up to 79.8%. Concurrently, experimental observations supported by numerical simulations demonstrate that HyGG beams exhibit superior turbulence resilience across multiple metrics compared to conventional Gaussian beams, particularly in their ability to withstand over 5 times stronger turbulence while maintaining similar intensity fluctuations. These dual investigations, on both turbulence mitigation and robust beam solutions, converge to form a unified strategy for enhancing free-space optical system performance. Collectively, our findings provide new insights into light–turbulence interactions and highlight the practical utility of vortex beams under atmospheric conditions.

Zhang, Boyu↗

How Humans Contribute to Safety

We have all heard, and much too often, how human error is the leading cause of accidents. What we haven’t been hearing is how humans produce safety far more often than reduce safety. Before we embark on developing technologies to replace the error-prone human, it behooves us to understand how humans produce safety lest we lose that primary source of resilience in our aviation system.

safety↗

Enhancing Unknown Waveform Detection by Learning Intra and Inter-domain Dependencies with Advanced Attention Fusion Mechanisms

Detection of unknown waveforms in mission-critical communications is a crucial area of interest for the Department of Energy (DoE). Traditional methods and recent deep learning-based approaches often assume that the training set includes all possible classes, which is impractical for detecting new waveforms. This limitation gives rise to the problem of open-set recognition (OSR), which involves correctly identifying known classes while detecting and rejecting unknown or unseen classes. To address this limitation, we propose a novel dual-domain complex-valued neural architecture that jointly processes time-domain and frequency-domain signal representations using transformer mechanisms. A transformer model is a deep learning architecture that uses self-attention mechanisms to process and learn relationships in sequential data. Our model employs a cosine similarity loss to extract domain-specific features and incorporates a transformer architecture in the latent space to weigh the importance of different features from the time and frequency domains. The transformer layer includes stacked self-attention and cross-attention modules to learn intra-domain and inter-domain dependencies, creating a more holistic signal representation. An attention-based fusion module intelligently combines the time and frequency-domain features using multi-head attention, enabling the network to learn the optimal feature for each domain in each input signal. Quantitative results demonstrate the impact of these architectural choices on overall performance, showing significant improvement after incorporating self and cross-attention modules and using complex attention fusion over simple weighted fusion. Our ongoing work will focus on addressing the limitations of threshold-based OSR methods by developing a novel generative framework that integrates a conditional diffusion probabilistic model (DPM). DPM is a generative framework that learns to synthesize complex data by reversing a gradual noising process using a neural network trained to denoise step-by-step. Our goal is to leverage the inherent strengths of DPMs for identifying unknown signals more robustly. One primary advantage of using a DPM is its ability to provide a more reliable anomaly score based on the model's reconstruction error, rather than relying solely on classifier confidence. Additionally, the iterative denoising process of DPMs makes this approach naturally resilient to low Signal-to-Noise Ratio (SNR) conditions, where traditional methods often fail. By implementing this generative framework, we aim to enhance the model's capability to accurately detect unknown waveforms and maintain performance in challenging environments.

99 - GENERAL AND MISCELLANEOUS↗

Resiliency in numerical algorithm design for extreme scale simulations

Here this work is based on the seminar titled ‘Resiliency in Numerical Algorithm Design for Extreme Scale Simulations’ held March 1–6, 2020, at Schloss Dagstuhl, that was attended by all the authors. Advanced supercomputing is characterized by very high computation speeds at the cost of involving an enormous amount of resources and costs. A typical large-scale computation running for 48 h on a system consuming 20 MW, as predicted for exascale systems, would consume a million kWh, corresponding to about 100k Euro in energy cost for executing 10 23 floating-point operations. It is clearly unacceptable to lose the whole computation if any of the several million parallel processes fails during the execution. Moreover, if a single operation suffers from a bit-flip error, should the whole computation be declared invalid? What about the notion of reproducibility itself: should this core paradigm of science be revised and refined for results that are obtained by large-scale simulation? Naive versions of conventional resilience techniques will not scale to the exascale regime: with a main memory footprint of tens of Petabytes, synchronously writing checkpoint data all the way to background storage at frequent intervals will create intolerable overheads in runtime and energy consumption. Forecasts show that the mean time between failures could be lower than the time to recover from such a checkpoint, so that large calculations at scale might not make any progress if robust alternatives are not investigated. More advanced resilience techniques must be devised. The key may lie in exploiting both advanced system features as well as specific application knowledge. Research will face two essential questions: (1) what are the reliability requirements for a particular computation and (2) how do we best design the algorithms and software to meet these requirements? While the analysis of use cases can help understand the particular reliability requirements, the construction of remedies is currently wide open. One avenue would be to refine and improve on system- or application-level checkpointing and rollback strategies in the case an error is detected. Developers might use fault notification interfaces and flexible runtime systems to respond to node failures in an application-dependent fashion. Novel numerical algorithms or more stochastic computational approaches may be required to meet accuracy requirements in the face of undetectable soft errors. These ideas constituted an essential topic of the seminar. The goal of this Dagstuhl Seminar was to bring together a diverse group of scientists with expertise in exascale computing to discuss novel ways to make applications resilient against detected and undetected faults. In particular, participants explored the role that algorithms and applications play in the holistic approach needed to tackle this challenge. This article gathers a broad range of perspectives on the role of algorithms, applications and systems in achieving resilience for extreme scale simulations. The ultimate goal is to spark novel ideas and encourage the development of concrete solutions for achieving such resilience holistically.

79 ASTRONOMY AND ASTROPHYSICS↗

PROTEUS: Machine Learning Driven Resilience for Extreme-scale Systems

The objective of this project is to design, develop, and evaluate scalable software to enhance resilience, data checkpointing, program restart, and analysis. The proposed tasks are to 1) develop scalable machine learning techniques to learn temporal change patterns in a scalable and in-situ manner, and to minimize data movement and maximize learning locally closest to data; 2) design a concise data representation and indexing mechanism to capture the distribution of changes in data that can guarantee point-wise user-defined tolerable errors while reducing the data storage requirements by an order of magnitude or more; 3) develop data reduction techniques as library modules; 4) exploit local SSD for minimizing data movement in storage hierarchy; 5) develop anomaly detection algorithms that can predict corruptions based on learning of emerging patterns; 6) develop software libraries to be incorporated within widely used data formats and APIs; and 7) evaluate the proposed software using DOE scientific applications. The outcomes of the proposed work are to satisfy many synergistic data reduction and resilience requirements for large-scale data intensive applications executed on extreme-scale computing systems. The developed mechanism for error-bound data approximation is directly applicable to existing scientific applications. Through machine learning from historical events and change distribution, this work will enable anomaly detection for DOE computer facility.

97 MATHEMATICS AND COMPUTING↗

To Create Safety Is Human

It is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations, are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

operations↗

Changing the Narrative About the Human Role in Accidents

It is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

operations↗

What Do People Do?

It is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

operations↗

So What Do People Actually Do?

I is often said that to err is human. It's true that failures can be traced to human limitation, but what's more important is that all successes, all safe operations are the result of human capabilities. This talk highlights the resilience people bring to aviation operations and discusses ways to change the common narrative that people are the creators of safety rather than only the source of error and failure.

safety↗

AGGREGATE: dAta-driven modelinG preservinG contRollable dEr for outaGe mAnagemenT and rEsiliency (Final Report)

The AGGREGATE project team successfully developed and validated various modules for outage management. Brief summaries of each module are provided to showcase their strength for outage management and restoration for a distribution system with a high penetration of connected distribution energy resources (DERs). In recent years, inverter-based DERs have been widely deployed in distribution system. A most of behind-the-meter (BTM) solar power generation is not visible to the utility. The data-driven DER and load estimation modules are using machine learning (ML) and artificial intelligence (AI) to manage this issue, which provides an opportunity for distribution system operators (DSOs) to operate systems and make decisions in real-time for a distribution system with a high penetration of DERs deployed. Also, the estimated DER and true load can be further leveraged in network aggregation and cold-load pick up estimation for reducing the computing complexity and providing for fast restoration. After load demand and DER power generations have been estimated, the information will support topology and state estimation (SE). The topology estimation module demonstrated the viability of mixed integer linear programming (MILP) formulation to estimate the most likely operational radial topology and outage sections using power flow measurements, historical/estimated load and DERs data and smart meter ping measurements. Formulation includes continuous (power flow, load and DERs data) and binary measurements (smart meter ping measurements) in a single formulation. Errors in continuous data and binary data are modeled as normal distribution and Bernoulli distribution, respectively. In the future distribution grid, the power injection from controllable DERs will be essential for efficient and resilient grid operation. However, determining the optimal DER injections and restoration actions is dependent on knowledge of the system states. State estimation (SE), already the cornerstone of transmission energy management systems, will become commonplace in distribution management systems as more measurements become available from deployment of automated metering infrastructure (AMI). Observability analysis is the first step in SE, as it determines the sufficiency of the available measurements for accurately estimating the current system states. A new type of pseudo-measurement called a Correlational Measurement (CM) is introduced in this module, to enhance the observability of the system to enable more accurate SE. CMs encapsulate knowledge of correlation between demand patterns for similar classes of loads as well as injection patterns for same-technology renewable DERs. During grid contingency scenarios, DERs have been traditionally disconnected, without any fault ride-through capabilities. However, with new regulations and better technology, it is feasible for these resources to contribute to the grid’s restoration after an adverse event and hence enhance resilience. The controllability module proposes a two-step restoration scheme for the power system restoration process by leveraging additional degrees of freedom in power electronics interfaced DERs for mitigating voltage problems. In a resilience mode without the utility system, the distribution grid relies on DERs to serve critical load. In such a severe event with multiple faults on the distribution feeders, actuation of various protective devices (PDs) divides the distribution system into electrical islands. The undetected actuated PDs due to fault current contributions from DERs can delay the restoration process, thereby reducing the system resilience. The Advanced Outage Management (AOM) and the Advanced Feeder Restoration (AFR) modules developed in this project provide improved system resilience with multiple DERs. AOM identifies the faulted sections and actuated PDs in a distribution system with DERs by incorporating smart meter data. The most credible outage scenario including fault locations, PD actuations, and fault indicator (FI) failures is identified by a set of binary integer linear programming incorporating hypotheses. The AFR module serves to restore a distribution system with available energy resources taking into consideration the availability of utility sources and DERs. By partitioning the system into islands, critical load will be served with the available generation resources within islands based on the solution of a MILP. When the utility systems become available, the optimal path will be determined by a spanning tree search algorithm that reconnects these islands back to substations and restores the remaining load. The transmission and distribution (T&D) co-simulation module was used to validate the effect of a control action performed on the distribution side assets as it propagates to the transmission side. This ensures that the control action performed results in a feasible operating point on both the transmission and the distribution system. In addition to validation, the team used the T&D co-simulation module to demonstrate how distribution system assets can be used to mitigate issues on the transmission system. Specifically, the team demonstrated that appropriate switching operations on the distribution side can alleviate the line overload condition on the transmission side without causing new operational constraint violations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Resilient Observer Design for Cyber-Physical Systems with Data-Driven Measurement Pruning

Resilient observer design for Cyber-Physical Systems (CPS) in the presence of adversarial false data injection attacks (FDIA) is an active area of research. The existing state-of-the-art algorithms tend to break down as more and more knowledge of the system is built into the attack model; also as the percentage of attacked nodes increases. From the view of optimization theory, the problem is often cast as a classical error correction problem for which a theoretical limit of has been established as the maximum percentage attacked nodes for which state recovery is guaranteed. Beyond this limit, the performance of -minimization based schemes, for instance, deteriorates rapidly. Similar performance degradation occurs for other types of resilient observers beyond certain percentages of attacked nodes. In order to increase the corresponding percentage of attacked nodes for which state recoveries can be guaranteed, researchers have begun to incorporate prior information into the underlying resilient observer design framework. For the most pragmatic cases, this prior information is often obtained through a data-driven machine learning process. Existing results have shown a strong positive correlation between the maximum attacked percentages that can be tolerated and the accuracy of the data-driven model. Motivated by these results, this chapter examines the case for pruning algorithms designed to improve the Positive Prediction Value (PPV) of the resulting prior information, given stochastic uncertainty characteristics of the underlying machine learning model. Theoretical quantification of the achievable improvement is given. Simulation results show that the pruning algorithm significantly increases the maximum correctable percentage of attacked nodes, even for machine learning model whose prediction power is comparable to the random flip of a coin.

Resilient Observer, Cyber-physical Systems, Data-D↗

Smart Data Mapping for Connecting Power System Model and Geospatial Data

Knowing the geospatial locations of power system model elements is the foundation for analyzing system vulnerability to natural hazards and connecting loads with end users and their communities. However, power system models and geospatial data for power grid assets may have been developed asynchronously without close coordination. Creating a direct mapping between the two may be a challenging task, considering heterogeneous data structures, target uses, historical legacies, and human errors. This work aims to build an automatic data mapping workflow to connect power system model elements and geospatial data for transmission network, and to support energy grid resilience studies for Puerto Rico. The primary steps in this workflow include constructing graphs using geospatial data, and aligning them to the transmission networks defined in the power system data. The results have been evaluated against existing manual mapping practices for part of the Puerto Rico Power Grid model to illustrate the performance of such auto-mapping solutions.

Resilience, geospatial data, grid transmission net↗

Magnetostrictive Ultrasonic Waveguide Transducer for In-Pile Thermometry

Real-time and reliable temperature measurement on nuclear fuels is crucial for the safe operation of existing pressurized-water reactors and future advanced nuclear reactors. Magnetostrictive materials deform when subjected to a magnetic field or exhibit magnetization variation when stressed. Here, based on these properties, this study prototyped an ultrasonic thermometer (UT) consisting of a magnetostrictive waveguide, a dc coil providing appropriate magnetic biasing, and an ac coil generating an acoustic impulse and detecting the resulting acoustic echoes. By tracking the time of flight between the excitation pulse and the echoes, the UT can potentially detect nuclear fuel cladding temperature from a long distance. In this study, magnetostrictive iron–gallium alloys, or Galfenol, were selected as the waveguide due to their large magnetostriction, superior temperature survivability, excellent radiation resilience, and high mechanical robustness. A multiphysics finite-element model considering electrical, magnetic, and mechanical dynamics in the magnetostrictive UT was then developed. The model exhibited an error of 0.24% in time-of-flight simulation and, therefore, enabled computer-aided design and guided signal processing. Between room temperature and 120 °C, the new Galfenol-based UT exhibits a linear sensitivity of 162.8×10 ₋6 °C ₋1 , which is 51.7% higher than a previous magnetostrictive UT based on iron–cobalt–vanadium alloys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Near-Term Reliability and Resilience (NTRR) (Final Report)

The Near-Term Reliability and Resiliency (NTRR) was awarded in December 2020 as an inter-lab project to examine the reliability and resilience of the electricity grid and natural gas transportation availability. The project builds on studies conducted by The North American Electric Reliability Corporation (NERC), the U.S. Department of Energy (DOE), and other non-governmental research and operational focused on reliability and resilience analyses challenges. The research was conceived to address near-term scenarios (within 10 years), when many local and regional policy transitions could begin to impact grid reliability, resilience, and supporting infrastructure availability. To integrate the natural gas interdependency, the team began with the generating capacity and demand projections from the 2020 NERC Long-Term Reliability Assessment and the Bulk Electric System (BES) transmission topologies defined in the Western Electricity Coordinating Council (WECC) Anchor Data Set, Eastern Interconnection Reliability Assessment Group Multi-Regional Modeling Working Group (ERAG/MMWG) Data Set, the team calculated baseline regional power sector gas demands from present electricity delivery year through the end of delivery year 2030/31 by applying security constrained economic dispatch. This demand was compiled along with demand projections for regional residential, commercial, and industrial natural gas demands from the most recent Energy Information Administration (EIA) Annual Energy Outlook Reference Case into Deloitte’s MarketBuilder® North American Gas Model. Through the application of these demands, MarketBuilder® was projected the topology of natural gas flows in the natural gas pipeline network across the interconnected North American system along with regional natural gas prices that may be seen by market participants in future years Additionally, contingencies and sensitivities focused on the built models of the Eastern Interconnection (EI) and Western Interconnection (WI). They address challenges from the following with the outcomes being an identification of performance under the extreme conditions and an identification of potential grid weaknesses that should be addressed to mitigate the reduced performance and improve the resilience and reliability of the specific regions as well as the National Grid: • Weather events including extreme heat, extreme cold, high wind, no wind, wind and solar forecasting errors, and wildfires. • Gas availability, factoring in supply disruption (contractual and physical), seasonal availability constraints, and infrastructure limitations; and • Transmission availability and congestion.

03 NATURAL GAS↗

Physics-based hybrid machine learning for critical heat flux prediction with uncertainty quantification

Critical heat flux (CHF) is a key quantity in nuclear system modeling due to its impact on heat transfer, safety margins, and reactor performance. This study develops and validates an uncertainty-aware hybrid modeling approach that combines machine learning with physics-based models to predict CHF in cases of dryout. The Biasi and Bowring empirical correlations were paired with three ML uncertainty quantification (UQ) techniques: deep neural network (DNN) ensembles, Bayesian neural networks (BNNs), and deep Gaussian processes (DGPs). A pure ML model without a base model was evaluated for comparison. Model performance was assessed under plentiful (7,350 points) and limited (9 points) training data scenarios using parity, uncertainty distributions, and calibration curves. Results show that the Biasi hybrid DNN ensemble achieved the best overall performance, with a mean absolute relative error of 1.846%, and well-calibrated uncertainty estimates. The BNN-based hybrids showed slightly higher error (2.14%) but superior uncertainty calibration. DGP models underperformed, with over 6% error and poor uncertainty calibration. All hybrid models outperformed pure machine learning configurations, demonstrating resistance against data scarcity. These findings indicate that hybrid modeling significantly improves predictive accuracy, interpretability, and resilience to data scarcity. The integration of uncertainty awareness provides actionable confidence in CHF predictions, which is vital for safety-critical decisions in nuclear applications. This hybrid approach offers a viable pathway for deploying ML models in reactor analysis tools while preserving domain knowledge and physical consistency.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Data-Informed Synthetic Networks of Water Distribution Systems for Resilience Analysis in Puerto Rico

The increasing potential of infrastructure disruptions calls for high-quality infrastructure models to be used in resilience analysis and decision making. Unfortunately, many utilities and communities do not have access to accurate and detailed models due to a lack of data and resources. Furthermore, security restrictions on sharing infrastructure models present roadblocks to research, analysis, and decision making. Recent advances in the development of synthetic water distribution models provide a potential solution to this problem. There is an opportunity to improve these methods by leveraging incomplete pipe datasets to aid synthetic network generation. To address this gap, we developed a methodology for synthetic network generation that incorporates partial pipe data using a modification of the minimum cost flow algorithm for network generation and pipe sizing. This methodology demonstrates how partial pipe data can be leveraged to improve site-specific synthetic network generation. For the study area of Mayagüez, Puerto Rico, a synthetic model generated using 50% of real pipe data matches the pressure of the validation system with an average error of 23.5 m of head, which improves upon the average error of 31.6 m of head produced by a synthetic model generated using no data of the real pipes. Additionally, synthetic networks are shown to replicate the pressure response under a disruption scenario of the validation network, suggesting potential use in resilience analysis.

resilience analysis↗

JANUS: Resilient and Adaptive Data Transmission for Enabling Timely and Efficient Cross-Facility Scientific Workflows

In modern science, the growing complexity of large-scale scientific projects has led to an increasing reliance on cross-facility scientific workflows, where resources and expertise from multiple institutions and geographic locations are leveraged to accelerate scientific discovery. These workflows often require transmitting huge amounts of scientific data through wide-area networks. Although high-speed networks like ESnet and transfer services such as Globus have improved data mobility, several challenges remain. The sheer volume of data can overwhelm network bandwidth, widely used transport protocols such as TCP suffer from inefficiencies due to retransmissions triggered by packet loss, and existing fault-tolerance mechanisms like erasure coding introduce substantial overhead. In this paper, we propose Janus, a resilient and adaptable data transmission approach designed for cross-facility scientific workflows. Unlike traditional TCP-based methods, Janus leverages UDP, integrates erasure coding for fault tolerance, and combines it with error-bounded lossy compression to reduce overhead. This novel design allows users to balance data transmission time and accuracy, optimizing transfer performance based on specific scientific requirements. Additionally, Janus dynamically adjusts erasure coding parameters in response to real-time network conditions, ensuring efficient data transfers even in fluctuating environments. We develop optimization models for determining ideal configurations and implement adaptive data transfer protocols to enhance reliability. Through extensive simulations and real-network experiments, we demonstrate that Janus significantly improves transfer efficiency while maintaining data fidelity.

Esaulov, Vladislav [Georgia State University, Atla↗