Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Resilience Framework”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Performance Metrics to Evaluate Utility Resilience Investments

In 2019, Sandia National Laboratories (Sandia) contracted Synapse Energy Economics (Synapse) to research the integration of community and electric grid resilience investment planning as part of the Designing Resilient Communities (DRC): A Consequence-Based Approach for Grid Investment project. Synapse produced a series of reports to explore the challenges and opportunities in several key areas, including benefit-cost analysis (BCA), performance metrics, microgrids, and regulatory mechanisms. This report focuses on BCA. BCA is an approach that electric utilities, electric utility regulators, and communities can use to evaluate the costs and benefits of a wide range of grid resilience investments in a comprehensive and consistent way. While BCA is regularly applied to some types of grid investments, application of BCA to grid resilience investments is in the early stages of development. Though resilience is increasingly cited in connection with grid investment proposals and plans, the resilience- related costs and benefits of grid resilience investments are typically not fully identified, infrequently quantified, and almost never monetized. Without complete assessments of costs and benefits, regulators can be hesitant to approve some types of grid resilience investments. This report provides the first application of the framework developed in the 2020 National Standard Practice Manual for Benefit-Cost Analysis of Distributed Energy Resources (NSPM for DERs) to grid resilience investments. We provide guidance on next steps for implementation to enable grid resilience investments to receive due consideration. We suggest developing BCA principles and standards for jurisdiction-specific BCA tests. We also recommend identifying the resilience impacts of the investments and quantification of these impacts by establishing utility performance metrics for resilience. Proactive integration of grid resilience investments into existing regulatory processes and practices can increase the capacity of jurisdictions to respond to and recover from the consequences of extreme events. 1 National Energy Screening Project. 2020. National Standard Practice Manual for Benefit-Cost Analysis of Distributed Energy Resources.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine-learning-assisted high-temperature reservoir thermal energy storage optimization

High-temperature reservoir thermal energy storage (HT-RTES) has the potential to become an indispensable component in achieving the goal of the net-zero carbon economy, given its capability to balance the intermittent nature of renewable energy generation. In this study, a machine-learning-assisted computational framework is presented to co-optimize the performance metrics of HT-RTES by combining physics-based simulation with stochastic hydrogeologic formation and thermal energy storage operation parameters, artificial neural network regression of the simulation data, and genetic algorithm-enabled multi-objective optimization. A doublet well configuration with a layered (aquitard-aquifer-aquitard) generic reservoir is simulated for cases of continuous operation and seasonal-cycle operation scenarios. Further, neural network-based surrogate models are developed for the two scenarios and applied to generate the Pareto fronts of the HT-RTES performance for four potential HT-RTES sites. The developed Pareto optimal solutions indicate the performance of HT-RTES is operation-scenario (i.e., fluid cycle) and reservoir-site dependent, and the performance metrics have competing effects for a given site and a given fluid cycle. The developed neural network models can be applied to identify suitable sites for HT-RTES, and the proposed framework sheds light on the design of resilient HT-RTES systems.

15 GEOTHERMAL ENERGY↗

Planning for a Resilient Home Electricity Supply System

Resilience of power systems is already a key issue that is getting frequent attention all over the world. It is useful to analyze resilience issues not only for bulk supply, but at all levels including at a customer level. This is because distributed energy resources can play a prominent role in enhancing resilience. Although the literature on planning models, tools and data for bulk supply and distribution systems have expanded in recent years, customer-centric planning, e.g., for an individual household, is yet to receive adequate attention. Although solar PV and battery storage at a household level have been analyzed, how these resources can be optimally combined, together with grid supply, from a resilience perspective is the focus of this study. The study demonstrates how a conceptual framework can be developed to show the trade-off between system costs and resilience including its dimensions such as duration, depth and frequency of service outages. A planning model is developed that incorporates multiple facets of resilience and individual customer preferences. The model considers power system resilience explicitly as a constraint. The model is implemented for a household level case study in Miami, Florida. The results show there are complex trade-offs among different dimensions of resilience. The study demonstrates how combined resilience metrics can be formulated and evaluated using the proposed least-cost planning model at a household level to optimize grid supply together with solar, battery storage and diesel generators. The model allows a planner to directly embed a resilience standard to drive the optimal supply mix. These concepts and the modeling construct can also be applied at other levels of planning, including community level and bulk supply system planning.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Machine Learning-Assisted High-Temperature Reservoir Thermal Energy Storage Optimization: Numerical Modeling and Machine Learning Input and Output Files

This data set includes the numerical modeling input files and output files used to synthesize data, and the reduced-order machine learning models trained from the synthesized data for reservoir thermal energy storage site identification. In this study, a machine-learning-assisted computational framework is presented to identify High-Temperature Reservoir Thermal Energy Storage (HT-RTES) site with optimal performance metrics by combining physics-based simulation with stochastic hydrogeologic formation and thermal energy storage operation parameters, artificial neural network regression of the simulation data, and genetic algorithm-enabled multi-objective optimization. A doublet well configuration with a layered (aquitard-aquifer-aquitard) generic reservoir is simulated for cases of continuous operation and seasonal-cycle operation scenarios. Neural network-based surrogate models are developed for the two scenarios and applied to generate the Pareto fronts of the HT-RTES performance for four potential HT-RTES sites. The developed Pareto optimal solutions indicate the performance of HT-RTES is operation-scenario (i.e., fluid cycle) and reservoir-site dependent, and the performance metrics have competing effects for a given site and a given fluid cycle. The developed neural network models can be applied to identify suitable sites for HT-RTES, and the proposed framework sheds light on the design of resilient HT-RTES systems. All the simulations and the neural network model were done by Idaho National Laboratory. A detailed description of the work was reported in publication linked below.

15 GEOTHERMAL ENERGY↗

SynthEsizing Novel H2 Sensors for Operational Resilience in Pipeline Infrastructure (SENSOR)

A flexible and extensible computational framework acts as a black-box materials discovery engine, capable of screening, predicting, and designing advanced materials with minimal manual intervention was developed. While developed for hydrogen sensing, the approach can be readily adapted to other materials challenges, offering a powerful tool for data-driven materials innovation.

08 HYDROGEN↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Tachyon: Intelligent Multi-Scale Modeling of Distributed Resilient Infrastructure and Workflows for Data Intensive HEP Analyses

The DOE High Energy Physics (HEP) program in Neutrino and Collider science drives data-intensive science and simulation on extreme-scale platforms. Modeling and optimizing the complex distributed components from experimental to leadership computing facilities are essential for HEP workflows to achieve required response times and resilience under various conditions. Tachyon proposes a framework for scalable modeling, simulation, and validation of key performance characteristics for the distributed infrastructure between FNAL and ALCF, along with associated HEP workflows.

Carothers, Chris [Rensselaer Poly.]↗

Hurricane-induced power outage risk under climate change is primarily driven by the uncertainty in projections of future hurricane frequency

Nine in ten major outages in the US have been caused by hurricanes. Long-term outage risk is a function of climate change-triggered shifts in hurricane frequency and intensity; yet projections of both remain highly uncertain. However, outage risk models do not account for the epistemic uncertainties in physics-based hurricane projections under climate change, largely due to the extreme computational complexity. Instead they use simple probabilistic assumptions to model such uncertainties. Here, we propose a transparent and efficient framework to, for the first time, bridge the physics-based hurricane projections and intricate outage risk models. We find that uncertainty in projections of the frequency of weaker storms explains over 95% of the uncertainty in outage projections; thus, reducing this uncertainty will greatly improve outage risk management. We also show that the expected annual fraction of affected customers exhibits large variances, warranting the adoption of robust resilience investment strategies and climate-informed regulatory frameworks.

54 ENVIRONMENTAL SCIENCES↗

Long-Duration Energy Storage Grid Integration-Valuation Framework and Incentive Gaps

Given these challenges and current modeling gaps on Long Duration Energy Storage (LDES), enhancing the structure and design of existing planning, operations, and organized wholesale markets can better characterize the value of LDES to the power system. To more thoroughly assess the gaps and barriers to LDES investment and readiness for integration into a future grid, we conducted stakeholder outreach through an online survey, interviews with individual independent system operators/regional transmission organizations, and a literature review. Based on this assessment, we identified a set of opportunities for LDES focused development, including a framework to quantify the contributions of LDES on resource adequacy, reliability, and resiliency. Specifically, we identify the potential demand for and benefits of an open-source, LDES-centric evaluation framework that can guide future planning, operations, market design, and policy reforms.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Organizational System Resilience to Disinformation: A Viable Systems Model Exploration

This paper explores the utility of organizational system modeling frameworks to provide valuable insight into information flows within organizations and subsequently the opportunities for increasing resilience against disinformation campaigns targeting the system's ability to utilize information within its decision making. Disinformation is a growing challenge for many organizations and in recent years has created delay in decision making. Here the paper has utilized the viable systems model (VSM) to characterize organizational systems and used this approach to outline potential subsystem requirements to promote resilience of the system. The results of this paper can support the development of simulations and models considering the human elements within the system as well as support the development of quantitative measures of resilience.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

Navigating Epistemic Uncertainty in the Management of Flash Droughts

Abstract: Flash droughts, characterized by their rapid onset, sharply contrast with the typically gradual development of traditional droughts. These events are triggered by a combination of low rainfall and high evaporation rates, driven by elevated temperatures, making them particularly challenging to predict and prepare for. As a relatively new concept, flash drought is not well understood, which introduces significant epistemic uncertainties regarding their nature and detection methods. This under-detection hinders planners' ability to effectively manage these events. Despite these uncertainties, flash droughts can have significant impacts, raising the question: how can decision-makers prepare for such events given the current knowledge gaps? To address this, we propose a methodological framework aimed at enhancing flash drought preparedness by guiding the selection of appropriate indicators based on their detection capabilities and the decision-makers' level of risk aversion. Our approach involves evaluating six different flash drought indicators and analyzing the level of agreement among them. Additionally, we consider the decision-makers' risk aversion, distinguishing between those who require consensus across all methods (risk-takers) and those who act based on a single method's indication (risk-averse). The insights gained from this study offer a pathway towards more informed decision-making processes regarding flash droughts, potentially mitigating their adverse effects through better preparedness and response strategies.

climate resilience↗

Cyber Conservative Operations

In an era of increasingly sophisticated and pervasive cyber threats, robust cyber resilience strategies are more critical than ever. This paper introduces the concept of Cyber Conservative Operations, a proactive approach designed to assist critical infrastructure owners and operators in managing risk and maintaining resilience in the face of imminent, yet not occurring, cyber events. By leveraging the principles of Cyber-Informed Engineering (CIE) and the energy sector’s practice of conservative operations, Cyber Conservative Operations offer a framework for planning and executing coordinated active defense and resilience-oriented actions ahead of cyber events. This approach aims to minimize the consequences of digitally-enabled hazards and ensure swift recovery from anticipated cyber threats. Cyber Conservative Operations enhance the ability of owners and operators to execute pre-planned defense and resilience actions, thereby reducing the impact of digitally-enabled hazards. These operations are crucial for addressing impacts that cannot be precisely quantified or forecasted and for preparing organizations for rapid recovery from imminent digital threats. The paper begins with a discussion of conservative operations as applied to the bulk power system (BPS), a current mechanism allowing BPS entities to enact defensive operating plans to mitigate impending grid unreliability. Building on this model, we present a concept for implementing cyber conservative operations at asset owner facilities. The paper concludes with two case studies of cyber conservative operations drawn from high-profile cyber events, illustrating the practical application and benefits of this proactive approach.

42 - ENGINEERING↗

Achieving Resilient In-Flight Performance for Advanced Air Mobility through Simplified Vehicle Operations

A research and development (R&D) approach is proposed for developing and validating concepts and technologies to achieve vehicle autonomy goals of Advanced Air Mobility (AAM) through Simplified Vehicle Operations (SVO). The approach applies resilience-engineering and human-automation teaming (HAT) principles to a framework for defining vehicle-based functions for the management of missions and flight trajectories, focusing initially on the en route flight domain. To achieve the SVO goal of reducing pilot training requirements and thereby increasing the pilot pool for AAM, while at the same time promoting ever-safer operations, a framework for identifying essential functions is proposed. In this framework, functions are first categorized by high-level functional purpose (mission management, flightpath management, tactical operations, and vehicle control) and then subcategorized by attributes of resilient-performing systems (abilities to monitor, respond, learn, and anticipate). The categorization by functional purpose provides structure within which HAT designs can be holistically explored and total levels of human vs. automation responsibility can be varied. The subcategorization by resilient-system attributes provides a mechanism for capturing safety-critical functions that may not be codified in current operational procedures and training curricula, particularly those where humans proactively enhance safety in currently undocumented ways. An R&D approach consisting of seven strategies is proposed in which automation engineering and human-factors communities can collaborate in the research, development, and design of an SVO roadmap to enable the ambitious objectives of AAM.

AAM↗

Strategies for the Design and Operation of Resilient Extraterrestrial Habitats

An Earth-independent permanent extraterrestrial habitat system must function as intended under continuous disruptive conditions, and with significantly limited Earth support and extended uncrewed periods. Designing for the demands that extreme environments such as wild temperature fluctuations, galactic cosmic rays, destructive dust, meteoroid impacts (direct or indirect), vibrations, and solar particle events, will place on long-term deep space habitats represents one of the greatest challenges in this endeavor. This context necessitates that we establish the know-how and technologies to build habitat systems that are resilient. Resilience is not simply robustness or redundancy: it is a system property that accounts for both anticipated and unanticipated disruptions via the design choices and maintenance processes, and adapts to them in operation. We currently lack the frameworks and technologies needed to achieve a high level of resilience in a habitat system. The Resilient Extra Terrestrial Habitats Institute (RETHi) has the mission of leveraging existing novel technologies to provide situational awareness and autonomy to enable the design of habitats that are able to adapt, absorb and rapidly recover from expected and unexpected disruptions. We are establishing both fully virtual and coupled physical-virtual simulation capabilities that will enable us to explore a wide range of potential deep space Smart Hab configurations and operating modes.

Space habitats↗

Strategies for the Design and Operation of Resilient Extraterrestrial Habitats

An Earth-independent permanent extraterrestrial habitat system must function as intended under continuous disruptive conditions, and with significantly limited Earth support and extended uncrewed periods. Designing for the demands that extreme environments such as wild temperature fluctuations, galactic cosmic rays, destructive dust, meteoroid impacts (direct or indirect), vibrations, and solar particle events, will place on long-term deep space habitats represents one of the greatest challenges in this endeavor. This context necessitates that we establish the know-how and technologies to build habitat systems that are resilient. Resilience is not simply robustness or redundancy: it is a system property that accounts for both anticipated and unanticipated disruptions via the design choices and maintenance processes and adapts to them in operation. We currently lack the frameworks and technologies needed to achieve a high level of resilience in a habitat system. The Resilient ExtraTerrestrial Habitats Institute (RETHi) has the mission of leveraging existing novel technologies to provide situational awareness and autonomy to enable the design of habitats that are able to adapt, absorb and rapidly recover from expected and unexpected disruptions. We are establishing both fully virtual and coupled physical-virtual simulation capabilities that will enable us to explore a wide range of potential deep space SmartHab configurations and operating modes.

Space habitats↗

Deciphering supramolecular and polymer-like behavior in metallogels: real-time insights into temperature-modulated gelation and rapid self-assembly dynamics

Bis(pyridyl) urea-based gelators, namely L2 and its isomeric mixture ( L1 + L2 ), are known to self-assemble into 1D architectures capable of inducing supramolecular gelation. Coordination with metal ions such as Ag( I ), Cu( II ), and Fe( III ) introduces structural reinforcement, enabling the formation of distinct 3D networks governed by metal-specific coordination geometries. Here, we present a comprehensive investigation into the temperature-responsive behavior (20–60 °C) of L2 and L1 + L2 , both in the absence and presence of Ag( I ), Dy( III ), Fe( III ), Cu( II ), and Ho( III ), using real-time small-angle neutron scattering (SANS). To probe long-term structural evolution/kinetics of self-assembly, real-time small-angle X-ray scattering (SAXS) was employed on L2 + Ag gels, complemented by differential scanning calorimetry (DSC) to evaluate thermal transitions. Our results reveal strikingly divergent gelation behaviors: L2 forms a highly rigid, covalent polymer-like network, while L1 + L2 exhibits remarkable thermal adaptability. Upon metal coordination, the assemblies exhibit pronounced crystallinity and exceptional thermal stability, as evidenced by persistent Bragg reflections and invariant d-spacings. Intriguingly, L2 : Fe (2 : 1) and L1 : L2 : Fe (0.5 : 0.5 : 1) in acetonitrile-d 3 (ACN-d 3 ) deviate from this trend, forming thermally labile amorphous gels. These systems show a complete loss of crystalline order, reduced Porod exponents—indicative of collapsed or branched fiber morphologies—and prominent melting and glass transition events in DSC. Fitting SANS and SAXS data to the correlation length model unveiled insightful nanostructural features. While most systems displayed minimal temperature-induced variation in mesh size or surface morphology, L2 : Ag in dimethyl sulfoxide-d 6 (DMSO-d 6 )/D 2 O and L2 : Fe (1 : 1) in ACN-d 3 exhibited a rare combination of thermally stable correlation lengths and increasing high- q exponents—strongly suggesting progressive fiber densification or surface smoothing within a robust gel framework. These findings highlight the tunability and structural resilience of supramolecular gels through precise control of ligand architecture, metal coordination, and temperature, offering valuable design principles for functional soft materials.

Pajoubpong, Jinnipha [Univ. of Cincinnati, OH (Uni↗

Understanding Mixed Precision GEMM with MPGemmFI: Insights into Fault Resilience

Emerging deep learning workloads urgently need fast general matrix multiplication (GEMM). Thus, one of the critical features of machine-learning-specific accelerators such as NVIDIA Tensor Cores, AMD Matrix Cores, and Google TPUs is the support of mixed-precision enabled GEMM. For DNN models, lower-precision FP data formats and computation offer acceptable correctness but significant performance, area, and memory footprint improvement. While promising, the mixed-precision computation on error resilience remains unexplored. To this end, we develop a fault injection framework that systematically injects fault into the mixed-precision computation results. We investigate how the faults affect the accuracy of machine learning applications. Based on the characteristics of error resilience, we offer lightweight error detection and correction solutions that significantly improve the overall model accuracy by 75% if the models experience hardware faults. The solutions can be efficiently integrated into the accelerator's pipelines.

Fang, Bo↗