Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Energy Resilience for Mission Assurance, Agile Co-simulation for Cyber Energy System Security (ACCESS) Model Advancements for Resilience Analysis (Version 3, July 2023)

Agile Co-simulation for Cyber Energy System Security (ACCESS) is a co-simulation platform developed by Lawrence Livermore National Laboratory (LLNL). The primary high-level use-case for ACCESS is to study existing or new cyber-physical critical infrastructure systems, with a heavy emphasis on 1) systems that utilize communication networks, and 2) studies that seek to understand cyber-related system impacts. ACCESS is currently used for several energy system resilience projects at LLNL. In the Energy Resilience for Mission Assurance (ERMA) project, ACCESS is used in the Mod eling for Metric Calculation task (specifically, subtask 4.3, Communications and Cyber Modeling) to model and simulate the cyber and communication system aspects of Defense Critical Electric Infrastructure (DCEI) systems, with a focus on computing specific communication system metrics that can impact system resilience and mission performance. Simulated communication system per formance will be fed back to other ERMA system components so that mission performance can be evaluated holistically. This report describes several enhancements to the ACCESS platform that were implemented during the execution of the ERMA project in support of reslience analysis. This includes the addition of new models and subsystems, enhancements to existing models, and integration with external systems. The remainder of this report is structured as follows. In Section 2, a brief background description of the ACCESS platform is provided, including an outline of ACCESS components, example use cases, and a set of communication network resilience metrics that can be computed with ACCESS. Section 3 describes the ACCESS model enhancements for ERMA in detail. Finally, Section 4 briefly outlines future integration opportunities between ACCESS and project participant capabilities identified during the progression of the project.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Demystifying the Resilience of Large Language Models: An End-to-End Perspective

Deep neural networks are known to be resilient to random bit-wise faults in their parameters. However, this resilience has primarily been established through evaluations of classification models. The extent to which this claim holds for large-language models remains underexplored. In this work, we conduct an extensive measurement study on the impact of random bitwise faults in commercial-scale language models. We perform an in-depth analysis of the resulting generation outputs. We first expose that these language models are not truly resilient to random bit-flips. While aggregate metrics such as accuracy may suggest resilience, an in-depth inspection of the generated outputs shows significant degradation in text quality. Our analysis also shows that tasks requiring more complex reasoning suffer more from performance and quality degradation. Moreover, we extend our analysis to models with augmented reasoning capabilities, such as Chain-of-Thought or Mixture of Experts architectures, and characterize their failure scenarios under random bit-flips.

Sun, Yu↗

Radiation-Induced Noise Resilience of Neuromorphic Architectures

Neuromorphic event-based networks use asynchronous time-dependent information to extract features from input data that can allow for edge-based distributed applications such as object recognition. The noise resilience properties of such networks, especially in the context of space applications, are yet to be explored. In this paper, we use the hierarchy of time surfaces (HOTS) algorithm, which is one of the neuromorphic algorithms, to understand the least and most resilient modules in a neuromorphic network. The HOTS algorithm relies on the computing of time surfaces that maps the temporal delays between neighboring pixels into normalized features that involve many computations that are also found in other neuromorphic networks such as exponential decays, distance computations, etcetera. We implemented HOTS on a Digilent PYNQ board with a Xilinx Zynq 7020 system on a chip, and we subjected the boards running the HOTS network inference to neutron radiation at the Los Alamos Neutron Science Center. Furthermore, we used simulation models from our previous similar experiments on the event-based sensor to create a neutron induced noise model to quantify the effect of this noise on the overall performance of the network. This experiment provides the preliminary measurements of the reliability of the HOTS algorithm and proposes methods to create a more reliable HOTS architecture in future spacecraft missions.

Engineering↗

Contact Graph Routing

Contact Graph Routing (CGR) is a dynamic routing system that computes routes through a time-varying topology of scheduled communication contacts in a network based on the DTN (Delay-Tolerant Networking) architecture. It is designed to enable dynamic selection of data transmission routes in a space network based on DTN. This dynamic responsiveness in route computation should be significantly more effective and less expensive than static routing, increasing total data return while at the same time reducing mission operations cost and risk. The basic strategy of CGR is to take advantage of the fact that, since flight mission communication operations are planned in detail, the communication routes between any pair of bundle agents in a population of nodes that have all been informed of one another's plans can be inferred from those plans rather than discovered via dialogue (which is impractical over long one-way-light-time space links). Messages that convey this planning information are used to construct contact graphs (time-varying models of network connectivity) from which CGR automatically computes efficient routes for bundles. Automatic route selection increases the flexibility and resilience of the space network, simplifying cross-support and reducing mission management costs. Note that there are no routing tables in Contact Graph Routing. The best route for a bundle destined for a given node may routinely be different from the best route for a different bundle destined for the same node, depending on bundle priority, bundle expiration time, and changes in the current lengths of transmission queues for neighboring nodes; routes must be computed individually for each bundle, from the Bundle Protocol agent's current network connectivity model for the bundle s destination node (the contact graph). Clearly this places a premium on optimizing the implementation of the route computation algorithm. The scalability of CGR to very large networks remains a research topic. The information carried by CGR contact plan messages is useful not only for dynamic route computation, but also for the implementation of rate control, congestion forecasting, transmission episode initiation and termination, timeout interval computation, and retransmission timer suspension and resumption.

Burleigh, Scott C.↗

Cyber-Power Co-Simulation for End-to-End Synchrophasor Network Analysis and Applications

The resiliency, reliability and security of the next generation cyber-power smart grid depend upon efficiently leveraging advanced communication and computing technologies. Also, developing real-time data-driven applications is critical to enable wide-area monitoring and control of the cyber-power grid given high-resolution data from Phasor Measurement Units (PMUs). North American Synchrophasor Initiative Network (NASPlnet) provides guidance for PMU data exchanges. With the advancement in networking and grid operation, it is necessary to evaluate the performance of different data flow architectures suggested by NASPInet and analyze the impact on applications. Therefore, we need a cyber-power co-simulation framework that supports very large-scale co-simulation capable of running in parallel, high-performance computing platforms and capturing real-life network behavior. This work presents an end-to-end automated and user-driven cyber-power co-simulation using NS3 to model communication networks, GridPACK to model the power grid, and HELICS as a co-simulation engine. Comparative analysis of latency in synchrophasor networks and a performance evaluation of a power system stabilizer application utilizing PMU data in an IEEE 39 bus test system is presented using this cosimulation testbed.

Mustafa, Hussain M.↗

Reinforcement Learning for Intentional Islanding in Resilient Power Transmission Systems

Intentional islanding is the process of identifying and deliberately decomposing the transmission network to form self-sustained islands from an endangered network during disruptions to improve resilience and security. Most existing intentional islanding models are offline resilience decision tools and hence do not provide outage responses in a timely manner. In this paper, a reinforcement learning (RL) based model for intentional islanding is developed, which offers real-time switching control, online deployability, and adaptability to varying system conditions. The intentional islanding process is formulated as a Markov decision process, where the optimal transmission switching policy is learned using the RL approach. The control policy is learned over an environment that encompasses a Power System Simulator for Engineering (PSS/E) model of the transmission network, facilitated by an interface to the standard openAI Gym framework. The proposed RL-based methodology aims to form stable and self-sustainable islands by ensuring voltage stability while reducing the power mismatch in the formed islands. A proximal policy optimization algorithm is designed, which is suitable for controlling the on/off status of the switches with multi-layer perceptron as value and actor networks. The effectiveness of the proposed framework in the self-recovery of the grid by island formation is applied on the modified IEEE 39-bus test network and validated by dynamic simulations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Assessing and Promoting Functional Resilience in Flight Crews During Exploration Missions

NASA plans to send humans to Mars in about 20 years. The NASA Human Research Program supports research to mitigate the major risks to human health and performance on extended missions. However, there will undoubtedly be unforeseen events on any mission of this nature - thus mitigation of known risks alone is not sufficient to ensure optimal crew health and performance. Research should be directed not only to mitigating known risks, but also to providing crews with the tools to assess and enhance resilience, as a group and individually. We can draw on ideas from complexity theory and network theory to assess crew and individual resilience. The entire crew or the individual crewmember can be viewed as a complex system that is composed of subsystems (individual crewmembers or physiological subsystems), and the interactions between subsystems are of crucial importance for overall health and performance. An understanding of the structure of the interactions can provide important information even in the absence of complete information on the component subsystems. This is critical in human spaceflight, since insufficient flight opportunities exist to elucidate the details of each subsystem. Enabled by recent advances in noninvasive measurement of physiological and behavioral parameters, subsystem monitoring can be implemented within a mission and also during preflight training to establish baseline values and ranges. Coupled with appropriate mathematical modeling, this can provide real-time assessment of health and function, and detect early indications of imminent breakdown. Since the interconnected web of physiological systems (and crewmembers) can be interpreted as a network in mathematical terms, we can draw on recent work that relates the structure of such networks to their resilience (ability to self-organize in the face of perturbation). There are many parameters and interactions to choose from. Normal variability is an established characteristic of a healthy physiological response. Healthy coupling has been investigated less extensively, but there are cases in which too tight or too loose coupling can be problematic. This might be in inter-individual behaviors, such as sleep cycles, coordination of work and meal times, and coupled motions during communication. Less apparent are couplings of physiological systems, nevertheless examples abound of coupled systems which might be monitored: cardio-respiratory rhythms; circadian rhythms, body temperature, and sleep; stress markers and cognition, sleep, and performance; profiles of biochemical markers related to immune function and nutritional status; sensorimotor aspects such as motion sickness, ataxia, reaction time, and manual control. Tools for resilience are then the means to measure and analyze these parameters, incorporate them into appropriate models of normal variability and interconnectedness, and recognize when parameters or their couplings are outside of normal limits. What to do when a problem is identified depends on its nature. Changes can be made to crew procedures, work pacing, interpersonal interactions, sleep cycles, meal timing and content, as guided by the model. Use and continued development of these methods could not only provide tools for resilience, but also meaningful autonomous work for the crew on an extended flight.

Shelhamer, Mark↗

Enhancing the Operational Resilience of Advanced Reactors with Digital Twins by Recurrent Neural Networks

Because of a lack of operational data and uncertainty in evaluation model for abnormal and accident scenarios, the established operating procedures can be biased in characterizing the reactor states and ensuring operational resilience. To reduce uncertainty associated with actual plant conditions, digital twin (DT) technology is suggested to support operator’s decision-making by effectively extracting and using knowledge of the current and future plant states from the knowledge base. This study first builds a knowledge base based on the characterization of issue space and the simulation tool. Next, this study discusses diagnosis and prognosis DTs for enhancing operational resilience by recovering the complete states of reactors and by predicting the future reactor behaviors. Finally, the decision-making module of the control system can determine the optimal control strategy that meets operational goals during loss-of-flow scenarios. To demonstrate and evaluate the DTs capability for supporting the operations of nuclear reactors, this study develops and assesses both the diagnosis and prognosis DTs in a nearly autonomous management and control system for an Experimental Breeder Reactor-II simulator during different loss-of-flow scenarios.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Optimal Network Reconfiguration and Scheduling With Hardware-in-the-Loop Validation for Improved Microgrid Resilience

With the increased occurrence of various major extreme weather events, power outages and prompt power system restorations have recently drawn more attention to the resilience and recovery of power systems. From the perspective of a more resilient power delivery at the distribution grid, system restoration using network topology reconfiguration together with optimal scheduling of distributed energy resources are adopted in this paper. The proposed optimization model aims at minimizing the total load shedding cost and other operational costs, in which linearized topological constraints borrowed from graph theory and linearized DistFlow models are respectively used to maintain the radial network topology and power flow balance after system contingencies. To demonstrate the applicability of the proposed strategy, a real-world case study of a networked three-microgrid system in Adjuntas, Puerto Rico, is used with the consideration of different independent/interconnected microgrid scenarios, contingencies, and fairness settings. Furthermore, hardware-in-the-loop testing is conducted for the same three-microgrid network, where the closely matched results with the simulated ones have validated the effectiveness of the proposed restoration strategy, which is now ready to move one step forward towards field deployment. Finally, to test the proposed restoration strategy in a larger networked system, the modified IEEE-33 bus test distribution system is considered, and the results show a more resilient power delivery for critical loads under three and four line outages.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Co-occurrence networks reveal more complexity than community composition in resistance and resilience of microbial communities

Plant response to drought stress involves fungi and bacteria that live on and in plants and in the rhizosphere, yet the stability of these myco- and micro-biomes remains poorly understood. We investigate the resistance and resilience of fungi and bacteria to drought in an agricultural system using both community composition and microbial associations. Here we show that tests of the fundamental hypotheses that fungi, as compared to bacteria, are (i) more resistant to drought stress but (ii) less resilient when rewetting relieves the stress, found robust support at the level of community composition. Results were more complex using all-correlations and co-occurrence networks. In general, drought disrupts microbial networks based on significant positive correlations among bacteria, among fungi, and between bacteria and fungi. Surprisingly, co-occurrence networks among functional guilds of rhizosphere fungi and leaf bacteria were strengthened by drought, and the same was seen for networks involving arbuscular mycorrhizal fungi in the rhizosphere. We also found support for the stress gradient hypothesis because drought increased the relative frequency of positive correlations.

59 BASIC BIOLOGICAL SCIENCES↗

Enhancing the Operational Resilience of Advanced Reactors with Digital Twins by Recurrent Neural Networks

Because of a lack of operation data during abnormal and accident scenarios, along with the existence of uncertainty in the evaluation model for transient and accident analysis, the established abnormal and emergency operating procedures can be biased in characterizing the reactor states and ensuring operational resilience. To improve state awareness and ensure operational flexibility for minimizing effects on the system due to anomaly, digital twin (DT) technology is suggested to support operator's decision-making by effectively extracting and using knowledge of the current and future plant states from the knowledge base. To demonstrate DT's capability for recovering the complete states of reactors and for predicting the future reactor behaviors, this paper develops and assesses both the diagnosis and prognosis DTs in a nearly autonomous management and control system for an Experimental Breeder Reactor-II simulator during different loss-of-flow scenarios.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Resilience Measurement Framework For Post-deployment Artificial Intelligence (ai) Integrated Systems

Resilience is largely defined as the ability to adapt or recover from adverse conditions, stresses, attacks, or compromises on systems that use or are enabled by digital resources. In Artificial Intelligence Management and Research for Advanced Networked Testbed Hub (AMARANTH), resilience is measured in the amount of time it took from the beginning of a testing period for the model to reach predictions outside of the original 95% confidence interval or using the Kullback-Leibler (KL) divergence theorem, the Population Stability Index (PSI), and traditional methods such as root mean squared error (RMSE) threshold. Artificial Intelligence (AI) model drift is of significant concern when deploying AI-integrated systems into critical and/or secure environments. Drift can impact resilience of the AI-integrated system post-deployment and requires consistent maintenance and upkeep to ensure the model is accurate and precise. To quantify model drift and predict the point when a model's drift becomes unacceptable, we describe using Kullback-Leibler (KL) divergence, Population Stability Index (PSI) and/or confidence interval width estimations to determine the point of failure and time to failure of a model post-deployment. Through simple code functions, the KL-divergence, PSI, confidence interval, and root mean squared (RMSE) point of failures can be used to derive when a model needs to be maintained as well as the impact of adversarial action through statistical means.

Yockey, Patience [Idaho National Laboratory (INL),↗

Optimal Siting of EV Fleet Charging Station Considering EV Mobility and Microgrid Formation for Enhanced Grid Resilience

Coordinating infrastructure planning for transportation and the power grid is essential for enhanced reliability and resilience during operation and disaster management. This paper presents a two-stage stochastic model to optimize the location of electric vehicle fleet charging stations (FEVCSs) to enhance the resilience of a distribution network. The first stage of this model deals with the decision to place an FEVCS at the most favorable and optimized location, whereas the second stage aims to minimize the weighted sum of the value of lost load in multiple potential scenarios with different faults. Indeed, the second stage is a joint grid restoration scheme with network reconfiguration and microgrid formation using available distributed generators and fleet electric vehicles. The proposed model is tested on a modified IEEE-33 node distribution network and a four-node transportation network. Case studies demonstrate the effectiveness of the proposed model.

25 ENERGY STORAGE↗

Plan evaluation for heat resilience: complementary methods to comprehensively assess heat planning in Tempe and Tucson, Arizona

Abstract Escalating impacts from climate change and urban heat are increasing the urgency for communities to equitably plan for heat resilience. Cities in the desert Southwest are among the hottest and fastest warming in the U.S., placing them on the front lines of heat planning. Urban heat resilience requires an integrated planning approach that coordinates strategies across the network of plans that shape the built environment and risk patterns. To date, few studies have assessed cities’ progress on heat planning. This research is the first to combine two emerging plan evaluation approaches to examine how networks of plans shape urban heat resilience through case studies of Tempe and Tucson, Arizona. The first methodology, Plan Quality Evaluation for Heat Resilience, adapts existing plan quality assessment approaches to heat. We assess whether plans meet 56 criteria across seven principles of high-quality planning and the types of heat strategies included in the plans. The second methodology, the Plan Integration for Resilience Scorecard™ (PIRS™) for Heat, focuses on plan policies that could influence urban heat hazards. We categorize policies by policy tool and heat mitigation strategy and score them based on their heat impact. Scored policies are then mapped to evaluate their spatial distribution and the net effect of the plan network. The resulting PIRS™ for Heat scorecard is compared with heat vulnerability indicators to assess policy alignment with risks. We find that both cities are proactively planning for heat resilience using similar plan and strategy types, however, there are clear and consistent opportunities for improvement. Combining these complementary plan evaluation methods provides a more comprehensive understanding of how plans address heat and a generalizable approach that communities everywhere could use to identify opportunities for improved heat resilience planning.

Environmental Sciences & Ecology↗

Statistical Analysis of Inter-Area Oscillations in the U.S. Eastern Interconnection: A 2017-2023 Perspective

Recent advancements and the accumulation of high-resolution, long-term phasor measurement unit (PMU) data have provided detailed insights into inter-area oscillations in power grids. This study conducts a comprehensive statistical analysis of inter-area oscillations within the United States Eastern Interconnection from 2017 to 2023. Utilizing data captured by the advanced wide-area Frequency Monitoring Network (FNET/GridEye), this investigation examines the occurrence patterns, dominant frequencies, damping ratios, and excitation mechanisms of these oscillations. Our analysis sheds light on the evolving statistical behaviors of inter-area oscillations, offering updated and critical information for grid operators and planners. The insights gained from this study can be instrumental in enhancing the operational resilience of the power network and guiding strategic developments in grid infrastructure to accommodate future challenges. Additionally, the study discusses emerging challenges associated with the modernization of the power grid, including increased renewable penetration, dynamic load variability, and cyber-physical vulnerabilities that complicate oscillation monitoring and control.

Inter-area oscillations↗

Fault Injection for TensorFlow Applications

As machine learning (ML) has seen increasing adoption in safety-critical domains (e.g., autonomous vehicles), the reliability of ML systems has also grown in importance. While prior studies have proposed techniques to enable efficient error-resilience (e.g., selective instruction duplication), a fundamental requirement for realizing these techniques is a detailed understanding of the application’s resilience. In this work, we present TensorFI 1 and TensorFI 2, high-level fault injection (FI) frameworks for TensorFlow-based applications. TensorFI 1 and 2 are able to inject both hardware and software faults in any general TensorFlow 1 and 2 program respectively. Both are configurable FI tools that are flexible, easy to use, and portable. They can be integrated into existing TensorFlow programs to assess their resilience for different fault types (e.g., bit-flips in particular operations or layers). We use the TensorFI 1 and TensorFI 2 to evaluate the resilience of 12 and 10 ML programs written in TensorFlow, including DNNs used in the autonomous vehicle domain. The results give us insights into why some of the models are more resilient. We also measure the performance overheads of the two injectors, and present 4 case studies, two for each tool, to demonstrate their utility.

97 MATHEMATICS AND COMPUTING↗

Data-driven Resilience Characterization of Control Dynamical Systems

In this paper, we define and quantify resiliency of a power network and propose data-driven algorithms for computing the same for the power grid. To do this, we use the Koopman operator framework to lift the controlled dynamical system to an abstract (possibly higher) dimensional space, where the evolution is linear. The linear system representation allows us to relate small time local controllability and observability of a general nonlinear control system to the controllability and observability of the lifted linear system. Finally, we define the resiliency of the underlying power grid in terms of the controllability and observability gramians of the lifted linear system. We illustrate the proposed approach to compute the resiliency metrics on time-series data obtained from a microgrid.

koopman operator, resilience, control↗