Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “network resilience”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Chapter 32 - Power Grid Resilience

The new energy paradigm is altering the current power grid trend from synchronous generator reliant system toward power-electronics-based distributed energy resources (DERs). The resilient operation of modern power grid is highly dependent upon cyber and physical reliable operation of DERs. Power grid resilience broadly refers to the ability of the grid to be robust to eventualities and singularities that may disrupt the continuity of reliable power flow. This could involve resilience related to the physical system but more recently, the focus on cyber resilience has been evolving rapidly especially given the burgeoning growth of distributed generation and power-electronics-based DERs under smart grid and given that such threats to reliable operation do not have to evolve localized to where the physical asset is. Commonly, in power grids dominated by DERS, control and energy management are structured at a multitime scale multilayer fashion: (i) primary control layer at microseconds time scale, (ii) secondary control layer at millisecond time scale, and (iii) tertiary control layer at seconds to minutes time scale. The cyber-related issues that may affect the power grid resilience can be introduced in through primary, secondary, and tertiary control layers. The failure of each of these control layers may impact the reliable and resilient operation of the overall network and cause widespread failures which leads to unintended blackouts. As such, this chapter focuses on cyber-related resilience issues in primary control layer, secondary control layer, tertiary control layer, wide area/utility control, and anomaly detection and resilient communication.

anomaly detection↗

Aggregate attack surface management for network discovery of operational technology

Interconnectivity has become a substratum of technology as the benefits of data-driven functionality are being realized in nearly all industries. Increased connectivity of Operational Technology (OT) exacerbates cyber risks because Industrial Control Systems (ICS) are becoming exposed to the Internet. These exposures are often done inadvertently through misconfigurations as additional network devices come online. Attack surface management (ASM) platforms can be used to identify vulnerabilities by performing external network discovery over the Internet using web spiders. These web spiders enable big data analytics of Internet of Things (IoT) devices as identifiable information of Internet-exposed equipment are archived in searchable databases that are made publicly available. There are a multitude of ASM service providers on the market. Here, this study was conducted to evaluate several commonly known tools to determine the aggregate attack surface of control systems. Queries were crafted by targeting commonly known manufacturers and communication protocols found in OT networks. Identified devices were that categorized based on technology types. Each query was replicated between several tools to target identical ICS equipment. Findings in this paper suggested a significant variance in the exposures discovered by each tool, but unique contributions were identified for each tool when a merged attack surface was derived. Therefore, all tools should be used in aggregate.

97 MATHEMATICS AND COMPUTING↗

A Markov framework for generalized post-event systems recovery modeling: From single to multihazards

State-dependent models can be used to represent the system recovery process as a series of stochastic transitions from lower to higher functional states. However, the applications of these models have been limited in scope and there is a lack of a generalized recovery modeling framework. A generalized framework would permit a robust forecasting of systems and system-of-systems recovery under multiple hazards, and more broadly, would contribute to community disaster preparedness. This paper develops a generalized post hazard-event recovery modeling framework based on state-dependent Markov-type processes. We then apply the proposed framework to solve a spectrum of problems that range from hind-casting single-system recovery following a single hazard event to forecasting post-event trajectories under multiple hazards and modeling the recovery of a system-of-systems. First, Markov chains are used to hind-cast the observed recovery for a portfolio of buildings affected by the 2014 South Napa, California, earthquake. Next, Markov processes are used to formulate a parametric post hazard-event recovery model, which can be updated using Bayesian statistics when relevant datasets become available. Semi-Markov processes are then used to develop a more general model of single hazard recovery, which accounts for the intensity of the loading and level of damage caused by the event. Semi-Markov processes with non-renewal features are then used to account for multihazard interactions in a post-event recovery model, and applied to a case study that involves a community in Charleston, South Carolina. Lastly, Markov-type processes are combined with Bayesian networks to model the recovery of residential, commercial, educational, and industrial buildings (system-of-systems) following a hazard event. Overall, these applications demonstrate the versatility of the Markov framework towards handling recovery problems with varying levels of complexity.

42 ENGINEERING↗

Primary Frequency Control Using Motor Drives for Short Term Grid Disturbances

Industrial motor systems make up a quarter of all electric sales in the United States. Variable speed drives (VSDs) can provide energy efficiency savings to the customer by regulating motor speed based on specific and varying needs. In addition to the benefits provided to the customer, VSDs can provide support to the grid through ancillary services. The Center for Ultra-Wide-Area Resilient Electric Energy Transmission Networks (CURENT) developed a power electronics converter-based grid emulator to allow testing of various power system architectures and demonstration of key technologies in monitoring, control, actuation, and visualization. This paper proposes using an active front-end VSD's connected motor load to provide frequency regulation to a large scale power grid. Each part of the emulator is described including motor and power electronics model and control. The proposed frequency regulation is implemented in VSDs and modeled in both a transmission system in EMTDC/PSCAD and verified on CURENT's hardware testbed.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Enhancing Active Distribution Systems Resilience by Fully Distributed Self-Healing Strategy

Distributed restoration can exploit smart grid technologies to enhance the resilience of active distribution networks toward a self-healing smart grid. However, the large number of decision variables, especially the binary ones for reconfiguration, bring challenges to developing scalable distributed distribution service restoration (DDSR) strategies. This paper proposes a fully distributed solution procedure based on the alternating direction method of multipliers (ADMM) for mixed-integer programming problems and applies to develop the DDSR framework. The method consists of relax-drive-polish phases, 1) relaxing binary variables, and applying the convex ADMM as a warm start; 2) driving the solutions toward Boolean values through a proximal operator; 3) fixing the obtained binding binary variables and solving the rest of the problem to polish results and achieve a high-quality suboptimal solution. Then, an autonomous clustering strategy and consensus ADMM are integrated with the proposed method to realize the fully distributed cluster-based framework of DDSR. This framework can first determine DER scheduling and switch status for reconfiguration to energize the out-of-service areas from local faults, and then provide the load restoration solution in a distributed manner for total blackouts in large-scale distribution networks. Furthermore, the effectiveness and scalability of the proposed DDSR framework are demonstrated through testing on the IEEE 123-node, IEEE 8500-node, and synthetic 100k-node test feeders.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Power Distribution Designing For Resilience Application

The number of power outages have been on the rise as extreme weather patterns have been occurring more frequently due to climate change. These outages have an economic impact in the billions of dollar. The modernization of the power grid and adoption of high penetration of distributed generation such as wind and solar have the potential to reduce the number and length of time of these power outages. However, planners and operators need a way to value the resilience the distributed assets provide to the grid. PowDDeR is a tool that allows them to evaluate the resilience each asset provides as well as how they provide resilience over the connected network.

Phillips, TylerB↗

Carnac for Emulytics (HPC Annual Report V.1.0)

Carnac, located at Sandia's California site, is an institutional cluster for Emulytics that provides security researchers with resources to model enterprise computer networks and evaluate how resilient they are from attacks. While multiple Emulytics cluster computers have been built at Sandia, Carnac is the first system that was developed as an institutional resource that can be shared among different groups with disparate requirements.

97 MATHEMATICS AND COMPUTING↗

NASA’s Lunar Communications and Navigation Architecture: Human Lunar Return

NASA’s Artemis missions will return humanity to the Moon, establishing a long-term presence there and opening more of the lunar surface to exploration than ever before. This rapid growth of lunar activity requires robust and resilient communications, navigation, and networking capabilities for crew safety, command and control of spacecraft, return of science data, and precise maneuvering of assets in space and on the lunar surface.

Michael J Zemba↗

Smart Data Mapping for Connecting Power System Model and Geospatial Data

Knowing the geospatial locations of power system model elements is the foundation for analyzing system vulnerability to natural hazards and connecting loads with end users and their communities. However, power system models and geospatial data for power grid assets may have been developed asynchronously without close coordination. Creating a direct mapping between the two may be a challenging task, considering heterogeneous data structures, target uses, historical legacies, and human errors. This work aims to build an automatic data mapping workflow to connect power system model elements and geospatial data for transmission network, and to support energy grid resilience studies for Puerto Rico. The primary steps in this workflow include constructing graphs using geospatial data, and aligning them to the transmission networks defined in the power system data. The results have been evaluated against existing manual mapping practices for part of the Puerto Rico Power Grid model to illustrate the performance of such auto-mapping solutions.

Resilience, geospatial data, grid transmission net↗

Resilience Design Patterns: A Structured Approach to Resilience at Extreme Scale (V.2.0)

Reliability is a serious concern for future extreme-scale high-performance computing (HPC) systems. Projections based on the current generation of HPC systems and technology roadmaps suggest the prevalence of very high fault rates in future systems. The errors resulting from these faults will propagate and generate various kinds of failures, which may result in outcomes ranging from result corruptions to catastrophic application crashes. Therefore, the resilience challenge for extreme-scale HPC systems requires coordination between various hardware and software technologies that are capable of handling a broad set of fault models at accelerated fault rates. Also, due to practical limits on power consumption in future HPC systems, they are likely to embrace innovative architectures, increasing the levels of hardware and software complexities. Therefore, the techniques that seek to improve resilience must navigate the complex trade-off space between resilience and the overheads to power consumption and performance. While the HPC community has developed various resilience solutions, application-level techniques as well as system-based solutions, the solution space of HPC resilience techniques remains fragmented. There are no formal methods to integrate the various HPC resilience techniques into composite solutions, nor are there methods to holistically evaluate the adequacy and efficacy of such solutions in terms of their protection coverage, and their performance & power efficiency characteristics. Additionally, few implementations of current resilience solutions are portable to newer architectures and software environments that will be deployed on future systems. We developed a new structured approach to the management of HPC resilience using the concept of resilience-based design patterns. In general, a design pattern is a repeatable solution to a commonly occurring problem. We identified the well-known solutions that are commonly used to deal with faults, errors and failures in HPC systems. In the initial design patterns specification (version 1.0), we described the various solutions, which address specific problems in the design of resilient HPC environments, in the form of patterns. Each pattern describes a problem caused by a fault, error or failure event in an HPC environment, and then describes the core of the solution of the problem in such a way that this solution may be adapted to different systems and implemented at different layers of the system stack. The catalog of these resilience design patterns provides designers with a collection of design elements. To construct complete resilience solutions using combinations of various patterns, we defined a framework that enhances HPC designers' understanding of the important constraints and the opportunities for the design patterns to be implemented and deployed at various layers of the system stack. The design framework is also useful for establishing interfaces and mechanisms to coordinate flexible fault management across hardware and software components, as well as to consider the trade-off between performance, resilience, and power consumption when constructing a solution. The resilience design patterns specification version 1.1 included more detailed explanations of the pattern solutions, the context in which the patterns are applicable, and the implications for hardware or software design. It also provided several additional examples and detailed case studies to demonstrate the use of patterns to build realistic solutions. In version 1.2 of the specification document, we have improved the pattern descriptions, including graphical representations of the pattern components. These improvements are largely based on critical comments, feedback and suggestions received from pattern experts and readers of the previous versions of the specification. The pattern classification has been modified to further clarify the relationships between pattern categories. This version of the specification also introduces a pattern language for resilience design patterns. The pattern language presents the patterns in the catalog as a network, revealing the relations among the resilience patterns. The language provides designers with the means to explore alternative techniques for handling a specific fault model that may have different efficiency and complexity characteristics. Using the pattern language also enables the design and implementation of comprehensive resilience solutions as a set of interconnected resilience patterns that can be instantiated across layers of the system stack. The overall goal of this work is to provide hardware and software designers, as well as the users and operators of HPC systems, a systematic methodology for the design and evaluation of resilience technologies in HPC systems that keep scientific applications running to a correct solution in a timely and cost-efficient manner despite frequent faults, errors, and failures of various types. Version 2.0 expands the resilience design pattern classification and catalog to include self-stabilization patterns and reliability, availability and performance models for each structural pattern.

97 MATHEMATICS AND COMPUTING↗

SDN-Based Smart Cyber Switching (SCS) for Cyber Restoration of a Digital Substation

In recent years, critical infrastructure and power grids have increasingly been targets of cyber-attacks, causing widespread and extended blackouts. Digital substations are particularly vulnerable to such cyber incursions, jeopardizing grid stability. This paper addresses these risks by proposing a cybersecurity framework that leverages software-defined networking (SDN) to bolster the resilience of substations based on the IEC- 61850 standard. The research introduces a strategy involving smart cyber switching (SCS) for mitigation and concurrent intelligent electronic device (CIED) for restoration, ensuring ongoing operational integrity and cybersecurity within a substation. The SCS framework improves the physical network’s behavior (i.e., leveraging commercial SDN capabilities) by incorporating an adaptive port controller (APC) module for dynamic port management and an intrusion detection system (IDS) to detect and counteract malicious IEC-61850-based sampled value (SV) and generic object-oriented system event (GOOSE) messages within the substation’s communication network. The framework’s effectiveness is validated through comprehensive simulations and a hardware-in-the-loop (HIL) testbed, demonstrating its ability to sustain substation operations during cyber-attacks and significantly improve the overall resilience of the power grid.

Liu, Chen-Ching (ORCID:0000000289417958)↗

The Interplay of Binary and Quantitative Structure on the Stability of Mutualistic Networks

Synopsis Understanding how the structure of biological systems impacts their resilience (broadly defined) is a recurring question across multiple levels of biological organization. In ecology, considerable effort has been devoted to understanding how the structure of interactions between species in ecological networks is linked to different broad resilience outcomes, especially local stability. Still, nearly all of that work has focused on interaction structure in presence-absence terms and has not investigated quantitative structure, i.e., the arrangement of interaction strengths in ecological networks. We investigated how the interplay between binary and quantitative structure impacts stability in mutualistic interaction networks (those in which species interactions are mutually beneficial), using community matrix approaches. We additionally examined the effects of network complexity and within-guild competition for context. In terms of structure, we focused on understanding the stability impacts of nestedness, a structure in which more-specialized species interact with smaller subsets of the same species that more-generalized species interact with. Most mutualistic networks in nature display binary nestedness, which is puzzling because both binary and quantitative nestedness are known to be destabilizing on their own. We found that quantitative network structure has important consequences for local stability. In more-complex networks, binary-nested structures were the most stable configurations, depending on the quantitative structures, but which quantitative structure was stabilizing depended on network complexity and competitive context. As complexity increases and in the absence of within-guild competition, the most stable configurations have a nested binary structure with a complementary (i.e., anti-nested) quantitative structure. In the presence of within-guild competition, however, the most stable networks are those with a nested binary structure and a nested quantitative structure. In other words, the impact of interaction overlap on community persistence is dependent on the competitive context. These results help to explain the prevalence of binary-nested structures in nature and underscore the need for future empirical work on quantitative structure.

Zoology↗

Transformations in Air Transportation Systems For the 21st Century

Globally, our transportation systems face increasingly discomforting realities: certain of the legacy air and ground infrastructures of the 20th century will not satisfy our 21st century mobility needs. The consequence of inaction is diminished quality of life and economic opportunity for those nations unable to transform from the 20th to 21st century systems. Clearly, new thinking is required regarding business models that cater to consumers value of time, airspace architectures that enable those new business models, and technology strategies for innovating at the system-of-networks level. This lecture proposes a structured way of thinking about transformation from the legacy systems of the 20th century toward new systems for the 21st century. The comparison and contrast between the legacy systems of the 20th century and the transformed systems of the 21st century provides insights into the structure of transformation of air transportation. Where the legacy systems tend to be analog (versus digital), centralized (versus distributed), and scheduled (versus on-demand) for example, transformed 21st century systems become capable of scalability through technological, business, and policy innovations. Where air mobility in our legacy systems of the 20th century brought economic opportunity and quality of life to large service markets, transformed air mobility of the 21st century becomes more equitable available to ever-thinner and widely distributed populations. Several technological developments in the traditional aircraft disciplines as well as in communication, navigation, surveillance and information systems create new foundations for 21st thinking about air transportation. One of the technological developments of importance arises from complexity science and modern network theory. Scale-free (i.e., scalable) networks represent a promising concept space for modeling airspace system architectures, and for assessing network performance in terms of robustness, resilience, and other metrics. The lecture offers an air transportation system topology and a scale-free network linkage graphic as framework for transportation system innovation. Successful outcomes of innovation in air transportation could lay the foundations for new paradigms for aircraft and their operating capabilities, air transportation system topologies, and airspace architectures and procedural concepts. These new paradigms could support scalable alternatives for the expansion of future air mobility to more consumers in more parts of the world.

Holmes, Bruce J.↗

Graph-Theoretic Approaches to Quantifying Power System Resiliency

Although gaining growing importance, the subject of power system resiliency still lacks a commonly acknowledged metric. As a contribution to solving this complication, in this paper we leverage the concepts of spanning trees and Fiedler value from graph theory to propose two topology-based indices for quantifying the resiliency of power systems. The proposed indices require least information and may be applied to any other flow network, such as water or gas pipeline networks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

SWARM: Reimagining scientific workflow management systems in a distributed world

Modern scientific workflows process massive amounts of data from diverse instruments and sensors, leveraging geographically distributed, heterogeneous compute and storage resources—from leadership-class systems to edge devices—connected by high-performance networks. The diversity of resources introduces challenges in harnessing their full potential, with resilience issues arising across applications, system software, networks, storage, and hardware. Today, workflow management systems (WMS) coordinate the execution of computation and data management tasks across target resources. However, WMS’s centralized nature makes them vulnerable to faults and scalability issues that may result in failures of entire computational campaigns. In conclusion, this paper introduces a novel agentic framework for workflow management, fully distributing and decentralizing the WMS functions and modeling them as swarm intelligence agents infused with advanced artificial intelligence solutions and traditional distributed computing algorithms that can make coordinated decisions in the presence of failures of the underlying cyberinfrastructure.

Swarm intelligence↗

Detection and imaging of chemicals and hidden explosives using terahertz time-domain spectroscopy and deep learning

Detecting concealed chemicals and explosives remains a critical challenge in global security. Terahertz time-domain spectroscopy (THz-TDS) offers a promising non-invasive and stand-off detection technique owing to its ability to penetrate optically opaque materials without causing ionization damage. While many chemicals exhibit distinct spectral features in the terahertz range, conventional terahertz-based detection methods often struggle in real-world environments, where variations in sample geometry, thickness, and packaging can lead to inconsistent spectral responses. In this study, we present a chemical imaging system that integrates THz-TDS with deep learning to enable accurate pixel-level identification and classification of different explosives. Operating in reflection mode and enhanced with plasmonic nanoantenna arrays, our THz-TDS system achieves a peak dynamic range of 96 dB and a detection bandwidth of 4.5 THz, supporting practical, stand-off operation. By analyzing individual time-domain pulses with deep neural networks, the system exhibits strong resilience to environmental variations and sample inconsistencies. Blind testing across eight chemicals—including pharmaceutical excipients and explosive compounds—resulted in an average classification accuracy of 99.42% at the pixel level. Notably, the system maintained an average accuracy of 88.83% when detecting explosives concealed under opaque paper coverings, demonstrating its robust generalization capability. These results highlight the potential of combining advanced terahertz spectroscopy with neural networks for highly sensitive and specific chemical and explosive detection in diverse and operationally relevant scenarios.

Imaging and sensing↗

Cybersecurity Assessment in DER-rich Distribution Operations: Criticality Levels and Impact Analysis

The integration of distributed energy resources (DERs) in distribution networks has become a pivotal strategy for achieving decarbonization, enhancing grid resilience, and optimizing grid efficiency. Remote monitoring and control op- erations of such resources rely on a network of sensors and communication infrastructure, exposing the system to potential cyber threats. Therefore, as the deployment of DERs increases, ensuring secure monitoring and control becomes an imperative challenge. This paper utilizes real-time feeder models, which are instrumental in developing cybersecurity testbeds tailored for hardware-in-loop (HIL) systems. These models enable users to simulate cyber attacks in a real-world environment and analyze the power distribution operations during vulnerabilities. Furthermore, we discuss several practical sets of grid parameters to identify critical levels of DERs and evaluate various scenarios that simulate cyber threats on sensitive DERs. The modified IEEE 123-bus model is used as the test case for demonstrating the proposed scenarios. The findings from this study provide valuable insights into the vulnerabilities and potential consequences of cyber attacks on DERs, allowing for better mitigation strategies and improved cyber resilience in future distribution networks.

Maharjan, Manisha↗