Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

A Model-Free Voltage Control Approach to Mitigate Motor Stalling and FIDVR for Smart Grids

Electric power networks are large and highly nonlinear dynamical systems that present unique challenges to control design. Though there is a large number of dynamic models for power system stability and control, many models are only useful with right assumptions and wrong for other tasks. Moreover, the dynamic behavior of the grid is increasingly complex under the banner of smart grids. These lead to the difficulty of developing appropriate dynamic modeling, and thus an efficient control strategy. To avoid such modeling challenges, here we present a novel dynamic voltage control strategy based on a model-free control (MFC) approach, requiring no modeling procedure. In particular, it focuses on fault-induced delayed voltage recovery (FIDVR) events, which require complex and accurate dynamic load models to replicate such events. This work utilizes MFC as an online controller to achieve the desired voltage stability under the FIDVR event. The proposed MFC strategy allows simple implementation and low computational cost for efficient mitigation of FIDVR. For benchmarking, a reasonably accurate dynamic performance model is explored. Simulation results with the IEEE 57 bus test network demonstrate the enhanced dynamic voltage profile for load buses having induction motors with the support of reactive power resources.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Cybersecurity Considerations for Grid-Connected Batteries with Hardware Demonstrations

The share of renewable and distributed energy resources (DERs), like wind turbines, solar photovoltaics and grid-connected batteries, interconnected to the electric grid is rapidly increasing due to reduced costs, rising efficiency, and regulatory requirements aimed at incentivizing a lower-carbon electricity system. These distributed energy resources differ from traditional generation in many ways including the use of many smaller devices connected primarily (but not exclusively) to the distribution network, rather than few larger devices connected to the transmission network. DERs being installed today often include modern communication hardware like cellular modems and WiFi connectivity and, in addition, the inverters used to connect these resources to the grid are gaining increasingly complex capabilities, like providing voltage and frequency support or supporting microgrids. To perform these new functions safely, communications to the device and more complex controls are required. The distributed nature of DER devices combined with their network connectivity and complex controls interfaces present a larger potential attack surface for adversaries looking to create instability in power systems. To address this area of concern, the steps of a cyberattack on DERs have been studied, including the security of industrial protocols, the misuse of the DER interface, and the physical impacts. These different steps have not previously been tied together in practice and not specifically studied for grid-connected storage devices. In this work, we focus on grid-connected batteries. We explore the potential impacts of a cyberattack on a battery to power system stability, to the battery hardware, and on economics for various stakeholders. We then use real hardware to demonstrate end-to-end attack paths exist when security features are disabled or misconfigured. Our experimental focus is on control interface security and protocol security, with the initial assumption that an adversary has gained access to the network to which the device is connected. We provide real examples of the effectiveness of certain defenses. This work can be used to help utilities and other grid-connected battery owners and operators evaluate the severity of different threats and the effectiveness of defense strategies so they can effectively deploy and protect grid-connected storage devices.

25 ENERGY STORAGE↗

An ODE-Enabled Distributed Transient Stability Analysis for Networked Microgrids

Networked microgrid (NMG) exhibits noteworthy resiliency and flexibility benefits for the mutual support from neighboring microgrids. With high penetration of distributed energy resources (DERs) and the associated controls, the transient stability analysis of NMGs is of critical significance. To address the issues of computation burdens and privacy in the centralized transient analysis, this paper devises an ordinary differential equation (ODE)-enabled distributed transient stability (DTS) methodology for NMGs. First, an ODE-based microgrid model is established to capture the dynamics in the droop control of DERs as well as network and load. Further, a distributed DTS is devised for the ODE representation of an NMG, allowing a privacy-preserving transient analysis of each microgrid while accurately reconstructing the frequency dynamics under droop controls in all DERs. In conclusion, extensive tests are performed to verify the validity of the ODE-based microgrid model through both dynamic response and eigenvalue analysis, and the efficacy of the DTS algorithm in simulating the large signal responses and the frequent fluctuations in NMG.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Aggregate data‐driven dynamic modeling of active distribution networks with DERs for voltage stability studies

Abstract Electric distribution networks increasingly host distributed energy resources based on power electronic converter (PEC) toward active distribution networks (ADN). Despite advances in computational capabilities, electromagnetic transient models are limited in scalability because of their reliance on exact data about the distribution system and each of its components. Similarly, the use of the DER_A model, which is intended to examine the combined dynamic behavior of many DERs, is limited by the difficulty in parameterization. There is a need for improved dynamic models of DERs for use in large power system simulations for stability analysis. This paper proposes an aggregate model‐free, data‐driven approach for deriving a dynamic partitioned model (DPM) of ADNs. Detailed residential distribution feeders were first developed, including PEC‐based DERs and composite load models (CMLDs), from which the aggregated DPM was derived. The performance was evaluated through various case studies and validated against the detailed ADN model and state‐of‐the‐art DER_A model with CMLD. The data‐driven DPM achieved a of over 90%, accurately representing the aggregated dynamic behavior of ADNs. Furthermore, the DPM significantly accelerated the simulation process with a computational speedup of 68 times compared to the detailed ADN and a 3.5 times speedup compared to the DER_A CMLD model.

42 ENGINEERING↗

Constructing a Testing Application for GWMS

Many of Fermilab’s High Energy Physics experiments require High Throughput Computing to carry out simulations, data reconstruction, and data analysis. Glidein Workflow Management System (GWMS) is a tool that distributes High Throughput Computing. The purpose of GWMS is to provide convenient access to Grid resources and sites. GWMS is coupled to HTCondor. HTCondor is the workload management system that is used for scheduling and job control. Users submit jobs to a local HTCondor queue and GWMS will make sure the job will run on one of the many remote resources.

Neely-Brown, LeRayah↗

Powered By: Transmission Planning

Presented as part of the NREL Grid Planning and Analysis Center's (GPAC) "Powered By" Series, this iteration focuses not on a specific tool but instead on a domain of expertise/capability being built out at NREL: transmission planning. Various sets of tools are integrated together in this presentation to present a combination of transmission focus areas, case studies, tools and future directions of transmission research.

capacity expansion↗

A Comprehensive Analysis of PINNs for Power System Transient Stability

The integration of machine learning in power systems, particularly in stability and dynamics, addresses the challenges brought by the integration of renewable energies and distributed energy resources (DERs). Traditional methods for power system transient stability, involving solving differential equations with computational techniques, face limitations due to their time-consuming and computationally demanding nature. This paper introduces physics-informed Neural Networks (PINNs) as a promising solution for these challenges, especially in scenarios with limited data availability and the need for high computational speed. PINNs offer a novel approach for complex power systems by incorporating additional equations and adapting to various system scales, from a single bus to multi-bus networks. Our study presents the first comprehensive evaluation of physics-informed Neural Networks (PINNs) in the context of power system transient stability, addressing various grid complexities. Additionally, we introduce a novel approach for adjusting loss weights to improve the adaptability of PINNs to diverse systems. Our experimental findings reveal that PINNs can be efficiently scaled while maintaining high accuracy. Furthermore, these results suggest that PINNs significantly outperform the traditional ode45 method in terms of efficiency, especially as the system size increases, showcasing a progressive speed advantage over ode45.

24 POWER TRANSMISSION AND DISTRIBUTION↗

STREAM: A Scalable Federated HPC Telemetry Platform

Obtaining and analyzing high performance computing (HPC) telemetry in real time is a complex task that can impact algo- rithmic performance, operating costs, and ultimately scientific outcomes. If your organization operates multiple HPC systems, filesystems, and clusters, telemetry streams can be synthesized in order to ease operational and analytics burden. In order to collect this telemetry, the Oak Ridge Leadership Computing Facility (OLCF) has deployed STREAM (Streaming Telemetry for Resource Events, Analytics, and Monitoring), which is a distributed and high-performance message bus based on Apache Kafka. STREAM collects center-wide performance information and must interface with many sources, including five HPE deployed supercomputers, each with their own Kafka cluster which is managed by HPCM. OLCF Supercomputers and their attached scratch filesystems currently send more than 300 million messages to over 200 topics producing around 1.3 Terabytes per day of telemetry data to STREAM. This paper describes the architectural principles that enable STREAM to be both resilient and highly performant while supporting multiple upstream Kafka clusters and other data sources. It also discusses the design challenges and decisions faced in adapting our existing system- monitoring infrastructure to support the first Exascale computing platform.

Adamson, Ryan↗

Orchestration of materials science workflows for heterogeneous resources at large scale

In the era of big data, materials science workflows need to handle large-scale data distribution, storage, and computation. Any of these areas can become a performance bottleneck. We present a framework for analyzing internal material structures (e.g., cracks) to mitigate these bottlenecks. We demonstrate the effectiveness of our framework for a workflow performing synchrotron X-ray computed tomography reconstruction and segmentation of a silica-based structure. Our framework provides a cloud-based, cutting-edge solution to challenges such as growing intermediate and output data and heavy resource demands during image reconstruction and segmentation. Specifically, our framework efficiently manages data storage, scaling up compute resources on the cloud. The multi-layer software structure of our framework includes three layers. A top layer uses Jupyter notebooks and serves as the user interface. A middle layer uses Ansible for resource deployment and managing the execution environment. A low layer is dedicated to resource management and provides resource management and job scheduling on heterogeneous nodes (i.e., GPU and CPU). At the core of this layer, Kubernetes supports resource management, and Dask enables large-scale job scheduling for heterogeneous resources. The broader impact of our work is four-fold: through our framework, we hide the complexity of the cloud’s software stack to the user who otherwise is required to have expertise in cloud technologies; we manage job scheduling efficiently and in a scalable manner; we enable resource elasticity and workflow orchestration at a large scale; and we facilitate moving the study of nonporous structures, which has wide applications in engineering and scientific fields, to the cloud. While we demonstrate the capability of our framework for a specific materials science application, it can be adapted for other applications and domains because of its modular, multi-layer architecture.

97 MATHEMATICS AND COMPUTING↗

CARILEC Resilient Energy Community CoP for Cybersecurity Workshop Series: Cybersecurity Assessment Tools [Slides]

For the last several years and in collaboration with CARILEC, USAID and NREL have been working to support cyber resilience at power sector utilities in Latin America and the Caribbean. Direct technical assistance with regional utilities has been a key component of USAID-NREL Partnership activities, and technical assistance has typically included a foundational cybersecurity assessment using NREL's Distributed Energy Resource Cybersecurity Framework (DER-CF) tool. The DER-CF allows organizations to benchmark and evaluate their cybersecurity posture across the areas of Governance, Technical Management, and Physical Security. To complement the activities of the newly created CAREC IT/OT and Cybersecurity Team, this webinar on cybersecurity assessment tools includes an overview of the DER-CF tool and a discussion with regional stakeholders and NREL experts on the DER-CF assessment process and other resources for cybersecurity assessments.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Toward equitable environmental exposure modeling through convergence of data, open, and citizen sciences: an example of air pollution exposure modeling amidst increasing wildfire smoke

Exposure modeling is critical in environmental epidemiology and human health but may face challenges (e.g., skewed data, unequal error, context-insensitive validation, and computational demands). Modeling decisions reflect the intended use of the models and the values that modelers prioritize. We aimed to provide a conceptual framework and machine learning (ML) modeling protocols that address these issues. With 500m-gridded hourly PM 2.5 and O 3 levels in Illinois before, during, and after the 2023 Canadian wildfire season as a motivating example, we conducted modeling experiments to evaluate modeling methods, guided by three domains we propose based on theories of science: 1) Data Diversity, leveraging open and citizen science data to enhance inclusivity, parsimony, and representativeness; 2) Equitable Accuracy, ensuring fairly distributed uncertainties across subpopulations; and 3) Sustainable Modeling, balancing accuracy with reducing computational demands to promote accessibility for under-resourced researchers. Here, we found that ML with publicly available data can achieve high accuracy. Depending on methods, performance may vary substantially, even with identical input data. Large but skewed data may reduce performance. Misuse of cross-validation protocols can underestimate prediction error; although we observed R 2 s of ∼98 %, the modeled estimates varied significantly, indicating the need for careful model validation. By using new modeling protocols including representativeness-considered training and validation data and a new loss function, we achieved high agreement between estimates and ground-based measurements (e.g., R 2 = ∼90 % for PM 2.5 ; ∼80 % for O 3 ), equally distributed errors across sociodemographic strata and urban–rural divides, and reduction in computation time—from several weeks or months to a few days.

Exposure assessment↗

Power and Communications Hardware-in-the-Loop CPS Architecture and Platform for DER Monitoring and Control Applications: Preprint

The rapid growth of distributed energy resources (DERs) has prompted increasing interest in the monitoring and control of DERs through hybrid smart grid communications. The deployment of communications and computation has transformed the traditional physical power grid into a smart cyber-physical system (CPS). To fully understand the interdependency between physical grid and cyber netowrks, this study designed a power and communications hardware-in-the-loop (PCommHIL) CPS architecture, which enables the flexible verification of DER monitoring and control with hybrid communications architectures and Internet protocols. Design, development and case study of a PCommHIL testbed for the DER coordination are discussed in detail, and the proposed platform integrates DER devices, Advanced Metering Infrastructures (AMIs), and a suite of hybrid communications networks for distribution automation applications. Case study on DER situational awareness and Volt-Var control validates the efficacy of this proposed PCommHIL platform with hybrid communications designs. Results show that the HAN communication technologies play a critical role in hybrid designs and it is the bottleneck for DER applications. High performance communication technologies are highly recommended to be applied in the HAN for enhanced monitoring and real-time control of DERs.

AMIs↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

$\mathrm{RADICAL}$-Pilot and $\mathrm{PMIx}$/$\mathrm{PRRTE}$: Executing Heterogeneous Workloads at Large Scale on Partitioned $\mathrm{HPC}$ Resources

Execution of heterogeneous workflows on high-performance computing (HPC) platforms present unprecedented resource management and execution coordination challenges for runtime systems. Task heterogeneity increases the complexity of resource and execution management, limiting the scalability and efficiency of workflow execution. Re-source partitioning and distribution of tasks execution over portioned re-sources promises to address those problems but we lack an experimental evaluation of its performance at scale. Here this paper provides a performance evaluation of the Process Management Interface for Exascale (PMIx) and its reference implementation PRRTE on the leadership-class HPC plat-form Summit, when integrated into a pilot-based runtime system called RADICAL-Pilot. We partition resources across multiple PRRTE Distributed Virtual Machine (DVM) environments, responsible for launching tasks via the PMIx interface. We experimentally measure the work-load execution performance in terms of task scheduling/launching rate and distribution of DVM task placement times, DVM startup and termination overheads on the Summit leadership-class HPC platform. Integrated solution with PMIx/PRRTE enables using an abstracted, standardized set of interfaces for orchestrating the launch process, dynamic process management and monitoring capabilities. It extends scaling capabilities allowing to overcome a limitation of other launching mechanisms (e.g., JSM/LSF). Explored different DVM setup configurations provide insights on DVM performance and a layout to leverage it. Our experimental results show that heterogeneous workload of 65,500 tasks on 2048 nodes, and partitioned across 32 DVMs, runs steady with resource utilization not lower than 52%. While having less concurrently executed tasks resource utilization is able to reach up to 85%, based on results of heterogeneous workload of 8200 tasks on 256 nodes and 2 DVMs.

97 MATHEMATICS AND COMPUTING↗

IRIS-DMEM: Efficient Memory Management for Heterogeneous Computing

This paper proposes an efficient data memory management approach for the Intelligent RuntIme System (IRIS) heterogeneous computing framework along with new data transfer policies. IRIS provides a task-based programming model for extreme heterogeneous computing (e.g., CPU, GPU, DSP, FPGA) with support for today's most important programming languages (e.g., OpenMP, OpenCL, CUDA, HIP, OpenACC). However, the IRIS framework either forces the programmer to introduce data transfer commands for each task or relies on suboptimal memory management for automatic and transparent data transfers. The work described here extends IRIS with novel heterogeneous memory handling and introduces novel data transfer policies by employing the Distributed data MEMory handler (DMEM) for efficient and optimal movement of data among the various computing resources. The proposed approach achieves performance gains of up to 7× for tiled LU factorization and tiled DGEMM (i.e., matrix multiplication) benchmarks. Moreover, this approach also reduces data transfers by up to 71% when compared to previous IRIS heterogeneous memory management handlers. This work compares the performance results of the IRIS framework's novel DMEM with the StarPU runtime and MAGMA math library for GPUs. Experiments show a performance gain of up to 1.95× over StarPU and 2.1× over MAGMA.

Miniskar, Narasinga Rao↗

A Microservices Architecture Toolkit for Interconnected Science Ecosystems

Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for inte- grating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architec- ture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability develop- ment, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.

Brim, Michael↗

Distributed Resources for the Earth System Grid Advanced Management (DREAM). Final Report

The DREAM project was funded more than 3 years ago to design and implement a next generation ESGF (Earth System Grid Federation) architecture which would be suitable for managing and accessing data and services resources on a distributed and scalable environment. In particular, the project intended to focus on the computing and visualization capabilities of the stack, which at the time were rather primitive. At the beginning, the team had the general notion that a better ESGF architecture could be built by modularizing each component, and redefining its interaction with other components by defining and exposing a well defined API. Although this was still the high level principle that guided the work, the DREAM project was able to accomplish its goals by leveraging new practices in IT that started just about 3 or 4 years ago: the advent of containerization technologies (specifically, Docker), the development of frameworks to manage containers at scale (Docker Swarm and Kubernetes), and their application to the commercial Cloud. Thanks to these new technologies, DREAM was able to improve the ESGF architecture (including its computing and visualization services) to a level of deployability and scalability beyond the original expectations.

54 ENVIRONMENTAL SCIENCES↗

A Comprehensive Analysis of PINNs for Power System Transient Stability

The integration of machine learning in power systems, particularly in stability and dynamics, addresses the challenges brought by the integration of renewable energies and distributed energy resources (DERs). Traditional methods for power system transient stability, involving solving differential equations with computational techniques, face limitations due to their time-consuming and computationally demanding nature. This paper introduces physics-informed Neural Networks (PINNs) as a promising solution for these challenges, especially in scenarios with limited data availability and the need for high computational speed. PINNs offer a novel approach for complex power systems by incorporating additional equations and adapting to various system scales, from a single bus to multi-bus networks. Our study presents the first comprehensive evaluation of physics-informed Neural Networks (PINNs) in the context of power system transient stability, addressing various grid complexities. Additionally, we introduce a novel approach for adjusting loss weights to improve the adaptability of PINNs to diverse systems. Our experimental findings reveal that PINNs can be efficiently scaled while maintaining high accuracy. Furthermore, these results suggest that PINNs significantly outperform the traditional ode45 method in terms of efficiency, especially as the system size increases, showcasing a progressive speed advantage over ode45.

97 MATHEMATICS AND COMPUTING↗