Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed Computing Resources”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Modeling Capabilities to Quantify the Benefits of DERMS (Cooperative Research and Development Final Report)

The funds-in under the CRADA will fund a team of National Renewable Energy Laboratory (NREL) researchers to participate in Energy I-Corps (formerly known as Lab-Corps). Energy I-Corps pairs teams of researchers with industry mentors for an intensive two-month training where the researchers define technology value propositions, conduct customer discovery interviews, and develop viable market pathways for their technologies. FedIMPACT, LLC and its affiliate IP Group, Inc., will evaluate the work completed at Energy I-Corps to determine whether it would like to pursue further commercialization and development of related technologies and background intellectual property.

24 POWER TRANSMISSION AND DISTRIBUTION↗

International Cybersecurity: Capabilities and Overview [Slides]

Innovations in clean energy technology are beginning to transform electric grids around the world. It is more important than ever to understand and improve the resiliency and security of the grid against natural and human disruptions as our energy systems become more distributed, intelligent, and interconnected. Through its advanced cybersecurity technical assistance portfolio, experts at the National Renewable Energy Laboratory (NREL) work with international governments to support secure and resilient deployment of renewable energy assets and address grid interconnection challenges. Cybersecurity technical assistance is tailored to the needs of our international partners.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Powered By CADET

The Capacity Expansion Decision Support for Distribution Networks (CADET) is a Python-based library and framework for creating electrical distribution system capacity planning tools for cost-effective, reliable power delivery. It enables the creation of modular, scalable, and extensible distribution capacity planning tools by providing a high-level optimization interface, parameter and options data managers, optimization constraint and objective libraries, generalized nomenclature, a system for tracking and modifying distribution network changes, optimization solution validation, and other capabilities. This webinar will describe 1) the motivation for creating CADET, 2) key designs, and 3) several use cases.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Modernizing GlideinWMS Factory Monitoring with Prometheus & Grafana

Large-scale scientific experiments like CMS and DUNE rely on the distributed workload management system GlideinWMS to efficiently utilize computing resources across heterogeneous computing environments. GlideinWMS currently records Factory statistics using Round Robin Databases (RRDBs), XML, and JSON files, and these statistics are displayed via custom monitoring Web pages, thereby limiting integration with modern observability platforms. This project investigates the use of Prometheus-based instrumentation to expose Factory metrics using OpenTelemetry principles. Factory statistics related to Glidein submission and job execution are exported as Prometheus metrics through the Prometheus Python Client Library and are served via an HTTP metrics endpoint. The collected metrics are inspected using the Prometheus web-based interface and are visualized through Grafana dashboards within the Landscape monitoring infrastructure at Fermilab. This project significantly streamlines the integration of modern monitoring technologies into GlideinWMS and establishes a framework for extending observability across additional system components.

Appiah, Gideon [Grambling State U.]↗

An Energy Service Interface for Distributed Energy Resources

Renewable energy resources, particularly wind and solar photovoltaic, are becoming significant contributors to electric power generation. These re-sources will contribute towards achieving sustainable electric power systems. However, renewable resources will dramatically increase the demand for flexible power system operations. This paper proposes an energy service interface that will allow aggregated distributed energy resources, such as residential loads and inverter-based systems, to participate in NERC-defined smart energy reliability services. Such cyber-physical systems will increase system flexibility by ensuring match between energy supply and energy demand.Aggregation and coordinated dispatch of millions of distributed energy resources will require development of large-scale computing networks. Several smart grid interface-enabling technologies, including IEEE 2030.5, Common Smart Inverter Profile, SunSpec Modbus, and CTA 2045, are discussed. Residential loads are categorized by their static and dynamic energy characteristics to identify services in which they can participate. The business model for the energy services interface as well as probabilistic modeling for resource estimation are highlighted as future considerations.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Netload Range Cost Curves for Coordinated Transmission-Distribution Planning Under DER Growth Uncertainty

The increasing penetration of distributed energy resources (DERs) requires better coordination between transmission and distribution (T&D) planning to ensure system security and cost efficiency. However, misaligned planning horizons, computational burdens, and privacy concerns hinder effective coordination, leading to either underutilized resources caused by overinvestments or reliability risks due to underinvestment. To address this challenge, we introduce netload range cost curves (NRCCs), a novel approach for managing long-term DER growth uncertainty through T&D coordination, while preserving existing data-sharing and regulatory structures. NRCCs provide pairs of (i) peak substation netload guarantees and (ii) corresponding distribution upgrade options and costs, enabling their seamless integration into transmission planning workflows. To compute NRCCs efficiently, we develop a transmission-aware distribution network planning (TADNP), which is subsequently integrated to an iterative computation procedure. These NRCCs are then embedded into an NRCC-informed transmission planning model to enable resource-efficient coordination. We illustrate our proposed approach with a case study based on realistic distribution and transmission systems in the San Francisco Bay Area, California. Our results indicate the possibility of dramatic savings in transmission investments by incorporating the proposed NRCC-integrated T&D coordination framework.

Li, Yujia↗

Continuous and Time-Domain Coherent Signal Conversion between Optical and Microwave Frequencies

A quantum network consisting of computational nodes connected by high-fidelity communication channels could expand information-processing capabilities significantly beyond those of classical networks. Superconducting qubits hold promise for scalable and high-fidelity quantum computation at microwave frequencies but must operate in an isolated cryogenic environment, obviating the potential for practical long-range communication. Quantum communication has, however, been demonstrated with optical photons. A fast efficient quantum-coherent interface between superconducting qubits and optical photons would provide a key resource for a large-scale quantum network or distributed quantum computer. Here, we describe the design and experimental operation of a device incorporating a silicon optomechanical nanobeam combined with an aluminum-nitride-based electromechanical transducer. We experimentally demonstrate classical continuous-wave operation of this device at room temperature with external conversion efficiencies of (2.5 +/- 0.4) x 10 -5 (microwave to optical) and (3.8 +/- 0.4) x 10 -5 (optical to microwave), corresponding to internal efficiencies of 2.4% and 3.7%, respectively. Finally, this device also has a larger bandwidth than previous efficient microwave-optical transducers, allowing us to operate in the time domain with 20-ns pulses.

74 ATOMIC AND MOLECULAR PHYSICS↗

WindWatts Computational Framework and Web UI [SWR-20-100]

This software provides a collection of algorithms, an API, and a functional Web UI for the Distributed Wind’s WindWatts project. DW WindWatts is a DOE WETO-funded project aimed at the development of tools to supporting the distributed wind industry, particularly with respect to siting and resource assessment. The computational framework and the back-end API are powering the easy to use front-end services available at: https://windwatts.nrel.gov. https://github.com/NREL/dw-tap-api; https://github.com/NREL/dw-tap; https://github.com/NREL/windwatts-data For reference, this software was previously known as DW TAP Computational Framework.

Phillips, Caleb↗

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

Hydropower Cybersecurity Value-at-Risk Framework

Hydropower remains one of the strongest forms of renewable energy generation methods. It is crucial to address the increasing risks associated with the rapid digitization. The push towards decarbonization also factors in the need to ensure security and resilience for grid-connected renewable energy resources. This report summarizes the U.S. Department of Energy's Water Power Technologies Office's effort to develop a cybersecurity valuation methodology that assists hydropower stakeholders in assessing risks associated with plan operations and gathers valuation guidance through a web-based application. The Hydropower Cybersecurity Value-at-Risk Framework delivers a platform for industry members to perform self-assessments and make informed decisions on their cybersecurity investments.

13 HYDRO ENERGY↗

Tools Assessing Performance

For the distributed wind industry, it can be challenging to accurately predict the performance and annual energy production of projects prior to their installation. The U.S. Department of Energy’s Tools Assessing Performance (TAP) project aims to improve wind resource characterization, thereby reducing the uncertainty of project performance and financing costs, increasing consumer confidence, and lowering the levelized cost of distributed wind energy. A collaborative effort among DOE National Laboratories, TAP will create a computational framework that provides the distributed wind community with access to newly developed wind resource data and modeling capabilities. These capabilities will allow users to perform timely and accurate performance assessments for distributed wind projects at locations across the United States.

wind, distributed, tools, performance↗

Real-Time Distributed Control of Smart Inverters for Network-level Optimization

The limitations of centralized optimization methods in managing electric power distribution systems operations have led to the distributed paradigm of computing and decision-making. Unfortunately, the existing distributed optimization algorithms are limited in their applicability to managing fast varying phenomena such as those resulting from highly variable Distributed Energy Resource (DER) generation patterns. They require a large number of communication rounds (in the order of 10 2 to 10 3 ) among the computing agents to solve one instance of the optimization problem. Related real-time distributed control methods are equally limited in their applications to power distribution systems with fast-changing DER generation; they require hundreds of rounds of communication and thus are slow in tracking the network-level optimal solutions. In this paper, we propose a novel distributed voltage controller that provides a fast-tracking of rapidly varying DER generation profiles while simultaneously converging to network-level optimal solutions within a few communication rounds. The proposed control algorithm leverages the radial topology of the system, which reduces the required communication rounds to reach the network-level optimum solution by order of magnitude. The novelty lies in carefully reducing the electrical network model from the perspective of each distributed controller and enabling appropriate data sharing among upstream and downstream nodes to achieve fast convergence. The simulation results demonstrate the effectiveness of the proposed approach in minimizing the feeder losses while maintaining the node voltage within the pre-specified limits.

voltage control, optimization, reactive power, inv↗

Towards Cross-Facility Workflows Orchestration through Distributed Automation

Modern science relies on end-to-end workflows that incorporate experimental instruments and utilize edge, cloud, or high-performance computing and storage resources. These components are geographically dispersed across various user facilities and interconnected through high-speed networks. In this paper, we present Zambeze, an automated distributed framework designed to facilitate this new class of cross-facility workflows. Utilizing swarm intelligence principles, Zambeze orchestrates science campaigns by managing distributed autonomous agents. These agents can offer a suite of services, including computing, storage, and data management. We demonstrate the feasibility of Zambeze through a real-world application involving electron microscopy, enhanced with Artificial Intelligence capabilities.

Skluzacek, Tyler↗

Perturbative readout-error mitigation for near-term quantum computers

Readout errors on near-term quantum computers can introduce significant error to the empirical probability distribution sampled from the output of a quantum circuit. These errors can be mitigated by classical postprocessing given the access of an experimental response matrix that describes the error associated with the measurement of each computational basis state. However, the resources required to characterize a complete response matrix and to compute the corrected probability distribution scale exponentially with the number of qubits, n . In this work, we modify standard matrix inversion techniques using perturbative approximations with significantly reduced complexity and bounded error when the likelihood of high-order bit-flip events is strongly suppressed. Given a characteristic error rate q , we discuss a method to recover the probability of the all-zeros bit string p 0 by sampling only a small subspace of the response matrix before inverting readout error, resulting in a relative speedup of poly [ 2 n / ( n w ) ] , which we motivate using a simplified error model for which the approximation incurs only O ( q w ) error for some integer w . We then provide a generalized technique to efficiently recover full output distributions with O ( q w ) error in the perturbative limit. These approximate techniques for readout-error correction may greatly accelerate near-term quantum computing applications.

97 MATHEMATICS AND COMPUTING↗

SuperLab 2.0 Showcase: Connecting Five Labs to Tackle Grid Complexity and Unlock Unique Grid Asset Potential

SuperLab 2.0 (5-Lab Demo) is a collaborative, national-scale experiment showcasing the coordination of geographically distributed energy assets in real time. The demonstration integrates 25 physical and digital assets, spanning wind, PV, batteries, electrolyzers, DC fast chargers, microgrid controllers, building automation systems, small modular reactor (SMR), control centers, and gas turbines, across five DOE national laboratories-NLR, INL, NETL, LBNL, and SNL. These assets are unified using Energy Sciences Network (ESnet), a low-latency, high-performance U.S. Department of Energy's (DOE) network, and controlled via a centralized energy controller hosted at NLR's ARIES facility. The demonstration validates the ability to stress-test hybrid energy systems under dynamic scenarios to de-risk advanced control strategies for greater resilience and flexibility. SuperLab 2.0 (5-Lab Demo) showcased a major advancement in federated national laboratory collaboration, enabling real-time, cross-laboratory experimentation to coordinate geographically dispersed distributed energy resources (DERs) using various communication protocols and networks. SuperLab 2.0 (5-Lab Demo) built on previous demonstrations conducted between NLR-PNNL and NLR-INL connecting diverse assets including distant protection devices, a SMR simulator, and a high temperature electrolyzer (HTE). Previous demos were based on a single connection between two labs with minimal coordination challenges. The 5-Lab demo with a centralized controller, distributed testbeds across different geographical locations, and use of protocols-based communication represents a scenario closer to real-world grid operations that coordinate resources across a region to meet system needs. This experiment studied how local DER controllers interact with a centralized energy controller during normal and abnormal events to maintain reliability. The SuperLab team across the five labs implemented a notional power system model equivalent of transmission and distribution lines, represented by the data networks interconnecting the labs. Each lab continuously exchanged local parameters (such as P and Q) from its Hardware-In-Loop (CHIL) and Power Hardware-In-Loop (PHIL) assets through centralized energy controller at NLR, enabling real-time interaction and coordination across sites. By leveraging ESnet as the communication backbone, the team successfully operated the distributed assets as a unified power system, with each bus represented by a different laboratory. This setup mirrors how assets interact in real-world power systems across dispersed locations with various protocols and latencies. At each lab site, assets were operated using their own local controllers which were coordinated through an overarching operation and control layer of centralized energy controller, equivalent to how an energy management system (EMS) orchestrates assets across a regional or national grid. SuperLab's federated connectivity utilized a Digital Real-Time Simulators (DRTS)-type gateway to connect Controller Hardware-In-Loop (CHIL) and PHIL assets between labs. To enable this federated connection through ESnet, a deterministic network was established where latency variations were consistent. This consistency allowed the development of digital filters for the power system assets across CHIL and PHIL interfaces to avoid unstable and unreliable grid conditions. This report provides an overview of the cross-laboratory configuration and offers insights into interconnecting geographically distributed research assets to test them as if they were co-located. This experiment represents a step toward linking nine DOE national laboratories, enabling nation-wide simulations that can address utility-driven challenges with grid resilience, flexibility, and modernization.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Preventive Power Outage Estimation Based on a Novel Scenario Clustering Strategy

The increasing occurrence of extreme weather events is challenging power grid operation. For extreme weather events, the system operator is responsible for estimating the power outages and scheduling the restoration resources. This paper proposes an outage evaluation framework to identify the possible unserved load profiles, vulnerable areas, and mobile energy adequacy. The outputs of an outage prediction model tool are used to generate numerous faulted line scenarios. Next, each scenario's nodal unserved load profile is obtained by solving a three-phase restoration model that considers repair crews and mobile energy resources (MERs). Then, a novel scenario clustering strategy is developed to cluster the unserved load profiles into multiple representative profiles which the system operator can focus on. Finally, case studies on a distribution system evaluate the damage caused by an extreme weather event and verify the effectiveness of the proposed scenario clustering strategy.

MATHEMATICS AND COMPUTING,POWER TRANSMISSION AND D↗