Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Distributed System and Computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Resilient Information Architecture Platform for Smart Grid (RIAPS)

A number of emerging trends will substantially alter the operation and control of the electric grid over the next several decades. These trends include ensuring resiliency under severe weather events, increasing integration of renewable electricity generation, supporting changing electricity demand patterns, and the improving cost effectiveness of distributed energy resources. To address these challenges, the future “Smart Grid” management will need to transition from centralized to coordinated distributed control paradigm. Reliable operation of the Smart Grid depends on distributed intelligence realized through software applications that run on distributed computing devices attached to the power system to collect data and collaboratively manage resources. However, much of the existing software for Smart Grid-enabled devices is either proprietary or developed with custom solutions, which limits interoperability among the heterogeneous devices and hinders the ability to manage system-level reliability, security, and resiliency requirements. Additionally, this approach makes Smart Grid applications hard to maintain, evolve, verify, and replace; resulting in high development and deployment costs. Further development of the Smart Grid requires a reusable software base-layer to move from hard-coded functionality to a plug-and-play architecture capable of managing system-level objectives and constraints in addition to providing consistent common services across heterogeneous devices and applications. Vanderbilt University, in collaboration with North Carolina State University and Washington State University has developed a foundation ‘software platform’ for developing and deploying robust, reliable, effective and secure software applications for the Smart Grid. The Resilient Information Architecture Platform for the Smart Grid (RIAPS) provides core services for building effective and powerful smart grid applications. It offers unique services for real-time data dissemination, fault tolerance, and coordination across apps distributed over the network.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Decentralized Distributed Proximal Policy Optimization (DD-PPO) for High Performance Computing Scheduling on Multi-User Systems

Resource allocation in High Performance Computing (HPC) environments presents a complex and multifaceted challenge for job scheduling algorithms. Beyond the efficient allocation of system resources, schedulers must account for and optimize multiple performance metrics, including job wait time and system throughput. Traditional heuristic-based scheduling algorithms increasingly struggle and lack the efficiency needed to meet the demands and address the complexity and scale of modern HPC systems. Consequently, recent research efforts have focused on leveraging advancements in Artificial Intelligence (AI) and Deep Learning (DL), particularly Reinforcement Learning (RL), to develop more adaptable and intelligent scheduling strategies. Previous RL-based scheduling approaches have explored a range of algorithms, from Deep Q-Networks (DQN) to Proximal Policy Optimization (PPO), and more recently, hybrid methods that integrate Graph Neural Networks (GNNs) with RL techniques. However, a common limitation across these methods is their reliance on relatively small datasets, with few methods being evaluated using large-scale, multi-million-job trace datasets representative of real-world HPC workloads. Moreover, existing RL schedulers face scalability issues due to centralized policy updates, which hinder training efficiency and performance when applied to large datasets. This study introduces a novel RL-based scheduler utilizing Decentralized Distributed Proximal Policy Optimization (DD-PPO) algorithm, which supports large-scale distributed training across multiple workers without requiring parameter synchronization at every step. By eliminating reliance on centralized updates to a shared policy, the DD-PPO scheduler enhances scalability, training efficiency, and sample utilization. Experimental validation using a large real-world dataset containing over 11.5 million job traces collected from petascale HPC systems over six years assesses the influence of dataset scale on training effectiveness and compares DD-PPO performance to traditional and advanced scheduling approaches. The experimental results demonstrate improved scheduling performance in comparison to both heuristic-based schedulers and existing RL-based scheduling algorithms.

AI↗

Shedding light on U.S. small and midsize data centers: Exploring insights from the CBECS survey

As demand for digital services accelerates, the energy and environmental footprint of data centers faces increasing scrutiny. While hyperscale cloud facilities have driven efficiency gains, small and midsize U.S. data centers remain a critical yet underexamined segment with significant untapped potential for energy savings. This study leverages data from the Commercial Buildings Energy Consumption Survey (CBECS) to analyze trends in server stocks, computing customers, cooling system adoption and efficiency, and geospatial distribution from 2012 to 2018. Findings reveal a sharp decline in small and midsize data centers, from 1.764 million to 1.398 million, with server counts dropping from 5.177 million to 4.262 million—aligning with the broader shift toward cloud computing. More than 40 % of servers in small data centers and 55 % in midsize data centers are housed in office buildings, and over half of all servers are concentrated in climate zones 5A (cold), 3A (mixed-humid), and 4A (mixed-humid), with the highest densities in metropolitan hubs. While direct expansion units remain the dominant cooling system, a clear transition toward more energy-efficient solutions, particularly air economizers, is evident. By integrating server and cooling system distributions, we estimate Power Usage Effectiveness (PUE) and Water Usage Effectiveness (WUE) for U.S. data centers by size and year. Results show that midsize data centers are more energy-efficient but more water-intensive due to the widespread use of water-cooled chillers. These findings highlight the trade-offs in cooling system selection and provide a critical foundation for policies aimed at enhancing efficiency in an evolving data center landscape.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

I/O in Machine Learning Applications on HPC Systems: A 360-degree Survey

Growing interest in Artificial Intelligence (AI) has resulted in a surge in demand for faster methods of Machine Learning (ML) model training and inference. This demand for speed has prompted the use of high performance computing (HPC) systems that excel in managing distributed workloads. Because data is the main fuel for AI applications, the performance of the storage and I/O subsystem of HPC systems is critical. In the past, HPC applications accessed large portions of data written by simulations or experiments or ingested data for visualizations or analysis tasks. ML workloads perform small reads spread across a large number of random files. This shift of I/O access patterns poses several challenges to modern parallel storage systems. In this paper, we survey I/O in ML applications on HPC systems, and target literature within a 6-year time window from 2019 to 2024. We define the scope of the survey, provide an overview of the common phases of ML, review available profilers and benchmarks, examine the I/O patterns encountered during offline data preparation, training, and inference, and explore I/O optimizations utilized in modern ML frameworks and proposed in recent literature. Lastly, we seek to expose research gaps that could spawn further R&D.

97 MATHEMATICS AND COMPUTING↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

Assessing the Impact of Cybersecurity Attacks on Energy Systems

This paper investigates the cyber resiliency of future power systems with high penetration of distributed energy resources using advanced distributed and (or) hierarchical control architectures. Specifically, we simulate cyberattacks on three prototypical use cases, and we identify attack scenarios that are the most damaging to the overall system performance. We show that these attacks can have a significant impact on grid operation. Results provide additional insights into the robustness of the system to the most common cyberattacks.

buildings↗

A Microservices Architecture Toolkit for Interconnected Science Ecosystems

Microservices architecture is a promising approach for developing reusable scientific workflow capabilities for inte- grating diverse resources, such as experimental and observational instruments and advanced computational and data management systems, across many distributed organizations and facilities. In this paper, we describe how the INTERSECT Open Architec- ture leverages federated systems of microservices to construct interconnected science ecosystems, review how the INTERSECT software development kit eases microservice capability develop- ment, and demonstrate the use of such capabilities for deploying an example multi-facility INTERSECT ecosystem.

Brim, Michael↗

Real-Time Distributed Control of Smart Inverters for Network-level Optimization

The limitations of centralized optimization methods in managing electric power distribution systems operations have led to the distributed paradigm of computing and decision-making. Unfortunately, the existing distributed optimization algorithms are limited in their applicability to managing fast varying phenomena such as those resulting from highly variable Distributed Energy Resource (DER) generation patterns. They require a large number of communication rounds (in the order of 10 2 to 10 3 ) among the computing agents to solve one instance of the optimization problem. Related real-time distributed control methods are equally limited in their applications to power distribution systems with fast-changing DER generation; they require hundreds of rounds of communication and thus are slow in tracking the network-level optimal solutions. In this paper, we propose a novel distributed voltage controller that provides a fast-tracking of rapidly varying DER generation profiles while simultaneously converging to network-level optimal solutions within a few communication rounds. The proposed control algorithm leverages the radial topology of the system, which reduces the required communication rounds to reach the network-level optimum solution by order of magnitude. The novelty lies in carefully reducing the electrical network model from the perspective of each distributed controller and enabling appropriate data sharing among upstream and downstream nodes to achieve fast convergence. The simulation results demonstrate the effectiveness of the proposed approach in minimizing the feeder losses while maintaining the node voltage within the pre-specified limits.

voltage control, optimization, reactive power, inv↗

Distributed Energy Resources Cybersecurity Framework & Risk Manager

Distributed Energy Resource Cybersecurity Framework (DER-CF) provides a holistic assessment for evaluating the cybersecurity posture of DER systems - filling an important gap that expands upon existing cybersecurity frameworks for more modern energy systems. The DER-CF is available as a written framework and an interactive Web tool. Distributed Energy Resource Risk Manager (DER-RM) process adheres closely to NIST's seven risk management steps: prepare, categorize, select, implement, assess, monitor, and authorize. This added feature is independent of the DER-CF's existing self-assessment and allows managers to focus on the RMF process.

cybersecurity↗

Engineering Privacy at the Edge: A Practical Guide to Differential Privacy in System Architectures

The rapid expansion of distributed and edge computing platforms—spanning autonomous vehicles, IoT sensors, and healthcare monitors—has heightened concerns about data privacy. Differential Privacy (DP) offers a rigorous mathematical framework to protect sensitive information while retaining analytical utility. This tutorial introduces the foundations of DP for both numerical and categorical datasets and extends the discussion to correlation-aware techniques tailored for structured and high-dimensional data. Hands-on demonstrations will begin with the PETINA (Privacy prEservaTIoN Algorithms) package for numerical data and continue with MIC-DP (Maximum Information Correlated Differential Privacy) for tabular data. Designed for researchers and practitioners in secure systems, embedded architectures, and AI accelerators, the tutorial emphasizes practical and scalable methods for integrating DP into real-world system designs.

Kotevska, Olivera [ORNL] (ORCID:0000000316772243)↗

Hierarchical off-diagonal low-rank approximation of Hessians in inverse problems, with application to ice sheet model initialization

Obtaining lightweight and accurate approximations of discretized objective functional Hessians in inverse problems governed by partial differential equations (PDEs) is essential to make both deterministic and Bayesian statistical large-scale inverse problems computationally tractable. The cubic computational complexity of dense linear algebraic tasks, such as Cholesky factorization, that provide a means to sample Gaussian distributions and determine solutions of Newton linear systems is a computational bottleneck at large-scale. These tasks can be reduced to log-linear complexity by utilizing hierarchical off-diagonal low-rank (HODLR) matrix approximations. In this work, we show that a class of Hessians that arise from inverse problems governed by PDEs are well approximated by the HODLR matrix format. In particular, we study inverse problems governed by PDEs that model the instantaneous viscous flow of ice sheets. In these problems, we seek a spatially distributed basal sliding parameter field such that the flow predicted by the ice sheet model is consistent with ice sheet surface velocity observations. Here, we demonstrate the use of HODLR Hessian approximation to efficiently sample the Laplace approximation of the posterior distribution with covariance further approximated by HODLR matrix compression. Computational studies are performed which illustrate ice sheet problem regimes for which the Gauss–Newton data-misfit Hessian is more efficiently approximated by the HODLR matrix format than the low-rank (LR) format. We then demonstrate that HODLR approximations can be favorable, when compared to global LR approximations, for large-scale problems by studying the data-misfit Hessian associated with inverse problems governed by the first-order Stokes flow model on the Humboldt glacier and Greenland ice sheet.

97 MATHEMATICS AND COMPUTING↗

Blueprinting Electrified Transit System Implementation

To achieve a more affordable and reliable transportation system, we need to smartly upgrade our power systems and install a large number of charging stations, but conventional planning methods are not up to the task. By applying advanced simulation and optimization tools, we can design a smarter, more cost-effective electric transportation network. The initial focus was on public transit systems, demonstrating how this approach can deliver broader economic, reliability, and air quality benefits nationwide.

24 POWER TRANSMISSION AND DISTRIBUTION↗

iDDS: intelligent distributed dispatch and scheduling for workflow orchestration

The intelligent distributed dispatch and scheduling (iDDS) service is a versatile workflow orchestration system designed for large-scale, distributed scientific computing. iDDS extends traditional workload and data management by integrating data-aware execution, conditional logic, and programmable workflows, enabling automation of complex and dynamic processing pipelines. Originally developed for the ATLAS experiment at the large hadron collider, iDDS has evolved into an experiment-agnostic platform that supports both template-driven workflows and a Function-as-a-Task model for Python-based orchestration. This paper presents the architecture and core components of iDDS, highlighting its scalability, modular message-driven design, and integration with systems such as PanDA and Rucio. We demonstrate its versatility through real-world use cases: fine-grained tape resource optimization for ATLAS, orchestration of large Directed Acyclic Graph (DAG) workflows for the Rubin Observatory, distributed hyperparameter optimization for machine learning applications, active learning for physics analyses, and AI-assisted detector design at the electron–ion collider. By unifying workload scheduling, data movement, and adaptive decision-making, iDDS reduces operational overhead and enables reproducible, high-throughput workflows across heterogeneous infrastructures. We conclude with current challenges and future directions, including interactive, cloud-native, and serverless workflow support.

97 MATHEMATICS AND COMPUTING↗

Multi-Area Model-Free State Estimation via Distributed Tensor Decomposition

This paper proposes a model-free method for distribution system state estimation based on tensor completion using canonical polyadic decomposition. In particular, we consider a setting where the network is divided into multiple areas. The measured physical quantities at buses located in the same area are processed by an area controller. A third-order tensor is constructed to collect these measured quantities. The measurements are analyzed locally to recover the full state information of the network. A closed-form iterative algorithm based on the alternating direction method of multipliers is developed to obtain the low-rank factors of the whole network state tensor where information exchange happens only between neighboring areas. To demonstrate the efficacy of the developed algorithm, numerical simulations are carried out using an IEEE test system.

alternating direction method of multipliers↗

Scalable Predictive Control and Optimization for Grid Integration of Large-Scale Distributed Energy Resources

Integrating a large number of distributed energy resources (DERs) into the power grid needs a scalable power balancing method. We formulate the power balancing problem as a look-ahead optimization problem to be solved sequentially by a power distribution system aggregator based on a model predictive control (MPC) framework. Solving large-scale look-ahead control problems requires proper configuration of the control steps. In this paper, to solve large-scale control problems, we propose a variable time granularity where control time steps nearby the current control step have finer resolutions. The aggregator objective includes maximization of power production revenue and minimization of power purchasing expense, renewable power curtailment, and mileage costs for energy storage and electric vehicle (EV) charging stations while satisfying system capacity and operational constraints. The control problem is formulated as a mixed-integer linear program (MILP) and solved using the XpressMP solver. We perform simulations considering a copper plate representation of a large distribution network consisting of 2507 devices (controllable DERs), including curtailable photovoltaics (PVs), energy storage batteries, EV charging stations, and buildings with heating, ventilation, and air conditioning units (HVACs). We show the effectiveness of the proposed approach in managing DERs interactively for maximum energy trading profit and local supply-demand power balancing. Finally, we demonstrate that the proposed method outperforms other benchmark controllers regarding computation time without compromising operational performance.

DER↗

Powered By ERAD [Slides]

Energy Resilience Analysis for Distribution Power System (ERAD) is a free, open-source Python toolkit for estimating the energy and service impacts of hazards like earthquakes and flooding. It uses a graph-based approach to capture high resolution connectivity among the grid, critical services, and customers and rapidly compute household level metrics and aggregated statistics across large distribution systems. It uses asset fragility curves that relate hazard severity to survival probability for power system equipment including cables, transformers, substations, etc. The tool is designed to be modular and extensible, allowing it to interface with third-party hazard simulators and integrate into broader resilience analysis workflows. ERAD enables researchers, students, communities, distribution utilities, and other stakeholders to understand hazard impacts and evaluate the effectiveness of different programs to improve energy resilience. The webinar was hosted by NLR researcher Aadil Latif.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Visibility-enhanced model-free deep reinforcement learning algorithm for voltage control in realistic distribution systems using smart inverters

Increasing integration of distributed solar photovoltaic (PV) into distribution networks could result in adverse effects on grid operation. Traditional model-based control algorithms require accurate model information that is difficult to acquire and thus are challenging to implement in practice. Here, this paper proposes a surrogate model-enabled grid visibility scheme to empower deep reinforcement learning (DRL) approach for distribution network voltage regulation using PV inverters with minimal system knowledge. In contrast to existing DRL methods, this paper presents and corroborates the adverse impact of missing load information on DRL performance and, based on this finding, proposes a surrogate model methodology to impute load information utilizing observable data. Additionally, a multi-fidelity neural network is utilized to construct the DRL training environment, chosen for its efficient data utilization and enhanced robustness to data uncertainty. The feasibility and effectiveness of the proposed algorithm are assessed by considering DRL testing across varying degrees of observable load information and diverse training environments on a realistic power system.

14 SOLAR ENERGY↗