Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed algorithm”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Operating a commercial building HVAC load as a virtual battery through airflow control

Virtual battery (VB) is an innovative method to model flexibility of building loads and effectively coordinate them with other resources at a system level. Unlike a real battery with a dedicated power conversion system for charging control, methods are required for operating building loads to deviate from the baseline to respond to grid signals. This paper presents a VB control for a commercial heating, ventilation, and air conditioning (HVAC) system to follow the desired power consumption in real-time by adjusting zonal airflow rates. The proposed method consists of two parts. At the system level, a mixed feedforward and feedback control is used to estimate the desired total airflow rate. At the zone level, two priority-based algorithms are then proposed to distribute the total airflow rate to individual zones. In particular, a zonal airflow limit estimation method is proposed using machine-learning techniques, in contrast to physics-based thermal models in existing studies, to more accurately capture zonal thermal dynamics and improve temperature control performance. An office building on the Pacific Northwest National Laboratory campus is implemented in EnergyPlus, and used to illustrate and validate the proposed control.

Wang, Jiyu↗

Hierarchical Optimal Power Flow with Improved Gradient Evaluation

Existing algorithms to solve alternating-current optimal power flow (AC-OPF) often exploit linear approximations to simplify system models and accelerate computations. In this paper, we improve a recent hierarchical OPF algorithm, which rested on primal-dual gradients evaluated in a linearized distribution power flow model. Specifically, we identify a risk of voltage violation arising from the model linearization, and propose a more accurate gradient evaluation method to eliminate that risk. We further develop a hierarchical primal-dual algorithm to solve OPF based on the proposed gradient evaluation method. Numerical results on IEEE networks show that our algorithm can enhance voltage safety with satisfactory computational efficiency.

distributed algorithm↗

Parallel sorting algorithm classification: is manual instrumentation necessary?

Understanding parallel algorithms is crucial for accelerating scientific simulations on complex, distributed memory, high-performance computers. Modern algorithm classification approaches learn semantics directly from source code to differentiate between algorithms, however, accessing source code is not always possible. We can learn about parallel algorithms from observing their performance, as programs running the same algorithms and using the same hardware should exhibit similar performance characteristics. We present an approach to learn algorithm classes from parallel performance data directly in order to classify algorithms without access to the source code. We extend previous work to enable classifying parallel sorting algorithms using automatic instrumentation instead of requiring manual region annotations in the source code. In this work, we design and demonstrate a study for classification of parallel sorting algorithms using parallel performance data collected from automatic instrumentation, and evaluate the performance of our new methodology on classification. We leverage Caliper to collect the performance data, Thicket for our exploratory data analysis (EDA), and PyTorch and Scikit-learn to evaluate the effectiveness of random forests, support vector machines (SVMs), decision trees, neural networks, and logistic regressions on parallel performance data. Additionally, we study noise in parallel performance data, whether the removal of noise and pre-processing of the data is necessary to accurately classify parallel sorting algorithms, and determine the effectiveness of features created from performance data. In conclusion, we demonstrate classification accuracy for these five different models of up to 97.7% across four different parallel algorithm classes.

Algorithm Classification↗

HTR-1.3 solver: Predicting electrified combustion using the hypersonic task-based research solver

Here this manuscript presents an updated open-source version of the Hypersonics Task-based Research (HTR) solver. The solver, whose main features are presented in Di Renzo et al. (2020) and Di Renzo & Pirozzoli (2021), is designed for direct numerical simulation of reacting flows at high Reynolds numbers. This new version extends the applications of the HTR solver to turbulent combustion in the presence of external electric fields. In particular, a new distributed Poisson solver compatible with heterogeneous architectures has been incorporated in the algorithm to compute the electric potential distribution in bi-periodic configurations. The drift fluxes of the electrically charged species are now included in the transport equations using a targeted essentially non-oscillatory scheme. A verification of these new features of the solver is provided using one-dimensional burner stabilized flames, whereas a three dimensional turbulent flame is utilized to discuss the scalability of the proposed numerical tool.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Understanding kinetics of defect annihilation in chemoepitaxy-directed self-assembly.

Directed self-assembly (DSA) of block copolymers (BCP) has attracted considerable interest from the semiconductor industry because it can achieve semiconductor-relevant structures with a relatively simple process and low cost. However, the self-assembling structures can become kinetically trapped into defective states, which greatly impedes the implementation of DSA in high-volume manufacturing. Understanding the kinetics of defect annihilation is crucial to optimizing the process and eventually eliminating defects in DSA. Such kinetic experiments, however, are not commonly available in academic laboratories. To address this challenge, we perform a kinetic study of chemoepitaxy DSA in a 300 mm wafer fab, where the complete defectivity information at various annealing conditions can be readily captured. Through extensive statistical analysis, we reveal the statistical model of defect annihilation in DSA for the first time. The annihilation kinetics can be well described by a power law model, indicating that all dislocations can be removed by sufficiently long annealing time. We further develop image analysis algorithms to analyze the distribution of dislocation size and configurations and discover that the distribution stays relatively constant over time. The defect distribution is determined by the role of the guiding stripe, which is found to stabilize the defects. Although this study is based on polystyrene-b-poly(methyl methacrylate) (PS-b-PMMA), we anticipate that these findings can be readily applied to other BCP platforms as well.

block copolymer↗

Non-parametric projections of the net-income distribution for all U.S. states for the Shared Socioeconomic Pathways

Abstract Income distributions are a growing area of interest in the examination of equity impacts brought on by climate change and its responses. Such impacts are especially important at subnational levels, but projections of income distributions at these levels are scarce. Here, we project U.S. state-level income distributions for the Shared Socioeconomic Pathways (SSPs). We apply a non-parametric approach, specifically a recently developed principal components algorithm to generate net income distributions for deciles across 50 U.S. states and the District of Columbia. We produce these projections to 2100 for three SSP scenarios in combination with varying projections of GDP per capita to represent a wide range of possible futures and uncertainties. In the generation of these scenarios, we also generated tax adjusted historical deciles by U.S. states, which we used for validating model performance. Our method thus produces income distributions by decile for each state, reflecting the variability in state income, population, and tax regimes. Our net income projections by decile can be used in both emissions- and impact-related research to understand distributional effects at various income levels and identify economically vulnerable populations.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Robust sampling for weak lensing and clustering analyses with the Dark Energy Survey

Recent cosmological analyses rely on the ability to accurately sample from high-dimensional posterior distributions. A variety of algorithms have been applied in the field, but justification of the particular sampler choice and settings is often lacking. Here, we investigate three such samplers to motivate and validate the algorithm and settings used for the Dark Energy Survey (DES) analyses of the first 3 yr (Y3) of data from combined measurements of weak lensing and galaxy clustering. We employ the full DES Year 1 likelihood alongside a much faster approximate likelihood, which enables us to assess the outcomes from each sampler choice and demonstrate the robustness of our full results. We find that the ellipsoidal nested sampling algorithm multinest reports inconsistent estimates of the Bayesian evidence and somewhat narrower parameter credible intervals than the sliced nested sampling implemented in polychord. We compare the findings from multinest and polychord with parameter inference from the Metropolis–Hastings algorithm, finding good agreement. We determine that polychord provides a good balance of speed and robustness for posterior and evidence estimation, and recommend different settings for testing purposes and final chains for analyses with DES Y3 data. Our methodology can readily be reproduced to obtain suitable sampler settings for future surveys.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Fast and efficient identification of anomalous galaxy spectra with neural density estimation

ABSTRACT Current large-scale astrophysical experiments produce unprecedented amounts of rich and diverse data. This creates a growing need for fast and flexible automated data inspection methods. Deep learning algorithms can capture and pick up subtle variations in rich data sets and are fast to apply once trained. Here, we study the applicability of an unsupervised and probabilistic deep learning framework, the probabilistic auto-encoder, to the detection of peculiar objects in galaxy spectra from the SDSS survey. Different to supervised algorithms, this algorithm is not trained to detect a specific feature or type of anomaly, instead it learns the complex and diverse distribution of galaxy spectra from training data and identifies outliers with respect to the learned distribution. We find that the algorithm assigns consistently lower probabilities (higher anomaly score) to spectra that exhibit unusual features. For example, the majority of outliers among quiescent galaxies are E+A galaxies, whose spectra combine features from old and young stellar population. Other identified outliers include LINERs, supernovae, and overlapping objects. Conditional modelling further allows us to incorporate additional information. Namely, we evaluate the probability of an object being anomalous given a certain spectral class, but other information such as metrics of data quality or estimated redshift could be incorporated as well. We make our code publicly available.

Böhm, Vanessa↗

MetallData

MetallData is an HPC platform for interactive data science applications at HPC-scales. It provides an ecosystem for persistent distributed data structures, including algorithms, interactivity and storage.

Pearce, RogerA↗

Fed-DeepONet: Stochastic Gradient-Based Federated Training of Deep Operator Networks

The Deep Operator Network (DeepONet) framework is a different class of neural network architecture that one trains to learn nonlinear operators, i.e., mappings between infinite-dimensional spaces. Traditionally, DeepONets are trained using a centralized strategy that requires transferring the training data to a centralized location. Such a strategy, however, limits our ability to secure data privacy or use high-performance distributed/parallel computing platforms. To alleviate such limitations, in this paper, we study the federated training of DeepONets for the first time. That is, we develop a framework, which we refer to as Fed-DeepONet, that allows multiple clients to train DeepONets collaboratively under the coordination of a centralized server. To achieve Fed-DeepONets, we propose an efficient stochastic gradient-based algorithm that enables the distributed optimization of the DeepONet parameters by averaging first-order estimates of the DeepONet loss gradient. Then, to accelerate the training convergence of Fed-DeepONets, we propose a moment-enhanced (i.e., adaptive) stochastic gradient-based strategy. Finally, we verify the performance of Fed-DeepONet by learning, for different configurations of the number of clients and fractions of available clients, (i) the solution operator of a gravity pendulum and (ii) the dynamic response of a parametric library of pendulums.

Moya, Christian↗

Protein Conformational States—A First Principles Bayesian Method

Automated identification of protein conformational states from simulation of an ensemble of structures is a hard problem because it requires teaching a computer to recognize shapes. We adapt the naïve Bayes classifier from the machine learning community for use on atom-to-atom pairwise contacts. The result is an unsupervised learning algorithm that samples a ‘distribution’ over potential classification schemes. We apply the classifier to a series of test structures and one real protein, showing that it identifies the conformational transition with >95% accuracy in most cases. A nontrivial feature of our adaptation is a new connection to information entropy that allows us to vary the level of structural detail without spoiling the categorization. This is confirmed by comparing results as the number of atoms and time-samples are varied over 1.5 orders of magnitude. Further, the method’s derivation from Bayesian analysis on the set of inter-atomic contacts makes it easy to understand and extend to more complex cases.

97 MATHEMATICS AND COMPUTING↗

Exploring the Landscape of Distributed Graph Clustering on Leadership Supercomputers

The rapid growth of large-scale datasets in fields like biology and social networks has driven the need for advanced graph analytics techniques. Community detection, a fundamental task in graph analytics, identifies closely connected groups of nodes within a network, providing valuable insights across various disciplines. This study focuses on two classic community detection methods, the Louvain algorithm and Markov Clustering (MCL), and evaluates the performance of two prominent distributed community detection algorithms: HiPDPL-GPU, our prior implementation, and HipMCL. We conduct experiments on GPU-accelerated heterogeneous HPC systems, Summit and Frontier, to assess their performance under varying conditions. Our objective is to identify the strengths and weaknesses of these algorithms in terms of scalability, and quality of solutions. We evaluate these algorithms on a diverse set of 70+ networks spanning 13 domains, with sizes ranging up to 4.2 billion edges. Our results demonstrate that HiPDPL-GPU consistently outperforms HipMCL, especially for large-scale networks. HiPDPL-GPU achieves significantly faster runtimes (47x to 1439x), higher modularity scores, and improved scalability. These findings highlight HiPDPL-GPU as a promising solution for efficient and effective large-scale graph analytics in diverse application domains, and provide insights into the feasibility of using MCL-based approaches for certain application domains.

Community detection, graph algorithms↗

Autonomous Intelligent Charging/Discharging of Electric Vehicles using Distributed Multi-Agent ADMM Framework for Grid Ancillary Services

The increasing popularity of Electric Vehicles (EVs) in the distribution grid along with technological advancement in EV electronics such as vehicle to grid (V2G) technique has enabled them to participate in grid ancillary services. To achieve this, the EVs need to establish a contract with third-party aggregators and connect to a charging unit, either residential or commercial. At any time they are connected, the EVs can decide to take part in the ancillary services program offered to them by the aggregators. If agreed, the aggregators will use the EVs as a power source capable of charging/discharging power according to the input signal, and in return, they will be compensated. This inter-temporal nature of charging/discharging is also transforming the traditional optimal power flow (OPF) problem into a dynamic OPF problem. This chapter aims at developing a multi-layer time-dependent optimization algorithm to utilize EV potential and provide ancillary services while maximizing its utilization function. Specifically, in the upper layer, an autonomous distributed ADMM algorithm is developed to optimize the cost for charging/discharging EVs while using them to regulate the voltage at each bus in the distribution grid. The distributed ADMM algorithm is also expanded to the lower layer where the individual EVs active and reactive power is controlled for voltage regulation while maintaining the desired state of the charge of the vehicle at the end of the charging period. Here, the effectiveness and performance improvement of the proposed multi-layer algorithm is illustrated through analytical analysis and simulation results.

Rahman, Towfiq↗

Autonomous Intelligent Charging/Discharging of Electric Vehicles using Distributed Multi-Agent ADMM Framework for Grid Ancillary Services

The increasing popularity of Electric Vehicles (EVs) in the distribution grid along with technological advancement in EV electronics such as vehicle to grid (V2G) technique has enabled them to participate in grid ancillary services. To achieve this, the EVs need to establish a contract with third-party aggregators and connect to a charging unit, either residential or commercial. At any time they are connected, the EVs can decide to take part in the ancillary services program offered to them by the aggregators. If agreed, the aggregators will use the EVs as a power source capable of charging/discharging power according to the input signal, and in return, they will be compensated. This inter-temporal nature of charging/discharging is also transforming the traditional optimal power flow (OPF) problem into a dynamic OPF problem. This chapter aims at developing a multi-layer time-dependent optimization algorithm to utilize EV potential and provide ancillary services while maximizing its utilization function. Specifically, in the upper layer, an autonomous distributed ADMM algorithm is developed to optimize the cost for charging/discharging EVs while using them to regulate the voltage at each bus in the distribution grid. The distributed ADMM algorithm is also expanded to the lower layer where the individual EVs active and reactive power is controlled for voltage regulation while maintaining the desired state of the charge of the vehicle at the end of the charging period. The effectiveness and performance improvement of the proposed multi-layer algorithm is illustrated through analytical analysis and simulation results.

Rahman, Towfiq↗

Electromagnetic Transient Simulation Algorithms for Evaluation of Large-Scale Extreme Fast Charging Systems (Distribution Grid Models)

The distribution and transmission grids are observing an increased penetration of power electronics in loads and generations. For example, there is increasing interest in integrating in extreme fast charging (XFC) systems for fast charging of electrical vehicles. As these systems are integrated, developing high-fidelity electromagnetic transient model of XFC systems in distribution grids and evaluating their interactions with the power grid would be of significant interest. This model will be utilized for design of XFC systems, to identify upgrades in distribution and/or transmission grids, for planning purposes by transmission planners or operators or owners, among others. It can also be utilized in operations for improved reliable performance of the grid and/or XFC station. The challenge with simulating these models is the high computational complexity introduced by the large number of states present in the system and the time-step needed to simulate the system. In this paper, advanced simulations algorithms are applied to reduce the computational complexity of simulating large-scale XFC systems. The algorithms include numerical stiffness-based segregation, time constant-based segregation, clustering and aggregation on differential algebraic equations (DAEs), and multi-order integration approaches. While the first three algorithms split the matrix that needs to be inverted from a large matrix to much smaller matrices, the final algorithm reduces the computational burden of applying higher-order integration approaches in the complete system. The comparison made in the previous sentence is with respect to use of homogeneous integration approaches used in conventional electromagnetic transient simulators like power systems computer aided design (PSCAD). The approaches mentioned here have resulted in speed-up of 36x in the simulation of a single distribution system with 15 XFCs.

Debnath, Suman↗

ConKer: An algorithm for evaluating correlations of arbitrary order

Context. High order correlations in the cosmic matter density have become increasingly valuable in cosmological analyses. However, computing these correlation functions is computationally expensive. Aims. We aim to circumvent these challenges by developing a new algorithm called ConKer for estimating correlation functions. Methods. This algorithm performs convolutions of matter distributions with spherical kernels using FFT. Since matter distributions and kernels are defined on a grid, it results in some loss of accuracy in the distance and angle definitions. We study the algorithm setting at which these limitations become critical and suggest ways to minimize them. Results. ConKer is applied to the CMASS sample of the SDSS DR12 galaxy survey and corresponding mock catalogs, and is used to compute the correlation functions up to correlation order n = 5. We compare the n = 2 and n = 3 cases to traditional algorithms to verify the accuracy of the new algorithm. We perform a timing study of the algorithm and find that three of the four distinct processes within the algorithm are nearly independent of the catalog size N , while one subdominant component scales as O ( N ). The dominant portion of the calculation has complexity of O ( N c 4/3 log N c ), where N c is the of cells in a three-dimensional grid corresponding to the matter density. Conclusions. We find ConKer to be a fast and accurate method of probing high order correlations in the cosmic matter density, then discuss its application to upcoming surveys of large-scale structure.

79 ASTRONOMY AND ASTROPHYSICS↗

Economic Dispatch With Distributed Energy Resources: Co-Optimization of Transmission and Distribution Systems

The increasing penetration of distributed energy resources (DERs) in the distribution networks has turned the conventionally passive load buses into active buses that can provide grid services for the transmission system. To take advantage of the DERs in the distribution networks, this letter formulates a transmission-and-distribution (T&D) systems co-optimization problem that achieves economic dispatch at the transmission level and optimal voltage regulation at the distribution level by leveraging large generators and DERs. A primal-dual gradient algorithm is proposed to solve this optimization problem jointly for T&D systems, and a distributed market-based equivalent of the gradient algorithm is used for practical implementation. Finally, the results are corroborated by numerical examples with the IEEE 39-Bus system connected with 7 different distribution networks.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Spatio-Temporal Deep Graph Network for Event Detection, Localization, and Classification in Cyber-Physical Electric Distribution System

This work proposes a deep graph learning framework to identify, locate, and classify power, cyber, and cyber power events at the distribution system level. The proposed algorithm jointly exploits spatial, temporal, and node-level cyber and physical data features. The developed graph neural network, together with a deep autoencoder, utilizes physical measurements from distribution level phasor measurement units and cyber data from communication network logs. The spatial structure of the synchrophasor measurements and network is incorporated through a weighted adjacency matrix. The temporal structure is incorporated by defining a spatial operation in the gated recurrent unit. This spatio-temporal learning element resides inside a power event detection, localization, and classification module that provides the degree of confidence for an event label. To accurately pinpoint the location of an event to the nearest bus equipped with a measurement unit, a combination of squared error and proximity score is utilized. Also included is a cyber event detection module that employs heteroskedasticity to analyze the significance of various cyber features during different types of attacks. Finally, a dual-bit cyber-power decision table determines the nature of the event. The proposed method is validated on two distribution systems modeled in OPAL-RT/Hypersim with limited phasor measurement units for different possible physical and cyber events. Further analyses include comparison with other state-of-the-art methods and validation in the presence of measurement noise. As a result, our method outperforms existing approaches and achieves an average detection accuracy of 97.97%, F1-score of 96.88%, precision of 96.53%, and recall of 98.57%.

24 POWER TRANSMISSION AND DISTRIBUTION↗