Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “distributed computation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

Applying the Risk Management Framework: The Distributed Energy Resource Risk Manager

As part of a multiyear effort, the National Renewable Energy Laboratory (NREL) has dedicated resources to understand and identify cybersecurity weaknesses in distributed energy resources (DERs) by performing assessments. Due to a lack of standardization and rapidly increasing adoption of DERs, there is a critical need to address cybersecurity needs for DER systems in an interactive way. Furthermore, federal agencies, which are required to obtain an authority to operate, are challenged by the complexities of including their DERs. To help meet this need, in early 2020, NREL released the Distributed Energy Resources Cybersecurity Framework (DERCF) and accompanying Web application. This process is supported by the Risk Management Framework (RMF) developed by the National Institute of Standards and Technology. This project, referred to as the DERCF RMF application, expands on the existing DERCF work to include methods that support walking a user through the seven RMF steps. The tool will be available for download at no cost from [link ]. The purpose of this paper is to describe the steps the DERCF team at NREL took to understand Steps 1-5 of the RMF process. Additionally, this document will identify future work on the first five steps as well as a plan for Steps 6 and 7.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Lattice QCD calculation of the pion generalized parton distribution

We present the results of a Lattice QCD computation of pion generalized parton distribution (GPD), employing perturbative matching up to next-to-next-to-leading order (NNLO). The computations are based on an ensemble of Nf=2+1 highly improved staggered quarks (HISQ) with a pion mass of 300 MeV and a lattice spacing of 0.04 fm. Centered on the zero-skewness limit, we utilize a recently proposed Lorentz-invariant definition of GPD, which is derived from Lorentz-invariant amplitudes. We analyze and compare these amplitudes in both Breit and non-Breit kinematic frames at comparable momentum transfers, validating their frame-independent nature. To obtain light-cone GPD, we integrate hybrid scheme renormalization with the large momentum effective theory (LaMET). Moreover, we determine the first three iso-vector generalized form factors (GFFs) of the pion using the ratio scheme renormalization and leading-twist factorization, achieving NNLO accuracy.

Ding, Heng-Tong↗

Large-scale frictionless jamming with power-law particle size distributions

Due to significant computational expense, discrete element method simulations of jammed packings of size-dispersed spheres with size ratios greater than 1:10 have remained elusive, limiting the correspondence between simulations and real-world granular materials with large size dispersity. Here, invoking a recently developed neighbor binning algorithm, we generate mechanically stable jammed packings of frictionless spheres with power-law size distributions containing up to nearly 4 000 000 particles with size ratios up to 1:100. By systematically varying the width and exponent of the underlying power laws, we analyze the role of particle size distributions on the structure of jammed packings. The densest packings are obtained for size distributions that balance the relative abundance of large-large and small-small particle contacts. Although the proportion of rattler particles and mean coordination number strongly depend on the size distribution, the mean coordination of nonrattler particles attains the frictionless isostatic value of six in all cases. The size distribution of nonrattler particles that participate in the load-bearing network exhibits no dependence on the width of the total particle size distribution beyond a critical particle size for low-magnitude exponent power laws. This signifies that only particles with sizes greater than the critical particle size contribute to the mechanical stability. However, for high-magnitude exponent power laws, all particle sizes participate in the mechanical stability of the packing.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Posterior Covariance Matrix Approximations

Here, the Davis equation of state (EOS) is commonly used to model thermodynamic relationships for high explosive (HE) reactants. Typically, the parameters in the EOS are calibrated, with uncertainty, using a Bayesian framework and Markov Chain Monte Carlo (MCMC) methods. However, MCMC methods are computationally expensive, especially for complex models with many parameters. This paper provides a comparison between MCMC and less computationally expensive Variational methods (Variational Bayesian and Hessian Variational Bayesian) for computing the posterior distribution and approximating the posterior covariance matrix based on heterogeneous experimental data. All three methods recover similar posterior distributions and posterior covariance matrices. This study demonstrates that for this EOS parameter calibration application, the assumptions made in the two Variational methods significantly reduce the computational cost but do not substantially change the results compared to MCMC.

97 MATHEMATICS AND COMPUTING↗

Multi-Level Optimal Power Flow Solver in Large Distribution Networks

Solving optimal power flow (OPF) problems for large distribution networks incurs high computational complexity. We consider a large multi-phase distribution network of tree topology with a deep penetration of active devices. We divide the network into collaborating areas featuring subtree topology and subareas featuring subsubtree topology. We design a multilevel implementation of the primal-dual gradient algorithm to solve the voltage regulation OPF problems while preserving nodal voltage information and topological information within areas and subareas. Numerical results on a 4,521-node system verify that the proposed algorithm can significantly improve the computational speed without compromising any optimality.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Multi-Level Optimal Power Flow Solver in Large Distribution Networks: Preprint

Solving optimal power flow (OPF) problem for large distribution networks incurs high computational complexity. We consider a large multi-phase distribution networks of tree topology with deep penetration of active devices. We divide the network into collaborating areas featuring subtree topology and subareas featuring subsubtree topology. We design a multi-level implementation of the primal-dual gradient algorithm for solving the voltage regulation OPF problems while preserving nodal voltage information and topological information within areas and subareas. Numerical results on a 4,521-node system verifies that the proposed algorithm can significantly improve computational speed without compromising any optimality.

41 EE - Solar Energy Technologies Office (EE-4S)↗

Evaluation of Weakly Informed Priors for FLEX Data

On behalf of the U. S. Nuclear Regulatory Commission (NRC), Idaho National Laboratory (INL) has reviewed the development and use of “weakly informed prior” (WIP) probability distributions to obtain industry-wide failure probability and rate distributions for portable “FLEX” equipment, as documented in the Pressurized Water Reactor Owner’s Group (PWROG) report, FLEX Equipment Data Collection and Analysis (PWROG 18043-P) by MM Degonish. The review includes a description of the methods used by the PWROG analysts. It also includes an implementation of those methods for comparison purposes and an implementation of two constrained-noninformative (CN) distributions for further WIP comparisons. The report documents several issues that the review identified in the development of the WIP distributions. They relate to the quality and representativeness of the current portable “FLEX” component data, the choice of informed distributions for related installed components, and the lack of verification for the factors chosen to weaken the information from the installed components. Two errors occurred in the computations: the posterior distribution of the WIP was used to modify the WIP itself, resulting in a lack of independence between the WIP and the data; and an error was made in the conversion of WIP mean and variance values to beta distributions. The INL reviewers think the WIP distributions presented in the report are untenable for use in risk assessments at the present time because of these issues.

99 GENERAL AND MISCELLANEOUS↗

Sensing Electrical Networks Securely & Economically (SENSE)

The growing adoption of distributed energy resources (DERs) like battery energy storage systems and roof top solar/PV and the rapid penetration of electric vehicles (EVs), the electric grid is undergoing a major transformation with elevated stress on legacy grid assets. Despite a lot of expenditure to address these challenges, both in dollars and manpower, utilities have not been able to receive the value that was promised. The gains have been most visible at the transmission and substation level, especially where the main objective was improving operational and economic efficiency for the utility. Improving visibility and control at a few select points enhances the existing and established paradigm of centralized command and control. With changing load patterns, load types and the overall transition to an “active grid”, the centralized control and coordination paradigm gets challenged. To address the challenges, a new architecture and mechanism is needed, one that supports decentralized control and decision making, extracting value streams at the grid edge, particularly as the changes are fueled by transitions occurring in the distribution system. To address this, a communications and data processing platform, “GAMMA” was developed and demonstrated through the project. At the heart of the platform, are distributed, intelligent edge nodes with sensing and compute capabilities, that can record and analyze information locally. They are embedded in sensors and actuators specific to different distribution system applications. Phase 1 of the project focused on developing novel sensor technology that can be used for monitoring utility pole top distribution transformers. The sensors were designed with the objective of being low-cost, communicating with the GAMMA cloud using novel “delay-tolerant” networking using Bluetooth and a secure mobile application. They were non-intrusive in nature so that they can be installed quickly in the field, resulting in overall low cost of deployment and operations. Following the successful completion of Phase 1, the team manufactured 100 units for a field demonstration in Phase 2. The field demonstration was carried out on two real feeder systems with the local utility partner. In total, 100 sensors were installed and operated over a period of 6 months in the state of Georgia. The platform is operational end to end, with the cloud infrastructure deployed on a distributed, serverless environment that can serve multiple data streams, an analytics engine and a portal to securely view the data from multiple assets. The data collected through the GAMMA Mobile Phone app showcased the viability of the novel delay tolerant networking architecture, and the data processing algorithms developed through the course of the project, were successful in extracting important information about the overall network, improving the utility’s visibility and situational awareness in the distribution feeder.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (distributed parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

graph algorithms, high performance comptuing↗

Distributed Multi-GPU Community Detection on Exascale Computing Platforms

Community detection is a fundamental operation in graph mining, and by uncovering hidden structures and patterns within complex systems it helps solve fundamental problems pertaining to social networks, such as information diffusion, epidemics, and recommender systems. Scaling graph algorithms for massive networks becomes challenging on modern distributed-memory multi-GPU (Graphics Processing Unit) systems due to limitations such as irregular memory access patterns, load imbalances, higher communication-computation ratios, and cross-platform support. We present a novel algorithm HiPDPL-GPU (Distributed Parallel Louvain) to address these challenges. We conduct experiments involving different partitioning techniques to achieve an optimized performance of HiPDPL-GPU on the two largest supercomputers: Frontier and Summit. Remarkably, HiPDPL-GPU processes a graph with 4.2 billion edges in less than 3 minutes using 1024 GPUs. Qualitatively, the performance of HiPDPL-GPU is similar or better compared to other state-of-the-art CPU- and GPU-based implementations. While prior GPU implementations have predominantly employed CUDA, our first-of-its-kind implementation for community detection is cross-platform, accommodating both AMD and NVIDIA GPUs.

Sattar, Naw Safrin↗

UPC++ v1.0 Programmer’s Guide, Revision 2021.9.0

UPC++ is a C++ library that provides Partitioned Global Address Space (PGAS) programming. It is designed for writing parallel programs that run efficiently and scale well on distributed-memory parallel computers. The PGAS model is single program, multiple-data (SPMD), with each separate constituent process having access to local memory as it would in C++. PGAS additionally provides one-sided Remote Memory Access (RMA) to a global address space, which is allocated in shared segments that are distributed over the processes. UPC++ also features Remote Procedure Call (RPC) communication, making it easy to move computation to operate on data that resides on remote processes. In UPC++, all communication operations are explicit, which encourages programmers to be aware of the cost of communication and data movement. Moreover, all communication operations are asynchronous by default, to enable programmers to write code that scales well even on hundreds of thousands of cores.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Building an Integrated Ecosystem of Computational and Observational Facilities to Accelerate Scientific Discovery

Future scientific discoveries will rely on flexible ecosystems that incorporate modern scientific instruments, high performance computing resources, parallel distributed data storage, and performant networks across multiple, independent facilities. In addition to connecting physical resources, such an ecosystem presents many challenges in logistics and accessibility, especially in orchestrating computations and experiments that span across leadership computing systems and experimental instruments. Past efforts have typically been application-specific or limited to interfaces for computing resources. This paper proposes a general framework for integrating computation resources and instrument operations, addressing challenges in code development/execution, data staging and collection, software stack, control mechanisms, resource authorization and governance, and hardware integration. We also describe a demonstration use case wherein a Bayesian optimization algorithm running on an edge computing resource guides a scanning probe microscope to autonomously and intelligently characterize a material sample. This science edge ecosystem framework will provide a blueprint for federating multi-institutional, disparate resources and orchestrating scientific workflows across them to enable next-generation discoveries.

Somnath, Suhas↗

Modular thermal energy storage system

A thermal energy storage (TES) system includes a plurality of closely packed TES modules, each TES module having a shell enclosing a plurality of sealed tubes that each contain a TES media. A computer-controlled flow control system includes a flow distributor, for example a flow distributor having a plenum configured to receive a heat transfer fluid (HTF), and a plurality of control valves controlled by the computer to controllably distribute the HTF from the plenum to the plurality of TES modules. Sensor data from the TES modules, for example temperature, pressure, and/or flow data, is provided to the computer. In some embodiments the plenum includes two or more compartments with separate HTF flow ports, which may be provided to the controller at different temperatures.

Wirz, Richard E.↗

True Load Balancing for Matricized Tensor Times Khatri-Rao Product

MTTKRP is the bottleneck operation in algorithms used to compute the CP tensor decomposition. For sparse tensors, utilizing the compressed sparse fibers (CSF) storage format and the CSF-oriented MTTKRP algorithms is important for both memory and computational efficiency on distributed-memory architectures. Existing intelligent tensor partitioning models assume the computational cost of MTTKRP to be proportional to the total number of nonzeros in the tensor. However, this is not the case for the CSF-oriented MTTKRP on distributed-memory architectures. We outline two deficiencies of nonzero-based intelligent partitioning models when CSF-oriented MTTKRP operations are performed locally: failure to encode processors' computational loads and increase in total computation due to fiber fragmentation. We focus on existing fine-grain hypergraph model and propose a novel vertex weighting scheme that enables this model encode correct computational loads of processors. We also propose to augment the fine-grain model by fiber nets for reducing the increase in total computational load via minimizing fiber fragmentation. In this way, the proposed model encodes minimizing the load of the bottleneck processor. In conclusion, parallel experiments with real-world sparse tensors on up to 1024 processors prove the validity of the outlined deficiencies and demonstrate the merit of our proposed improvements in terms of parallel runtimes.

97 MATHEMATICS AND COMPUTING↗

MILK : a Python scripting interface to MAUD for automation of Rietveld analysis

Modern diffraction experiments ( e.g. in situ parametric studies) present scientists with many diffraction patterns to analyze. Interactive analyses via graphical user interfaces tend to slow down obtaining quantitative results such as lattice parameters and phase fractions. Furthermore, Rietveld refinement strategies ( i.e. the parameter turn-on-off sequences) tend to be instrument specific or even specific to a given dataset, such that selection of strategies can become a bottleneck for efficient data analysis. Managing multi-histogram datasets such as from multi-bank neutron diffractometers or caked 2D synchrotron data presents additional challenges due to the large number of histogram-specific parameters. To overcome these challenges in the Rietveld software Material Analysis Using Diffraction ( MAUD ), the MAUD Interface Language Kit ( MILK ) is developed along with an updated text batch interface for MAUD . The open-source software MILK is computer-platform independent and is packaged as a Python library that interfaces with MAUD . Using MILK , model selection ( e.g. various texture or peak-broadening models), Rietveld parameter manipulation and distributed parallel batch computing can be performed through a high-level Python interface. A high-level interface enables analysis workflows to be easily programmed, shared and applied to large datasets, and external tools to be integrated with MAUD . Through modification to the MAUD batch interface, plot and data exports have been improved. The resulting hierarchical folders from Rietveld refinements with MILK are compatible with Cinema: Debye–Scherrer , a tool for visualizing and inspecting the results of multi-parameter analyses of large quantities of diffraction data. In this manuscript, the combined Python scripting and visualization capability of MILK is demonstrated with a quantitative texture and phase analysis of data collected at the HIPPO neutron diffractometer.

97 MATHEMATICS AND COMPUTING↗

Communication Lower Bounds and Optimal Algorithms for Symmetric Matrix Computations

In this article, we focus on the communication costs of three symmetric matrix computations: (i) multiplying a matrix with its transpose, known as a symmetric rank-k update (SYRK) (ii) adding the result of the multiplication of a matrix with the transpose of another matrix and the transpose of that result, known as a symmetric rank-2k update (SYR2K) (iii) performing matrix multiplication with a symmetric input matrix (SYMM). All three computations appear in the Level 3 Basic Linear Algebra Subroutines (BLAS) and have wide use in applications involving symmetric matrices. We establish communication lower bounds for these kernels using sequential and distributed-memory parallel computational models, and we show that our bounds are tight by presenting communication-optimal algorithms for each setting. Our lower bound proofs rely on applying a geometric inequality for symmetric computations and analytically solving constrained nonlinear optimization problems. As a result, the symmetric matrix and its corresponding computations are accessed and performed according to a triangular block partitioning scheme in the optimal algorithms.

Al Daas, Hussam [Rutherford Appleton Laboratory, D↗

Tensor network simulations of quasi-GPDs in the massive Schwinger model

Generalized parton distribution functions (GPDs) are off-diagonal light-cone matrix elements that encode the internal structure of hadrons in terms of quark and gluon degrees of freedom. In this work, we present the first nonperturbative study of quasi-GPDs in the massive Schwinger model, quantum electrodynamics in 1+1 dimensions (QED 2 ), within the Hamiltonian formulation of lattice field theory. Quasidistributions are spatial correlation functions of boosted states, which approach the relevant light-cone distributions in the luminal limit. Using tensor networks, we prepare the first excited state in the strongly coupled regime and boost it to close to the light-cone on lattices of up to 400 lattice sites. We compute both quasiparton distribution functions and, for the first time, quasi-GPDs, and study their convergence for increasingly boosted states. In addition, we perform analytic calculations of GPDs in the two-particle Fock-space approximation and in the Reggeized limit, providing qualitative benchmarks for the tensor network results. Our analysis establishes computational benchmarks for accessing partonic observables in low-dimensional gauge theories, offering a starting point for future extensions to higher dimensions, non-Abelian theories, and quantum simulations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

MixPI: Mixed-time slicing path integral software for quantized molecular dynamics simulations

We introduce the MixPI software to implement path integral molecular dynamics (PIMD) simulations for the study of condensed phase systems where nuclear quantum effects (NQEs) are important. In contrast to existing PIMD simulation software, MixPI enables the implementation of mixed quantum–classical path integral simulations where only a subset of system degrees of freedom (dofs) are treated quantum mechanically in an extended phase space while the remaining dofs are described classically. We expect this software to be particularly useful for simulations of electron and proton transfer in condensed phase systems, as well as for the study of biological and material systems where only a handful of dofs contribute significantly to the observed NQEs. We demonstrate the use of MixPI in two different systems. The first is a simple water model where we implement a set of mixed quantum–classical simulations to compute average energy and radial distribution functions. We use these simulations to benchmark the effectiveness of MixPI and to demonstrate how it enables systematic investigation into the origin of observed NQEs. We then compute radial distribution functions for a system where MixPI is essential: a solvated metal (M 2+ ) cation described using an explicit quantized electron localized on an M 3+ ion in water.

chemical physics↗