Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel and distributed computing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

PV Hosting Capacity Estimation: Experiences with Scalable Framework

Hosting capacity is an indication of the amount of solar photovoltaics (PV) that can be hosted in a distribution system without additional changes to infrastructure or oper-ations. This paper presents a framework for estimating the PV hosting capacity at scale. First, we analyze computational, modeling and other key challenges of performing relevant, large-scale simulations, provided along with the experiences and lessons learned. Then, we develop two open-source Python-based software tools to conduct repeatable distribution analyses: the Distribution Integration Solution Cost Options (DISCO) for configuring and analyzing simulations and the Job Automation and Deployment Engine (JADE) for parallelizing jobs on high-performance computing clusters. A case study of hosting capacity estimation for the SMART-DS San Francisco (SFO) 2000+ synthetic feeders, is used to demonstrate the capability of the developed DISCO+JADE framework and tools. The framework and tools can help utilities assess the overall hosting capacity of their service territory, which can help them better plan for the overall upgrade costs to integrate more PV in the future. The experiences are shared to aid the tool users and researchers to conduct relevant studies and research.

distributed energy resources↗

Situational Awareness of Grid Anomalies (SAGA)

The modern power industry becomes more vulnerable to cyber events due to the growing interconnectivity, interdependence, and complexity of the electric power grid. High-fidelity modeling and simulation tools that support the preventative risk analysis on potential cyber-relevant events is essential for ensuring the situational awareness of the system operator as it provides an inexpensive and risk-free environment to test the system responses under various cyber-relevant events and hereby can support research on cyber anomaly detection, optimal protective resource allocation, and mitigation measures. In this webinar, we will share NREL's cybersecurity research capabilities by highlighting the development of a scalable cyber-physical event test bed and demonstration with real hardware in the loop. The developed cyber-physical event test bed is backboned by an integrated transmission, distribution, and communication dynamic co-simulation framework and a plug-and-play cyber event generation module. It is designed to be modular and compatible with parallel computing, and thereby supports large-scale system simulations at an affordable computation cost. The test bed can capture millisecond-to-minutes dynamic frequency and voltage responses under cyber events from the bulk transmission system to the active distribution systems and distributed energy resources at the grid edge.

co-simulation↗

Enabling Modular Autonomous Feedback‐Loops in Materials Science through Hierarchical Experimental Laboratory Automation and Orchestration

Abstract Materials acceleration platforms (MAPs) operate on the paradigm of integrating combinatorial synthesis, high‐throughput characterization, automatic analysis, and machine learning. Within a MAP, one or multiple autonomous feedback loops may aim to optimize materials for certain functional properties or to generate new insights. The scope of a given experiment campaign is defined by the range of experiment and analysis actions that are integrated into the experiment framework. Herein, the authors present a method for integrating many actions within a hierarchical experimental laboratory automation and orchestration (HELAO) framework. They demonstrate the capability of orchestrating distributed research instruments that can incorporate data from experiments, simulations, and databases. HELAO interfaces laboratory hardware and software distributed across several computers and operating systems for executing experiments, data analysis, provenance tracking, and autonomous planning. Parallelization is an effective approach for accelerating knowledge generation provided that multiple instruments can be effectively coordinated, which the authors demonstrate with parallel electrochemistry experiments orchestrated by HELAO. Efficient implementation of autonomous research strategies requires device sharing, asynchronous multithreading, and full integration of data management in experimental orchestration, which to the best of the authors’ knowledge, is demonstrated for the first time herein.

36 MATERIALS SCIENCE↗

Implementation of a practical Markov chain Monte Carlo sampling algorithm in PyBioNetFit

Abstract Summary Bayesian inference in biological modeling commonly relies on Markov chain Monte Carlo (MCMC) sampling of a multidimensional and non-Gaussian posterior distribution that is not analytically tractable. Here, we present the implementation of a practical MCMC method in the open-source software package PyBioNetFit (PyBNF), which is designed to support parameterization of mathematical models for biological systems. The new MCMC method, am, incorporates an adaptive move proposal distribution. For warm starts, sampling can be initiated at a specified location in parameter space and with a multivariate Gaussian proposal distribution defined initially by a specified covariance matrix. Multiple chains can be generated in parallel using a computer cluster. We demonstrate that am can be used to successfully solve real-world Bayesian inference problems, including forecasting of new Coronavirus Disease 2019 case detection with Bayesian quantification of forecast uncertainty. Availability and implementation PyBNF version 1.1.9, the first stable release with am, is available at PyPI and can be installed using the pip package-management system on platforms that have a working installation of Python 3. PyBNF relies on libRoadRunner and BioNetGen for simulations (e.g. numerical integration of ordinary differential equations defined in SBML or BNGL files) and Dask.Distributed for task scheduling on Linux computer clusters. The Python source code can be freely downloaded/cloned from GitHub and used and modified under terms of the BSD-3 license (https://github.com/lanl/pybnf). Online documentation covering installation/usage is available (https://pybnf.readthedocs.io/en/latest/). A tutorial video is available on YouTube (https://www.youtube.com/watch?v=2aRqpqFOiS4&t=63s). Supplementary information Supplementary data are available at Bioinformatics online.

59 BASIC BIOLOGICAL SCIENCES↗

On the Efficient Evaluation of the Exchange Correlation Potential on Graphics Processing Unit Clusters

The predominance of Kohn–Sham density functional theory (KS-DFT) for the theoretical treatment of large experimentally relevant systems in molecular chemistry and materials science relies primarily on the existence of efficient software implementations which are capable of leveraging the latest advances in modern high-performance computing (HPC). With recent trends in HPC leading toward increasing reliance on heterogeneous accelerator-based architectures such as graphics processing units (GPU), existing code bases must embrace these architectural advances to maintain the high levels of performance that have come to be expected for these methods. In this work, we purpose a three-level parallelism scheme for the distributed numerical integration of the exchange-correlation (XC) potential in the Gaussian basis set discretization of the Kohn–Sham equations on large computing clusters consisting of multiple GPUs per compute node. In addition, we purpose and demonstrate the efficacy of the use of batched kernels, including batched level-3 BLAS operations, in achieving high levels of performance on the GPU. We demonstrate the performance and scalability of the implementation of the purposed method in the NWChemEx software package by comparing to the existing scalable CPU XC integration in NWChem.

97 MATHEMATICS AND COMPUTING↗

EQSIM—A multidisciplinary framework for fault-to-structure earthquake simulations on exascale computers, part II: Regional simulations of building response

The existing observational database of the regional-scale distribution of strong ground motions and measured building response for major earthquakes continues to be quite sparse. As a result, details of the regional variability and spatial distribution of ground motions, and the corresponding distribution of risk to buildings and other infrastructure, are not comprehensively understood. Utilizing high-performance computing platforms, emerging high-resolution, physics-based ground motion simulations can now resolve frequencies of engineering interest and provide detailed synthetic ground motions at high spatial density. This provides an opportunity for new insight into the distribution of infrastructure seismic demands and risk. In the work presented herein, the EQSIM fault-to-structure computational framework described in a companion paper, McCallen et al., is employed to investigate the regional-scale response of buildings to large earthquakes. A representative M = 7.0 strike-slip event is used to explore the distribution and amplitude of building demand, and comparisons are made between building response computed with fault-to-structure simulations and building response computed with existing measured near-fault earthquake records. New information on the distribution and variability of building response from high-performance parallel simulations is described and analyzed, and favorable first comparisons between building response predicted with both fault-to-structure simulations and real ground motions records are presented.

58 GEOSCIENCES↗

ATLAS Data Analysis using a Parallel Workflow on Distributed Cloud-based Services with GPUs

A new type of parallel workflow is developed for the ATLAS experiment at the Large Hadron Collider, that makes use of distributed computing combined with a cloud-based infrastructure. This has been developed for a specific type of analysis using ATLAS data, one popularly referred to as Simulation-Based Inference (SBI). The JAX library is used for the parts of the workflow to compute gradients as well as accelerate program execution using just-in-time compilation, which becomes essential in a full SBI analysis and can also offer significant speed-ups in more traditional types of analysis.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Laser-induced slip casting as an additive manufacturing approach for silicon carbide

Here, this work presents processing silicon carbide (SiC) with the laser-induced slip casting (LIS) additive manufacturing (AM). SiC was stabilized in water with polyethyleneimine (PEI) dispersant, and SiC slurries were made with rheology for LIS printing. High-density ceramic parts were printed, followed by single-step binder burnout and sintering. The printed parts achieved 93–95 % of theoretical density. X-ray computed tomography (XCT) revealed a small distribution of flaws exceeding 100 microns. The mechanical properties were measured in both parallel and perpendicular to the printing layers, and the orientation with layers perpendicular to the bending moment resulted in higher strength compared to the parallel direction. Porosity resulting from processing and large inclusions of boron carbide (B4C) were the root cause of failure in the measured samples. Despite these defects through this effort, this new approach demonstrates promise for green forming of SiC with densities greater than 95 % theoretical and tensile strengths above 250 MPa.

Additive Manufacturing↗

Distributed approximate minimal Steiner trees with millions of seed vertices on billion-edge graphs

In this report, we present a parallel 2-approximation Steiner minimal tree algorithm and its MPI-based distributed implementation. In place of expensive distance computations between all pairs of seed vertices, the solution we employ exploits a cheaper Voronoi cell computation. Our design leverages asynchronous processing and message prioritization to accelerate convergence of distance computations, and harnesses vertex and edge centric processing to offer fast time-to-solution. We demonstrate scalability and performance using real-world graphs with up to 128 billion edges and 512 compute nodes, and show the ability to find Steiner trees with up to one million seed vertices. Using 12 data instances, we present comparison with the state-of-the-art exact solver, SCIP-Jack, and two sequential 2-approximate algorithms. We empirically show that, on average, the total distance of the Steiner tree identified by our solution is 1.1290 times greater than the Steiner minimal tree – well within the theoretical approximation bound of 2.

97 MATHEMATICS AND COMPUTING↗

Virtual Time III, Part 2: Combining Conservative and Optimistic Synchronization

This is Part 2 of a trio of works intended to provide a unifying framework in which conservative and optimistic synchronization for parallel discrete event simulations can be freely and transparently combined in the same logical process on an event-by-event basis. Here, in this article, we continue the outline of an approach called Unified Virtual Time (UVT) that was introduced in Part 1, showing in detail via two extended examples how conservative synchronization can be refactored and combined with optimistic synchronization in the UVT framework. We describe UVT versions of both a basic time windowing algorithm called Unified Simple Time Windows and a refactored version of the Chandy-Misra-Bryant Null Message algorithm called Unified CMB.

97 MATHEMATICS AND COMPUTING↗

A Robust Parallel Distributed State Estimation for Large Scale Distribution Systems

The growing need and interest in real-time monitoring of large distribution networks motivated by the rapid population of renewable sources, EVs and etc. demand a computationally efficient state estimation framework. Furthermore, this paper presents an improved computational framework for implementing a robust state estimator using a multi-core processor. The main contribution of the paper is the proposed computational framework along with two partitioning strategies which enable fast and robust state estimation for large scale radial and/or meshed distribution systems. Formulation of the proposed method and its implementation are described in detail. Performance of the estimator is tested by simulations first using a small 84-bus radial distribution system. Then the method’s scalability is demonstrated by simulations on two very large scale distribution networks one configured radially and the other meshed each containing over 12,500 buses.

42 ENGINEERING↗

How fast can one resize a distributed file system?

Efficient resource utilization becomes a major concern as large-scale distributed computing infrastructures keep growing in size. Malleability, the possibility for resource managers to dynamically increase or decrease the amount of resources allocated to a job, is a promising way to save energy and costs. However, state-of-the-art parallel and distributed storage systems have not been designed with malleability in mind. The reason is mainly the supposedly high cost of data transfers required by resizing operations. Nevertheless, as network and storage technologies evolve, old assumptions about potential bottlenecks can be revisited. In this study, we evaluate the viability of malleability as a design principle for a distributed storage system. We specifically model the minimal duration of the commission and decommission operations. To show how our models can be used in practice, we evaluate the performance of these operations in HDFS, a relevant state-of-the-art distributed file system. We show that the existing decommission mechanism of HDFS is good when the network is the bottleneck, but can be accelerated by up to a factor 3 when storage is the limiting factor. We also show that the commission in HDFS can be substantially accelerated. With the highlights provided by our model, we suggest improvements to speed both operations in HDFS. We discuss how the proposed models can be generalized for distributed file systems with different assumptions and what perspectives are open for the design of efficient malleable distributed file systems.

97 MATHEMATICS AND COMPUTING↗

Toward designing effective exascale scientific computing workflows: experiences and best practices

Many fields within scientific computing have embraced advances in big-data analysis and machine learning, which often requires the deployment of large, distributed and complicated workflows that may combine training neural networks, performing simulations, running inference, and performing database queries and data analysis in asynchronous, parallel and pipelined execution frameworks. Such a shift has brought into focus the need for scalable, efficient workflow management solutions with reproducibility, error and provenance handling, traceability, and checkpoint-restart capabilities, among other needs. Here, we discuss challenges and best-practices for deploying exascale-generation computational science workflows on resources at the Oak Ridge Leadership Computing Facility (OLCF). We present our experiences with large-scale deployment of distributed workflows on the Summit supercomputer, including for bioinformatics and computational biophysics, materials science, and deep learning model optimization. We also present problems and solutions created by working within a Python-centric software base on traditional HPC systems, and discuss steps that will be required before the convergence of HPC, AI, and data science can be fully realized. Our results point to a wealth of exciting new possibilities for harnessing this convergence to tackle new scientific challenges.

Coletti, Mark↗

Structural Simluation Toolkit (SST) v.11.0

The SST provides a parallel framework to perform system simulation of computer architectures to determine their performance and power consumption. Additionally, the SST contains basic models of a computer processor, and interconnect and can connect to an external memory simulator (DRAMSim II). The SST framework provides a simple interface by which other computer simulation models can be combined under a common parallel discrete event-based simulation environment. This allows design exploration of future architectures, analysis of how current computer programs will function on future architectures. The SST provides a parallel discrete event simulation framework, including partitioning and object distribution over MPI. It also provides a mechanism by which components can report their power consumption for analysis.

Rodrigues, ArunF.↗

Neutron Absorber Plate Characterization Plan for Criticality Experiments Design

After being used in nuclear installations, depleted fuel can still be highly reactive and must be handled securely to prevent any radiological or criticality concerns. In particular, spent fuel from use in nuclear power reactors must be stored and transported in specifically designed containers using neutron absorber materials to prevent criticality. Various neutron absorber material types exist and are manufactured by various entities, as thoroughly described in the Handbook of Neutron Absorber Materials for Spent Nuclear Fuel Storage and Transportation Applications written by EPRI. Presently, one of the most modern and most widely used types of neutron absorber material contains particles of boron carbide, or B 4 C, embedded in aluminum matrix: Boralcan, manufactured by Rio Tinto. It is very important for the community to know as much as possible about such neutron absorber materials. Therefore, in the recent years, a US Department of Energy National Nuclear Security Administration–Nuclear Criticality Safety Program funded project initiated design of an experiment that places Boralcan neutron-absorbing plates in an established critical assembly using low-enriched uranium fuel at the Sandia Pulsed Reactor Facility/Critical Experiments (SPRF/CX) apparatus at Sandia National Laboratories. The goal of the experiment is to produce high-quality benchmark data to submit to the International Criticality Safety Benchmark Evaluation Project (ICSBEP), for use in validating calculational tools and nuclear data by criticality safety analysts. The project, named IER-554, is currently in its final design stage, following a successful preliminary design. In the work documented in the design study, ten critical configurations using Boralcan neutron absorber plates were designed, and the experiment was proven to be feasible, with a predicted low k eff uncertainty around 100 pcm. An overview of the modeled cutout of the critical assembly with a Boralcan plate is shown in Figure 1, representing one of the configurations planned for the critical experiments. Before the plates are inserted in the critical assembly, it is necessary to know more about their composition and uniformity. This summary focuses on the plate characterization plans. Each plate will undergo (1) neutron transmission measurements at different locations to determine the 10 B areal density and (2) an in-depth x-ray computed tomography (XCT) examination to obtain the exact Sizes and distribution of the B4C powder particles inside the plates. In parallel, plate modeling studies are performed with a goal to determine the validity of the currently used approximation of modeling the neutron absorber plates as a homogeneous mixture of Aluminum 1100 alloy and B4C— instead of explicitly modeling the B4C particles. By using the experimental 10 B areal density measurements, and the exact size and location of the B4C particles obtained by XCT, a plate model can theoretically be built that reproduces the plate with extremely high fidelity. The results of this modeling study could increase the confidence of the criticality safety community in its modeling methods when using this type of neutron absorber material, and the industry could use these validations to change the boron loading credit limits from the U.S. Nuclear Regulatory Commission standard review plan for dry cask storage of spent nuclear fuel. The modeling calculations are performed with SCALE 6.3.0 using the KENO V.a sequence for criticality calculations with the ENDF/B-VIII.0 continuous-energy cross section library.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Stable parallel training of Wasserstein conditional generative adversarial neural networks

In this work, we propose a stable, parallel approach to train Wasserstein conditional generative adversarial neural networks (W-CGANs) under the constraint of a fixed computational budget. Differently from previous distributed GANs training techniques, our approach avoids inter-process communications, reduces the risk of mode collapse and enhances scalability by using multiple generators, each one of them concurrently trained on a single data label. The use of the Wasserstein metric also reduces the risk of cycling by stabilizing the training of each generator. We illustrate the approach on the CIFAR10, CIFAR100, and ImageNet1k datasets, three standard benchmark image datasets, maintaining the original resolution of the images for each dataset. Performance is assessed in terms of scalability and final accuracy within a limited fixed computational time and computational resources. To measure accuracy, we use the inception score, the Fréchet inception distance, and image quality. An improvement in inception score and Fréchet inception distance is shown in comparison to previous results obtained by performing the parallel approach on deep convolutional conditional generative adversarial neural networks as well as an improvement of image quality of the new images created by the GANs approach. Weak scaling is attained on both datasets using up to 2000 NVIDIA V100 GPUs on the OLCF supercomputer Summit.

97 MATHEMATICS AND COMPUTING↗

Neglecting Model Parametric Uncertainty Can Drastically Underestimate Flood Risks

Abstract Floods drive dynamic and deeply uncertain risks for people and infrastructures. Uncertainty characterization is a crucial step in improving the predictive understanding of multi‐sector dynamics and the design of risk‐management strategies. Current approaches to estimate flood hazards often sample only a relatively small subset of the known unknowns, for example, the uncertainties surrounding the model parameters. This approach neglects the impacts of key uncertainties on hazards and system dynamics. Here we mainstream a recently developed method for Bayesian inference to calibrate a computationally expensive distributed hydrologic model. We compare three different calibration approaches: (a) stepwise line search, (b) precalibration or screening, and (c) the Fast Model Calibrations (FaMoS) approach. FaMoS deploys a particle‐based approach that takes advantage of the massive parallelization afforded by modern high‐performance computing systems. We quantify how neglecting parametric uncertainty and data discrepancy can drastically underestimate extreme flood events and risks. Precalibration improves prediction skill score over a stepwise line search. The Bayesian calibration improves the uncertainty characterization of model parameters and flood risk projections.

54 ENVIRONMENTAL SCIENCES↗

Optimizing the Weather Research and Forecasting Model with OpenMP Offload and Codee

Currently, the Weather Research and Forecasting model (WRF) utilizes shared memory (OpenMP) and distributed memory (MPI) parallelisms. To take advantage of GPU resources on the Perlmutter supercomputer at NERSC, we port parts of the computationally expensive routine Fast Spectral Bin Microphysics (FSBM) to NVIDIA GPUs using OpenMP device offloading directives. To facilitate this process, we explore a workflow for optimization which uses both runtime profilers and a static code inspection tool Codee to refactor the subroutine. We observe an 2.24x overall speedup for the CONUS-12km storm test case.

Wichitrnithed, Chayanon (Namo) [Odin Institute]↗