Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “massive parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

Massively parallel, computationally guided design of a proenzyme

Proteins have shown promise as therapeutics and diagnostics, but their effectiveness is limited by our inability to spatially target their activity. To overcome this limitation, we developed a computationally guided method to design inactive proenzymes or zymogens, which are activated through cleavage by a protease. Since proteases are differentially expressed in various tissues and disease states, including cancer, these proenzymes could be targeted to the desired microenvironment. We tested our method on the therapeutically relevant protein carboxypeptidase G2 (CPG2). We designed Pro-CPG2s that are inhibited by 80 to 98% and are partially to fully reactivatable following protease treatment. The developed methodology, with further refinements, could pave the way for routinely designing protease-activated protein-based therapeutics and diagnostics that act in a spatially controlled manner. Confining the activity of a designed protein to a specific microenvironment would have broad-ranging applications, such as enabling cell type-specific therapeutic action by enzymes while avoiding off-target effects. While many natural enzymes are synthesized as inactive zymogens that can be activated by proteolysis, it has been challenging to redesign any chosen enzyme to be similarly stimulus responsive. Here, we develop a massively parallel computational design, screening, and next-generation sequencing-based approach for proenzyme design. For a model system, we employ carboxypeptidase G2 (CPG2), a clinically approved enzyme that has applications in both the treatment of cancer and controlling drug toxicity. Detailed kinetic characterization of the most effectively designed variants shows that they are inhibited by ∼80% compared to the unmodified protein, and their activity is fully restored following incubation with site-specific proteases. Introducing disulfide bonds between the pro- and catalytic domains based on the design models increases the degree of inhibition to 98% but decreases the degree of restoration of activity by proteolysis. A selected disulfide-containing proenzyme exhibits significantly lower activity relative to the fully activated enzyme when evaluated in cell culture. Structural and thermodynamic characterization provides detailed insights into the prodomain binding and inhibition mechanisms. The described methodology is general and could enable the design of a variety of proproteins with precise spatial regulation.

59 BASIC BIOLOGICAL SCIENCES↗

Massively Parallel Bayesian Model Calibration and Uncertainty Quantification with Applications to Nuclear Fuels and Materials

The U.S. Department of Energy (DOE)’s Nuclear Energy Advanced Modeling and Simulation (NEAMS) program aims to develop predictive capabilities by applying computational methods to the analysis and design of advanced reactor and fuel cycle systems. This program has been providing engineering-scale support for the development of BISON, a high-fidelity and high-resolution fuel performance tool. Fuel behavior in a nuclear reactor is governed by a complex network of mechanisms interacting with various other physics aspects in the reactor system. Any model developed to represent the fuel behavior will likely be idealized resulting in uncertainties in their predictions compared to the observed data. As such, this report was motivated by the need to identify the sources of uncertainties and quantify and propagate them through the fuel model outputs. Such quantification of uncertainties will establish a level of model trustworthiness, identify approaches to improve the model trustworthiness, and even guide optimal experiment design for maximal information gain. To accomplish the uncertainty quantification for computational models, this report has relied on the Bayesian framework which provides probabilistic treatment of models their inputs and outputs. The current state-of-the-art on performing Bayesian Uncertainty Quantification (UQ) for nuclear engineering models using High Performance Computing (HPC) resources have been reviewed. Implementation of capabilities for massively parallel Bayesian UQ in Multiphysics Object-Oriented Simulation Environment (MOOSE) is discussed. Several verification cases are discussed to verify the accuracy of the quantified uncertainties using the developed computational capabilities in MOOSE. Then, the problem of quantifying the uncertainties in TRI-Structural isOtropic (TRISO) fuel silver release is addressed. For the first time, the uncertainties arising from the TRISO Fission Gas Release (FGR) model due to model inadequacy and experimental noise are quantified. Also, the Bayesian capabilities are applied to the calibration of the MATPRO creep model, a widely used model in several fuel assessment cases. The impact of the prediction uncertainties in the MATPRO model on the fuel cladding behavior as part of the TRIBULATION assessment case (which is an integral effects case) is investigated. This report concludes with a discussion on the future work for the UQ for computational models.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Massively parallel quantum chemical density matrix renormalization group method

Here, we present, to the best of our knowledge, the first attempt to exploit the supercomputer platform for quantum chemical density matrix renormalization group (QCDMRG) calculations. We have developed the parallel scheme based on the in-house MPI global memory library, which combines operator and symmetry sector parallelisms, and tested its performance on three different molecules, all typical candidates for QC-DMRG calculations. In case of the largest calculation, which is the nitrogenase FeMo cofactor cluster with the active space comprising 113 electrons in 76 orbitals and bond dimension equal to 6000, our parallel approach scales up to approximately 2000 CPU cores.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Massively parallel transport sweeps on meshes with cyclic dependencies

When solving the first-order form of the linear Boltzmann equation, a common misconception is that the matrix-free computational method of “sweeping the mesh”, used in conjunction with the Discrete Ordinates method, is too complex or does not scale well enough to be implemented in modern high performance computing codes. This has led to considerable efforts in the development of matrix-based methods that are computationally expensive and is partly driven by the requirements placed on modern spatial discretizations. In particular, modern transport codes are required to support higher order elements, a concept that invariably adds a lot of complexity to sweeps because of the introduction of cyclic dependencies with curved mesh cells. In this article we will present a comprehensive implementation of sweeping, to a piecewise-linear DFEM spatial discretization with particular focus on handling cyclic dependencies and possible extensions to higher order spatial discretizations. We find that these methods are implemented in a new C++ simulation framework called Chi-Tech (). We present some typical simulation results with some performance aspects that one can expect during real world simulations, we also present a scaling study to >100k processes where Chi-Tech maintains greater than 80% efficiency solving a total of 87.7 trillion angular flux unknowns for a 116 group simulation.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Portfolio Approach to Massively Parallel Bayesian Optimization

One way to reduce the time of conducting optimization studies is to evaluate designs in parallel rather than just one-at-a-time. For expensive-to-evaluate black-boxes, batch versions of Bayesian optimization have been proposed. They work by building a surrogate model of the black-box to simultaneously select multiple designs via an infill criterion. Still, despite the increased availability of computing resources that enable large-scale parallelism, the strategies that work for selecting a few tens of parallel designs for evaluations become limiting due to the complexity of selecting more designs. It is even more crucial when the black-box is noisy, necessitating more evaluations as well as repeating experiments. Here we propose a scalable strategy that can keep up with massive batching natively, focused on the exploration/exploitation trade-off and a portfolio allocation. We compare the approach with related methods on noisy functions, for mono and multi-objective optimization tasks. These experiments show orders of magnitude speed improvements over existing methods with similar or better performance.

97 MATHEMATICS AND COMPUTING↗

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

2.0 - MOOSE: Enabling massively parallel multiphysics simulation

The last 2 years have been a period of unprecedented growth for the MOOSE community and the software itself. The number of monthly visitors to the website has grown from just over 3,000 to now averaging 5,000. In addition, over 1,800 pull requests have been merged since the beginning of 2020, and the new discussions forum has averaged 600 unique visitors per month. The previous publication has been cited over 200 times since it was published 2 years ago. This paper serves as an update on some of the key additions and changes to the code and ecosystem over the last 2 years, as well as recognizing contributions from the community.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

3.0 - MOOSE: Enabling massively parallel multiphysics simulations

The development of MOOSE has kept accelerating since the last release, with over 2,100 pull requests merged over the last 30 months that involved nearly fifty contributors across close to a dozen institutions internationally. The growth in MOOSE's capabilities and downstream applications is reflected in the growth of the community. User support provided on the GitHub discussions forum has steadily increased to nearly 50 daily interactions. New simulation projects, notably to model advanced nuclear reactor and fusion devices, are driving a significant expansion of the capabilities. This paper reports on these developments, with several major released features, new physics modules, and key improvements to the user experience and simulation workflow.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗