Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “massive parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

2.0 - MOOSE: Enabling massively parallel multiphysics simulation

The last 2 years have been a period of unprecedented growth for the MOOSE community and the software itself. The number of monthly visitors to the website has grown from just over 3,000 to now averaging 5,000. In addition, over 1,800 pull requests have been merged since the beginning of 2020, and the new discussions forum has averaged 600 unique visitors per month. The previous publication has been cited over 200 times since it was published 2 years ago. This paper serves as an update on some of the key additions and changes to the code and ecosystem over the last 2 years, as well as recognizing contributions from the community.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

3.0 - MOOSE: Enabling massively parallel multiphysics simulations

The development of MOOSE has kept accelerating since the last release, with over 2,100 pull requests merged over the last 30 months that involved nearly fifty contributors across close to a dozen institutions internationally. The growth in MOOSE's capabilities and downstream applications is reflected in the growth of the community. User support provided on the GitHub discussions forum has steadily increased to nearly 50 daily interactions. New simulation projects, notably to model advanced nuclear reactor and fusion devices, are driving a significant expansion of the capabilities. This paper reports on these developments, with several major released features, new physics modules, and key improvements to the user experience and simulation workflow.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

4.0 MOOSE: Enabling massively parallel Multiphysics simulation

Approaching 18 years of existence, MOOSE—the Multiphysics Object-Oriented Simulation Environment—is being developed at a higher pace than ever before. With significant support from four research institutions across the globe, and dozens of new contributors, the capabilities of the framework are being expanded to meet modeling challenges in a wide variety of fields from nuclear system design, to geomechanics, to material science. This includes new development in equation discretization techniques, solver methods, meshing capabilities, application deployment, and user interface improvements. Applications built on MOOSE benefit from all these improvements.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Implementation of Relativistic Coupled Cluster Theory for Massively Parallel GPU-Accelerated Computing Architectures

In this paper, we report reimplementation of the core algorithms of relativistic coupled cluster theory aimed at modern heterogeneous high-performance computational infrastructures. The code is designed for parallel execution on many compute nodes with optional GPU coprocessing, accomplished via the new ExaTENSOR back end. The resulting ExaCorr module is primarily intended for calculations of molecules with one or more heavy elements, as relativistic effects on the electronic structure are included from the outset. In the current work, we thereby focus on exact two-component methods and demonstrate the accuracy and performance of the software. The module can be used as a stand-alone program requiring a set of molecular orbital coefficients as the starting point, but it is also interfaced to the DIRAC program that can be used to generate these. We therefore also briefly discuss an improvement of the parallel computing aspects of the relativistic self-consistent field algorithm of the DIRAC program.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Massively Parallel Implementation of the CCSD(T) Method Using the Resolution-of-the-Identity Approximation and a Hybrid Distributed/Shared Memory Parallelization Model

In this work, a parallel algorithm is described for the coupled-cluster singles and doubles method augmented with a perturbative correction for triple excitations [CCSD(T)] using the resolution-of-the-identity (RI) approximation for two-electron repulsion integrals (ERIs). The algorithm bypasses the storage of four-center ERIs by adopting an integral-direct strategy. The CCSD amplitude equations are given in a compact quasi-linear form by factorizing them in terms of amplitude-dressed three-center intermediates. A hybrid MPI/OpenMP parallelization scheme is employed, which uses the OpenMP-based shared memory model for intranode parallelization and the MPI-based distributed memory model for internode parallelization. Parallel efficiency has been optimized for all terms in the CCSD amplitude equations. Two different algorithms have been implemented for the rate-limiting terms in the CCSD amplitude equations that entail and -scaling computational costs, where N O and N V denote the number of correlated occupied and virtual orbitals, respectively. One of the algorithms assembles the four-center ERIs requiring N V 4 and N O 2 N V 2 -scaling memory costs in a distributed manner on a number of MPI ranks, while the other algorithm completely bypasses the assembling of quartic memory-scaling ERIs and thus largely reduces the memory demand. It is demonstrated that the former memory-expensive algorithm is faster on a few hundred cores, while the latter memory-economic algorithm shows a better strong scaling in the limit of a few thousand cores. The program is shown to exhibit a near-linear scaling, in particular for the compute-intensive triples correction step, on up to 8000 cores. The performance of the program is demonstrated via calculations involving molecules with 24–51 atoms and up to 1624 atomic basis functions. As the first application, the complete basis set (CBS) limit for the interaction energy of the π-stacked uracil dimer from the S66 data set has been investigated. This work reports the first calculation of the interaction energy at the CCSD(T)/aug-cc-pVQZ level without local orbital approximation. The CBS limit for the CCSD correlation contribution to the interaction energy was found to be -8.01 kcal/mol, which agrees very well with the value -7.99 kcal/mol reported by Schmitz, Hättig, and Tew [ Phys. Chem. Chem. Phys. 2014 , 16 , 22167-22178]. The CBS limit for the total interaction energy was estimated to be -9.64 kcal/mol.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Low cost and massively parallel force spectroscopy with fluid loading on a chip

Current approaches for single molecule force spectroscopy are typically constrained by low throughput and high instrumentation cost. Herein, a low-cost, high throughput technique is demonstrated using microfluidics for multiplexed mechanical manipulation of up to ~4000 individual molecules via molecular fluid loading on-a-chip (FLO-Chip). The FLO-Chip consists of serially connected microchannels with varying width, allowing for simultaneous testing at multiple loading rates. Molecular force measurements are demonstrated by dissociating Biotin-Streptavidin and Digoxigenin-AntiDigoxigenin interactions along with unzipping of double stranded DNA of varying sequence under different dynamic loading rates and solution conditions. Rupture force results under varying loading rates and solution conditions are in good agreement with prior studies, verifying a versatile approach for single molecule biophysics and molecular mechanobiology. FLO-Chip enables straightforward, rapid, low-cost, and portable mechanical testing of single molecules that can be implemented on a wide range of microscopes to broaden access and may enable new applications of molecular force spectroscopy.

42 ENGINEERING↗

MEUMAPPS (Microstructure Evolution Using Massively Parallel Phase-field Simulations)

The software “MEUMAPPS” is a h high-performance computing code used to simulate microstructure evolution associated with diffusional solid-state transformations in structural alloys. Understanding microstructure evolution during thermo-mechanical processing of structural alloys is the first step towards designing and processing alloys for specific applications by meeting property requirements demanded by the application. In this instance, the code is an important component in predicting processing-structure linkages during additive manufacturing of structural alloys as a function of processing parameters and alloy composition used in an additive manufacturing process. The code provides a detailed three-dimensional distribution of different types of phases / constituents and their morphologies, as well as the compositions of these phases that make up the microstructure. The microstructure is simulated by solving the governing partial differential equations using a Fourier Spectral Method.

Radhakrishnan, Balasubram↗

A massively parallel time-domain coupled electrodynamics–micromagnetics solver

We present a high-performance coupled electrodynamics–micromagnetics solver for full physical modeling of signals in microelectronic circuitry. The overall strategy couples a finite-difference time-domain approach for Maxwell’s equations to a magnetization model described by the Landau–Lifshitz–Gilbert equation. The algorithm is implemented in the Exascale Computing Project software framework, AMReX, which provides effective scalability on manycore and GPU-based supercomputing architectures. Furthermore, the code leverages ongoing developments of the Exascale Application Code, WarpX, which is primarily being developed for plasma wakefield accelerator modeling. Our temporal coupling scheme provides second-order accuracy in space and time by combining the integration steps for the magnetic field and magnetization into an iterative sub-step that includes a trapezoidal temporal discretization for the magnetization. The performance of the algorithm is demonstrated by the excellent scaling results on NERSC multicore and GPU systems, with a significant (59×) speedup on the GPU using a node-by-node comparison. We demonstrate the utility of our code by performing simulations of an electromagnetic waveguide and a magnetically tunable filter.

97 MATHEMATICS AND COMPUTING↗

Nyx: A Massively Parallel AMR Code for Computational Cosmology

Nyx is a highly parallel, adaptive mesh, finite-volume N-body compressible hydrodynamics solver for cosmological simulations. It has been used to simulate different cosmological scenarios with a recent focus on the intergalactic medium and Lyman alpha forest. Together, Nyx, the compressible astrophysical simulation code, Castro, and the low Mach number code MAESTROeX, make up the AMReX-Astrophysics Suite of open-source, adaptive mesh, performance-portable astrophysical simulation codes. Other examples of cosmological simulation research codes include Enzo, Enzo-P/Cello, RAMSES, ART, FLASH, Cholla, as well as Gadget, Gasoline, Arepo, Gizmo, and SWIFT.

79 ASTRONOMY AND ASTROPHYSICS↗

DeepHyper: A Python Package for Massively Parallel Hyperparameter Optimization in Machine Learning

Machine learning models are increasingly applied across scientific disciplines, yet their effectiveness often hinges on heuristic decisions—such as data transformations, training strategies, and model architectures—that are not learned by the models themselves. Automating the selection of these heuristics and analyzing their sensitivity is crucial for building robust and efficient learning workflows. DeepHyper addresses this challenge by democratizing hyperparameter optimization, providing accessible tools to streamline and enhance machine learning workflows from a laptop to the largest supercomputer in the world. Building on top of hyperparameter optimization, it unlocks new capabilities around ensembles of models for improved accuracy and uncertainty quantification. All of these organized around efficient parallel computing.

ensemble↗

C HIMERA : A Massively Parallel Code for Core-collapse Supernova Simulations

Herein we provide a detailed description of the C HIMERA code, a code developed to model core collapse supernovae (CCSNe) in multiple spatial dimensions. The CCSN explosion mechanism remains the subject of intense research. Progress to date demonstrates that it involves a complex interplay of neutrino production, transport, and interaction in the stellar core, three-dimensional stellar core fluid dynamics and its associated instabilities, nuclear burning, and the fundamental physics of the neutrino–stellar core weak interactions and the equations of state of all stellar core constituents—particularly, the nuclear equation of state associated with core nucleons, both free and bound in nuclei. C HIMERA , by incorporating detailed neutrino transport, realistic neutrino–matter interactions, three-dimensional hydrodynamics, realistic nuclear, leptonic, and photonic equations of state, and a nuclear reaction network, along with other refinements, can be used to study the role of neutrino radiation, hydrodynamic instabilities, and a variety of input physics in the explosion mechanism itself. It can also be used to compute observables such as neutrino signatures, gravitational radiation, and the products of nucleosynthesis associated with CCSNe. The code contains modules for neutrino transport, multidimensional compressible hydrodynamics, nuclear reactions, a variety of neutrino interactions, equations of state, and modules to provide data for post-processing observables such as the products of nucleosynthesis, and gravitational radiation. C HIMERA is an evolving code, being updated periodically with improved input physics and numerical refinements. We detail here the current version of the code, from which future improvements will stem, which can in turn be described as needed in future publications.

79 ASTRONOMY AND ASTROPHYSICS↗

Simulating coupled surface–subsurface flows with ParFlow v3.5.0: capabilities, applications, and ongoing development of an open-source, massively parallel, integrated hydrologic model

Surface flow and subsurface flow constitute a naturally linked hydrologic continuum that has not traditionally been simulated in an integrated fashion. Recognizing the interactions between these systems has encouraged the development of integrated hydrologic models (IHMs) capable of treating surface and subsurface systems as a single integrated resource. IHMs are dynamically evolving with improvements in technology, and the extent of their current capabilities are often only known to the developers and not general users. This article provides an overview of the core functionality, capability, applications, and ongoing development of one open-source IHM, ParFlow. ParFlow is a parallel, integrated, hydrologic model that simulates surface and subsurface flows. ParFlow solves the Richards equation for three-dimensional variably saturated groundwater flow and the two-dimensional kinematic wave approximation of the shallow water equations for overland flow. The model employs a conservative centered finite-difference scheme and a conservative finite-volume method for subsurface flow and transport, respectively. ParFlow uses multigrid-preconditioned Krylov and Newton–Krylov methods to solve the linear and nonlinear systems within each time step of the flow simulations. The code has demonstrated very efficient parallel solution capabilities. ParFlow has been coupled to geochemical reaction, land surface (e.g., the Common Land Model), and atmospheric models to study the interactions among the subsurface, land surface, and atmosphere systems across different spatial scales. This overview focuses on the current capabilities of the code, the core simulation engine, and the primary couplings of the subsurface model to other codes, taking a high-level perspective.

58 GEOSCIENCES↗

Explicit and implicit solution of the Navier-Stokes equations on a massively parallel computer

The design, implementation, and performance of a two-dimensional time-accurate Navier-Stokes solver for the CM2 supercomputer are described. The program uses a single processor for each grid point. Two different time-stepping methods have so far been implemented: an explicit third-order Runge-Kutta method and an implicit approximation-factorization method. The CM2 results are checked against those of a mature well-vectorized Cray 2 program, both for correctness and performance. The code is found to be correct, and the performance in some cases is up to several times that of the Cray 2.

Levit, Creon↗