Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 559 records · Page 31

ExaCA: A performance portable exascale cellular automata application for alloy solidification modeling

Modeling the as-solidified grain structures that form during alloy processing is a critical component in understanding process-property relationships, particularly for additive manufacturing (AM) where grain structure is very sensitive to processing conditions. While cellular automata (CA)-based models have proven able to predict aspects of microstructure for several alloys and AM process conditions, long run times and large resource sets required limit the utility and the problem size to which existing CA models can be applied. As part of the ExaAM project, an initiative within the Exascale Computing Project (ECP) to develop, test, and optimize an exascale-capable coupled and self-consistent model of AM parts, we developed ExaCA (https://github.com/LLNL/ExaCA) for the liquid–solid phase transformation in the wake of AM melt pools. The CA-based code is parallelized using MPI and the Kokkos programming model, the latter enabling simulation on both CPUs and GPUs within a single-source implementation. Here, we detail the steps taken to transform a baseline, MPI-based CA code into one that is performant on CPUs and GPUs. Performance testing of ExaCA on Summit (a pre-exascale machine at Oak Ridge National Laboratory) was used to quantify CPU–GPU speedup comparing with equal numbers of nodes. Testing showed comparable CPU performance to the MPI-only CA code and a 5-20x speedup when running AM-based test problems using GPUs. The improved performance of CA through GPU utilization and the performance portable nature of ExaCA will enable accurate part-scale modeling by harnessing the power of current and future generations of high performance computing resources. Future work will include improving the strong scaling of ExaCA on GPUs by reducing load imbalance associated with the locality of the problem, and continuing performance optimization across exascale hardware.

36 MATERIALS SCIENCE↗

UQpy: A general purpose Python package and development environment for uncertainty quantification

In this paper, we present the UQpy software toolbox, an open-source Python package for general uncertainty quantification (UQ) in mathematical and physical systems. The software serves as both a user-ready toolbox that includes many of the latest methods for UQ in computational modeling and a convenient development environment for Python programmers advancing the field of UQ. The paper presents an introduction to the software's architecture and existing capabilities, divided in the code in a set of modules centered around different UQ tasks such as sampling methods, generation of random processes and random fields, probabilistic inverse modeling, reliability analysis, surrogate modeling, and active learning. The paper also highlights the importance of the RunModel module, which is used to drive simulations in the uncertainty analyses performed in UQpy. This module conveniently allows the user to define computational models directly in Python, or to run simulations from a third-party software in serial or in parallel. To illustrate the various capabilities, two examples are tracked throughout the paper and analyzed repeatedly for various UQ tasks. The first is a Python model solving a nonlinear structural dynamics problem, used to illustrate UQpy's capabilities in sampling and forward propagation of high dimensional random vectors (stochastic processes), and probabilistic inference. The second model is a third-party Abaqus finite element model solving the thermomechanical response of a beam structure. This example is used to illustrate UQpy's capabilities in variance reduction sampling techniques, reliability analysis, surrogate modeling and active learning techniques.

97 MATHEMATICS AND COMPUTING↗

High-Resolution Comonomer Sequencing of Blocky Brominated Syndiotactic Polystyrene Copolymers Using 13C NMR Spectroscopy and Computer Simulations

This work demonstrates the first high-resolution comonomer sequencing of Blocky brominated syndiotactic polystyrene (sPS-co-sPS-Br) copolymers based on pentad assign- ments of the quaternary carbon region of the nuclear magnetic resonance spectrum. Copolymers containing p-bromostyrene (Br-Sty) units were prepared in matched sets using postpolymerization bromination methods carried out in the heterogeneous gel state (Blocky) and homogeneous solution state (Random). Quantitative information from the quaternary carbon spectra, heteronuclear multiple bond correlation spectroscopy, electronic structure calculations, and simulated statistically random copolymers was correlated to confirm the carbon resonance assignments for all 20 possible pentad comonomer sequences. Using the experimental pentad sequence prevalences, a computer code was developed to simulate chains with microstructures typical of each sample as a means to visually represent the copolymer blockiness with quantitative precision. Based on the microstructure and distribution of run lengths in these chains, the simulations revealed that the Blocky copolymers contain a high degree of blockiness. By comparing the run lengths in the simulated chains to the average number of styrene units in a crystalline segment of sPS (found by small-angle X-ray scattering), copolymer crystallizability was predicted. For the simulated Blocky B-21% (21 mol % Br-Sty) chain, the probability of randomly selecting a styrene unit in a crystallizable block was 25.8%, while that in the simulated Random R-18% was zero, in excellent agreement with the experimental crystallization behavior measured by differential scanning calorimetry. Additionally, these predictions confirmed that the simulated chains accurately represent the ensemble of chains in their respective copolymer samples. Furthermore, each simulated Blocky chain contained one or more long sPS blocks that paralleled the measured 38-40 styrene units spanning a crystalline segment within the sPS/CCl4 gel. This finding affirmed that the long sPS segments originated from the precise lamellar structure within the heterogeneous gel morphology (i.e., block length is correlated with lamellar thickness). Overall, the ability to tailor the copolymer microstructure through control of the semicrystalline gel morphology opens the door to synthesizing ordered copolymers by postpolymerization functionalization processes with unprecedented levels of compositional control.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Automated calculation and convergence of defect transport tensors

Defect diffusion is a key process in materials science and catalysis, but as migration mechanisms are often too complex to enumerate a priori, calculation of transport tensors typically have no measure of convergence and require significant end-user intervention. These two bottlenecks prevent high-throughput implementations essential to propagate model-form uncertainty from interatomic interactions to predictive simulations. In order to address these issues, we extend a massively parallel accelerated sampling scheme, autonomously controlled by Bayesian estimators of statewide sampling completeness, to build atomistic kinetic Monte Carlo models on a state-space irreducible under exchange and space group symmetries. Focusing on isolated defects, we derive analytic expressions for drift and diffusion coefficients, providing a convergence metric by calculating the Kullback–Leibler divergence across the ensemble of diffusion processes consistent with the sampling uncertainty. The autonomy and efficacy of the method is demonstrated on surface trimers in tungsten and Hexa-interstitials in magnesium oxide, both of which exhibit complex, correlated migration mechanisms.

36 MATERIALS SCIENCE↗

Scalable quantum computational science: A perspective from block-encodings and polynomial transformations

Significant developments made in quantum hardware and error correction recently have been driving quantum computing toward practical utility. However, gaps remain between abstract quantum algorithmic development and practical applications in computational sciences. In this perspective article, we propose several properties that scalable quantum computational science methods should possess. We further discuss how block-encodings and polynomial transformations can potentially serve as a unified framework with the desired properties. Recent advancements on these topics are presented, including the construction and assembly of block-encodings, and various generalizations of quantum signal processing (QSP) algorithms to perform polynomial transformations. The scalability of QSP methods on parallel and distributed quantum architectures is also highlighted. Promising applications in simulation and observable estimation in chemistry, physics, and optimization problems are presented. We hope this perspective serves as a gentle introduction to state-of-the-art quantum algorithms for the computational science community and inspires future development of scalable quantum computational science methodologies that bridge theory and practice.

Bayesian inference↗

A New Integrated Analysis Suite for Fast-Ion Study in KSTAR

Here, an integrated workflow for fast-ion analysis was developed by adapting the One Modeling Framework for Integrated Task (OMFIT) workflow manager to support a standard and unified analysis platform for KSTAR users. The newly established analysis suite offers a graphical user interface–based workflow to enable users to readily access and handle experimental data archived in various data formats and servers. Further, users can analyze the data by importing modules designed for conducting certain tasks, such as profile fitting, equilibrium reconstruction, and postprocessing of tokamak data. The procedures for preparing the inputs for fast-ion simulations are streamlined by a common workflow manager, which enables the parallel processing of various tasks to efficiently analyze large fast-ion datasets. The OMFIT platform comprises a flexible Python-based application that enables users to freely manipulate the Python scripts for applications that are unavailable in the standard workflow. The framework also offers mapping tools to translate the output data into the Integrated Modeling and Analysis Suite format to maintain application compatibility for future ITER burning plasma experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Simulation study of particle transport by weakly coherent mode in the Alcator C-Mod tokamak

A simulation study has been conducted of the physical mechanisms behind the weakly coherent mode (WCM) and its produced particle transport in the I-mode edge plasmas by using the BOUT++ code. The WCM is identified in our simulations by its poloidal and radial distributions as well as its frequency and wavenumber spectra. Its produced radial particle flux is calculated and compared with the experimental value. The good agreement indicates that the WCM is an important particle transport channel in the I-mode pedestal. It is found that the WCM can transport particles across the strong outer shear layer of the Er well established in the formation of I-mode, based on which a possible explanation is provided why I-mode does not feature a density pedestal. The key point lies in the change of the cross-phase between the electric potential and density fluctuations induced by the E × B Doppler shift. In the strong shear layer, although the electric potential fluctuation is significantly suppressed, the cross-phase is close to π/2, resulting in a strong drive of the density fluctuation and particle transport. To identify the physical nature of the WCM, a linear dispersion relation for drift Alfvén modes is derived in the slab geometry. A drift Alfvén wave instability is found to have similar dependence to the simulated linear instability behind the WCM on the resistivity and the parallel electron pressure gradient and thermal force terms in the parallel Ohm's law.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

MrHyDE v.1.0

SAND2024-01324O MrHyDE, which stands for Multi-resolution Hybridized Differential Equations, is a general-purpose C++ package for the solution of coupled multiphysics and multiscale systems on massively parallel computing systems. MrHyDE is designed to enable moving beyond forward simulation for multiscale applications which includes optimization, control, uncertainty quantification, and stochastic inversion. The framework provides interfaces to several packages within the Trilinos framework and leverages automatic differentiation to enable adjoint capabilities for large-scale, gradient-based optimization. MrHyDE provides automated multiscale capabilities through a subgrid model interface and multiscale Dirichlet-to-Neumann maps. For extreme-scale applications, MrHyDE provides in situ data-compression algorithms to reduce memory requirements while maintaining performance. MrHyDE is a general-purpose, computational framework for the solution of multiscale and multiphysics applications. It uses a combination of structure-preserving, physics-compatible discretizations, fully implicit methods, multi-resolution schemes, or fully explicit methods. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

SciDAC↗

Attribution of North American Subseasonal Precipitation Prediction Skill

The skill of NOAA’s official monthly U.S. precipitation forecasts (issued in the middle of the prior month) has historically been low, having shown modest skill over the southern United States, but little or no skill over large portions of the central United States. The goal of this study is to explain the seasonal and regional variations of the North American subseasonal (weeks 3–6) precipitation skill, specifically the reasons for its successes and its limitations. The performances of multiple recent-generation model reforecasts over 1999–2015 in predicting precipitation are compared to uninitialized simulation skill using the atmospheric component of the forecast systems. This parallel analysis permits attribution of precipitation skill to two distinct sources: one due to slowly evolving ocean surface boundary states and the other to faster time-scale initial atmospheric weather states. A strong regionality and seasonality in precipitation forecast performance is shown to be analogous to skill patterns dictated by boundary forcing constraints alone. The correspondence is found to be especially high for the North American pattern of the maximum monthly skill that is achieved in the reforecast. The boundary forcing of most importance originates from tropical Pacific SST influences, especially those related to El Niño–Southern Oscillation. Furthermore, we discuss physical constraints that may limit monthly precipitation skill and interpret the performance of existing models in the context of plausible upper limits.

54 ENVIRONMENTAL SCIENCES↗

A fast particle-based approach for calibrating a 3-D model of the Antarctic ice sheet

We consider the scientifically challenging and policy-relevant task of understanding the past and projecting the future dynamics of the Antarctic ice sheet. The Antarctic ice sheet has shown a highly nonlinear threshold response to past climate forcings. Triggering such a threshold response through anthropogenic greenhouse gas emissions would drive drastic and potentially fast sea level rise with important implications for coastal flood risks. Previous studies have combined information from ice sheet models and observations to calibrate model parameters. These studies have broken important new ground but have either adopted simple ice sheet models or have limited the number of parameters to allow for the use of more complex models. These limitations are largely due to the computational challenges posed by calibration as models become more computationally intensive or when the number of parameters increases. Here, we propose a method to alleviate this problem: a fast sequential Monte Carlo method that takes advantage of the massive parallelization afforded by modern high-performance computing systems. We use simulated examples to demonstrate how our sample-based approach provides accurate approximations to the posterior distributions of the calibrated parameters. The drastic reduction in computational times enables us to provide new insights into important scientific questions, for example, the impact of Pliocene era data and prior parameter information on sea level projections. These studies would be computationally prohibitive with other computational approaches for calibration such as Markov chain Monte Carlo or emulation-based methods. We also find considerable differences in the distributions of sea level projections when we account for a larger number of uncertain parameters. For example, based on the same ice sheet model and data set, the 99th percentile of the Antarctic ice sheet contribution to sea level rise in 2300 increases from 6.5 m to 13.1 m when we increase the number of calibrated parameters from three to 11. With previous calibration methods, it would be challenging to go beyond five parameters. Here, this work provides an important next step toward improving the uncertainty quantification of complex, computationally intensive and decision-relevant models.

54 ENVIRONMENTAL SCIENCES↗

Thermal quench of open field plasma intercepting with recycling walls

When a fusion plasma suddenly intercepts a solid surface (wall or pellet), thermal collapse is distinctly kinetic & has novel physics. VPIC simulations & theory revealed the fascinating dynamics of four propagating fronts controlling parallel electron temperature cooling.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A Model for Pair Production Limit Cycles in Pulsar Magnetospheres

Abstract It was recently proposed that the electric field oscillation as a result of self-consistent e ± pair production may be the source of coherent radio emission from pulsars. Direct particle-in-cell simulations of this process have shown that the screening of the parallel electric field by this pair cascade manifests as a limit cycle, as the parallel electric field is recurrently induced when pairs produced in the cascade escape from the gap region. In this work, we develop a simplified time-dependent kinetic model of e ± pair cascades in pulsar magnetospheres that can reproduce the limit-cycle behavior of pair production and electric field screening. This model includes the effects of a magnetospheric current, the escape of e ± , as well as the dynamic dependence of pair production rate on the plasma density and energy. Using this simple theoretical model, we show that the power spectrum of electric field oscillations averaged over many limit cycles is compatible with the observed pulsar radio spectrum.

Astronomy & Astrophysics↗

Dynamic Phasor Modeling of Multi-Converter Systems

Generalized form of Dynamic Phasor (DP)-based modeling of multi-converter systems containing high order harmonics has not been proposed in the literature yet due to the complexity and large number of variables that should be included in the model. In this work a generalized form of a system-level model comprising high number of single-phase converters is presented. The proposed method can be applied to any number of single-phase voltage source converters (VSIs). Test cases of this modeling method containing 5 to 100 parallel connected single-phase voltage source inverters are modeled and simulation times are recorded. The results of the proposed method are compared and validated with conventional average models as well as detailed switching models. A systematic approach for comparing the accuracy and timestep between dynamic phasor modeling method and detailed switching model is illustrated. Advantage of DP models over conventional average models for stability assessment are also discussed at the end of the paper.

Xue, Yaosuo↗

Study of Inverter Control Strategies on the Stability of Microgrids Toward 100% Renewable Penetration: Preprint

This paper investigates microgrid transient stability with mixed generation - synchronous generator (SG), grid-forming (GFM) and grid-following (GFL) inverters - under increasing penetration levels toward a 100% renewable generation microgrid. Specifically, the dynamics of a microgrid with an SG and GFL inverter(s), an SG with GFM inverter(s), and an SG with GFM and GFL inverters under each penetration are evaluated with an electromagnetic transient study with two critical dynamic events: unplanned islanding and switching in a pumped induction motor load. Analysis and simulation results indicate that the microgrid with GFL inverters running in parallel with the SG can provide a faster power response than the GFM inverters to compensate for the deviations of the frequency and voltage. The scenario with the mixed SG, GFM, and GFL inverter has the best transient and steady-state stability toward 100% inverter-based resource (IBR) penetration. This comprehensive study provides helpful references for microgrid engineers to understand the microgrid stability when facing various choice of installing IBRs (GFL, GFM, or mixed).

droop control↗

Study of Inverter Control Strategies on the Stability of Microgrids Toward 100% Renewable Penetration

This paper investigates microgrid transient stability with mixed generation - synchronous generator (SG), grid-forming (GFM) and grid-following (GFL) inverters - under increasing penetration levels toward a 100% renewable generation microgrid. Specifically, the dynamics of a microgrid with an SG and GFL inverter(s), an SG with GFM inverter(s), and an SG with GFM and GFL inverters under each penetration are evaluated with an electromagnetic transient study with two critical dynamic events: unplanned islanding and switching in a pumped induction motor load. Analysis and simulation results indicate that the microgrid with GFL inverters running in parallel with the SG can provide a faster power response than the GFM inverters to compensate for the deviations of the frequency and voltage. The scenario with the mixed SG, GFM, and GFL inverter has the best transient and steady-state stability toward 100% inverter-based resource (IBR) penetration. This comprehensive study provides helpful references for microgrid engineers to understand the microgrid stability when facing various choice of installing IBRs (GFL, GFM, or mixed).

droop control↗

Separatrix-to-Wall Simulations of Impurity Transport with a Fully Three-Dimensional Wall in DIII-D

A novel multi-code workflow to interpret collector probe deposition patterns in DIII-D has been developed. The components of the workflow consist of a detailed computer-aided design (CAD) file of the vessel wall and the scrape-off layer (SOL) codes MAFOT, OSM, DIVIMP and 3DLIM. A special-purpose toolkit enables passing the output of these codes between each other to provide a full-SOL picture of impurity transport. A demonstration of the workflow is described to support evidence of near-SOL tungsten parallel accumulation during trace W impurity experiments on DIII-D. Iteration between simulated deposition patterns in 3DLIM and DIVIMP predicts a region of elevated W density near the separatrix about halfway between the outboard midplane and the top of the plasma. Furthermore, this workflow will be used to better interpret collector probe experiments on DIII-D.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel Algorithms for Computing the Tensor-Train Decomposition

The tensor-train (TT) decomposition expresses a tensor in a data-sparse format used in molecular simulations, high-order correlation functions, and optimization. In this paper, we propose four parallelizable algorithms that compute the TT format from various tensor inputs: (1) Parallel-TTSVD for traditional format, (2) PSTT and its variants for streaming data, (3) Tucker2TT for Tucker format, and (4) TT-fADI for solutions of Sylvester tensor equations. We provide theoretical guarantees of accuracy, parallelization methods, scaling analysis, and numerical results. For example, for a d-dimension tensor in $\mathbb{R}$ $n\times∙∙∙$$\times$$n$ a two-sided sketching algorithm PSTT2 is shown to have a memory complexity of $O(n^{[d/2]})$, improving upon $O(n^{d—1})$ from previous algorithms.

97 MATHEMATICS AND COMPUTING↗

Microfabricated ion trap chip with an integrated microwave antenna

An ion trap chip, which may be used for quantum information processing and the like, includes an integrated microwave antenna. The antenna is formed as a radiator connected by one of its ends to the center trace of a microwave transmission line and connected by its other end to a current return path through a ground trace of the microwave transmission line. The radiator includes several parallel, coplanar radiator traces connected in series. The radiator traces are connected such that they all carry electric current in the same direction, so that collectively, they simulate a single, unidirectionally flowing sheet of current. In embodiments, induced currents in underlying metallization planes are suppressed by parallel slots that extend in a direction perpendicular to the radiator traces.

Nordquist, Christopher↗