Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

Fast Muon Capture Monitoring in Mu2e with the CAPHRI Detector

The Mu2e experiment at Fermilab will search for the charged lepton flavor violating (CLFV) process of a neutrinoless muon-to-electron conversion in the field of an aluminum nucleus. Reaching the experiment’s target sensitivity requires precise normalization of the physics signal through accurate monitoring of the muon capture rate on the stopping target. For this purpose, the Calorimeter Precise High-Resolution Intensity detector (CAPHRI) has been developed. The detector is composed of four LYSO crystals installed in the upstream disk of the Mu2e calorimeter and read out with the standard calorimeter readout. CAPHRI measures the muon capture rate by detecting the characteristic 1.8~MeV gamma emission line of the $^{27}Al(\mu^−, \nu n \gamma) ^{26}Mg$ nuclear reaction. The fast, precise response enables injection-by-injection monitoring of proton beam intensity fluctuations. We report on the commissioning and performance characterization of CAPHRI. The response of each channel is calibrated at two SiPM overvoltages using both the intrinsic self-emission of the LYSO crystals and cosmic ray signals. In parallel, Monte Carlo simulations are used to evaluate the detector acceptance and the expected signal-to-background ratio under realistic running conditions. Preliminary results show a crystal light yield consistent with expectations and a channel inter-calibration at the 2--4% level. Simulation studies indicate that the detector acceptance and background rejection satisfy the requirements for physics operations, with about 1000 detected events per beam injection at a beam power of 1.5~kW. These results demonstrate that CAPHRI is an effective tool for beam monitoring and signal normalization in Mu2e.

Ciccarella, V. [Frascati; U. Rome La Sapienza (mai↗

Towards multiscale modeling of ocean surface turbulent mixing using coupled MPAS-Ocean v6.3 and PALM v5.0

A multiscale modeling approach for studying the ocean surface turbulent mixing is explored by coupling an ocean general circulation model (GCM) MPAS-Ocean with the Parallelized Large Eddy Simulation Model (PALM). The coupling approach is similar to the superparameterization approach that has been used to represent the effects of deep convection in atmospheric GCMs. However, the focus of this multiscale modeling approach is on the small-scale turbulent mixing and their interactions with the larger-scale processes in the ocean, so that a more flexible coupling strategy is used. To reduce the computational cost, a customized version of PALM is ported on the general-purpose graphics processing unit (GPU) with OpenACC, achieving 10–16 times overall speedup as compared to running on a single CPU. Even with the GPU-acceleration technique, a superparameterization-like approach to represent the ocean surface turbulent mixing in GCMs using embedded high fidelity and three-dimensional large eddy simulations (LESs) over the global ocean is still computationally intensive and infeasible for long simulations. However, running PALM regionally on selected MPAS-Ocean grid cells is shown to be a promising approach moving forward. The flexible coupling between MPAS-Ocean and PALM allows further exploration of the interactions between the ocean surface turbulent mixing and larger-scale processes, as well as future development and improvement of ocean surface turbulent mixing parameterizations for GCMs.

54 ENVIRONMENTAL SCIENCES↗

Packaging a 650V/400A GaN Half-bridge Power Module with Ultra-low Parasitics for Electric Vehicle Drive Applications

This paper proposes a compact and efficient half-bridge power module with three 650 V / 150 A GaN dies in parallel. The power module incorporates a main power printed circuit board (PCB), an interface PCB, and a flex PCB to achieve low parasitics in both power loop and gate-side connection, resolving the issue of high parasitics typically encountered with wire bonding in high-current applications. Additionally, the interface PCB decouples the design constraints between the power loop and the gate loops. The proposed design is optimized with a vertical loop configuration to reduce power loop inductance through magnetic flux cancellation. Finite element analysis indicates that the power loop inductance is 0.58 nH at 100 MHz, while the maximum die junction temperature reaches 131 °C under an ambient temperature of 65 °C and a load current of 385 A. The proposed multi-piece PCB structure reduces the inductance of the drive circuit to minimize EMI and to mitigate false triggering. At the same time, it reduces impedance mismatches across different driver circuits, thereby achieving dynamic current sharing in multi-chip parallel configurations. Under simulation conditions of 400 V / 385 A, the current imbalance among chips was limited to 5 A. A 400 V / 385 A double-pulse test was conducted to experimentally validate the performance of the proposed power module.

30 DIRECT ENERGY CONVERSION↗

Simulation of 24,000 Electron Dynamics: Real-Time Time-Dependent Density Functional Theory (TDDFT) with the Real-Space Multigrids (RMG)

Here, we present the theory, implementation, and benchmarking of a real-time time-dependent density functional theory (RT-TDDFT) module within the RMG code, designed to simulate the electronic response of molecular systems to external perturbations. Our method offers insights into nonequilibrium dynamics and excited states across a diverse range of systems, from small organic molecules to large metallic nanoparticles. Benchmarking results demonstrate excellent agreement with established TDDFT implementations and showcase the superior stability of our time integration algorithm, enabling long-term simulations with minimal energy drift. The scalability and efficiency of RMG on massively parallel architectures allow for simulations of complex systems, such as plasmonic nanoparticles with thousands of atoms. Future extensions, including nuclear and spin dynamics, will broaden the applicability of this RT-TDDFT implementation, providing a powerful toolset for studies of photoactive materials, nanoscale devices, and other systems where real-time electronic dynamics is essential.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

HPC for optimizing process parameters to control material evolution in seamless induction hardening of wind turbine main shaft bearings

Work proposed in this project focused on understanding the effect of martensitic transformation in the steel on the potential for cracking during seamless induction hardening (SIH) as a function of process conditions to allow the process to optimally scale up. Large-scale, three-dimensional phase-field simulations of martensitic transformation were performed using MEUMAPPS-SS (Microstructure Evolution Using Massively Parallel Phase-field Simulations – Solid State) code developed at Oak Ridge National Laboratory. The simulations were guided by location-specific thermal history generated by experimental measurements of time-temperature history generated at The Timken Company. The simulations were able to capture the morphological evolution of the martensite variants in an Fe-1.0C-1.5Cr steel based on the Nishiyama-Wasserman (NW) orientation relationship. The simulations were also able to quantify the stress-state at the interface between impinging martensite variants. The simulations indicated that the magnitude of the various stress and strain components were dependent on the sizes of the impinging plates with a reduction in these quantities with reduced plate size in agreement with experimental findings. The results obtained from the simulations will be used to guide the optimization of the alloy thermal conditions to eliminate quench cracking during SIH of bearing steels.

17 WIND ENERGY↗

The role of the spatial heterogeneity and correlation length of surface wettability on two-phase flow in a CO 2 -water-rock system

This study characterized and modeled heterogeneous surface wettability in sandstone and investigated the role of spatial heterogeneity and correlation length of surface wettability on relative permeability in a supercritical CO 2 (scCO 2 )-brine-rock system. Understanding the role of wettability heterogeneity on relative permeability is essential to geological CO 2 sequestration, oil and gas recovery, and contaminated groundwater remediation. Although numerous studies have attempted to understand the influences of surface wettability, capillary number (Ca), and viscosity ratio, the role of the spatial variation and correlation length of surface wettability on two-phase flow in three-dimensional (3D) porous media has not been unraveled due to the challenges in the measurement and representation of realistic rock surface wettability. In this work, we conducted in-situ measurements of surface contact angle (CA) in a Bentheimer sandstone after CO 2 flooding using micro-computed tomography (micro-CT), and found that the pore-scale CA distribution on rock surfaces followed a log-normal distribution associated with a spatial correlation length. Based on the statistical information from CT scanning, a Gaussian random field was used to model CA distributions that had desired standard deviations and spatial correlation lengths, which were then adjusted within a certain range of values for sensitivity analyses to study their combined effects on the two-phase flow in the porous medium using the lattice Boltzmann (LB) method. The LB two-phase flow simulation was accelerated using hybrid, multicore parallel computing to overcome the challenges in simulating multiphase flow in a large 3D domain having 800 × 800 × 600 nodes. The simulation results showed that the surface wettability heterogeneity (i.e., standard deviation of CA) had a lesser effect on the relative permeability of the wetting fluid (water) but a more significant impact on the relative permeability of the non-wetting fluid (scCO 2 ). The Corey model was used to fit the LB-simulated relative permeability curves of water and scCO 2 and showed that the variations in the relative permeability curves for both water and scCO 2 increased as the standard deviation and spatial correlation length of CA increased. This study illustrated that the assumption of homogeneous surface wettability may cause errors in multiphase flow simulations. Furthermore, the impacts of both the standard deviation and spatial correlation length of CAs should be accounted for. This is the first study that explored the spatial correlation lengths associated with CA distributions on sandstone surfaces and comprehensively investigated the roles of both spatial variation and correlation length of CA on two-phase flow properties in 3D porous media. The optimized LB multiphase flow model was proved a powerful tool to study the interplays and combined effects of these statistical parameters, which had critical applications in numerous natural and engineering processes that involved multiphase flow in porous media.

58 GEOSCIENCES↗

Toward a Fully Integrated Multiphysics Simulation Framework for Fusion Blanket Design

Fusion is an attractive clean-energy solution, thanks to its various advantages, such as reduced radioactivity, little high-level nuclear waste, ample fuel supplies, and increased safety. However, the harsh operating environment introduced by a complex fusion plasma system makes design and integration of fusion blankets incredibly challenging and time-consuming. This work focuses on developing a fully integrated multiphysics simulation framework based on an advanced open-source platform—the Multiphysics Object-Oriented Simulation Environment (MOOSE)—to alleviate the difficulties in fusion blanket design and integration. MOOSE is a massively parallel finite element/volume multiphysics simulation platform that has been widely adopted within the nuclear fission community. Even though fission and fusion are fundamentally different, they involve similar multiphysics phenomena. A fully integrated open-source multiphysics simulation framework tailored for the fusion blanket design will be implemented by leveraging the well-established multiphysics capabilities in MOOSE. Once successfully developed, this fully integrated framework will rapidly evaluate a blanket design concept and offer insights for subsequent iterations. As the first step, we will mainly aim to integrate neutronics analysis, system thermal hydraulics simulation, and full 3-D heat transfer calculations. The efficacy of the integrated framework will be verified using an innovative solid ceramic blanket design. While the project’s final goal is to enable a fully integrated multiphysics simulation platform for various fusion blanket concepts, here this work, as a preliminary step, will mainly focus on a solid ceramic breeder helium-cooled blanket.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Architecture for Co-Simulation of Transportation and Distribution Systems with Electric Vehicle Charging at Scale in the San Francisco Bay Area

This work describes the Grid-Enhanced, Mobility-Integrated Network Infrastructures for Extreme Fast Charging (GEMINI) architecture for the co-simulation of distribution and transportation systems to evaluate EV charging impacts on electric distribution systems of a large metropolitan area and the surrounding rural regions with high fidelity. The current co-simulation is applied to Oakland and Alameda, California, and in future work will be extended to the full San Francisco Bay Area. It uses the HELICS co-simulation framework to enable parallel instances of vetted grid and transportation software programs to interact at every model timestep, allowing high-fidelity simulations at a large scale. This enables not only the impacts of electrified transportation systems across a larger interconnected collection of distribution feeders to be evaluated, but also the feedbacks between the two systems, such as through control systems, to be captured and compared. The findings are that with moderate passenger EV adoption rates, inverter controls combined with some distribution system hardware upgrades can maintain grid voltages within ANSI C.84 range A limits of 0.95 to 1.05 p.u. without smart charging. However, EV charging control may be required for higher levels of charging or to reduce grid upgrades, and this will be explored in future work.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Performance Analysis of Speculative Parallel Adaptive Local Timestepping for Conservation Laws

Stable simulation of conservation laws, such as those used to model fluid dynamics and plasma physics applications, requires the satisfaction of the so-called Courant-Friedrichs-Lewy condition. By allowing regions of the mesh to advance with different timesteps that locally satisfy this stability constraint, significant work reduction can be attained when compared to a time integration scheme using a single timestep size. However, parallelizing this algorithm presents considerable difficulty. Since the stability condition depends on the state of the system, dependencies become dynamic and potentially non-local. In this article, we present an adaptive local timestepping algorithm using an optimistic (Timewarp-based) parallel discrete event simulation. We introduce waiting heuristics to limit misspeculation and a semi-static load balancing scheme to eliminate load imbalance as parts of the mesh require finer or coarser timesteps. Last, we outline an interface for separating the physics of the specific conservation law from the temporal integration allowing for productive adoption of our proposed algorithm. We present a misspeculation study for three conservation laws, demonstrating both the productivity of the local timestepping API, for which 74% of the lines of code are reused across different conservation laws, and the robustness of the waiting heuristics—at most 1.5% of element updates are rolled back. Our performance studies demonstrate up to a 2.8× speedup versus a baseline unoptimized local timestepping approach, a 4x improvement in per-node throughput compared to an MPI parallelization of synchronous timestepping, and scalability up to 3,072 cores on NERSC’s Cori Haswell partition.

97 MATHEMATICS AND COMPUTING↗

Parallel Randomized Tucker Decomposition Algorithms

The Tucker tensor decomposition is a natural extension of the singular value decomposition (SVD) to multiway data. Here, we propose to accelerate Tucker tensor decomposition algorithms by using randomization and parallelization. We present two algorithms that scale to large data and many processors, significantly reduce both computation and communication cost compared to previous deterministic and randomized approaches, and obtain nearly the same approximation errors. The key idea in our algorithms is to perform randomized sketches with Kronecker-structured random matrices, which reduces computation compared to unstructured matrices and can be implemented using a fundamental tensor computational kernel. We provide probabilistic error analysis of our algorithms and implement a new parallel algorithm for the structured randomized sketch. Our experimental results demonstrate that our combination of randomization and parallelization achieves accurate Tucker decompositions much faster than alternative approaches. We observe up to a 16X speedup over the fastest deterministic parallel implementation on 3D simulation data.

Tucker decompositions↗

Code modernization strategies for short-range non-bonded molecular dynamics simulations

Modern HPC systems are increasingly relying on greater core counts and wider vector registers. Thus, applications need to be adapted to fully utilize these hardware capabilities. One class of applications that can benefit from this increase in parallelism are molecular dynamics simulations. In this paper, we describe our efforts at modernizing the ESPResSo++ simulation package for molecular dynamics by restructuring its particle data layout for efficient memory accesses and applying vectorization techniques to benefit the calculation of short-range non-bonded forces, which results in an overall three times speedup and serves as a baseline for further optimizations. We also implement fine-grained parallelism for multi-core CPUs through HPX, a C++ runtime system which uses lightweight threads and an asynchronous many-task approach to maximize concurrency. Our goal is to evaluate the performance of an HPX-based approach compared to the bulk-synchronous MPI-based implementation. This requires the introduction of an additional layer to the domain decomposition scheme that defines the task granularity. On spatially inhomogeneous systems, which impose a corresponding load-imbalance in traditional MPI-based approaches, we demonstrate that by choosing an optimal task size, the efficient work-stealing mechanisms of HPX can overcome the overhead of communication resulting in an overall 1.4 times speedup compared to the baseline MPI version.

97 MATHEMATICS AND COMPUTING↗

NESSi : The N on- E quilibrium S ystems S imulation package

The nonequilibrium dynamics of correlated many-particle systems is of interest in connection with pump–probe experiments on molecular systems and solids, as well as theoretical investigations of transport properties and relaxation processes. Nonequilibrium Green’s functions are a powerful tool to study interaction effects in quantum many-particle systems out of equilibrium, and to extract physically relevant information for the interpretation of experiments. Here, we present the open-source software package NESSi (The Non-Equilibrium Systems Simulation package) which allows to perform many-body dynamics simulations based on Green’s functions on the L-shaped Kadanoff–Baym contour. NESSi contains the library libcntr which implements tools for basic operations on these nonequilibrium Green’s functions, for constructing Feynman diagrams, and for the solution of integral and integro-differential equations involving contour Green’s functions. The library employs a discretization of the Kadanoff–Baym contour into time points and a high-order implementation of integration routines. The total integrated error scales up to $\mathcal{O}(N^{-7})$, which is important since the numerical effort increases at least cubically with the simulation time. A distributed-memory parallelization over reciprocal space allows large-scale simulations of lattice systems. We provide a collection of example programs ranging from dynamics in simple two-level systems to problems relevant in contemporary condensed matter physics, including Hubbard clusters and Hubbard or Holstein lattice models. The libcntr library is the basis of a follow-up software package for nonequilibrium dynamical mean-field theory calculations based on strong-coupling perturbative impurity solvers.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Polarized fusion and potential in situ tests of fuel polarization survival in a tokamak plasma

Abstract The use of spin-polarized fusion fuels would provide a significant boost towards the ignition of a burning plasma. The cross section for D + T → α + n, would be increased by 1.5 if the fuels were injected with parallel polarization. Furthermore, our simulations demonstrate additional non-linear power gains in large-scale machines such as ITER, due to increased alpha heating. Such benefits require the survival of spin polarizations for periods comparable to the particle confinement time. During the 1980s, calculations predicted that polarizations could survive a plasma environment, although concerns persisted regarding the cumulative impacts of wall recycling. In that era, technical challenges prevented direct tests and left the large scale fueling of a power reactor beyond reach. Over the last decades, this situation has changed dramatically. Detailed simulations of ITER have predicted negligible wall recycling in a high-power reactor, and recent advances in laser-driven sources project the capability of producing large quantities of ∼100% polarized D and T. The remaining crucial step is an in-situ demonstration of polarization survival in a plasma. For this, we outline a measurement strategy using the isospin-mirror reaction, D + 3 He → α + p. Polarized 3 He avoids the complexities of handling tritium, while encompassing the same spin-physics. We evaluate two methods of delivering deuterium, using dynamically polarized Lithium-Deuteride (with vector polarization P V D of 70%) or frozen-spin Hydrogen-Deuteride (with P V D of 40%), together with a method of injecting optically-pumped 3 He (with 65% polarization). Pellets of these materials all have long polarization decay times (∼6 min for LiD at 2 K, ∼2 months for HD at 2 K, and ∼3 d for 3 He at 77 K), all far greater than a plasma shot in a research tokamak such as DIII-D (∼20 s). Both species can be propelled from a single cryogenic injection gun. We review plasma requirements and strategies for detecting polarization survival. Polarization alters both fusion yields and the angular distribution of fusion products, and each of these provides a potential signal. In this paper we simulate a selection of shots with similar characteristics in a future high-T ion H plasma, and find ratios of yields from shots with fuel spins parallel and antiparallel reaching 1.3 (HD + 3 He) to 1.6 (LiD + 3 He) over a wide range of poloidal angles. (A companion paper finds sensitivity to fusion product angular distributions as reflected in the pitch angles of protons and alphas reaching the plasma facing wall.)

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Examination of Semi-Analytical Solution Methods in the Coarse Operator of Parareal Algorithm for Power System Simulation

With continuing advances in high-performance parallel computing platforms, parallel algorithms have become powerful tools for development of faster than real-time power system dynamic simulations. In particular, it has been demonstrated in recent years that parallel-in-time (Parareal) algorithms have the potential to achieve such an ambitious goal. Here, the selection of a fast and reasonably accurate coarse operator of the Parareal algorithm is crucial for its effective utilization and performance. This paper examines semi-analytical solution (SAS) methods as the coarse operators of the Parareal algorithm and explores performance of the SAS methods to the standard numerical time integration methods. Two promising time-power series-based SAS methods were considered; Adomian decomposition method and Homotopy analysis method with a windowing approach for improving the convergence. Numerical performance case studies on 10-generator 39-bus system and 327-generator 2383-bus system were performed for these coarse operators over different disturbances, evaluating the number of Parareal iterations, computational time, and stability of convergence. All the coarse operators tested with different scenarios have converged to the same corresponding true solution (if they are convergent) and the SAS methods provide comparable computational speed, while having more stable convergence to the true solution in many cases.

97 MATHEMATICS AND COMPUTING↗

An Algorithmic and Software Pipeline for Very Large Scale Scientific Data Compression with Error Guarantees

Efficient data compression is becoming increasingly critical for storing scientific data because many scientific applications produce vast amounts of data. This paper presents an end-to-end algorithmic and software pipeline for data compression that guarantees both error bounds on primary data (PD) and derived data, known as Quantities of Interest (QoI).We demonstrate the effectiveness of the pipeline by compressing fusion data generated by a large-scale fusion code, XGC, which produces tens of petabytes of data in a single day. We demonstrate that the compression is conducted by setting aside computational resources known as staging nodes, and does not impact the simulation performance. For efficient parallel I/O, the pipeline uses ADIOS2, which many codes such as XGC already use for their parallel I/O. We show that our approach can compress the data by two orders of magnitude while guaranteeing high accuracy on both the PD and the QoIs. Further, the amount of resources required by compression is a few percent of the resources required by simulation while ensuring that the compression time for each stage is less than the corresponding simulation time.This pipeline consists of three main steps. The first step decomposes the data using domain decomposition into small subdomains. Each subdomain is then compressed independently to achieve a high level of parallelism. The second step uses existing techniques that guarantee error bounds on the primary data for each subdomain. The third step uses a post-processing optimization technique based on Lagrange multipliers to reduce the QoI errors for data corresponding to each subdomain. The Lagrange multipliers generated can be further quantized or truncated to increase the compression level. All of the above characteristics of our approach make it highly practical to apply on-the-fly compression while guaranteeing errors on QoIs that are critical to the scientists.

Banerjee, Tania↗

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations [Slides]

A separate MCNP6.3 V&V document reports on all the default calculations for all test suites. This report does not include the subcritical multiplication benchmark suite. After some additional clean-up and finalizing the post-processing and documentation steps, the subcritical multiplication benchmark suite will be released in the next version of our vnvstats repository. We tested the new HDF5 PTRAC feature in MCNP6.3 and found encouraging outcomes. Identical results coming out of the simulation with respect to the legacy PTRAC results. The overall runtime for all simulations is reduced by ~20% with the new HDF5 PTRAC capability. We consider giving the new HDF5 PTRAC features a try and using it for all subcritical multiplication and any other relevant (PTRAC) calculations.

97 MATHEMATICS AND COMPUTING↗

Verification and Performance Impact of the New Parallel MCNP6.3 Particle Track Output Capability for Subcritical Multiplication Simulations

The MCNP6® code, version 6.3, has several new features that are intended to ultimately replace legacy features that are now marked for deprecation. One of these features is the new particle track output (PTRAC) format and capability, where the legacy PTRAC capability still exists alongside the modern PTRAC capability in MCNP6.3. While the MCNP6.3 code has been extensively verified and validated for many applications, the PTRAC feature is not exercised in any of the typical verification and validation (V&V) applications studied during the course of a typical MCNP code release. The primary goal of this paper is to verify that the legacy and modern PTRAC feature produces equivalent results for subcritical multiplication benchmarks previously studied. In the process of verifying that the simulated benchmark results are equivalent, the computational performance is compared between the legacy and modern PTRAC uses. In addition to verification of the update, which is important to the community as a whole, this effort also supports advances in the simulation of recent subcritical neutron noise measurements that require higher computational effort per second of real-time measurement than that of systems typically measured.

97 MATHEMATICS AND COMPUTING↗