Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 343 records · Page 19

The impact of non-local parallel electron transport on plasma-impurity reaction rates in tokamak scrape-off layer plasmas

Abstract Plasma-impurity reaction rates are a crucial part of modelling tokamak scrape-off layer (SOL) plasmas. To avoid calculating the full set of rates for the large number of important processes involved, a set of effective rates are typically derived which assume Maxwellian electrons. However, non-local parallel electron transport may result in non-Maxwellian electrons, particularly close to divertor targets. Here, the validity of using Maxwellian-averaged rates in this context is investigated by computing the full set of rate equations for a fixed plasma background from kinetic and fluid SOL simulations. We consider the effect of the electron distribution as well as the impact of the electron transport model on plasma profiles. Results are presented for lithium, beryllium, carbon, nitrogen, neon and argon. It is found that electron distributions with enhanced high-energy tails can result in significant modifications to the ionisation balance and radiative power loss rates from excitation, on the order of 50%–75% for the latter. Fluid electron models with Spitzer-Härm or flux-limited Spitzer-Härm thermal conductivity, combined with Maxwellian electrons for rate calculations, can increase or decrease this error, depending on the impurity species and plasma conditions. Based on these results, we also discuss some approaches to experimentally observing non-local electron transport in SOL plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Bayesian model updating with finite element vs surrogate models: Application to a miter gate structural system

Bayesian finite element (FE) model updating using direct model evaluations of large-scale high-fidelity FE models is extremely computationally expensive. Surrogate models can be used as fast emulators of FE models to accelerate the model calibration process. The physics/mechanics-based FE models are still the underpinning behind the surrogate models. Here, this paper evaluates the loss in accuracy and the gain in computational time while performing Bayesian model updating by using surrogate model evaluations compared to using direct FE model evaluations. This evaluation is crucial before entirely relying on surrogate models in model updating for structural health monitoring (SHM) and damage prognosis (DP) purposes. This paper also demonstrates Bayesian updating and surrogate model construction of large-scale high-fidelity FE models of infrastructure systems. In this regard, the miter gate structural system is considered as the testbed structure. Three predominant damage modes (loss of contact between gate and wall, loss of thickness due to corrosion, and loss of tension in the diagonal rods) are considered for model updating purposes. Bayesian model updating is performed using direct FE evaluations by leveraging parallel computing. Two types of surrogates, namely polynomial chaos expansion (PCE) and Gaussian process regression (GPR), are developed for the miter gate. Model updating is performed again using the trained surrogate models, and the updating results are compared with their counterparts obtained using the direct FE evaluation results. The posterior distribution of the FE model parameters obtained using the trained surrogates are sufficiently accurate with respect to the posterior obtained utilizing the direct FE evaluations. In addition, an approximate 4-fold decrease in the computational time was observed when using surrogate model evaluations instead of direct FE evaluations for model updating.

42 ENGINEERING↗

A novel approach to increase accuracy in remotely sensed evapotranspiration through basin water balance and flux tower constraints

Remote sensing-derived evapotranspiration (RSET) products capture the spatiotemporal variations of evapotranspiration (ET) from field to basin scales with unprecedented details. However, their accuracy varies across RSET estimation methods and diverse hydroclimate regions. While ET modeling efforts to account for biophysical processes and controlling parameters have made good progress in recent years, a parallel approach of integrating in-situ ET with RSET could reduce biases in RSET products. Basin water balance ET (WBET) and flux tower ET are widely applied to evaluate RSET accuracy, yet such ET measurements are rarely used for RSET bias corrections, especially for large area applications. To address this issue, we propose a novel approach: the water balance equivalence (WABE) method, which generates spatially continuous WBET for correcting biases in RSET products. The WABE method computes synthetic WBET by integrating observed WBET and flux tower-derived FLUXCOM ET, which fills the spatial gaps of observed WBET and generates a spatially continuous WBET dataset. Synthetic WBET (2002–2015 annual average) of eight-digit hydrologic unit code (HUC8) basins across the conterminous United States (CONUS), constituting 44 % (887 out of 2035 basins) of CONUS basins, was determined within 2.0 % (RMSE = 12 %) of observed WBET at CONUS and between 1–12 % (RMSE = 3–33 %) across 18 regions in CONUS. With WABE-based bias corrections, the overall annual bias of RSET decreased from 10 % (RMSE = 34 %) to 6 % (RMSE = 26 %) across 37 flux tower sites. The WABE method offers a new approach for RSET accuracy improvement and shows great promise for large area implementations with a potential to yield substantial benefits for building accurate basin water budgets and water management decisions.

Khand, Kul↗

How eigenmode self-interaction affects zonal flows and convergence of tokamak core turbulence with toroidal system size

Self-interaction is the process by which a microinstability eigenmode that is extended along the direction parallel to the magnetic field interacts non-linearly with itself. This effect is particularly significant in gyrokinetic simulations accounting for kinetic passing electron dynamics and is known to generate stationary $E\times B$ zonal flow shear layers at radial locations near low-order mode rational surfaces (Weikl et al. Phys. Plasmas, vol. 25, 2018, 072305). Here, we find that self-interaction, in fact, plays a very significant role in also generating fluctuating zonal flows, which is critical to regulating turbulent transport throughout the radial extent. Unlike the usual picture of zonal flow drive in which microinstability eigenmodes coherently amplify the flow via modulational instabilities, the self-interaction drive of zonal flows from these eigenmodes are uncorrelated with each other. It is shown that the associated shearing rate of the fluctuating zonal flows therefore reduces as more toroidal modes are resolved in the simulation. In simulations accounting for the full toroidal domain, such an increase in the density of toroidal modes corresponds to an increase in the toroidal system size, leading to a finite system size effect that is distinct from the well-known profile shearing effect.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Global Impacts of Marine Methanethiol Emissions and Chemistry in the Atmosphere

Oceanic emissions of dimethyl sulfide (DMS) have long been known to influence aerosol particle composition, cloud condensation nuclei (CCN) concentration, and Earth’s radiative budget. However, the impact of oceanic emissions of methanethiol (MeSH), a sulfur compound produced by the same oceanic precursor as DMS, has been relatively less explored. The gas-phase oxidation of MeSH has a higher effective yield of SO 2 and a shorter oxidative lifetime compared to DMS, highlighting the relevance of this pathway for the modeled representation of particle formation, growth, and CCN abundance in the marine atmosphere. Here, we use the global chemical transport model GEOS-Chem to explore possible scenarios representative of specific environmental conditions and MeSH emission schemes based on previous experimental studies. We further implement and test previously reported chemical mechanisms for MeSH oxidation, along with additional improvements, highlighting key uncertainties and sensitivities for regional and global sulfur budgets. We place our results in the context of recent modeling updates to DMS chemistry and cloud processing, which further impact SO 2 production in the marine atmosphere in parallel with MeSH oxidation. Within the overall marine sulfur budget, our findings highlight that MeSH plays a significant role in SO 2 production in the marine atmosphere, contributing to regional surface layer concentration increases of up to 40–60%. These results point to the importance of MeSH for efforts aimed at improving the modeled representation of sulfur spatiotemporal patterns relevant to air quality predictions and climate impact assessments.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Electronic coupling in square planar La 4 Ni 3 O 8

A study of a dd excitation in La4Ni 3 O 8 (La-438) using x-ray absorption scattering (XAS) and resonant inelastic x-ray scattering (RIXS) at the Ni K-edge is presented here. The incident energy dependence of this dd excitation shows a maximum at the 1s→4p π transition. Its intensity at the main edge is proportional to the amount of incident x-ray polarization parallel to the c-axis. These observations suggest that the RIXS process underlying this excitation includes a strong Ni 3d-Ni 4p Coulomb interaction and excludes the “4p-as-spectator” approximation. The dominant Ni 3d Coulomb interaction is with Ni 4p π with limited or no interaction with the Ni 4p σ . An insulating gap closing is observed as a function of temperature.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Fast neural Poincaré maps for toroidal magnetic fields

Poincaré maps for toroidal magnetic fields are routinely employed to study gross confinement properties in devices built to contain hot plasmas. In most practical applications, evaluating a Poincaré map requires numerical integration of a magnetic field line, a process that can be slow and that cannot be easily accelerated using parallel computations. We propose a novel neural network architecture, the HénonNet, and show that it is capable of accurately learning realistic Poincaré maps from observations of a conventional field-line-following algorithm. After training, such learned Poincaré maps evaluate much faster than the field-line integration method. Moreover, the HénonNet architecture exactly reproduces the primary physics constraint imposed on field-line Poincaré maps: flux preservation. Furthermore, this structure-preserving property is the consequence of each layer in a HénonNet being a symplectic map. We demonstrate empirically that a HénonNet can learn to mock the confinement properties of a large magnetic island by using coiled hyperbolic invariant manifolds to produce a sticky chaotic region at the desired island location. This suggests a novel approach to designing magnetic fields with good confinement properties that may be more flexible than ensuring confinement using KAM tori.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Controlling colloidal crystals via morphing energy landscapes and reinforcement learning

We report a feedback control method to remove grain boundaries and produce circular shaped colloidal crystals using morphing energy landscapes and reinforcement learning–based policies. We demonstrate this approach in optical microscopy and computer simulation experiments for colloidal particles in ac electric fields. First, we discover how tunable energy landscape shapes and orientations enhance grain boundary motion and crystal morphology relaxation. Next, reinforcement learning is used to develop an optimized control policy to actuate morphing energy landscapes to produce defect-free crystals orders of magnitude faster than natural relaxation times. Morphing energy landscapes mechanistically enable rapid crystal repair via anisotropic stresses to control defect and shape relaxation without melting. This method is scalable for up to at least N = 10 3 particles with mean process times scaling as N 0.5 . Further scalability is possible by controlling parallel local energy landscapes (e.g., periodic landscapes) to generate large-scale global defect-free hierarchical structures.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

FuseIM: Fusing Probabilistic Traversals for Influence Maximization on Exascale Systems

Probabilistic breadth-first traversals (BPTs) are used in many network science and graph machine learning applications. In this paper, we are motivated by the application of BPTs in stochastic diffusion-based graph problems such as influence maximization. These applications heavily rely on BPTs to implement a Monte-Carlo sampling step for their approximations. Given the large sampling complexity, stochasticity of the diffusion process, and the inherent irregularity in real-world graph topologies, efficiently parallelizing these BPTs remains significantly challenging. In this paper, we present a new algorithm to fuse massive number of concurrently executing BPTs with random starts on the input graph. Our algorithm is designed to fuse BPTs by combining separate traversals into a unified frontier on distributed multi-GPU systems. To show the general applicability of the fused BPT technique, we have incorporated it into two state-of-the-art influence maximization parallel implementations (gIM and Ripples). Our experiments on up to 4K nodes of the OLCF Frontier supercomputer (32,768 GPUs and 196K CPU cores) show strong scaling behavior, and that fused BPTs can improve the performance of these implementations up to 34x (for gIM) and ~360x (for Ripples).

Neff, Reece W.↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗

MAUD Interface Tool Kit (MILK)

Materials Analysis Using Diffraction (MAUD) is an open source Rietveld refinement program which fits diffraction models to diffraction spectra allowing quantitative diffraction analysis. The scripting interface developed here using python facilitates the choice of parameters to refine during the Rietveld fitting process and provides a simple MPI framework for running MAUD instances in parallel. The scripting language facilitates batch, custom, and reproducible Rietveld refinements of large datasets.

Savage, Daniel↗

System and method of storing and analyzing information

A system and method of storing and analyzing information is disclosed. The system includes a compiler layer to convert user queries to data parallel executable code. The system further includes a library of multithreaded algorithms, processes, and data structures. The system also includes a multithreaded runtime library for implementing compiled code at runtime. The executable code is dynamically loaded on computing elements and contains calls to the library of multithreaded algorithms, processes, and data structures and the multithreaded runtime library.

Feo, John T.↗

FULL RANGE TUNE SCAN STUDIES USING GRAPHICS PROCESSING UNITS WITH CUDA IN EIC BEAM-BEAM SIMULATIONS

The hadron beam in the Electron-Ion Collider (EIC) suffers high order betatron and synchro-betatron resonances. In this paper, we present a weak-strong full range (0.0 ~ 0.5) fractional tune scan with a step size as small as 0.001. Multiple Graphics Processing Units (GPUs) are used to speed up the simulation. A code parallelized with MPI and CUDA is implemented. The good tune region from weak-strong scan is further checked by the self-consistent strong-strong simulation. This study provides beam dynamics guidance in choosing proper working points for the future EIC.

43 PARTICLE ACCELERATORS↗

The MPACT 2020 Milestone: Safeguards and Security by Design of Future Nuclear Fuel Cycle Facilities.

The Materials Protection, Accounting, and Control Technologies (MPACT) campaign, within the U.S. Department of Energy Office of Nuclear Energy, has developed a Virtual Facility Distributed Test Bed for safeguards and security design for future nuclear fuel cycle facilities. The purpose of the Virtual Test Bed is to bring together experimental and modeling capabilities across the U.S. national laboratory and university complex to provide a one-stop-shop for advanced Safeguards and Security by Design (SSBD). Experimental testing alone of safeguards and security technologies would be cost prohibitive, but testbeds and laboratory processing facilities with safeguards measurement opportunities, coupled with modeling and simulation, provide the ability to generate modern, efficient safeguards and security systems for new facilities. This Virtual Test Bed concept has been demonstrated using a generic electrochemical reprocessing facility as an example, but the concept can be extended to other facilities. While much of the recent work in the MPACT program has focused on electrochemical safeguards and security technologies, the laboratory capabilities have been applied to other facilities in the past (including aqueous reprocessing, fuel fabrication, and molten salt reactors as examples). This paper provides an overview of the Virtual Test Bed concept, a description of the design process, and a baseline safeguards and security design for the example facility. Parallel papers in this issue go into more detail on the various technologies, experimental testing, modeling capabilities, and performance testing.

98 NUCLEAR DISARMAMENT, SAFEGUARDS, AND PHYSICAL P↗

Kink mechanism in Cu/Nb nanolaminates explored by $\mathcal{in}$ $\mathcal{situ}$ pillar compression

We report Nano metallic laminates (NMLs) exhibit different failure modes depending on the loading conditions due to their mechanical anisotropies. Kinking is a typical failure mode in many NMLs compressed along a layer-parallel direction. However, a detailed description of the microstructure evolution during kink band (KB) formation and an in-depth understanding of the formation mechanisms are lacking. In this work, the KB process is investigated in Cu/Nb NMLs by in situ micro pillar compression in the scanning electron microscope (SEM) along a layer-parallel direction. Post-mortem S/TEM and transmission Kikuchi diffraction (TKD) analyses show that kink banding leads to significant microstructure changes characterized by an accumulation of geometrically necessary dislocations (GNDs) and of tilt geometrically necessary boundaries (GNBs) near KB boundaries (KBBs). The distinct microstructure evolution implies that KB formation is facilitated by the inhomogeneous microstructures resulting in constrained deformation modes. Specifically, dislocations active on slip planes nearly parallel to the interfaces make a major contribution to kink evolution after the onset of kinking. Once layer-parallel slip systems are activated, preexisting lattice dislocations and dislocations nucleating from interfaces will accumulate as GNDs near KBBs via the stochastic storage of lattice dislocations that have certain Burgers vectors. GNDs can further transform into GNBs via cross-slip and climb driven processes near the KBB. Furthermore, GNBs near KBBs can grow by incorporating more GNDs or by coalescence to accommodate the KB evolution. We further hypothesize that microstructural perturbations and their ensuing stresses can initiate KB formation in Cu/Nb NMLs.

36 MATERIALS SCIENCE↗

Scalable quantum computational science: A perspective from block-encodings and polynomial transformations

Significant developments made in quantum hardware and error correction recently have been driving quantum computing toward practical utility. However, gaps remain between abstract quantum algorithmic development and practical applications in computational sciences. In this perspective article, we propose several properties that scalable quantum computational science methods should possess. We further discuss how block-encodings and polynomial transformations can potentially serve as a unified framework with the desired properties. Recent advancements on these topics are presented, including the construction and assembly of block-encodings, and various generalizations of quantum signal processing (QSP) algorithms to perform polynomial transformations. The scalability of QSP methods on parallel and distributed quantum architectures is also highlighted. Promising applications in simulation and observable estimation in chemistry, physics, and optimization problems are presented. We hope this perspective serves as a gentle introduction to state-of-the-art quantum algorithms for the computational science community and inspires future development of scalable quantum computational science methodologies that bridge theory and practice.

Bayesian inference↗

Effect of sintering temperature on adhesion of spray-on piezoelectric transducers

Conventionally sol-gel spray-on transducers require a high-temperature (> 700 ◦C) sintering process; however, this process can affect the microstructure of the substrate material. For mechanical elbows and valves utilized for fluid transport in the energy sector, the components are designed to have a specific microstructure, and deviations from these specifications can create weak points in the system. For this reason it is important to investigate how the temperature of the deposition process affects the substrate. This paper investigates the effect of high-temperature and low-temperature (< 150 ◦C) processing conditions on the surface composition of the substrate. Furthermore, the resultant transducers from high- and low-temperature fabrication processes are compared to determine if a low-temperature processing method is feasible. For these studies a sol-gel spray-on process is employed to deposit piezoelectric ceramics onto a stainless-steel 316L substrate. Energy-dispersive X-ray spectroscopy is utilized to determine the composition of the substrate surface before and after transducer deposition. Results indicate that the high-temperature processing conditions may alter the surface composition of the metal due to a diffusion of the metal into the ceramic, which results in a metal surface that is bonded to the ceramic. Furthermore, it is shown that low-temperature processing of spray-on transducers is a viable method for transducer fabrication where the resultant transducers meet the industry minimum requirement of 30 dB signalto-noise ratio. In parallel simulation calculations, finite-element method (FEM) studies were performed to model the adhesive strength of the low-temperature processed transducer to the substrate surface. Comparisons between the simulations and experiments suggest that the bond strength is much greater than the commercial gel bonds and closer to hardened epoxy glue bonds. These results indicate that spray-on transducers fabricated under lowtemperature processing conditions are a viable solution for leave-in-place monitoring of structures.

M. Sinding, Kyle↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗