Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Enabling kilometer-scale E3SM land model simulation over North America: A new integrated framework solution

This study introduces a novel framework designed to enhance the performance, scalability, and portability of the kilometer-scale E3SM Land Model (km-ELM) within the E3SM modeling infrastructure. By seamlessly integrating cutting-edge data tools, we address existing challenges such as slow performance, limited scalability, and difficulties in software integration in current data-driven ELM simulation over large geographic areas. Our innovative approach leverages the KiloCraft data toolkit to generate unified inputs for simulations ranging from a single-cite case, to a 72,083-cell regional case to a continental configuration encompassing 21.6 million land grid cells at a 1 km × 1 km resolution. We conduct extensive strong- and weak-scaling experiments on three state-of-the-art supercomputers, utilizing up to 100,800 CPU cores across 2400 compute nodes to evaluate end-to-end metrics including wall-clock time, simulation-years-per-day (SYPD), initialization costs, and I/O throughput. Our results reveal the land (LND) component’s efficient scaling, demonstrating near-ideal weak scaling and strong-scaling parallel efficiencies reaching up to 87% at 50,400 cores. We confirm portability and reproducibility through bitwise-equivalent outputs across different machines using identical inputs over supported machines. Notably, at extreme scales, we identify I/O as a critical bottleneck and that leads to effective solution with the SCORPIO/ADIOS stack. Collectively, these findings validate the deployment of km-ELM at a continental scale with high parallel efficiency and provide essential guidance on configuration, decomposition, and I/O settings for optimized kilometer-scale land simulations in E3SM. This work emphasizes the innovative design and practical solutions that enhance the operational capabilities of km-ELM, focusing on software performance and scalability while leaving detailed scientific evaluations of simulated land processes for future investigations.

E3SM land model (ELM), km-ELM, scalability, perfor↗

High Order Wall-Modeled Large-Eddy Simulation on Mixed Unstructured Meshes

In the present study, an algebraic equilibrium wall model previously developed for hexahedral elements is extended to handle mixed meshes including prismatic, tetrahedral, and pyramidal elements in the context of a discontinuous high-order method. This extension is needed for complex geometries, for which high-order mixed elements (e.g., tetrahedral and pyramidal elements) are often necessary near solid walls to avoid meshing challenges. Various design decisions are discussed to achieve the best performance on massively parallel CPU/GPU architectures for a production-level high-order large-eddy simulation solver based on the flux reconstruction/correction procedure via reconstruction method, hpMusic. The extension to other elements is first evaluated using a benchmark channel flow problem at various Reynolds numbers. After that, flow over the NASA high-lift Common Research Model (CRM-HL) from the 4th AIAA High-lift Prediction Workshop is computed to further test the new implementation. Computational results at the third- and fourth-order accuracies are compared with experimental data.

Engineering↗

Dynamic Phasor Modeling of Three Phase Voltage Source Inverters

With the increase in the development and implementation of distributed energy resources, application of parallel connected inverters is increasing and as a result having accurate modeling and simulation tools that can help in the design, analysis and stability assessment of the grid is of great importance. Development of fundamental methods that can achieve accurate, reliable, and computationally efficient results can be very beneficial. Application of Dynamic phasor (DP) modeling method has been limited to study of limited harmonics and small combination of interconnected converters due to the complexity associated with developing models that describe larger systems. In this paper, application of DP modeling method is expanded to model any number of parallel connected three phase voltage source inverter (VSI) with inclusion of a wider harmonic content including fundamental, subharmonics, inter-harmonics, switching frequency and their sidebands. Results achieved from this modeling method is compared with conventional average model as well as detailed switching model and the effect of inclusion of wider harmonic content on accuracy of DP modeling method is demonstrated.

Xue, Yaosuo↗

An Adaptive-Mesh-Refinement Based Computational Tool for Simulating Catalysis at Mesoscale

In this work, we present a computational tool for mesoscale applications using open-source exascale- computing compatible adaptive-mesh-refinement (AMR) library, AMReX [2]. AMReX is software library that enables development of application solvers with block-structured Cartesian AMR. Our tool has capabilities to include realistic geometry representation, chemical species transport, reactions and thermodynamics that are critical for capturing mesoscale physics. A significant achievement is the ability of our solver to automatically import electron microscopy data in the form of a stereolithography (STL) or pixelated file format (mrc, tiff) without undergoing the tedious task of unstructured mesh generation. This feature allows for rapid simulation of catalyst particles with complex morphologies using an immersed-boundary formulation. The use of AMR allows for higher resolutions at catalyst surface interfaces, which in turn provides an accurate description of surface reactions and transport. Our solver uses a hybrid distributed and shared memory parallelism (OpenMP/GPU-based) with which strong scaling up to 10,000 processors for realistic catalyst particle simulations have been demonstrated.

BIOMASS FUELS,MATHEMATICS AND COMPUTING↗

Quantum Simulators and Applications on Quantum Framework

Simulating quantum circuits is essential for validating quantum algorithms. However, no single simulator consistently performs best - efficiency depends on circuit structure, entanglement, and depth. In this work, we integrate Qiskit-Aer (state-vector and matrix product state) and QTensor, a tree-tensor-network based simulator, into the Quantum Framework (QFw), a modular platform that supports multiple quantum backends via a unified interface. We also enable distributed quantum approximate optimization algorithm (DQAOA) application compatibility with QFw, allowing sub-problems to be solved in parallel at scale. We then benchmark DQAOA and TFIM (transverse field Ising model) circuits across supported simulators, showing how performance varies significantly with problem type. All simulations are deployed on the Frontier supercomputer using QFw's MPI-based orchestration for distributed, multinode execution. These results underscore the need for simulatoragnostic infrastructure to enable systematic evaluation and highperformance scaling of quantum workloads. QFw provides a practical and extensible path toward reproducible quantum algorithm development across diverse application domains.

Chundury, Srikar [ORNL] (ORCID:0009000183359259)↗

GLEAM: Galaxy Line Emission & Absorption Modeling

We present Galaxy Line Emission & Absorption Modeling (gleam), a Python tool for fitting Gaussian models to emission and absorption lines in large samples of 1D extragalactic spectra. gleam is tailored to work well in batch mode without much human interaction. With gleam, users can uniformly process a variety of spectra, including galaxies and active galactic nuclei, in a wide range of instrument setups and signal-to-noise regimes. gleam also takes advantage of multiprocessing capabilities to process spectra in parallel. With the goal of enabling reproducible workflows for its users, gleam employs a small number of input files, including a central, user-friendly configuration in which fitting constraints can be defined for groups of spectra and overrides can be specified for edge cases. For each spectrum, gleam produces a table containing measurements and error bars for the detected spectral lines and continuum and upper limits for nondetections. For visual inspection and publishing, gleam can also produce plots of the data with fitted lines overlaid. In the present paper, we describe gleam’s main features, the necessary inputs, expected outputs, and some example applications, including thorough tests on a large sample of optical/infrared multi-object spectroscopic observations and integral field spectroscopic data. gleam is developed as an open-source project hosted at https://github.com/multiwavelength/gleam and welcomes community contributions.

79 ASTRONOMY AND ASTROPHYSICS↗

A Practical Framework for Simulating Time-Resolved Spectroscopy Based on a Real-Time Dyson Expansion

Time-resolved spectroscopy is a powerful tool for probing electron dynamics in molecules and solids, revealing transient phenomena on subfemtosecond time scales. The interpretation of experimental results is often enhanced by parallel numerical studies, which can provide insight and validation for experimental hypotheses. However, developing a theoretical framework for simulating time-resolved spectra remains a significant challenge. The most suitable approach involves the many-body nonequilibrium Green's function formalism, which accounts for crucial dynamical many-body correlations during time evolution. While these dynamical correlations are essential for observing emergent behavior in time-resolved spectra, they also render the formalism prohibitively expensive for large-scale simulations. Substantial effort has been devoted to reducing this computational cost─through approximations and numerical techniques─while preserving the key dynamical correlations. The ultimate goal is to enable first-principles simulations of time-dependent systems ranging from small molecules to large, periodic, multidimensional solids. Here, in this perspective, we outline key challenges in developing practical simulations for time-resolved spectroscopy, with a particular focus on Green's function methodologies. We highlight a recent advancement toward a scalable framework: the real-time Dyson expansion (RT-DE) [Phys. Rev. Lett. 2024, 133, 226902]. We introduce the theoretical foundation of RT-DE and discuss strategies for improving scalability, which have already enabled simulations of system sizes beyond the reach of previous fully dynamical approaches. We conclude with an outlook on future directions for extending RT-DE to first-principles studies of dynamically correlated, nonequilibrium systems.

Reeves, Cian C. [Univ. of California, Santa Barbar↗

Radio-Frequency Resonances and Damping in Metallic Magnetic Calorimeter Sensors

Metallic magnetic calorimeters (MMCs) are particle detectors that combine ultra-high energy resolution with a predictable and smooth response based on the physics of paramagnetism. For best energy resolution, MMCs are read out with dc SQUID preamplifiers. Since the ac Josephson effect also makes dc SQUIDs broadband RF sources in the 1–100 GHz range, the SQUID can potentially excite RF modes of the MMC sensor, with negative consequences. Further, the importance of this possibility is magnified in direct-coupled MMCs, where the MMC sensor is part of the SQUID loop to maximize performance. For these reasons, the RF behavior of MMC sensors must be investigated. In this report, we present the results of exploratory RF simulations of MMC sensor modes and damping, and we assess three approaches to damp the parallel-meander direct-coupled MMC without excessive noise increase.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

HOSS: an implementation of the combined finite-discrete element method

Nearly thirty years since its inception, the combined finite-discrete element method (FDEM) has made remarkable strides in becoming a mainstream analysis tool within the field of Computational Mechanics. FDEM was developed to effectively “bridge the gap” between two disparate Computational Mechanics approaches known as the finite and discrete element meth-ods. At Los Alamos National Laboratory (LANL) researchers developed the Hybrid Optimization Software Suite (HOSS) as a hybrid multi-physics platform, based on FDEM, for the simulation of solid material behavior complemented with the latest technological enhancements for full fluid–solid interaction. Furthermore, in HOSS, several newly developed FDEM algorithms have been implemented that yield more accurate material deformation formulations, inter-particle interaction solvers, and fracture and fragmentation solutions. Additionally, an explicit computational fluid dynamics solver and a novel fluid–solid interaction algorithms have been fully integrated (as opposed to coupled) into the HOSS’ solid mechanical solver, allowing for the study of an even wider range of problems. Advancements such as this are leading HOSS to become a tool of choice for multi-physics problems. Finally, HOSS has been successfully applied by a myriad of researchers for analysis in rock mechanics, oil and gas industries, engineering application (structural, mechanical and biomedical engineering), mining, blast loading, high velocity impact, as well as seismic and acoustic analysis. This paper intends to summarize the latest development and application efforts for HOSS.

58 GEOSCIENCES↗

Grand Unification of Quantum Algorithms

Quantum algorithms offer significant speed-ups over their classical counterparts for a variety of problems. The strongest arguments for this advantage are borne by algorithms for quantum search, quantum phase estimation, and Hamiltonian simulation, which appear as subroutines for large families of composite quantum algorithms. A number of these quantum algorithms have recently been tied together by a novel technique known as the quantum singular value transformation (QSVT), which enables one to perform a polynomial transformation of the singular values of a linear operator embedded in a unitary matrix. In the seminal GSLW’19 paper on the QSVT [Gilyén et al., ACM STOC 2019], many algorithms are encompassed, including amplitude amplification, methods for the quantum linear systems problem, and quantum simulation. Here, we provide a pedagogical tutorial through these developments, first illustrating how quantum signal processing may be generalized to the quantum eigenvalue transform, from which the QSVT naturally emerges. Paralleling GSLW’19, we then employ the QSVT to construct intuitive quantum algorithms for search, phase estimation, and Hamiltonian simulation, and also showcase algorithms for the eigenvalue threshold problem and matrix inversion. This overview illustrates how the QSVT is a single framework comprising the three major quantum algorithms, suggesting a grand unification of quantum algorithms.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

A Scalable Semi-Implicit Barotropic Mode Solver for the MPAS-Ocean

A scalable semi-implicit barotropic mode solver for the ocean component of the model for prediction across scales has been implemented as a competitor to an existing explicit-subcycling scheme to allow faster and more stable simulations while not sacrificing accuracy. The semi-implicit solver adopts the pipelined preconditioned bi-conjugate gradient stabilization algorithm as an iterative solver in conjunction with the restricted additive Schwarz preconditioner that accelerates the convergence rate of the iterative solver. The preconditioner is constructed from a linearized barotropic system that also reorders the system for optimal performance, while the semi-implicit solver deals with the fully nonlinear barotropic system that requires reassembly of the coefficient matrix for every time step. Several numerical experiments, from simple one-dimensional tests to three-dimensional real-world tests, demonstrate that the semi-implicit solver has almost the same accuracy and better parallel scalability compared with the existing scheme while allowing faster and more stable simulations. Furthermore, the semi-implicit solver accelerates the barotropic mode up to 2.9 times faster than the existing scheme on 16,320 processors, leading to an overall runtime speedup of 1.9.

97 MATHEMATICS AND COMPUTING↗

Sparse Linear Solvers for Large-scale Electromagnetic Transient Simulations

Linear solvers form the basis for electromagnetic transient (EMT) simulations. There is a need to speed up EMT simulations as larger regions are analyzed using EMT simulations. For the same, the performance of linear solvers plays an important role. Exploiting the sparsity of the matrices generated in EMT simulations could assist with speed-up. Scalability is also crucial as power grids expand, demanding solutions capable of accommodating the increasing system size. Recent studies from the North American Electric Reliability Corporation (NERC) increasingly emphasize that EMT simulation models of the power grid will grow larger with the inclusion of power electronics components. Parallelisms in sparsity patterns exploit modern central processing units (CPUs), multi-core CPUs, and graphics processing units (GPUs) architectures in sparse solver designs. Therefore, this paper explores publicly available existing linear solvers and investigates their efficiency in large-scale power grid simulations. A large-scale power grid is developed by increasing the size of the IEEE 39 bus test system to up to 39000 bus systems.

Hsu, Kuan-Chieh↗

Scoping study of lower hybrid current drive for CFETR

The paper assesses the applicability of lower hybrid current drive (LHCD) for two potential operating scenarios for the China Fusion Engineering Test Reactor (CFETR): the “hybrid” scenario in which some of the plasma current is sustained by the Ohmic transformer, and the fully non-inductive “steady state” scenario. Here, the πScope workflow engine was used to set up a large number of ray tracing/Fokker-Planck simulations (> 10 4 ) with parametric scans in the antenna poloidal position and launched parallel refractive index (n || ) for both the hybrid and steady state scenarios. Modeling predicts efficient off-axis current drive (1.3 MA for 20 MW launched power) with a peak near ρ of 0.6-0.65 for waves launched from the high field side (HFS). Waves launched from the low field side (LFS) damp at larger radius (ρ > 0.73) with similar efficiency to HFS launch. Stability analysis of the CFETR scenarios favors current drive profiles peaked near the mid-radius, suggesting that HFS launch is preferable due to the current drive location. The effect of wave scattering from density blobs in the edge/scrape-off-layer region was assessed through rotation of the perpendicular wavenumber at the ray origin. Simulations show that this effect can be quite large both in efficiency and damping location, however by adjusting the launched n|| much of the unperturbed performance can be recovered.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Exascale models of stellar explosions: Quintessential multi-physics simulation

The ExaStar project aims to deliver an efficient, versatile, and portable software ecosystem for multi-physics astrophysics simulations run on exascale machines. The code suite is a component-based multi-physics toolkit, built on the capabilities of current simulation codes (in particular Flash-X and Castro), and based on the massively parallel adaptive mesh refinement framework AMReX. It includes modules for hydrodynamics, advanced radiation transport, thermonuclear kinetics, and nuclear microphysics. The code will reach exascale efficiency by building upon current multi- and many-core packages integrated into an orchestration system that uses a combination of configuration tools, code translators, and a domain-specific asynchronous runtime to manage performance across a range of platform architectures. The target science includes multi-physics simulations of astrophysical explosions (such as supernovae and neutron star mergers) to understand the cosmic origin of the elements and the fundamental physics of matter and neutrinos under extreme conditions.

79 ASTRONOMY AND ASTROPHYSICS↗

Simulating of magnetic reconnection in solar flares

Magnetic reconnection is an astrophysical process where neighboring magnetic field lines running anti-parallel are reconfigured. Observations of solar flares are what pushed for serious research regarding magnetic reconnection as it seemed to be the underlying mechanism. The reconfiguration of the field lines results in an explosive release of energy as magnetic field energy is converted into plasma kinetic and thermal energies. Our work begins with Athena++, an astrophysical magnetohydrodynamic (MHD) simulation code, and an existing magnetic reconnection setup.

79 ASTRONOMY AND ASTROPHYSICS↗

Prediction and uncertainty quantification of shale well performance using multifidelity Monte Carlo

Uncertainty quantification is an integral component of reservoir management, especially considering the inherent uncertainty in subsurface systems. While a standard practice to estimate the uncertainty, Monte Carlo (MC) simulation is computationally intense when the sampling population comprises high-fidelity simulations. Alternatively, the Multi-fidelity Monte Carlo (MFMC) simulation overcomes this computational intensity by integrating low- and high-fidelity simulations. Our goal is to minimize the number of expensive high-fidelity simulations while maintaining accuracy and using numerous fast and cheap low-fidelity simulations to efficiently sample to input parameter space of interest. We selected gas production from unconventional wells to demonstrate the potential speedups and accuracy of the MFMC approach. The model fidelity usually determines the trade-off between accuracy and efficiency. While the high-fidelity model is more accurate, the low-fidelity model is more efficient. Our high-fidelity simulation includes reservoir simulations of a hydraulically fractured well. On the other hand, our low-fidelity model comprises the parallel-plate flow model. We used differential programming to efficiently solve the 1D flow model, where automatic differentiation is used to efficiently compute the gradients. We matched the production profile of high-fidelity simulations with our low-fidelity simulations. Then, we used a support vector regression to map the high- and low-fidelity input parameters. The mapping function is essential to tune the low-dimensional parameter space of the low-fidelity model to the high-dimensional parameter space of the high-fidelity model. We found that we can use a combination of 9 high fidelity and 10,000 low fidelity simulations to efficiently and accurately simulate pressure management. This method is at least two orders of magnitude faster than only using high-fidelity simulations. Finally, from a broader perspective, MFMC could efficiently estimate the uncertainty of various systems and models, integrating low- and high-fidelity models.

04 OIL SHALES AND TAR SANDS↗

Direct numerical simulations for hybrid rocket boundary layers: Performance modeling and scaling

This paper presents a comprehensive performance and scaling analysis of direct numerical simulations for reacting boundary layers, focusing on slab burner configurations. Using a PETSc-based finite volume CFD framework, the study evaluates the scalability and computational cost of flow, chemistry, and radiation evaluations across 2D and 3D simulations. Polymethyl methacrylate (PMMA) is the fuel with pure O 2 as the oxidizer, modeled using a detailed chemical kinetics mechanism with 113 species and 660 reactions. A ray-tracing-based radiation solver, designed for distributed memory applications, is implemented to model radiation heat transfer. Parallel scalability is analyzed for the coupled flow, chemistry, and radiation heat transfer processes. Weak and strong scaling studies are conducted on up to 15,000 computational ranks, revealing robust performance when flow cells exceed 200 per rank. Chemistry evaluations dominate the computational cost in large 3D simulations, accounting for approximately 40% of the total runtime, while flow processes contribute around 35%, and radiation solver contributions remain below 10% due to reduced evaluation frequencies. GPU accelerated chemistry evaluation, implemented with Zero-RK, demonstrates significant promise, achieving up to a 4x speedup for workloads exceeding 30,000 cells per GPU. However, diminishing returns are observed for smaller workloads due to CPU-GPU communication overhead. This study identifies key challenges, including memory bottlenecks and the effects of domain partitioning on flow scalability, while highlighting the potential of GPU-accelerated chemistry to reduce computational costs. In conclusion, these findings provide realizable run configurations for 2D, 3D, and GPU-accelerated cases, offering insights for optimizing reactive flow solvers.

CFD Scalability↗

Pilgrim Hot Springs: GEOPHIRES Inputs and Outputs for Direct-Use Geothermal District Heating and Cooling

This dataset includes files for a techno-economic analysis conducted using the GEOPHIRES simulator to examine the feasibility of expanding a larger district heating site in a remote location: Pilgrim Hot Springs, Alaska. Files included here are GEOPHIRES inputs and outputs for five different scenarios with varying demand, cycle, and system design characteristics to analyze. Also included is the link to the GEOPHIRES GitHub, as well as a link to the dataset that contains the energy modelling used to determine the heating demand for the district. For a list of the differences between scenarios, see the included "Input Overview.txt" file. Fields included in the input files are: subsurface technical parameters, surface technical parameters, financial parameters, capital and O&M parameters, as well as simulation parameters. The output files are case reports that summarize all equipment, reservoir characteristics, costs, and heating profiles.

15 GEOTHERMAL ENERGY↗