Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

MAESTROeX

MAESTROeX is a massively parallel, finite-volume C++/F90 solver for low Mach number astrophysical flows. The code utilizes a low Mach number equation set allowing for more efficient, long-time integration of highly subsonic flows compared to compressible approaches. The recommended range of applicability is for flows where the Mach number does not exceed~0.1.

Harpole, Alice↗

TINES - Time Integration, Newton and Eigen Solver v. 1.0

SAND2021-1505 O. TINES is an open source software providing math infrastructure for solving many stiff time ordinary differential equations (ODEs) and/or differential algebraic equations (DAEs) using a batch hierarchical parallelism. The code is written using a parallel programming model (i.e., Kokkos) to future-proof the next generation parallel computing platforms such as GPU accelerators. This code is developed to support Exascale Catalytic Chemistry (ECC) Project. The code provides fundamental math helpers that can aid other research projects. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525.

Kim, Kyungjoo↗

Modeling transient edge plasma transport with dynamic recycling

The work presents numerical simulation studies of the role that dynamic plasma recycling on the main wall and divertor target surfaces plays in transient edge plasma transport phenomena, such as edge localized modes (ELMs). The studies are performed by coupling the edge plasma transport code UEDGE [Rognlien et al., J. Nucl. Mater. 196–198, 347 (1992)] and the wall reaction–diffusion transport code FACE [Smirnov et al., Fusion Sci. Technol. 71, 75 (2017)]. The two-dimensional, time-dependent, two-way coupling of the codes, in a realistic tokamak geometry, is accomplished using the Integrated Plasma Simulator framework [Elwasif et al., in 18th Euromicro Conference on Parallel, Distributed and Network-Based Processing (PDP 2010), Pisa, Italy (IEEE, 2010), pp. 419–427] for all modeled material plasma boundaries. The simulations show that dynamic plasma recycling has substantially different characteristics on the main wall and on the divertor plates. It is demonstrated that during an ELM cycle the outer wall can dynamically absorb and release a number of particles comparable to that expelled by the ELM from the core plasma, by far exceeding the dynamic retention capacity of the divertor surfaces. The resulting evolution of the edge and divertor plasma conditions during an ELM cycle is analyzed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Uncertainty Visualization of Critical Points of 2D Scalar Fields for Parametric and Nonparametric Probabilistic Models

This paper presents a novel end-to-end framework for closed-form computation and visualization of critical point uncertainty in 2D uncertain scalar fields. Critical points are fundamental topological descriptors used in the visualization and analysis of scalar fields. The uncertainty inherent in data (e.g., observational and experimental data, approximations in simulations, and compression), however, creates uncertainty regarding critical point positions. Uncertainty in critical point positions, therefore, cannot be ignored, given their impact on downstream data analysis tasks. Here, in this work, we study uncertainty in critical points as a function of uncertainty in data modeled with probability distributions. Although Monte Carlo (MC) sampling techniques have been used in prior studies to quantify critical point uncertainty, they are often expensive and are infrequently used in production-quality visualization software. We, therefore, propose a new end-to-end framework to address these challenges that comprises a threefold contribution. First, we derive the critical point uncertainty in closed form, which is more accurate and efficient than the conventional MC sampling methods. Specifically, we provide the closed-form and semianalytical (a mix of closed-form and MC methods) solutions for parametric (e.g., uniform, Epanechnikov) and nonparametric models (e.g., histograms) with finite support. Second, we accelerate critical point probability computations using a parallel implementation with the VTK-m library, which is platform portable. Finally, we demonstrate the integration of our implementation with the ParaView software system to demonstrate near-real-time results for real datasets.

97 MATHEMATICS AND COMPUTING↗

Three-Dimensional Grid Visualization for Planning Activities: A Dubai Case Study

National Laboratory of the Rockies (NLR), in collaboration with the Dubai Electricity and Water Authority (DEWA) and Infra-X, has undertaken the Energy Visualization Analysis Project. The aim of this project is to enhance analytical and 3D visualization capabilities for distribution network planning and renewable energy integration. As modern grid continues to evolve with large-scale solar PV deployment and emerging distributed energy resources (DERs), the ability to effectively analyze, visualize, and communicate complex grid behaviors has become increasingly critical. The project focuses on developing empirical use cases based on real distribution feeder data and engineering workflows, ensuring the outcomes are directly aligned with operational environment. Through time-series power flow simulations and nodal hosting capacity analysis, the study quantifies the impacts of high PV penetration on voltage and thermal limits within representative 11 kV feeders. These analyses identify specific nodes and conditions where DER integration challenges arise. Furthermore, a Battery Energy Storage System (BESS) optimization algorithm was applied to determine the optimal size and placement of storage systems that can mitigate network constraints and enhance hosting capacity. The comparative results between base-case and BESS-augmented scenarios clearly demonstrate improvements in network stability and load management efficiency. In parallel, the NLR team developed an immersive 3D visualization framework, enabling interactive exploration of grid simulations using commodity head-mounted display (HMD) systems. This framework transforms conventional 2D simulation data into spatially intuitive visual environments - allowing engineers to analyze feeder conditions, PV hosting potential, and BESS effects in real time. This report represents the first foundational phase in establishing a visualization-driven analytical ecosystem. It provides a methodological foundation for data integration, visualization architecture, and simulation-based decision support, paving the way for large-scale adoption of immersive visualization across DEWA's Smart Grid Initiative, R&D activities, and future network resilience studies.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Intelligent Partitioning based Fully Parallel AC Security-Constrained Optimal Power Flow

Today’s power grid is becoming more diverse and integrated with high-level distributed energy resources and smart control technologies that is creating a new set of grid management challenges in terms of large-scale, nonlinear, and non-convex problem modeling, complex and time-consuming computation, as well as difficult uncertainty handling. This project focused on solving a challenging multi-period security-constrained generation scheduling problem, which is of great importance for maximizing the social welfare of real-time dispatch, day-ahead market, as well as weekly planning of power systems. Our developed software explored parallel optimization algorithms for complex and realistic power system models, and develop fast, efficient, and robust grid optimization solutions on the high-performance computing platform that will enable increased grid economics, flexibility, resilience, as well as energy security in the United States.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A fully implicit, asymptotic-preserving, semi-Lagrangian algorithm for the time dependent anisotropic heat transport equation

In this paper, we extend the operator-split asymptotic-preserving, semi-Lagrangian algorithm for time dependent anisotropic heat transport equation proposed in Chacón et al. (2014) [18] to use a fully implicit time integration with backward differentiation formulas. The proposed implicit method can deal with arbitrary heat-transport anisotropy ratios $\mathcal{X}$∥ /$ \mathcal{X}$⟂ $\ggg$ 1 (with $\mathcal{X}$∥, $ \mathcal{X}$⟂ the parallel and perpendicular heat diffusivities, respectively) in complicated magnetic field topologies in an accurate and efficient manner. Further, the implicit algorithm is second-order accurate temporally and demonstrates an accurate treatment at boundary layers (e.g., island separatrices), which was not ensured by the operator-split implementation. The condition number of the resulting algebraic system is independent of the anisotropy ratio, and is inverted with preconditioned GMRES. We propose a simple preconditioner that renders the finite-dimensional linear operator compact, resulting in mesh-independent convergence rates for topologically simple magnetic fields, and convergence rates scaling as ~ (NΔt) 1/4 (with N the total mesh size and Δt the timestep) in topologically complex magnetic-field configurations. We demonstrate the accuracy and performance of the approach with test problems of varying complexity, including an analytically tractable boundary-layer problem in a straight magnetic field, and a topologically complex magnetic field featuring magnetic islands with extreme anisotropy ratios $\mathcal{X}$∥ /$ \mathcal{X}$⟂ = 10 10 ) .

97 MATHEMATICS AND COMPUTING↗

Heterogeneous techniques for rescaling energy deposits in the CMS Phase-2 endcap calorimeter

We present the porting to heterogeneous architectures of the algorithm used for applying linear transformations of raw energy deposits in the CMS High Granularity Calorimeter (HGCAL). This is the first heterogeneous algorithm to be fully integrated with HGCAL’s reconstruction chain. After introducing the latter and giving a brief description of the structural components of HGCAL relevant for this work, the role of the linear transformations in the calibration is reviewed. The many ways in which parallelization is achieved are described, and the successful validation of the heterogeneous algorithm is covered. Detailed performance measurements are presented, including throughput and execution time for both CPU and GPU algorithms, therefore establishing the corresponding speedup. We finally discuss the interplay between this work and the porting of other algorithms in the existing reconstruction chain, as well as integrating algorithms previously ported but not yet integrated.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Oxide-nitride heteroepitaxy for low-loss dielectrics in superconducting quantum circuits

Superconducting qubits show great promise for the realization of fault-tolerant quantum computing, but lossy, amorphous dielectrics limit current technology. Identifying highly crystalline and stoichiometric dielectrics with intrinsically low microwave loss is therefore a central materials challenge, yet experimentally validated platforms remain scarce. In this work, we integrate a crystalline dielectric into a heteroepitaxial TiN/$γ$-Al$_2$O$_3$/TiN trilayer grown via pulsed laser deposition. Correlative high-resolution imaging, diffraction, and spectroscopy measurements confirm the single-crystal quality and chemical integrity of all layers, with minimal defects and limited anion interdiffusion across the oxide-nitride interfaces. Using microwave lumped-element resonators with parallel-plate capacitors, we report the first direct measurement of the dielectric loss of epitaxial $γ$-Al$_2$O$_3$, for which we find a low intrinsic two-level system loss, $δ_{\text{TLS}}^0 = (2.8 \pm 0.1) \times 10^{-5}$. These results establish heteroepitaxial oxides on transition metal nitrides as an attractive materials platform for superconducting quantum circuits, particularly for integration into compact device architectures such as merged-element transmons and microwave kinetic inductance detectors.

Garcia-Wetten, David A. [Northwestern U.]↗

3D Coded SUMMA: Communication-Efficient and Robust Parallel Matrix Multiplication

In this paper, we propose a novel fault-tolerant parallel matrix multiplication algorithm called 3D Coded SUMMA that achieves higher failure-tolerance than replication-based schemes for the same amount of redundancy. This work bridges the gap between recent developments in coded computing and fault-tolerance in high-performance computing (HPC). The core idea of coded computing is the same as algorithm-based fault-tolerance (ABFT), which is weaving redundancy in the computation using error-correcting codes. In particular, we show that MatDot codes, an innovative code construction for parallel matrix multiplications, can be integrated into three-dimensional SUMMA (Scalable Universal Matrix Multiplication Algorithm [30]) in a communication-avoiding manner. To tolerate any two node failures, the proposed 3D Coded SUMMA requires ~50% less redundancy than replication, while the overhead in execution time is only about 5–10%.

97 MATHEMATICS AND COMPUTING↗

Spinbox: tools for many-body quantum systems in a Monte Carlo context

Spinbox is a piece of software that facilitates quantum mechanical calculations relevant to Monte Carlo simulation of atomic nuclei. At the front lines of research on the nuclear many-body problem are a large number of supercomputer-scale simulation codes. These codes produce valuable results but can be hard to understand, especially for those without intimate knowledge of the relevant theoretical methods. Thus, tools that fill pedagogical roles are extremely valuable. Spinbox makes it easy for one to replicate and analyze the computational processes relevant to a Quantum Monte Carlo (QMC) simulation that may be difficult to understand/debug/analyze due to the scale of the corresponding simulation software. Spinbox is written in Python using other state-of-the-art Python modules for numerical calculations. While a number of Python libraries exist that are suited to general quantum many-body calculations, the motivation of Spinbox is quite particular. In Diffusion Monte Carlo methods (DMC, GFMC, AFDMC), the central calculation is the imaginary-time propagation of individual samples of the many-body wavefunction. Although quantum wavefunctions generally must be described by a probability distribution over a basis, DMC imbues particles (within one sample) with classical spatial coordinates. This method is unusual, so other Python packages are typically not set up to do this easily. Furthermore, the software has built-in options for nuclear systems assuming isospin symmetry, which can be set up with other libraries but is a nontrivial process to do so. Features: - numerical representation of samples of the many-body wavefunctions, including tensor-product states (used in AFDMC) - numerical representation of many-body operators, including tensor-product operators: general, spin, imaginary-time propagation, etc. - the correct associated arithmetic and algebra, implemented as class methods - classes for representing realistic nuclear two- and three-body Hamiltonians (e.g. Argonne V18, Illinois NNN) - large-scale parallel integration over random variables, crucial for the AFDMC method My goal is to make this package open source so that anyone may use it and contribute to it, particularly other researchers doing AFDMC calculations

Fox, Jordan↗

Characterization of a Pixelated Cadmium Telluride Detector System Using a Polychromatic X-Ray Source and Gold Nanoparticle-Loaded Phantoms for Benchtop X-Ray Fluorescence Imaging

In this paper, the imaging dose and scan time have been considered as the two major constraints for routine benchtop x-ray fluorescence computed tomography (XFCT) imaging. One way to address this issue is to acquire x-ray fluorescence (XRF) signals in parallel through a 2D array of single-crystal detectors or a pixelated detector along with the cone-beam x-ray source. To identify a detector system suitable for this purpose, a commercially available, fully spectroscopic cadmium telluride (CdTe) pixelated detector, HEXITEC (High-Energy X-ray Imaging Technology), was tested under the experimental conditions optimized for benchtop XFCT imaging of gold nanoparticles (GNPs). Specifically, two different parallel-hole stainless steel collimators were fabricated and coupled with the detector for seamless integration into our existing benchtop cone-beam XFCT system. After the detector deployment, this benchtop XFCT system was used to detect XRF photons from GNP-loaded phantoms. A pixel-merging algorithm was introduced to enhance the sensitivity of XRF photon detection thereby minimizing the scan time. The effect of pixel-level charge sharing correction algorithms was investigated within the context of benchtop XFCT imaging. The detector energy resolution, in terms of the full width at half maximum (FWHM) values at different gold K-shell XRF energies, was also determined. Of the two charge sharing correction algorithms examined, the charge sharing addition gave better sensitivity than the charge sharing discrimination (csd). On the other hand, under the current experimental conditions, the energy resolution of the HEXITEC detector was the best with the csd and estimated to be 1.56 keV FWHM at 66-69 keV photon energy. Overall, despite some degradation of the detector energy resolution (compared with typical single crystal CdTe detectors), the HEXITEC detector enabled parallel data acquisition under the experimental conditions typical of benchtop XFCT imaging and operated well within our benchtop XFCT setup.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Virtual Time III, Part 1: Unified Virtual Time Synchronization for Parallel Discrete Event Simulation

Algorithms for synchronization of parallel discrete event simulation have historically been divided between conservative methods that require lookahead but not rollback, and optimistic methods that require rollback but not lookahead. In this paper we present a new approach in the form of a framework called Unified Virtual Time (UVT) that unifies the two approaches, combining the advantages of both within a single synchronization theory. Whenever timely lookahead information is available, a logical process (LP) executes conservatively using an irreversible event handler. When lookahead information is not available the LP does not block, as it would in a classical conservative execution, but instead executes optimistically using a reversible event handler. The switch from conservative to optimistic synchronization and back is decided on an event-by-event basis by the simulator, transparently to the model code. UVT treats conservative synchronization algorithms as optional accelerators for an underlying optimistic synchronization algorithm, enabling the speed of conservative execution whenever it is applicable, but otherwise falling back on the generality of optimistic execution. We describe UVT in a novel way, based on fundamental invariants, monotonicity requirements, and synchronization rules. UVT permits zero-delay messages and pays careful attention to tie-handling using superposition. We prove that under fairly general conditions a UVT simulation always makes progress in virtual time. This is Part 1 of a trio of papers describing the UVT framework for PDES, mixing conservative and optimistic synchronization and integrating throttling control.

97 MATHEMATICS AND COMPUTING↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Stochastic evaluation of fourth-order many-body perturbation energies

A scalable, stochastic algorithm evaluating the fourth-order many-body perturbation (MP4) correction to energy is proposed. Three hundred Goldstone diagrams representing the MP4 correction are computer generated and then converted into algebraic formulas expressed in terms of Green’s functions in real space and imaginary time. They are evaluated by the direct (i.e., non-Markov, non-Metropolis) Monte Carlo (MC) integration accelerated by the redundant-walker and control-variate algorithms. The resulting MC-MP4 method is efficiently parallelized and is shown to display O(n 5.3 ) size-dependence of cost, which is nearly two ranks lower than the O(n 7 ) dependence of the deterministic MP4 algorithm. Furthermore, it evaluates the MP4/aug-cc-pVDZ energy for benzene, naphthalene, phenanthrene, and corannulene with the statistical uncertainty of 10 mE h (1.1% of the total basis-set correlation energy), 38 mE h (2.6%), 110 mE h (5.5%), and 280 mE h (9.0%), respectively, after about 10 9 MC steps.

74 ATOMIC AND MOLECULAR PHYSICS↗

Determining the tilt of the Raman laser beam using an optical method for atom gravimeters

The tilt of a Raman laser beam is a major systematic error in precision gravity measurement using atom interferometry. The conventional approach to evaluating this tilt error involves modulating the direction of the Raman laser beam and conducting time-consuming gravity measurements to identify the error minimum. In this work, we demonstrate a method to expediently determine the tilt of the Raman laser beam by transforming the tilt angle measurement into characterization of parallelism, which integrates the optical method of aligning the laser direction, commonly used in freely falling corner-cube gravimeters, into an atom gravimeter. A position-sensing detector (PSD) is utilized to quantitatively characterize the parallelism between the test beam and the reference beam, thus measuring the tilt precisely and rapidly. After carefully positioning the PSD and calibrating the relationship between the distance measured by the PSD and the tilt angle measured by the tiltmeter, we achieved a statistical uncertainty of less than 30 µrad in the tilt measurement. Furthermore, we compared the results obtained through this optical method with those from the conventional tilt modulation method for gravity measurement. The comparison validates that our optical method can achieve tilt determination with an accuracy level of better than 200 µrad, corresponding to a systematic error of 20 µGal in g measurement. This work has practical implications for real-world applications of atom gravimeters.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Dynamic Response of a Semiactive Suspension System with Hysteretic Nonlinear Energy Sink Based on Random Excitation by means of Computer Simulation

This paper aims to investigate the property and behavior of the hysteretic nonlinear energy sink (HNES) coupled to a half vehicle system which is a nine-degree-of-freedom, nonlinear, and semiactive suspension system in order to improve the ride comfort and increase the stability in shock mitigation by using the computer simulation method. The HNES model is a semiactive suspension device, which comprises the famous Bouc–Wen (B-W) model employed to describe the force produced by both the purely hysteretic spring and linear elastic spring of potentially negative stiffness connected in parallel, for the half vehicle system. Nine nonlinear motion equations of the half vehicle system are derived in terms of the seven displacements and the two dimensionless hysteretic variables, which are integrated numerically by employing the direct time integration method for studying both the variables of vertical displacements, velocities, accelerations, chassis pitch angle, and the ride comfort and driver safety, respectively, based on the bump and random road inputs of the pseudoexcitation method as excitation signal. Simulation results show that, compared with the HNES model and the magnetorheological (MR) model coupled to the half vehicle system, the ride comfort and stability have been evidently improved. A successful validation process has been performed, which indicated that both the ride comfort and driver safety properties of the HNES model coupled to half vehicle significantly improved.

Chen, Hui↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗