Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,279 records · Page 71

Plastic Parallel Pathways Platform - 4P Model

Global momentum is building towards a circular economy capable of keeping plastics in use and out of waste streams. Given that 79% of all plastic produced since 1950 has accumulated in landfills or the natural environment,rapid implementation of various end-of-life (EoL) management technologies will be needed to reach this target. However, it can be challenging to develop an effective plastic EoL strategy when the available options - chemical or molecular recycling, energy recovery, upcycling, downcycling, closed-loop (plastic-to-plastic) or open-loop (plastic-to-x) recycling, among others - can generate products ranging from low-grade to virgin-quality plastic and from fuels to value-added chemicals. We present a flexible material flow model capable of analyzing the effects of both plastic-to-plastic and plastic-to-x EoL management strategies on the U.S. PET economy. This Plastic Parallel Pathways Platform (4P) assesses the environmental impacts, costs, and circularity of a PET system in which waste is managed through six potential EoL pathways: landfill, incineration with energy recovery, pyrolysis to fuel oil, upcycling to glass fiber reinforced plastic (GFRP), mechanical recycling to low-grade PET, and chemical recycling (glycolysis) to bottle-grade PET. We compare the pathways across multiple metrics using multi-criteria decision analysis (MCDA) and then use a brute force algorithm to predict an optimal combination of EoL pathways to minimize greenhouse gas (GHG) emissions and costs and maximize circularity. This work highlights the need to implement a diverse portfolio of EoL strategies in parallel to enable a PET economy that meets environmental, economic, and circularity requirements simultaneously.

downcycling↗

Parallel performance of algebraic multigrid domain decomposition

Algebraic multigrid (AMG) is a widely used scalable solver and preconditioner for large-scale linear systems resulting from the discretization of a wide class of elliptic PDEs. While AMG has optimal computational complexity, the cost of communication has become a significant bottleneck that limits its scalability as processor counts continue to grow on modern machines. This article examines the design, implementation, and parallel performance of a novel algorithm, algebraic multigrid domain decomposition (AMG-DD), designed specifically to limit communication. The goal of AMG-DD is to provide a low-communication alternative to standard AMG V-cycles by trading some additional computational overhead for a significant reduction in communication cost. Numerical results show that AMG-DD achieves superior accuracy per communication cost compared with AMG, and speedup over AMG is demonstrated on a large GPU cluster.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING↗

Solving larger maximum clique problems using parallel quantum annealing

Quantum annealing has the potential to find low energy solutions of NP-hard problems that can be expressed as quadratic unconstrained binary optimization problems. However, the hardware of the quantum annealer manufactured by D-Wave Systems, which we consider in this work, is sparsely connected and moderately sized (on the order of thousands of qubits), thus necessitating a minor-embedding of a logical problem onto the physical qubit hardware. The combination of relatively small hardware sizes and the necessity of a minor-embedding can mean that solving large optimization problems is not possible on current quantum annealers. In this research, we show that a hybrid approach combining parallel quantum annealing with graph decomposition allows one to solve larger optimization problem accurately. We apply the approach to the Maximum Clique problem on graphs with up to 120 nodes and 6395 edges.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Parallelized POD-based suboptimal economic model predictive control of a state-constrained Boussinesq approximation

Motivated by an energy efficient building application, we want to optimize a quadratic cost functional subject to the Boussinesq approximation of the Navier-Stokes equations and to bilateral state and control constraints. Since the computation of such an optimal solution is numerically costly, we design an efficient strategy to compute a sub-optimal (but applicationally acceptable) solution with significantly reduced computational effort. We employ an economic Model Predictive Control (MPC) strategy to obtain a feedback control. The MPC sub-problems are based on a linear-quadratic optimal control problem subjected to mixed control and state constraints and a convection-diffusion equation, reduced with proper orthogonal decomposition. Finally, to solve each sub-problem, we apply a primal-dual active set strategy. The method can be fully parallelized, which enables the solution of large problems with real-world parameters.

97 MATHEMATICS AND COMPUTING↗

Spatiotemporal parallelization of an analytical heat conduction model for additive manufacturing via a hybrid OpenMP + MPI approach

The ability to do thermal simulations for entire additive manufacturing builds is a key computational problem facing the additive manufacturing community; however, complex numerical models considering multiple physical phenomena currently do not have the capacity for simulations at this scale. To this end, conduction only analytic models offer a viable approach due to the massive drop in computational expense. In this work, we extend an existing implementation which uses a governing equation which can be evaluated at any point in space and time. This implementation already utilizes OpenMP with a spatial decompositions scheme stemming from a melt pool tracking algorithm. Furthermore, we then combine this with a parallel in time (PinT) approach to make the problem highly parallelizable. The new scheme, which uses MPI for internode communication and OpenMP for intranode communication, is shown to scale very well across multiple computational nodes. This approach results in the ability to simulate the 3D solidification conditions for entire layers of additively manufactured parts in minutes making part scale thermal simulations more practical.

36 MATERIALS SCIENCE↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Parallel computing for power system climate resiliency: Solving a large-scale stochastic capacity expansion problem with mpi-sppy

Here we propose a nodal stochastic generation and transmission expansion planning model that incorporates the output from high-resolution global climate models through load and generation availability scenarios. We implement our model in Pyomo and perform computational studies on a realistically-sized test case of the California electric grid in a high performance computing environment. We propose model reformulations and algorithm tuning to efficiently solve this large problem using a variant of the Progressive Hedging Algorithm. We utilize the parallelization capabilities and overall versatility of mpi-sppy, exploiting its hub-and-spoke architecture to concurrently obtain inner and outer bounds on an optimal expansion plan. Initial results show that instances with 360 representative days on a system with over 8,000 buses can be solved to within 5% of optimality in under 4 h of wall clock time, a first step towards solving a large-scale power system expansion planning problem across a wide range of climate-informed operational scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Microgrid energy scheduling under uncertain extreme weather: Adaptation from parallelized reinforcement learning agents

Microgrids are useful solutions for integrating renewable energy resources and providing seamless green electricity to minimize carbon footprint. In recent years, extreme weather events happened often worldwide and caused significant economic and societal losses. Such events bring uncertainties to the microgrid energy scheduling problems and increase the challenges of microgrid operation. Traditional optimization approaches suffer from the inaccuracy of the uncertain microgrid model and the unseen events. Existing reinforcement learning (RL) - based approaches are also hampered by the limited generalization and the increasing computational burden when stochastic formulations are required to accommodate the uncertainties. This paper proposes a new parallelized reinforcement learning (PRL) method based on the probabilistic events to handle the microgrid energy uncertainties. Specifically, several local learning agents are employed to interact with pertinent microgrid environments in a distributed manner and report outcomes to the global agent, which will optimize microgrid energy resources online during extreme events. The stochastic microgrid energy optimization problem is reformulated to include all possible scenarios with probabilities. The advantage estimate functions of learning agents are designed with a backward sweep to transfer the outcomes to the value function updating process. Two simulation studies, stochastic optimization and online testing, are performed to compare with several existing RL approaches. Results substantiate that the proposed PRL method can achieve up to 20% optimization performance improvement with 4 and 28 times less computation cost than Q-learning with experience replay and multi-agent Q-learning approaches, respectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

HPLC-Parallel accelerator and molecular mass spectrometry analysis of 14 C-labeled amino acids

Accelerator mass spectrometry (AMS) is the method of choice for quantitation of low amounts of 14 C-labeled biomolecules. Despite exquisite sensitivity, an important limitation of AMS is its inability to provide structural information about the analyte. This limitation is not critical when the labeled compounds are well-characterized prior to AMS analysis. However, analyte identity is important in other experiments where, for example, a compound is metabolized and the structures of its metabolites are not known. We previously described a moving wire interface that enables direct AMS measurement of liquid sample in the form of discrete drops or HPLC eluent without the need for individual fraction collection, termed liquid sample-AMS (LS-AMS). Here, we now report the coupling of LS-AMS with a molecular mass spectrometer, providing parallel accelerator and molecular mass spectrometry (PAMMS) detection of analytes separated by liquid chromatography. The repeatability of the method was examined by performing repeated injections of 14 C-labeled tryptophan, and relative standard deviations of the 14 C peak areas were ≤10.57% after applying a normalization factor based on a standard. Five 14 C-labeled amino acids were separated and detected to provide simultaneous quantitative AMS and structural MS data, and AMS results were compared with solid sample-AMS (SS-AMS) data using Bland-Altman plots. To demonstrate the utility of the workflow, yeast cells were grown in a medium with 14 C-labeled tryptophan. The cell extracts were analyzed by PAMMS, and 14 C was detected in tryptophan and its metabolite kynurenine.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Rate-induced aging effects on Parallel-Plate Avalanche Counter (PPAC) caused by heavy ion beams

The Facility for Rare Isotope Beams (FRIB) is one of the premier scientific user facilities for nuclear science with radioactive beams, capable of producing most (approximately 80%) of the isotopes expected to exist, from oxygen to uranium, at energies up to 200 MeV/u. With the increase in beam power from the present 10 kW to the planned 400 kW, FRIB experiments are about to enter a new era. An unprecedented rate capability as well as stable performance of all the planned instrumentation intended for beam diagnostics and beam tuning is required at the expected high beam intensities (> 1 MHz). A summary of aging phenomena at high heavy-ion beam rates observed in the Advanced Rare Isotope Separator (ARIS) detectors for beam diagnostics, including Parallel Plate Avalanche Counters (PPAC) and plastic scintillation for time-of-flight measurements, is discussed. Current research and development project to mitigate rate-induced aging are presented.

Aging effects↗

Complex Fluid‐Driven Fractures Caused by Crack‐Parallel Stress

Abstract Managing fluid‐driven fracture networks is crucial for subsurface resource utilization, yet the current understanding of the key controlling factors remains insufficient. While geologic discontinuities have been shown to significantly influence fracture network complexity, this study identifies another major contributor. We conducted a new set of experiments using a transparent true triaxial cell, which enabled video recording of the temporal evolution of fluid‐driven fracture paths. Using pseudo‐2D samples without macroscale structural discontinuities, we observed multiple occurrences of hydraulic fracture curving and branching under anisotropic boundary stresses. We proposed a theoretical model demonstrating that the stress parallel to the crack line in the solid matrix near the crack tip (i.e., the T ‐stress) accounts for the observed fracture curving behavior. This finding suggests that T ‐stress is an additional mechanism contributing to the complexity of fluid‐driven fracture networks in the subsurface, besides the geologic discontinuities.

58 GEOSCIENCES↗

The impact of non-local parallel electron transport on plasma-impurity reaction rates in tokamak scrape-off layer plasmas

Abstract Plasma-impurity reaction rates are a crucial part of modelling tokamak scrape-off layer (SOL) plasmas. To avoid calculating the full set of rates for the large number of important processes involved, a set of effective rates are typically derived which assume Maxwellian electrons. However, non-local parallel electron transport may result in non-Maxwellian electrons, particularly close to divertor targets. Here, the validity of using Maxwellian-averaged rates in this context is investigated by computing the full set of rate equations for a fixed plasma background from kinetic and fluid SOL simulations. We consider the effect of the electron distribution as well as the impact of the electron transport model on plasma profiles. Results are presented for lithium, beryllium, carbon, nitrogen, neon and argon. It is found that electron distributions with enhanced high-energy tails can result in significant modifications to the ionisation balance and radiative power loss rates from excitation, on the order of 50%–75% for the latter. Fluid electron models with Spitzer-Härm or flux-limited Spitzer-Härm thermal conductivity, combined with Maxwellian electrons for rate calculations, can increase or decrease this error, depending on the impurity species and plasma conditions. Based on these results, we also discuss some approaches to experimentally observing non-local electron transport in SOL plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Role of perturbed parallel magnetic field effects in predicting turbulent transport in NSTX

This study presents analysis of gyrokinetic simulations on the National Spherical Torus Experiment (NSTX) to investigate the effects of electromagnetic fields on plasma turbulence and transport. The simulations, performed with varying levels of fidelity using the gyrokinetic CGYRO code, include electrostatic (ES), single-field electromagnetic (EM1), and two-field electromagnetic (EM2) models. A detailed comparison across the simulation database reveals that electromagnetic effects increase both predicted growth rates and quasilinear fluxes, with EM2 simulations producing stronger turbulence than ES and EM1 cases. Quasilinear modeling using QLGYRO demonstrates that while the perturbed parallel magnetic field (δB ∥ ) does not drastically affect the total flux at experimental gradients, it leads to a shift in the dominant instability, altering mode structures from microtearing to kinetic ballooning modes (KBMs). The proximity of the plasma profiles to the KBM threshold is explored, with the experimental conditions being near the onset of KBM-driven transport. The KBM, with its large growth rates, is identified as a potential driver of electron temperature flattening, as it can rapidly transport heat across flux surfaces. Performing stability analysis shows core-localized unstable a low- mode that could contribute to the flattening at the early times of the discharge. TGYRO predictive modeling, incorporating both TGLF and QLGYRO, indicates that the inclusion of δB ∥ significantly improves the accuracy of temperature profile predictions in NSTX high-beta plasmas, although challenges remain in modeling the sharp flux discontinuities caused by KBM-driven instabilities.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel Algorithms for Efficient Computation of High-Order Line Graphs of Hypergraphs

This paper considers structures of systems beyond dyadic (pairwise) interactions and investigates mathematical modeling of multi-way interactions and connections as hypergraphs, where captured relationships among system entities are set-valued. To date, in most situations, entities in a hypergraph are considered connected as long as there is at least one common ``neighbor''. However, minimal commonality sometimes discards the ``strength'' of connections and interactions among groups. To this end, considering the ``width'' of a connection, referred to as the \emph{$s$-overlap} of neighbors, provides more meaningful insights into how closely the communities or entities interact with each other. In addition, $s$-overlap computation is the fundamental kernel to construct the line graph of a hypergraph, a low-order approximation of the hypergraph which can carry significant information about the original hypergraph. Subsequent stages of a data analytics pipeline then can apply highly-tuned graph algorithms on the line graph to reveal important features. Given a hypergraph, computing the $s$-overlaps by exhaustively considering all pairwise entities can be computationally prohibitive. To tackle this challenge, we develop efficient algorithms to compute $s$-overlaps and the corresponding line graph of a hypergraph. We propose several heuristics to avoid execution of redundant work and improve performance of the $s$-overlap computation. Our parallel algorithm, combined with these heuristics, is orders of magnitude (more than $10\times$) faster than the naive algorithm in all cases and the SpGEMM algorithm with filtration in most cases (especially with large $s$ value).

hypergraph algorithms, graph algorithms, parallel ↗

HYPPO: A Surrogate-Based Multi-Level Parallelism Tool for Hyperparameter Optimization

We present a new software, HYPPO, that enables the automatic tuning of hyperparameters of various deep learning (DL) models. Unlike other hyperparameter optimization (HPO) methods, HYPPO uses adaptive surrogate models and directly accounts for uncertainty in model predictions to find accurate and reliable models that make robust predictions. Using asynchronous nested parallelism, we are able to significantly alleviate the computational burden of training complex architectures and quantifying the uncertainty. HYPPO is implemented in Python and can be used with both TensorFlow and PyTorch libraries. We demonstrate various software features on time-series prediction and image classification problems as well as a scientific application in computed tomography image reconstruction. Finally, we show that (1) we can reduce by an order of magnitude the number of evaluations necessary to find the most optimal region in the hyperparameter space and (2) we can reduce by two orders of magnitude the throughput for such HPO process to complete.

adaptation models↗