Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

In-depth analysis on parallel processing patterns for high-performance Dataframes

The Data Science domain has expanded monumentally in both research and industry communities during the past decade, predominantly owing to the Big Data revolution. Artificial Intelligence (AI) and Machine Learning (ML) are bringing more complexities to data engineering applications, which are now integrated into data processing pipelines to process terabytes of data. Typically, a significant amount of time is spent on data preprocessing in these pipelines, and hence improving its efficiency directly impacts the overall pipeline performance. The community has recently embraced the concept of Dataframes as the de-facto data structure for data representation and manipulation. However, the most widely used serial Dataframes today (R, pandas) experience performance limitations while working on even moderately large data sets. We believe that there is plenty of room for improvement by taking a look at this problem from a high-performance computing point of view. In a prior publication, we presented a set of parallel processing patterns for distributed dataframe operators and the reference runtime implementation, Cylon. In this paper, we are expanding on the initial concept by introducing a cost model for evaluating the said patterns. Furthermore, we evaluate the performance of Cylon on the ORNL Summit supercomputer.

97 MATHEMATICS AND COMPUTING↗

Microgrid energy scheduling under uncertain extreme weather: Adaptation from parallelized reinforcement learning agents

Microgrids are useful solutions for integrating renewable energy resources and providing seamless green electricity to minimize carbon footprint. In recent years, extreme weather events happened often worldwide and caused significant economic and societal losses. Such events bring uncertainties to the microgrid energy scheduling problems and increase the challenges of microgrid operation. Traditional optimization approaches suffer from the inaccuracy of the uncertain microgrid model and the unseen events. Existing reinforcement learning (RL) - based approaches are also hampered by the limited generalization and the increasing computational burden when stochastic formulations are required to accommodate the uncertainties. This paper proposes a new parallelized reinforcement learning (PRL) method based on the probabilistic events to handle the microgrid energy uncertainties. Specifically, several local learning agents are employed to interact with pertinent microgrid environments in a distributed manner and report outcomes to the global agent, which will optimize microgrid energy resources online during extreme events. The stochastic microgrid energy optimization problem is reformulated to include all possible scenarios with probabilities. The advantage estimate functions of learning agents are designed with a backward sweep to transfer the outcomes to the value function updating process. Two simulation studies, stochastic optimization and online testing, are performed to compare with several existing RL approaches. Results substantiate that the proposed PRL method can achieve up to 20% optimization performance improvement with 4 and 28 times less computation cost than Q-learning with experience replay and multi-agent Q-learning approaches, respectively.

24 POWER TRANSMISSION AND DISTRIBUTION↗

HPLC-Parallel accelerator and molecular mass spectrometry analysis of 14 C-labeled amino acids

Accelerator mass spectrometry (AMS) is the method of choice for quantitation of low amounts of 14 C-labeled biomolecules. Despite exquisite sensitivity, an important limitation of AMS is its inability to provide structural information about the analyte. This limitation is not critical when the labeled compounds are well-characterized prior to AMS analysis. However, analyte identity is important in other experiments where, for example, a compound is metabolized and the structures of its metabolites are not known. We previously described a moving wire interface that enables direct AMS measurement of liquid sample in the form of discrete drops or HPLC eluent without the need for individual fraction collection, termed liquid sample-AMS (LS-AMS). Here, we now report the coupling of LS-AMS with a molecular mass spectrometer, providing parallel accelerator and molecular mass spectrometry (PAMMS) detection of analytes separated by liquid chromatography. The repeatability of the method was examined by performing repeated injections of 14 C-labeled tryptophan, and relative standard deviations of the 14 C peak areas were ≤10.57% after applying a normalization factor based on a standard. Five 14 C-labeled amino acids were separated and detected to provide simultaneous quantitative AMS and structural MS data, and AMS results were compared with solid sample-AMS (SS-AMS) data using Bland-Altman plots. To demonstrate the utility of the workflow, yeast cells were grown in a medium with 14 C-labeled tryptophan. The cell extracts were analyzed by PAMMS, and 14 C was detected in tryptophan and its metabolite kynurenine.

38 RADIATION CHEMISTRY, RADIOCHEMISTRY, AND NUCLEA↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Rate-induced aging effects on Parallel-Plate Avalanche Counter (PPAC) caused by heavy ion beams

The Facility for Rare Isotope Beams (FRIB) is one of the premier scientific user facilities for nuclear science with radioactive beams, capable of producing most (approximately 80%) of the isotopes expected to exist, from oxygen to uranium, at energies up to 200 MeV/u. With the increase in beam power from the present 10 kW to the planned 400 kW, FRIB experiments are about to enter a new era. An unprecedented rate capability as well as stable performance of all the planned instrumentation intended for beam diagnostics and beam tuning is required at the expected high beam intensities (> 1 MHz). A summary of aging phenomena at high heavy-ion beam rates observed in the Advanced Rare Isotope Separator (ARIS) detectors for beam diagnostics, including Parallel Plate Avalanche Counters (PPAC) and plastic scintillation for time-of-flight measurements, is discussed. Current research and development project to mitigate rate-induced aging are presented.

Aging effects↗

Complex Fluid‐Driven Fractures Caused by Crack‐Parallel Stress

Abstract Managing fluid‐driven fracture networks is crucial for subsurface resource utilization, yet the current understanding of the key controlling factors remains insufficient. While geologic discontinuities have been shown to significantly influence fracture network complexity, this study identifies another major contributor. We conducted a new set of experiments using a transparent true triaxial cell, which enabled video recording of the temporal evolution of fluid‐driven fracture paths. Using pseudo‐2D samples without macroscale structural discontinuities, we observed multiple occurrences of hydraulic fracture curving and branching under anisotropic boundary stresses. We proposed a theoretical model demonstrating that the stress parallel to the crack line in the solid matrix near the crack tip (i.e., the T ‐stress) accounts for the observed fracture curving behavior. This finding suggests that T ‐stress is an additional mechanism contributing to the complexity of fluid‐driven fracture networks in the subsurface, besides the geologic discontinuities.

58 GEOSCIENCES↗

The impact of non-local parallel electron transport on plasma-impurity reaction rates in tokamak scrape-off layer plasmas

Abstract Plasma-impurity reaction rates are a crucial part of modelling tokamak scrape-off layer (SOL) plasmas. To avoid calculating the full set of rates for the large number of important processes involved, a set of effective rates are typically derived which assume Maxwellian electrons. However, non-local parallel electron transport may result in non-Maxwellian electrons, particularly close to divertor targets. Here, the validity of using Maxwellian-averaged rates in this context is investigated by computing the full set of rate equations for a fixed plasma background from kinetic and fluid SOL simulations. We consider the effect of the electron distribution as well as the impact of the electron transport model on plasma profiles. Results are presented for lithium, beryllium, carbon, nitrogen, neon and argon. It is found that electron distributions with enhanced high-energy tails can result in significant modifications to the ionisation balance and radiative power loss rates from excitation, on the order of 50%–75% for the latter. Fluid electron models with Spitzer-Härm or flux-limited Spitzer-Härm thermal conductivity, combined with Maxwellian electrons for rate calculations, can increase or decrease this error, depending on the impurity species and plasma conditions. Based on these results, we also discuss some approaches to experimentally observing non-local electron transport in SOL plasmas.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Role of perturbed parallel magnetic field effects in predicting turbulent transport in NSTX

This study presents analysis of gyrokinetic simulations on the National Spherical Torus Experiment (NSTX) to investigate the effects of electromagnetic fields on plasma turbulence and transport. The simulations, performed with varying levels of fidelity using the gyrokinetic CGYRO code, include electrostatic (ES), single-field electromagnetic (EM1), and two-field electromagnetic (EM2) models. A detailed comparison across the simulation database reveals that electromagnetic effects increase both predicted growth rates and quasilinear fluxes, with EM2 simulations producing stronger turbulence than ES and EM1 cases. Quasilinear modeling using QLGYRO demonstrates that while the perturbed parallel magnetic field (δB ∥ ) does not drastically affect the total flux at experimental gradients, it leads to a shift in the dominant instability, altering mode structures from microtearing to kinetic ballooning modes (KBMs). The proximity of the plasma profiles to the KBM threshold is explored, with the experimental conditions being near the onset of KBM-driven transport. The KBM, with its large growth rates, is identified as a potential driver of electron temperature flattening, as it can rapidly transport heat across flux surfaces. Performing stability analysis shows core-localized unstable a low- mode that could contribute to the flattening at the early times of the discharge. TGYRO predictive modeling, incorporating both TGLF and QLGYRO, indicates that the inclusion of δB ∥ significantly improves the accuracy of temperature profile predictions in NSTX high-beta plasmas, although challenges remain in modeling the sharp flux discontinuities caused by KBM-driven instabilities.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallel Algorithms for Efficient Computation of High-Order Line Graphs of Hypergraphs

This paper considers structures of systems beyond dyadic (pairwise) interactions and investigates mathematical modeling of multi-way interactions and connections as hypergraphs, where captured relationships among system entities are set-valued. To date, in most situations, entities in a hypergraph are considered connected as long as there is at least one common ``neighbor''. However, minimal commonality sometimes discards the ``strength'' of connections and interactions among groups. To this end, considering the ``width'' of a connection, referred to as the \emph{$s$-overlap} of neighbors, provides more meaningful insights into how closely the communities or entities interact with each other. In addition, $s$-overlap computation is the fundamental kernel to construct the line graph of a hypergraph, a low-order approximation of the hypergraph which can carry significant information about the original hypergraph. Subsequent stages of a data analytics pipeline then can apply highly-tuned graph algorithms on the line graph to reveal important features. Given a hypergraph, computing the $s$-overlaps by exhaustively considering all pairwise entities can be computationally prohibitive. To tackle this challenge, we develop efficient algorithms to compute $s$-overlaps and the corresponding line graph of a hypergraph. We propose several heuristics to avoid execution of redundant work and improve performance of the $s$-overlap computation. Our parallel algorithm, combined with these heuristics, is orders of magnitude (more than $10\times$) faster than the naive algorithm in all cases and the SpGEMM algorithm with filtration in most cases (especially with large $s$ value).

hypergraph algorithms, graph algorithms, parallel ↗

HYPPO: A Surrogate-Based Multi-Level Parallelism Tool for Hyperparameter Optimization

We present a new software, HYPPO, that enables the automatic tuning of hyperparameters of various deep learning (DL) models. Unlike other hyperparameter optimization (HPO) methods, HYPPO uses adaptive surrogate models and directly accounts for uncertainty in model predictions to find accurate and reliable models that make robust predictions. Using asynchronous nested parallelism, we are able to significantly alleviate the computational burden of training complex architectures and quantifying the uncertainty. HYPPO is implemented in Python and can be used with both TensorFlow and PyTorch libraries. We demonstrate various software features on time-series prediction and image classification problems as well as a scientific application in computed tomography image reconstruction. Finally, we show that (1) we can reduce by an order of magnitude the number of evaluations necessary to find the most optimal region in the hyperparameter space and (2) we can reduce by two orders of magnitude the throughput for such HPO process to complete.

adaptation models↗

Virtual Time III, Part 1: Unified Virtual Time Synchronization for Parallel Discrete Event Simulation

Algorithms for synchronization of parallel discrete event simulation have historically been divided between conservative methods that require lookahead but not rollback, and optimistic methods that require rollback but not lookahead. In this paper we present a new approach in the form of a framework called Unified Virtual Time (UVT) that unifies the two approaches, combining the advantages of both within a single synchronization theory. Whenever timely lookahead information is available, a logical process (LP) executes conservatively using an irreversible event handler. When lookahead information is not available the LP does not block, as it would in a classical conservative execution, but instead executes optimistically using a reversible event handler. The switch from conservative to optimistic synchronization and back is decided on an event-by-event basis by the simulator, transparently to the model code. UVT treats conservative synchronization algorithms as optional accelerators for an underlying optimistic synchronization algorithm, enabling the speed of conservative execution whenever it is applicable, but otherwise falling back on the generality of optimistic execution. We describe UVT in a novel way, based on fundamental invariants, monotonicity requirements, and synchronization rules. UVT permits zero-delay messages and pays careful attention to tie-handling using superposition. We prove that under fairly general conditions a UVT simulation always makes progress in virtual time. This is Part 1 of a trio of papers describing the UVT framework for PDES, mixing conservative and optimistic synchronization and integrating throttling control.

97 MATHEMATICS AND COMPUTING↗

Parallel transport sweeps on two-dimensional cartesian and hexagonal grids

This paper aims to provide a proof of concept for parallel transport sweeps on two-dimensional hexagonal grids for the discrete ordinates transport equation. While the method is an extension of the popular and well-established Koch-Baker-Alcoulffe (KBA) algorithm, there are significant differences between the cartesian and hexagonal grid and thereafter sweep. The most important is the three-way connectivity of hexagons within the grid which creates greater dependencies between the elements. The KBA method in structured orthogonal grids was first implemented in the DRAGON5 code and the method is first described here. The differences in implementation for the hexagonal grid are also described. Benchmark results are also presented, showing roughly 10 times speedup in computational times with roughly 100 processors, in both cases. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

VerifyIO: Ensuring Correctness of Consistency Semantics in Parallel I/O

Abstract—High-performance computing (HPC) applications generate and consume substantial amounts of data, typically managed by parallel file systems. These applications access file systems either through the POSIX interface or by using highlevel I/O libraries. While the POSIX consistency model remains dominant in HPC, emerging file systems and popular I/O libraries increasingly adopt alternative consistency models that relax semantics in various ways, creating significant challenges for correctness and portability. This paper addresses these challenges by proposing a trace-driven I/O consistency verification workflow, implemented in our open-source tool, VerifyIO, which collects execution traces, detects data conflicts, and verifies proper synchronization against specified consistency models. Our extensive evaluation of 91 test case executions across three widely used I/O libraries with four I/O consistency models reveals critical consistency issues at both application and implementation levels.

Consistency Semantics↗

Revealing Parallel Inter‐ and Intra‐Ligand Charge Transfer Dynamics in [Ru(L) 2 (dppz)] 2+ Molecular Lightswitch with N K‐Edge X‐Ray Absorption Spectroscopy

In photoactive metal complexes the localization of photoexcited charges dictates the site of chemical reactivity, but few studies measure the charge redistribution in these systems with spatial precision. Herein, we track the inter- and intra-ligand charge transfer processes that underpin light-driven charge separation in the well-studied “molecular lightswitch” [Ru(bpy) 2 dppz] 2+ (aqueous [Ruthenium II (2,2′-bipyridine)2(dipyrido[3,2-a:2′,3′-c]phenazine)] 2+ [Cl − ] 2 ) by probing the electronic structure of ligand nitrogen atoms in real-time using ultrafast X-ray absorption spectroscopy and first principles calculations. We confirm the localization of excited electron density on the phenazine N atoms of dppz and we newly identify two parallel electron transfer pathways to populate this state. Sub-70 fs electron transfer to the phenazine portion of dppz is observed and attributed to intra-ligand electron transfer following Ru-to-dppz metal-to-ligand charge transfer (MLCT) excitation. This fast charge transfer was not reported in prior ultrafast studies. The slower (ca. 2 ps) charge transfer reported extensively in time-resolved optical absorption and emission studies is reassigned here to inter-ligand electron “hopping” between nearly isoenergetic ligand moieties following Ru-to-bpy MLCT excitation. In conclusion, the results demonstrate much faster charge separation than previously identified in this well-studied system, highlighting how extended azaacene ligand motifs promote the competitive charge transfer processes needed to drive light-driven electron transfer chemistry.

Donor-acceptor systems↗

Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarks

Abstract Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree‐/rule‐based ML methods to cope with the bias‐variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single‐object and multiple‐objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores.

Wu, Xingfu↗

Strong parallel evidence of selection during switchgrass sward establishment in hybrid and lowland ecotypes

Switchgrass sward establishment results in up to 90% seedling mortality. The degree of selection during sward establishment has not been reported using modern genetic methods. Pooled leaf samples were sequenced from replicated swards of 46 half-sib families from two breeding groups (lowland and hybrid) before and through 3 years of stand establishment. Pooled allele frequencies were then assessed using fixation indices (Fst) and an independent data set was used to predict the polygenic impact of establishment selection on two traits (heading date and winter survivorship). Last, the DNA pools were assigned survival rankings to predict the sward survival genomically estimated breeding values within the training data set. Strong and parallel selection occured in both breeding groups. Five genomic regions exceeded the significant threshold of 99.9% in >10 families, indicating consistent selection across families and breeding groups. Polygenic trait predictions determined that establishment selection was partially associated with winter survivorship but resulted in variable heading date alterations. The genomewide variation is consistent with selection for a small number of related parental lines. This study observed strong selection for a small number of hybrid and coastal ecotype individuals which are promising germplasm sources for improved sward survival. This confirms prior reports of sward selection during grassland establishment and highlights the strength of pooled DNA sequencing for survival traits.

54 ENVIRONMENTAL SCIENCES↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Anticipating gelation and vitrification with medium amplitude parallel superposition (MAPS) rheology and artificial neural networks

Abstract Anticipating qualitative changes in the rheological response of complex fluids (e.g., a gelation or vitrification transition) is an important capability for processing operations that utilize such materials in real-world environments. One class of complex fluids that exhibits distinct rheological states are soft glassy materials such as colloidal gels and clay dispersions, which can be well characterized by the soft glassy rheology (SGR) model. We first solve the model equations for the time-dependent, weakly nonlinear response of the SGR model. With this analytical solution, we show that the weak nonlinearities measured via medium amplitude parallel superposition (MAPS) rheology can be used to anticipate the rheological aging transitions in the linear response of soft glassy materials. This is a rheological version of a technique called structural health monitoring used widely in civil and aerospace engineering. We design and train artificial neural networks (ANNs) that are capable of quickly inferring the parameters of the SGR model from the results of sequential MAPS experiments. The combination of these data-rich experiments and machine learning tools to provide a surrogate for computationally expensive viscoelastic constitutive equations allows for rapid experimental characterization of the rheological state of soft glassy materials. We apply this technique to an aging dispersion of Laponite ® clay particles approaching the gel point and demonstrate that a trained ANN can provide real-time detection of transitions in the nonlinear response well in advance of incipient changes in the linear viscoelastic response of the system.

Lennon, Kyle R. (ORCID:0000000212515461)↗