Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,351 records · Page 75

VerifyIO: Ensuring Correctness of Consistency Semantics in Parallel I/O

Abstract—High-performance computing (HPC) applications generate and consume substantial amounts of data, typically managed by parallel file systems. These applications access file systems either through the POSIX interface or by using highlevel I/O libraries. While the POSIX consistency model remains dominant in HPC, emerging file systems and popular I/O libraries increasingly adopt alternative consistency models that relax semantics in various ways, creating significant challenges for correctness and portability. This paper addresses these challenges by proposing a trace-driven I/O consistency verification workflow, implemented in our open-source tool, VerifyIO, which collects execution traces, detects data conflicts, and verifies proper synchronization against specified consistency models. Our extensive evaluation of 91 test case executions across three widely used I/O libraries with four I/O consistency models reveals critical consistency issues at both application and implementation levels.

Consistency Semantics↗

Revealing Parallel Inter‐ and Intra‐Ligand Charge Transfer Dynamics in [Ru(L) 2 (dppz)] 2+ Molecular Lightswitch with N K‐Edge X‐Ray Absorption Spectroscopy

In photoactive metal complexes the localization of photoexcited charges dictates the site of chemical reactivity, but few studies measure the charge redistribution in these systems with spatial precision. Herein, we track the inter- and intra-ligand charge transfer processes that underpin light-driven charge separation in the well-studied “molecular lightswitch” [Ru(bpy) 2 dppz] 2+ (aqueous [Ruthenium II (2,2′-bipyridine)2(dipyrido[3,2-a:2′,3′-c]phenazine)] 2+ [Cl − ] 2 ) by probing the electronic structure of ligand nitrogen atoms in real-time using ultrafast X-ray absorption spectroscopy and first principles calculations. We confirm the localization of excited electron density on the phenazine N atoms of dppz and we newly identify two parallel electron transfer pathways to populate this state. Sub-70 fs electron transfer to the phenazine portion of dppz is observed and attributed to intra-ligand electron transfer following Ru-to-dppz metal-to-ligand charge transfer (MLCT) excitation. This fast charge transfer was not reported in prior ultrafast studies. The slower (ca. 2 ps) charge transfer reported extensively in time-resolved optical absorption and emission studies is reassigned here to inter-ligand electron “hopping” between nearly isoenergetic ligand moieties following Ru-to-bpy MLCT excitation. In conclusion, the results demonstrate much faster charge separation than previously identified in this well-studied system, highlighting how extended azaacene ligand motifs promote the competitive charge transfer processes needed to drive light-driven electron transfer chemistry.

Donor-acceptor systems↗

Utilizing ensemble learning for performance and power modeling and improvement of parallel cancer deep learning CANDLE benchmarks

Abstract Machine learning (ML) continues to grow in importance across nearly all domains in modeling to learn from data. Often a tradeoff exists between a model's ability to minimize bias and variance. In this article, we utilize ensemble learning to combine linear, nonlinear, and tree‐/rule‐based ML methods to cope with the bias‐variance tradeoff and result in more accurate models. We use the datasets collected for two parallel cancer deep learning CANDLE benchmarks, NT3 and P1B2, to build performance and power models based on hardware performance counters using single‐object and multiple‐objects ensemble learning to identify the most important counters for improvement on the Cray XC40 Theta at Argonne National Laboratory. Based on the insights from these models, we improve the performance and energy of P1B2 and NT3 by optimizing the deep learning environments TensorFlow, Keras, Horovod, and Python under the huge page size of 8 MB. Experimental results show that ensemble learning not only produces more accurate models but also provides more robust performance counter ranking. We achieve up to 61.15% performance improvement and up to 62.58% energy saving for P1B2 and up to 55.81% performance improvement and up to 52.60% energy saving for NT3 on up to 24,576 cores.

Wu, Xingfu↗

Strong parallel evidence of selection during switchgrass sward establishment in hybrid and lowland ecotypes

Switchgrass sward establishment results in up to 90% seedling mortality. The degree of selection during sward establishment has not been reported using modern genetic methods. Pooled leaf samples were sequenced from replicated swards of 46 half-sib families from two breeding groups (lowland and hybrid) before and through 3 years of stand establishment. Pooled allele frequencies were then assessed using fixation indices (Fst) and an independent data set was used to predict the polygenic impact of establishment selection on two traits (heading date and winter survivorship). Last, the DNA pools were assigned survival rankings to predict the sward survival genomically estimated breeding values within the training data set. Strong and parallel selection occured in both breeding groups. Five genomic regions exceeded the significant threshold of 99.9% in >10 families, indicating consistent selection across families and breeding groups. Polygenic trait predictions determined that establishment selection was partially associated with winter survivorship but resulted in variable heading date alterations. The genomewide variation is consistent with selection for a small number of related parental lines. This study observed strong selection for a small number of hybrid and coastal ecotype individuals which are promising germplasm sources for improved sward survival. This confirms prior reports of sward selection during grassland establishment and highlights the strength of pooled DNA sequencing for survival traits.

54 ENVIRONMENTAL SCIENCES↗

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249↗

Anticipating gelation and vitrification with medium amplitude parallel superposition (MAPS) rheology and artificial neural networks

Abstract Anticipating qualitative changes in the rheological response of complex fluids (e.g., a gelation or vitrification transition) is an important capability for processing operations that utilize such materials in real-world environments. One class of complex fluids that exhibits distinct rheological states are soft glassy materials such as colloidal gels and clay dispersions, which can be well characterized by the soft glassy rheology (SGR) model. We first solve the model equations for the time-dependent, weakly nonlinear response of the SGR model. With this analytical solution, we show that the weak nonlinearities measured via medium amplitude parallel superposition (MAPS) rheology can be used to anticipate the rheological aging transitions in the linear response of soft glassy materials. This is a rheological version of a technique called structural health monitoring used widely in civil and aerospace engineering. We design and train artificial neural networks (ANNs) that are capable of quickly inferring the parameters of the SGR model from the results of sequential MAPS experiments. The combination of these data-rich experiments and machine learning tools to provide a surrogate for computationally expensive viscoelastic constitutive equations allows for rapid experimental characterization of the rheological state of soft glassy materials. We apply this technique to an aging dispersion of Laponite ® clay particles approaching the gel point and demonstrate that a trained ANN can provide real-time detection of transitions in the nonlinear response well in advance of incipient changes in the linear viscoelastic response of the system.

Lennon, Kyle R. (ORCID:0000000212515461)↗

An asynchronous parallel high-throughput model calibration framework for crystal plasticity finite element constitutive models

Crystal plasticity finite element model (CPFEM) is a powerful numerical simulation in the integrated computational materials engineering toolboxes that relates microstructures to homogenized materials properties and establishes the structure–property linkages in computational materials science. However, to establish the predictive capability, one needs to calibrate the underlying constitutive model, verify the solution and validate the model prediction against experimental data. Bayesian optimization (BO) has stood out as a gradient-free efficient global optimization algorithm that is capable of calibrating constitutive models for CPFEM. Here in this paper, we apply a recently developed asynchronous parallel constrained BO algorithm to calibrate phenomenological constitutive models for stainless steel 304 L, Tantalum, and Cantor high-entropy alloy.

304L stainless steel↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

Dynamic analysis of fully constrained Cable-Driven Parallel Robots for automated prefabricated component installation

This paper presents a dynamic analysis and validation framework to assess a fully constrained six-anchor Cable-Driven Parallel Robot (CDPR) for automated installation of prefabricated facade components. Compared with conventional eight-anchor systems, the six-anchor configuration simplifies setup and reduces cost, but it also reduces control authority, shrinks the wrench-feasible workspace, and tightens orientation limits. Consequently, it is unclear a priori whether dynamically feasible trajectories exist to move the end effector from pickup to the facade. A constrained trajectory optimization is formulated to enforce the system dynamics, cable-tension bounds, and pose/velocity limits, and the framework is evaluated in simulation at three levels: (i) an idealized reference model, (ii) a lab-scale prototype model incorporating measured anchor misalignments and identified damping, and (iii) a full-scale three-story building model with load decomposition for structural feasibility checks. Across these scenarios, the analysis shows that optimal, constraint-satisfying trajectories exist that move the end effector from pickup to installation while maintaining a near-plumb, level orientation at the final pose. Collectively, this multi-scale dynamic analysis and validation framework supports the deployment readiness of the six-anchor CDPR and provides a prototype-based sensitivity case study of how measured anchor placement deviations affect feasibility.

CDPR↗

A novel xylosylated fucoglucuronan in Penium reveals structural parallels to rhamnogalacturonan-I and its broad evolutionary footprint in lower plants

Green algae inhabit aquatic environments across the planet and play a crucial role in sustaining the global ecosystem. Ancestors of some Charophytes adapted to terrestrial conditions and eventually evolved into land plants. Extant green algae have inherited traits from their ancestors and evolved into their current morphological and chemical forms, as reflected by their cell walls with distinct shapes and compositions. To illuminate the evolution of plant cell walls and bridge the gap between green algae and land plants, we investigated the charophyte Penium margaritaceum, a close relative of terrestrial plants. We discovered a previously unknown polysaccharide in both its culture medium and cell wall. This polysaccharide, termed xylosylated fucoglucuronan (XFG), possesses a rhamnogalacturonan-I (RG-I)-like backbone composed of repeating [-3-α-Fucp-(1,4)-α-GlcpA-] disaccharides that are extensively xylosylated and acetylated. Surveying approximately 20 non-vascular plants revealed that XFG and RG-I (or related structures) first emerge in certain Chlorophyceae and subsequently co-occur throughout lineages along the evolutionary trajectory to bryophytes, thereby bridging aquatic green algae to early land plants. The striking structural parallels between XFG, RG-I, and ulvan suggest a shared evolutionary origin, offering new insight into how plant cell walls adapted during the transition from marine to freshwater environments and ultimately to land.

Algae↗

Parallels between enzyme catalysis, electrocatalysis, and photoelectrosynthesis

Catalysts are central to accelerating chemistry in biology and technology. In biochemistry, the relationship between the velocity of an enzymatic reaction and the concentration of chemical substrates is described via the Michaelis-Menten model. Additionally, the modeling and benchmarking of synthetic molecular electrocatalysts are also well developed. However, such efforts have not been as rigorously extended to photoelectrosynthetic reactions, where, in addition to chemical substrates and charge carriers, light is a required reagent. In this perspective, we draw parallels between concepts involving enzyme catalytic efficiency, the benchmarking of molecular electrocatalysts, and the performance of photoelectrosynthetic assemblies, while highlighting key differences, assumptions, and limitations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE↗

Parallel three-dimensional simulations of quasi-static elastoplastic solids

Hypo-elastoplasticity is a flexible framework for modeling the mechanics of many hard materials under small elastic deformation and large plastic deformation. Under typical loading rates, most laboratory tests of these materials happen in the quasi-static limit, but there are few existing numerical methods tailor-made for this physical regime. Here, we extend to three dimensions a recent projection method for simulating quasi-static hypo-elastoplastic materials. The method is based on a mathematical correspondence to the incompressible Navier–Stokes equations, where the projection method of Chorin (1968) is an established numerical technique. We develop and utilize a three-dimensional parallel geometric multigrid solver employed to solve a linear system for the quasi-static projection. Our method is tested through simulation of three-dimensional shear band nucleation and growth, a precursor to failure in many materials. As an example system, we employ a physical model of a bulk metallic glass based on the shear transformation zone theory, but the method can be applied to any elastoplasticity model. We consider several examples of three-dimensional shear banding, and examine shear band formation in physically realistic materials with heterogeneous initial conditions under both simple shear deformation and boundary conditions inspired by friction welding.

97 MATHEMATICS AND COMPUTING↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

CSPlib: A performance portable parallel software toolkit for analyzing complex kinetic mechanisms

Computational singular perturbation (CSP) is a method to analyze dynamical systems. It targets the decoupling of fast and slow dynamics using an alternate linear expansion of the right-hand side of the governing equations based on eigenanalysis of the associated Jacobian matrix. This representation facilitates diagnostic analysis, detection and control of stiffness, and the development of simplified models. For this work, we have implemented CSP in a C++ open-source library CSPlib using the Kokkos parallel programming model to address portability across diverse heterogeneous computing platforms, i.e., multi/many-core CPUs and GPUs. We describe the CSPlib implementation and present its computational performance across different computing platforms using several test problems. Specifically, we test the CSPlib performance for a constant pressure ignition reactor model on different architectures, including IBM Power 9, Intel Xeon Skylake, and NVIDIA V100 GPU. The size of the chemical kinetic mechanism is varied in these tests. As expected, the Jacobian matrix evaluation, the eigensolution of the Jacobian matrix, and matrix inversion are the most expensive computational tasks. When considering the higher throughput characteristic of GPUs, GPUs performs better for small matrices with higher occupancy rate. CPUs gain more advantages from the higher performance of well-tuned and optimized linear algebra libraries such as OpenBLAS.

97 MATHEMATICS AND COMPUTING↗

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A library of calcium mineral reference spectra recorded by parallel imaging using NEXAFS spectromicroscopy

Calcium minerals are ubiquitous in geology and life chemistry. Understanding the phase and chemical state of calcium minerals is important for numerous processes including materials chemistry, hard tissue biogenesis and geological processes. Photoemission spectroscopies such as near edge X-ray absorption fine structure (NEXAFS) and scanning transmission X-ray microscopy have been instrumental in identifying and characterizing calcium minerals in all these areas. In this work, we have recorded reference spectra for a range of different calcium minerals including a series of calcium carbonates, calcium oxalates and calcium phosphates. While collections of reference spectra for several calcium minerals can be found in the literature, these spectra have been reported in different contexts using a variety of instruments. We, here, report a comprehensive list of references recorded in parallel in a single experiment by imaging an array of calcium minerals using a NEXAFS microscope. We present reference NEXAFS spectra at the calcium L-, carbon K- and oxygen K-edges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗