Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

An asynchronous parallel high-throughput model calibration framework for crystal plasticity finite element constitutive models

Crystal plasticity finite element model (CPFEM) is a powerful numerical simulation in the integrated computational materials engineering toolboxes that relates microstructures to homogenized materials properties and establishes the structure–property linkages in computational materials science. However, to establish the predictive capability, one needs to calibrate the underlying constitutive model, verify the solution and validate the model prediction against experimental data. Bayesian optimization (BO) has stood out as a gradient-free efficient global optimization algorithm that is capable of calibrating constitutive models for CPFEM. Here in this paper, we apply a recently developed asynchronous parallel constrained BO algorithm to calibrate phenomenological constitutive models for stainless steel 304 L, Tantalum, and Cantor high-entropy alloy.

304L stainless steel↗

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

MAPPRAISER: A massively parallel map-making framework for multi-kilo pixel CMB experiments

Forthcoming cosmic microwave background (CMB) polarized anisotropy experiments have the potential to revolutionize our understanding of the Universe and fundamental physics. The sought-after, tale-telling signatures will be however distributed over voluminous data sets which these experiments will collect. These data sets will need to be efficiently processed and unwanted contributions due to astrophysical, environmental, and instrumental effects characterized and efficiently mitigated in order to uncover the signatures. This poses a significant challenge to data analysis methods, techniques, and software tools which will not only have to be able to cope with huge volumes of data but to do so with unprecedented precision driven by the demanding science goals posed for the new experiments. A keystone of efficient CMB data analysis is solvers of very large linear systems of equations. Such systems appear in very diverse contexts throughout CMB data analysis pipelines, however they typically display similar algebraic structures and can therefore be solved using similar numerical techniques. Linear systems arising in the so-called map-making problem are one of the most prominent and common ones. In this work we present a massively parallel, flexible and extensible framework, comprised of a numerical library, MIDAPACK, and a high level code, MAPPRAISER, which provide tools for solving efficiently such systems. Here, the framework implements iterative solvers based on conjugate gradient techniques: enlarged and preconditioned using different preconditioners. We demonstrate the framework on simulated examples reflecting basic characteristics of the forthcoming data sets issued by ground-based and satellite-borne instruments, executing it on as many as 16,384 compute cores. The software is developed as an open source project freely available to the community at: https://github.com/B3Dcmb/midapack.

79 ASTRONOMY AND ASTROPHYSICS↗

Dynamic analysis of fully constrained Cable-Driven Parallel Robots for automated prefabricated component installation

This paper presents a dynamic analysis and validation framework to assess a fully constrained six-anchor Cable-Driven Parallel Robot (CDPR) for automated installation of prefabricated facade components. Compared with conventional eight-anchor systems, the six-anchor configuration simplifies setup and reduces cost, but it also reduces control authority, shrinks the wrench-feasible workspace, and tightens orientation limits. Consequently, it is unclear a priori whether dynamically feasible trajectories exist to move the end effector from pickup to the facade. A constrained trajectory optimization is formulated to enforce the system dynamics, cable-tension bounds, and pose/velocity limits, and the framework is evaluated in simulation at three levels: (i) an idealized reference model, (ii) a lab-scale prototype model incorporating measured anchor misalignments and identified damping, and (iii) a full-scale three-story building model with load decomposition for structural feasibility checks. Across these scenarios, the analysis shows that optimal, constraint-satisfying trajectories exist that move the end effector from pickup to installation while maintaining a near-plumb, level orientation at the final pose. Collectively, this multi-scale dynamic analysis and validation framework supports the deployment readiness of the six-anchor CDPR and provides a prototype-based sensitivity case study of how measured anchor placement deviations affect feasibility.

CDPR↗

A novel xylosylated fucoglucuronan in Penium reveals structural parallels to rhamnogalacturonan-I and its broad evolutionary footprint in lower plants

Green algae inhabit aquatic environments across the planet and play a crucial role in sustaining the global ecosystem. Ancestors of some Charophytes adapted to terrestrial conditions and eventually evolved into land plants. Extant green algae have inherited traits from their ancestors and evolved into their current morphological and chemical forms, as reflected by their cell walls with distinct shapes and compositions. To illuminate the evolution of plant cell walls and bridge the gap between green algae and land plants, we investigated the charophyte Penium margaritaceum, a close relative of terrestrial plants. We discovered a previously unknown polysaccharide in both its culture medium and cell wall. This polysaccharide, termed xylosylated fucoglucuronan (XFG), possesses a rhamnogalacturonan-I (RG-I)-like backbone composed of repeating [-3-α-Fucp-(1,4)-α-GlcpA-] disaccharides that are extensively xylosylated and acetylated. Surveying approximately 20 non-vascular plants revealed that XFG and RG-I (or related structures) first emerge in certain Chlorophyceae and subsequently co-occur throughout lineages along the evolutionary trajectory to bryophytes, thereby bridging aquatic green algae to early land plants. The striking structural parallels between XFG, RG-I, and ulvan suggest a shared evolutionary origin, offering new insight into how plant cell walls adapted during the transition from marine to freshwater environments and ultimately to land.

Algae↗

Parallels between enzyme catalysis, electrocatalysis, and photoelectrosynthesis

Catalysts are central to accelerating chemistry in biology and technology. In biochemistry, the relationship between the velocity of an enzymatic reaction and the concentration of chemical substrates is described via the Michaelis-Menten model. Additionally, the modeling and benchmarking of synthetic molecular electrocatalysts are also well developed. However, such efforts have not been as rigorously extended to photoelectrosynthetic reactions, where, in addition to chemical substrates and charge carriers, light is a required reagent. In this perspective, we draw parallels between concepts involving enzyme catalytic efficiency, the benchmarking of molecular electrocatalysts, and the performance of photoelectrosynthetic assemblies, while highlighting key differences, assumptions, and limitations.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE↗

DFT-FE 1.0: A massively parallel hybrid CPU-GPU density functional theory code using finite-element discretization

In this work, we present DFT-FE 1.0, building on DFT-FE 0.6 [Comput. Phys. Commun. 246, 106853 (2020)], to conduct fast and accurate large-scale density functional theory (DFT) calculations (reaching ~ 100,000 electrons) on both many-core CPU and hybrid CPU-GPU computing architectures. This work involves improvements in the real-space formulation—via an improved treatment of the electrostatic interactions that substantially enhances the computational efficiency—as well high-performance computing aspects, including the GPU acceleration of all the key compute kernels in DFT-FE. We demonstrate the accuracy by comparing the ground-state energies, ionic forces and cell stresses on a wide-range of benchmark systems against those obtained from widely used DFT codes. Further, we demonstrate the numerical efficiency of our implementation, which yields ~ 20× CPU-GPU speed-up by using GPU acceleration on hybrid CPU-GPU nodes. Notably, owing to the parallel-scaling of the GPU implementation, we obtain wall-times of 80–140 seconds for full ground-state calculations, with stringent accuracy, on benchmark systems containing ~ 6, 000 – 15,000 electrons.

pseudopotential↗

CSPlib: A performance portable parallel software toolkit for analyzing complex kinetic mechanisms

Computational singular perturbation (CSP) is a method to analyze dynamical systems. It targets the decoupling of fast and slow dynamics using an alternate linear expansion of the right-hand side of the governing equations based on eigenanalysis of the associated Jacobian matrix. This representation facilitates diagnostic analysis, detection and control of stiffness, and the development of simplified models. For this work, we have implemented CSP in a C++ open-source library CSPlib using the Kokkos parallel programming model to address portability across diverse heterogeneous computing platforms, i.e., multi/many-core CPUs and GPUs. We describe the CSPlib implementation and present its computational performance across different computing platforms using several test problems. Specifically, we test the CSPlib performance for a constant pressure ignition reactor model on different architectures, including IBM Power 9, Intel Xeon Skylake, and NVIDIA V100 GPU. The size of the chemical kinetic mechanism is varied in these tests. As expected, the Jacobian matrix evaluation, the eigensolution of the Jacobian matrix, and matrix inversion are the most expensive computational tasks. When considering the higher throughput characteristic of GPUs, GPUs performs better for small matrices with higher occupancy rate. CPUs gain more advantages from the higher performance of well-tuned and optimized linear algebra libraries such as OpenBLAS.

97 MATHEMATICS AND COMPUTING↗

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

A library of calcium mineral reference spectra recorded by parallel imaging using NEXAFS spectromicroscopy

Calcium minerals are ubiquitous in geology and life chemistry. Understanding the phase and chemical state of calcium minerals is important for numerous processes including materials chemistry, hard tissue biogenesis and geological processes. Photoemission spectroscopies such as near edge X-ray absorption fine structure (NEXAFS) and scanning transmission X-ray microscopy have been instrumental in identifying and characterizing calcium minerals in all these areas. In this work, we have recorded reference spectra for a range of different calcium minerals including a series of calcium carbonates, calcium oxalates and calcium phosphates. While collections of reference spectra for several calcium minerals can be found in the literature, these spectra have been reported in different contexts using a variety of instruments. We, here, report a comprehensive list of references recorded in parallel in a single experiment by imaging an array of calcium minerals using a NEXAFS microscope. We present reference NEXAFS spectra at the calcium L-, carbon K- and oxygen K-edges.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Flow regime and Reynolds number variation effects on the mixing behavior of parallel flows

The hydraulic single-phase mixing of three parallel rectangular channels is experimentally investigated at various Reynolds numbers (Re) and flow regime combinations. Particle Image Velocimetry results for seven mixing cases are presented and discussed with varying Re combinations ranging from 1,824 to 20,844. While all cases result in the same Re ratio of ~0.69 between the inner and outer flows, two cases represent multi-regime mixing with the inner-outer regime pair of laminar-transitional and transitional-turbulent, while the other 5 cases are all characteristic of turbulent mixing with varying levels of turbulence. The outer channels initially share characteristics with a backward facing step. The center channel is found to initially behave like a slot jet, but then sees a significant increase in velocity decay. This inner flow velocity decay increased dramatically in the laminar-transitional mixing case, whose centerline velocity decay was ~6 times larger than the decay in the turbulent mixing cases. Second order statistics revealed a consistent mixing layer thickness of ~0.1 hydraulic diameters for all the cases but showed more intense shearing in the multi-regime mixing cases. The combined point and thereby the mixing layer length is determined using centerline velocity decay profiles, which show a much more aggressive mixing in multi-regime flows. Multi-regime mixing demonstrated superior characteristics relative to turbulent mixing due to a more dramatic velocity decay in the inner flow and a shorter mixing length. The contributions of this work include communicating the benefits of multi-regime mixing and providing detailed characterization efforts that can serve future efforts for validating computational models. Here this research also lays the groundwork for future studies aimed at achieving high levels of mixing without a severe penalty in pressure drop.

42 ENGINEERING↗

Tusas: A fully implicit parallel approach for coupled phase-field equations

In this study, we develop a fully-coupled, fully-implicit approach for phase-field modeling of solidification in metals and alloys. Predictive simulation of solidification in pure metals and metal alloys remains a significant challenge in the field of materials science, as microstructure formation during the solidification process plays a critical role in the properties and performance of the solid material. Our simulation approach consists of a finite element spatial discretization of the fully-coupled nonlinear system of partial differential equations at the microscale, which is treated implicitly in time with a preconditioned Jacobian-free Newton-Krylov method. The approach is algorithmically scalable as well as efficient due to an effective preconditioning strategy based on algebraic multigrid and block factorization. We implement this approach in the open-source Tusas framework, which is a general, flexible tool developed in C++ for solving coupled systems of nonlinear partial differential equations. The performance of our approach is analyzed in terms of algorithmic scalability and efficiency, while the computational performance of Tusas is presented in terms of parallel scalability and efficiency on emerging heterogeneous architectures. We demonstrate that modern algorithms, discretizations, and computational science, and heterogeneous hardware provide a robust route for predictive phase-field simulation of microstructure evolution during additive manufacturing.

97 MATHEMATICS AND COMPUTING↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

A parallel variable population multi-objective optimizer for accelerator beam dynamics optimization

The simultaneous optimization of multiple objective functions is needed in many particle accelerator applications. In this paper, we present a parallel evolution based multi-objective optimizer that uses a variable population from generation to generation and an external storage to save good solutions. Two heuristic optimization methods, one uses the unified differential evolution and the other uses the real-coded genetic algorithm, are included in the optimizer to generate next generation candidate solutions, and are compared in the test examples. Finally, as an application, we applied this optimizer to the beam dynamics design optimization of a photoinjector and attained the optimal front solutions after 200 generations with the unified differential evolution offspring production scheme.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A parallel strategy for density functional theory computations on accelerated nodes

Using the Löwdin orthonormalization of tall-skinny matrices as a proxy-app for wavefunction-based Density Functional Theory solvers, we investigate a distributed memory parallel strategy focusing on Graphics Processing Unit (GPU)-accelerated nodes as available on some of the top ranked supercomputers at the present time. Here we present numerical results in the strong limit regime, as it is particularly relevant for First-Principles Molecular Dynamics. We also examine how matrix product-based iterative solvers provide a competitive alternative to dense eigensolvers on GPUs, allowing to push the strong scaling limit of these computations to a larger number of distributed tasks. Our strategy, which relies on replicated Gram matrices and efficient collective communications using the NCCL library, leads to a time-to-solution under 0.5 s for the Löwdin orthonormalization of a tall-skinny matrix of 3000 columns on Summit at Oak Ridge Leadership Facility (OLCF). Given the similarity in computational operations between one iteration of a DFT solver and this proxy-app, this shows the possibility of solving accurately the DFT equations well under a minute for 3000 electronic wave functions, and thus perform First-Principles molecular dynamics of physical systems much larger than traditionally solved on CPU systems.

97 MATHEMATICS AND COMPUTING↗

The heat transfer coefficient associated with a moving packed bed of silica particles flowing through parallel plates

Concentrating Solar Power (CSP) with thermal energy storage has the potential to be a renewable energy technology with long duration, inexpensive energy storage. Higher temperature operation increases the conversion efficiency and reduces the cost of energy storage. Several emerging CSP designs utilize particles as the solar receiver due to their high temperature stability and low cost. In some designs, the particles are also used as the thermal energy storage (TES) media. In either configuration, an energy transfer is required between the hot particles and the working fluid in the power cycle in a Particle-to-Fluid Heat Exchanger (PtFHX). Understanding the heat transfer between a moving packed bed and a stationary surface is critical to the successful design of a PtFHX. In this paper, a test facility is described in which a moving packed bed of silica sand with particle size 100–600 μm is introduced into the channel formed by two parallel plates, one of which is heated. The effective static thermal conductivity of the particles used for the test are separately measured over the entire range of test temperatures. The inlet and outlet bulk temperatures of the particle flow are measured as are the surface temperatures at several axial locations along the centerline of the plate. The result is the measurement of heat transfer coefficient as a function of temperature for several velocities. The uncertainty of the measurements is presented and the results are compared to model results found in the literature.

14 SOLAR ENERGY↗