Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

ExTreeM: Scalable Augmented Merge Tree Computation via Extremum Graphs

Over the last decade merge trees have been proven to support a plethora of visualization and analysis tasks since they effectively abstract complex datasets. Here, this paper describes the ExTreeM-Algorithm: A scalable algorithm for the computation of merge trees via extremum graphs. The core idea of ExTreeM is to first derive the extremum graph G of an input scalar field f defined on a cell complex K, and subsequently compute the unaugmented merge tree of f on G instead of K; which are equivalent. Any merge tree algorithm can be carried out significantly faster on G, since K in general contains substantially more cells than G. To further speed up computation, ExTreeM includes a tailored procedure to derive merge trees of extremum graphs. The computation of the fully augmented merge tree, i.e., a merge tree domain segmentation of K, can then be performed in an optional post-processing step. All steps of ExTreeM consist of procedures with high parallel efficiency, and we provide a formal proof of its correctness. Our experiments, performed on publicly available datasets, report a speedup of up to one order of magnitude over the state-of-the-art algorithms included in the TTK and VTK-m software libraries, while also requiring significantly less memory and exhibiting excellent scaling behavior.

97 MATHEMATICS AND COMPUTING↗

How to form a wormhole

Abstract We provide a simple but very useful description of the process of wormhole formation. We place two massive objects in two parallel universes (modeled by two branes). Gravitational attraction between the objects competes with the resistance coming from the brane tension. For sufficiently strong attraction, the branes are deformed, objects touch and a wormhole is formed. Our calculations show that more massive and compact objects are more likely to fulfill the conditions for wormhole formation. This implies that we should be looking for wormholes either in the background of black holes and compact stars, or massive microscopic relics. Our formation mechanism applies equally well for a wormhole connecting two objects in the same universe.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-order harmonics and the reverse of the squaring up process in the triangular-lattice magnet HoPdAl 4 ⁢Ge 2

We uncover high-order harmonics and the reverse of the squaring-up process, in terms of analyzing the evolution of magnetic orders in a centrosymmetric layered triangular-lattice magnet HoPdAl 4 ⁢Ge 2 , based on a detailed study of the crystal structure, magnetic susceptibility, magnetization, heat capacity, and magnetic structure. Temperature dependencies of magnetic susceptibility and heat capacity show two magnetic transitions at T N = 10.5 K and T t = 5.5 K. Below T N , Ho 3+ spins order antiferromagnetically as a transverse spin-density wave with the propagation vector k 1 = (001.5 – δ) with δ ≈ 0.18. Upon further cooling through T t , the high-order harmonics with k n = (001.5 – nδ), n = 3, 5, 7 develop, suggesting a squaring-up process. It is surprising that the squaring-up process does not continue down to 0 K but reverses the trend below 3 K. Magnetic-field-induced metastable transitions were observed in M⁡(H) curves with the fields applied both parallel and perpendicular to the triangular-lattice plane. Neutron-diffraction results suggest that the magnetization process in HoPdAl 4 ⁢Ge 2 involves the conversion of k n = (001.5 – n⁢δ) with n = 1, 3, 5, 7, to a ferromagnetic k F = (000) component. It is worth noting that the already weak seventh harmonic magnetic peak is enhanced by applying a small magnetic field in the ab plane at 1.5 K or warming up to 3 K, accompanied by a slight decrease of δ.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Exploiting user activeness for data retention in HPC systems

HPC systems typically rely on the fixed-lifetime (FLT) data retention strategy, which only considers temporal locality of data accesses to parallel file systems. However, our extensive analysis based on the leadership-class HPC system traces suggests that the FLT approach often fails to capture the dynamics in users' behavior and leads to undesired data purge. In this study, we propose an activeness-based data retention (ActiveDR) solution, which advocates considering the data retention approach from a holistic activeness-based perspective. By evaluating the frequency and impact of users' activities, ActiveDR prioritizes the file purge process for inactive users and rewards active users with extended file lifetime on parallel storage. Our extensive evaluations based on the traces of the prior Titan supercomputer show that, when reaching the same purge target, ActiveDR achieves up to 37% file miss reduction as compared to the current FLT retention methodology.

Zhang, Wei↗

Data-Driven Optimization of the Processing Window for 316H Components Fabricated Using Laser Powder Bed Fusion

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing with a specific focus on laser powder bed fusion (LPBF). As an initial case study, the program has selected 316H stainless steel (SS) as an initial material around which to develop a code case development strategy. This strategy involves two parallel approaches: (1) an equivalency approach whereby round-robin testing across multiple collaborating laboratories demonstrates repeatability in processing and direct comparisons with conventional wrought 316H material and (2) a revolutionary approach to code qualification combining in situ data collection and high-fidelity modeling to capture, predict, and bound the performance of LPBF 316HSS components. As part of this campaign, this work package has initiated an extensive process optimization campaign across three laboratories, each printing variations of LPBF 316HSS using three different LPBF units (Concept Laser, EOS, and Renishaw). In FY23, ORNL has focused on unique experimental designs spanning wide ranges in energy inputs and turning knobs such as scan speed, laser power, hatch spacing, layer thickness, spot size, scan rotation, and more. On the Concept Laser M2, 72 different combinations of processing variables were investigated with duplicate samples and different powder compositions. In total, 252 samples were printed with combined in situ sensing data. A parallel design of experiments was conducted on the Renishaw AM400 with an additional 390 printed specimens for analysis. All 642 miniature specimens, each with unique features included in each print to capture geometry-related heterogeneity, were subjected to high-throughput x-ray computed tomography (XCT) analysis to enable the downselection of specific processing parameters of interest. Then, using electrical discharge machining (EDM), miniature tensile specimens were extracted for mechanical testing and microscopy investigations. From the analysis performed in FY23, it was found that powder composition drastically affects the resulting microstructure and mechanical performance of 316SS. Specifically, changing from 316L to 316HSS powder results in a wide range of grain sizes with varying degrees of preferred grain orientation, which increases as a function of energy density. It was also found that due to stored heat in thin fin–type features, large microstructural differences can be seen within one part printed with one set of processing parameters. These variations in microstructure features, including grain size, the nanoscale dislocation structure, and grain texture, will all affect the irradiation performance and high-temperature mechanical performance of LPBF 316HSS parts. Two sets of concept laser processing parameters, spanning both refined and columnar grain structures, were scaled to print larger 316H builds for campaign testing (high-temperature creep and irradiation). In addition, at least two optimized processing parameter sets were identified for the Renishaw AM400 for round-robin testing in FY24 with Argonne National Laboratory. Future work includes printing samples using identical parameters identified by partner institutions, providing material for corrosion and high-temperature mechanical testing, and continuing evaluations of heterogeneity in larger printed parts.

36 MATERIALS SCIENCE↗

Data-Driven Optimization of the Processing Window for 316H Components Fabricated Using Laser Powder Bed Fusion

The Advanced Materials and Manufacturing Technologies Program is focused on accelerating the development and deployment of advanced materials and components fabricated via additive manufacturing with a specific focus on laser powder bed fusion (LPBF). As an initial case study, the program has selected 316H stainless steel (SS) as an initial material around which to develop a code case development strategy. This strategy involves two parallel approaches: (1) an equivalency approach whereby round-robin testing across multiple collaborating laboratories demonstrates repeatability in processing and direct comparisons with conventional wrought 316H material and (2) a revolutionary approach to code qualification combining in situ data collection and high-fidelity modeling to capture, predict, and bound the performance of LPBF 316HSS components. As part of this campaign, this work package has initiated an extensive process optimization campaign across three laboratories, each printing variations of LPBF 316HSS using three different LPBF units (Concept Laser, EOS, and Renishaw). In FY23, ORNL has focused on unique experimental designs spanning wide ranges in energy inputs and turning knobs such as scan speed, laser power, hatch spacing, layer thickness, spot size, scan rotation, and more. On the Concept Laser M2, 72 different combinations of processing variables were investigated with duplicate samples and different powder compositions. In total, 252 samples were printed with combined in situ sensing data. A parallel design of experiments was conducted on the Renishaw AM400 with an additional 390 printed specimens for analysis. All 642 miniature specimens, each with unique features included in each print to capture geometry-related heterogeneity, were subjected to high-throughput x-ray computed tomography (XCT) analysis to enable the downselection of specific processing parameters of interest. Then, using electrical discharge machining (EDM), miniature tensile specimens were extracted for mechanical testing and microscopy investigations. From the analysis performed in FY23, it was found that powder composition drastically affects the resulting microstructure and mechanical performance of 316SS. Specifically, changing from 316L to 316HSS powder results in a wide range of grain sizes with varying degrees of preferred grain orientation, which increases as a function of energy density. It was also found that due to stored heat in thin fin–type features, large microstructural differences can be seen within one part printed with one set of processing parameters. These variations in microstructure features, including grain size, the nanoscale dislocation structure, and grain texture, will all affect the irradiation performance and high-temperature mechanical performance of LPBF 316HSS parts. Two sets of concept laser processing parameters, spanning both refined and columnar grain structures, were scaled to print larger 316H builds for campaign testing (high-temperature creep and irradiation). In addition, at least two optimized processing parameter sets were identified for the Renishaw AM400 for round-robin testing in FY24 with Argonne National Laboratory. Future work includes printing samples using identical parameters identified by partner institutions, providing material for corrosion and high-temperature mechanical testing, and continuing evaluations of heterogeneity in larger printed parts.

36 MATERIALS SCIENCE↗

Rapid in situ diversification rates in Rhamnaceae explain the parallel evolution of high diversity in temperate biomes from global to local scales

Summary The macroevolutionary processes that have shaped biodiversity across the temperate realm remain poorly understood and may have resulted from evolutionary dynamics related to diversification rates, dispersal rates, and colonization times, closely coupled with Cenozoic climate change. We integrated phylogenomic, environmental ordination, and macroevolutionary analyses for the cosmopolitan angiosperm family Rhamnaceae to disentangle the evolutionary processes that have contributed to high species diversity within and across temperate biomes. Our results show independent colonization of environmentally similar but geographically separated temperate regions mainly during the Oligocene, consistent with the global expansion of temperate biomes. High global, regional, and local temperate diversity was the result of high in situ diversification rates, rather than high immigration rates or accumulation time, except for Southern China, which was colonized much earlier than the other regions. The relatively common lineage dispersals out of temperate hotspots highlight strong source‐sink dynamics across the cosmopolitan distribution of Rhamnaceae. The proliferation of temperate environments since the Oligocene may have provided the ecological opportunity for rapid in situ diversification of Rhamnaceae across the temperate realm. Our study illustrates the importance of high in situ diversification rates for the establishment of modern temperate biomes and biodiversity hotspots across spatial scales.

Plant Sciences↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N. One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an “intelligent” mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗

Improving Performance of M-to-N Processing and Data Redistribution in In Transit Analysis and Visualization

In an in transit setting, a parallel data producer, such as a numerical simulation, runs on one set of ranks M, while a data consumer, such as a parallel visualization application, runs on a different set of ranks N: One of the central challenges in this in transit setting is to determine the mapping of data from the set of M producer ranks to the set of N consumer ranks. This is a challenging problem for several reasons, such as the producer and consumer codes potentially having different scaling characteristics and different data models. The resulting mapping from M to N ranks can have a significant impact on aggregate application performance. In this work, we present an approach for performing this M-to-N mapping in a way that has broad applicability across a diversity of data producer and consumer applications. We evaluate its design and performance with a study that runs at high concurrency on a modern HPC platform. By leveraging design characteristics, which facilitate an ''intelligent'' mapping from M-to-N, we observe significant performance gains are possible in terms of several different metrics, including time-to-solution and amount of data moved.

Loring, Burlen↗

Systems and methods for tensor scheduling

A technique for efficient scheduling of operations in a program for parallelized execution thereof using a multi-processor runtime environment having two or more processors includes constraining the type or number of loop optimization transforms that may be explored such that memory and processing capacity available for the scheduling task are not exceeded, while facilitating a tradeoff between memory locality, parallelization, and/or data communication between memory modules of the multi-processor runtime environment.

Meister, Benoit J.↗

Adaptive, Active Learning, and Multifidelity Monte Carlo Methods in the MOOSE Stochastic Tools Module

MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.

97 MATHEMATICS AND COMPUTING↗

Direct Air Reactive Capture and Conversion for Utility-Scale Energy Storage (Final Report)

This final report for FEW0277 summarizes the work performed over the project performance period of October 2021 – March 2025. This project was funded under the “Reactive Capture and Conversion R&D” lab call released in FY2021. The goal of the project was to develop dual-function materials and process for capturing CO 2 from the atmosphere and converting it into CH 4 . The work was organized into four parallel tracks in 1) direct air capture materials synthesis and characterization, 2) catalysts for CO 2 conversion, 3) mechanistic investigations via ab initio simulations, and 4) process modeling, technoeconomic analysis, and lifecycle assessment. The project was split into two budget periods. The first budget period focused on development of amine-based materials, due to their known performance for CO 2 direct air capture and their potential to act synergistically with metal catalysts to enable a low-temperature methanation pathway. The second budget period focused on development of alkali-based materials and a simulated-moving-bed process for high conversion catalytic reduction of captured CO 2 to CH 4 . All project milestones were completed during the project performance period and are summarized in this report. Our work resulted in publication of eight peer-reviewed manuscripts, one patent application, and numerous presentations given at domestic and international conferences and invited academic department seminars.

03 NATURAL GAS↗

Rapid Commissioning of Large Machine Tools Using Finite Element-Based Correction of Geometric Errors

Large computer numerical control (CNC) machine tools derive their stiffness from monolithic cast iron bases or weldments that are sometimes integral to machine motion systems like box ways or guideways. However, the sheer size of castings and even floor flatness deviations result in dimensional errors in these systems, which manifest as machine motion errors. Typical geometric alignment processes rely on an iterative approach, where measurements are taken to assess alignment (straightness, squareness, and parallelism), followed by adjustment of the machine supports (fixators or leveling pads), which can take weeks even for an experienced operator. Conversely, a novel method is proposed to shorten the correction time by eliminating the trial-and-error process in favor of a more deterministic approach guided by a finite element (FE) method. A feasibility study is conducted on a CNC polymer hybrid machine, with a steel weldment frame, supported by six leveling pads. An FE model of the frame is utilized to obtain recommended leveling pad adjustments, based on measurement of machine errors taken using a laser tracker. After a single adjustment cycle, measurements reveal that geometric errors of the machine tool are reduced from 2.22 mm of flatness deviation to 0.32 mm, achieving an 85.6% reduction. Furthermore, the entire process including measurement, adjustment, and assessment is completed in just 6 h by two operators who are not professional service engineers. In conclusion, this methodology demonstrates feasibility for scaling up, especially to large, high-precision CNC machine tools with bases mounted by fixators, offering the capability for bidirectional adjustment.

42 ENGINEERING↗

Parallel Implicit Hydrodynamics for High Explosive Burn Calculations

High explosives in hostile environments will require calculational capabilities that model processes, which evolve on timescales from minutes to nanoseconds. Eventually the HE will begin to move metal. To handle this temporal evolution an implicit hydrodynamics coupled to the chemical release of the HE energy is required. In addition, the use of chemical kinetics to model the transition from, the initially, slow heating of a confined high explosive through to deflagration and on to detonation requires many computational zones to model high explosive engineering systems. This requirement means that a fully parallel implicit hydrodynamics is essential. In this paper we present the calculation of a nonlinear matrix equation for the advanced particle pressure that has been made parallel and implemented in our AMR code, BABBO. This new parallel implicit hydrodynamics has been applied to a cookoff problem, as well as, one and two dimensional shock problems. Results are presented and discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Parametric, Data-Driven, Non-Intrusive Reduced-Order Model Framework for Crystal Plasticity Simulations of Voids

The influence of the internal structure at micrometer length scales on the deformation of polycrystalline materials can be effectively captured using crystal plasticity finite element methods (CPFEM). However, the complexity and nonlinearity of the deformation equations CPFEM solves demand significant computational power and resources to achieve accurate predictions, limiting its broader application. To address this challenge, we have identified a reduced-order representation of the complex data in order to establish a computationally efficient reduced-order models (ROM) and drastically reduce the computational expense of CPFEM. Specifically, in this work, we developed a parametric, data-driven, and non-intrusive ROM framework for CPFEM using proper orthogonal decomposition (POD) and sparse variational Gaussian process (SVGP) regression for single-crystal microstructures under tensile loading conditions. The developed protocol enables one to compress field into a latent/low-dimensional space described by principal component analysis (PCA) via the singular value decomposition (SVD) algorithm. As a result, the high-dimensional data are reduced to a significantly smaller amount of dimensions with POD bases and POD coefficients. Furthermore, we deployed an ensemble of SVGPs—extended from the classical Gaussian process (GP) regression for scalability and handling big data—in a massively parallel manner to train and predict latent POD coefficients using known POD bases from a set of previously obtained simulations results. Lastly, using the predicted POD coefficients, we reconstructed the full-field results and showed reasonable agreement compared with the true values obtained from running CPFEM. The developed framework is validated with a set of CPFEM simulations of a single embedded void in single-crystal aluminum alloy. While the framework is broadly applicable, this work specifically focuses on single-crystal microstructures, a single load case (e.g., tensile), and a specific void geometry (spherical).

Anisotropy↗

Untangling the threads of cellulose mercerization

Naturally occurring plant cellulose, our most abundant renewable resource, consists of fibers of long polymer chains that are tightly packed in parallel arrays in either of two crystal phases collectively referred to as cellulose I. During mercerization, a process that involves treatment with sodium hydroxide, cellulose goes through a conversion to another crystal form called cellulose II, within which every other chain has remarkably changed direction. We designed a neutron diffraction experiment with deuterium labelling in order to understand how this change of cellulose chain direction is possible. Here we show that during mercerization of bacterial cellulose, chains fold back on themselves in a zigzag pattern to form crystalline anti-parallel domains. This result provides a molecular level understanding of one of the most widely used industrial processes for improving cellulosic materials.

59 BASIC BIOLOGICAL SCIENCES↗

Parallelized domain decomposition for multi-dimensional Lagrangian random walk mass-transfer particle tracking schemes

Lagrangian particle tracking schemes allow a wide range of flow and transport processes to be simulated accurately, but a major challenge is numerically implementing the inter-particle interactions in an efficient manner. This article develops a multi-dimensional, parallelized domain decomposition (DDC) strategy for mass-transfer particle tracking (MTPT) methods in which particles exchange mass dynamically. We show that this can be efficiently parallelized by employing large numbers of CPU cores to accelerate run times. In order to validate the approach and our theoretical predictions we focus our efforts on a well-known benchmark problem with pure diffusion, where analytical solutions in any number of dimensions are well established. In this work, we investigate different procedures for “tiling” the domain in two and three dimensions (2-D and 3-D), as this type of formal DDC construction is currently limited to 1-D. An optimal tiling is prescribed based on physical problem parameters and the number of available CPU cores, as each tiling provides distinct results in both accuracy and run time. We further extend the most efficient technique to 3-D for comparison, leading to an analytical discussion of the effect of dimensionality on strategies for implementing DDC schemes. Increasing computational resources (cores) within the DDC method produces a trade-off between inter-node communication and on-node work. For an optimally subdivided diffusion problem, the 2-D parallelized algorithm achieves nearly perfect linear speedup in comparison with the serial run-up to around 2700 cores, reducing a 5 h simulation to 8 s, while the 3-D algorithm maintains appreciable speedup up to 1700 cores.

97 MATHEMATICS AND COMPUTING↗

TunIO: An AI-powered Framework for Optimizing HPC I/O

I/O operations are a known performance bottleneck of HPC applications. To achieve good performance, users often employ an iterative multistage tuning process to find an optimal I/O stack configuration. However, an I/O stack contains multiple layers, such as high-level I/O libraries, I/O middleware, and parallel file systems, and each layer has many parameters. These parameters and layers are entangled and influenced by each other. The tuning process is time-consuming and complex. In this work, we present TunIO, an AI-powered I/O tuning framework that implements several techniques to balance the tuning cost and performance gain, including tuning the high-impact parameters first. Furthermore, TunIO analyzes the application source code to extract its I/O kernel while retaining all statements necessary to perform I/O. It utilizes a smart selection of high-impact configuration parameters of the given tuning objective. Finally, it uses a novel Reinforcement Learning (RL)-driven early stopping mechanism to balance the cost and performance gain. Experimental results show that TunIO leads to a reduction of up to ≈73% in tuning time while achieving the same performance gain when compared to H5Tuner. It achieves a significant performance gain/cost of 208.4 MBps/min (I/O bandwidth for each minute spent in tuning) over existing approaches under our testing.

Rajesh, Neeraj↗