Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Portage: A Modular Data Remap Library for Multiphysics Applications on Advanced Architectures

Portage is a scalable and extensible remap library for numerical simulations. It supports state-of-the-art remap schemes for meshes and particles in 2D and 3D up to a second-order accuracy. Portage ensures critical properties such as local/global conservation and bounds preservation for mesh remap. It enables multi-material field remap through a dedicated plugin, and leverages the hybrid parallelism exposed by advanced architectures using multi-processing and multi-threading.

97 MATHEMATICS AND COMPUTING↗

Efficient Treatment of Large Active Spaces through Multi-GPU Parallel Implementation of Direct Configuration Interaction

In this study, we have extended our graphical processing unit (GPU)-accelerated direct configuration interaction program to multiple devices, reducing iteration times for configuration spaces of 165 million determinants to only 3 s using NVIDIA P100 GPUs. Similar improvements in the one- and two-particle reduced density matrix formation allow for fast analytical energy gradients and electronic properties. Our parallel algorithm enables the calculation of arbitrarily large configuration spaces (limited only by available system memory), with iteration times of 13 min for an active space of 18 electrons in 18 orbitals (2.4 billion determinants) using six consumer grade NVIDIA 1080Ti GPUs. These advances enable routine molecular dynamics simulations, geometry optimizations, and absorption spectrum calculations for molecules with large configuration spaces, a task that has heretofore required massive computational effort. In this work, we demonstrate the utility of our program by generating the absorption spectrum for diphenyl acetylene at the floating occupation molecular orbital complete active space configuration interaction level of theory. Lastly, several active spaces were investigated to assess the dependence of spectral features on orbital space dimension.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Analysis of Infrastructures for Processing Plastic Waste using Pyrolysis-Based Chemical Upcycling Pathways

Modern mechanical recycling infrastructure for plastic is capable of processing only a small subset of waste plastics, reinforcing the need for parallel disposal methods such as landfilling and incineration. Emerging pyrolysis-based chemical technologies can "upcycle" plastic waste into high-value polymer and chemical products and process a broader range of waste plastics. In this work, we study the economic and environmental benefits of deploying an upcycling infrastructure in the continental United States for producing low-density polyethylene (LDPE) and polypropylene (PP) from post-consumer mixed plastic waste. Our analysis aims to determine the market size that the infrastructure can create, the degree of circularity that it can achieve, the prices for waste and derived products it can propagate, and the environmental benefits of diverting plastic waste from landfill and incineration facilities it can produce. We apply a computational framework that integrates techno-economic analysis, life cycle assessment, and value chain optimization. Our results demonstrate that the infrastructure generates an economy of nearly 20 billion USD and positive prices for plastic waste, opening opportunities for compensation to residents who provide plastic waste. Our analysis also indicates that the infrastructure can achieve a plastic-to-plastic degree of circularity of 34% and remains viable under various external factors (including technology efficiencies, capital investment budgets, and polymer market values). Finally, we present significant environmental benefits of upcycling over alternative landfill and incineration waste disposal methods, and comment on ongoing work expanding our modeling methodology to other chemical upcycling pathway case studies, including hydroformylation of specific plastics to chemicals.

Interdisciplinary↗

Investigation of turbulent inflow specification in Euler–Lagrange simulations of mid-field spray

The process of atomization of a liquid jet by a parallel high-speed gas stream results in a spray, whose downstream development is of considerable interest to several applications. The round jet spray can be spatially divided into (i) a near-field (near-nozzle) region of liquid atomization and (ii) a downstream mid-field region of fully-dispersed droplets. In order to accurately model mid-field droplet dispersion, this work aims at developing a rigorous and robust injection model for Euler–Lagrange spray simulations. Results from experiments are used to obtain the relevant droplet number density, size distribution, and mean and standard deviation velocity distributions of the injection model, systematically in a step-by-step process. Two-phase large eddy simulations are performed by stochastically generating the Lagrangian droplets at the inlet of the mid-field region. Number flux, diameter distribution, mean velocity, and other time-averaged statistics at several downstream locations are shown to agree well with the corresponding experimental data.

42 ENGINEERING↗

Process Design and Techno-Economic Analysis of the Modular Staged Pressurized Oxy-Combustion (SPOC) Power Plant for Biomass

This work describes the process design and techno-economic analysis (TEA) of the modular stage pressurized oxy-combustion (SPOC) power plant for biomass firing and coal-biomass co-firing. The SPOC process was modelled using Aspen Plus®, and largely based on a previous model designed by this group for SPOC coal firing. To enable comparison with current National Energy Technology Laboratory (NETL) Bio-Energy Carbon Capture and Storage (BECCS) studies, a 550 MWe SPOC power plant with a supercritical Rankine cycle (241 bar, 593°C, and 593°C), and 90% carbon capture was modeled, and hybrid poplar biomass was chosen. Two cases were evaluated, namely 100% biomass (carbon negative) and 25% biomass co-firing (carbon neutral), and the 100% Powder River Basin coal firing case was chosen for comparison purposes. In the SPOC process, oxygen is produced via a cryogenic air separation unit (ASU) and the heat generated from the compression of air is integrated into the steam cycle and utilized for boiler feed water regeneration. Unique to the SPOC process, the boilers are arranged in a series-parallel configuration, with minimized flue gas recirculation. The flue gas is cooled and scrubbed in the direct-contact cooler (DCC) column, and the water leaving the bottom of the DCC is at a sufficiently high temperature that it can be used for boiler feed water heating, improving plant thermal efficiency. The SPOC efficiencies were above the BECCS cases with capture, and no efficiency penalty on the SPOC plant was observed with an increase of biomass in the mix mostly due to the higher oxygen content in biomass that resulted in lower oxygen requirement from the ASU, and the higher moisture in biomass that due to the key benefit of the SPOC process can be partially recovered as latent heat.

Magalhaes, Duarte↗

Unbalanced Parallel I/O: An Often-Neglected Side Effect of Lossy Scientific Data Compression

Lossy compression techniques have demonstrated promising results in significantly reducing the scientific data size while guaranteeing the compression error bounds. However, one important yet often neglected side effect of lossy scientific data compression is its impact on the performance of parallel I/O. Our key observation is that the compressed data size is often highly skewed across processes in lossy scientific compression. To understand this behavior, we conduct extensive experiments where we apply three lossy compressors MGARD, ZFP, and SZ, which are specifically designed and optimized for scientific data, to three real-world scientific applications Gray-Scott simulation, WarpX, and XGC. Our analysis result demonstrates that the size of the compressed data is always skewed even if the original data is evenly decomposed among processes. Such skewness widely exists in different scientific applications using different compressors as long as the information density of the data varies across processes. We then systematically study how this side effect of lossy scientific data compression impacts the performance of parallel I/O. We observe that the skewness in the sizes of the compressed data often leads to I/O imbalance, which can significantly reduce the efficiency of I/O bandwidth utilization if not properly handled. In addition, writing data concurrently to a single shared file through MPI-IO library is more sensitive to the unbalanced I/O loads. Therefore, we believe our research community should pay more attention to the unbalanced parallel I/O caused by lossy scientific data compression.

Wang, Xinying↗

High-throughput synthesis of high-entropy alloys via parallelized electric field assisted sintering

Materials discovery and design is an expensive and time-consuming process, though necessary to advance many engineering fields. In this work, a novel tooling design is utilized in conjunction with electric field assisted sintering (EFAS) to effectively create a new high-throughput synthesis technique: parallelized EFAS. Through this technique, a wide range of material compositions and geometries can be synthesized in parallel as isolated samples or as part of contiguous arrays. Multiple tooling designs are explored to examine both the flexibility and limitations of the technique. A series of increasing complex alloys is produced simultaneously using in situ alloying, beginning with pure Ni and adding equimolar constituents up to the septenary high-entropy alloy AlCoCrCuFeMnNi. Microstructural characterization reveals each sample is effectively fully dense and chemically homogenous while exhibiting phases in agreement with CALPHAD predictions. Scalability of parallelized EFAS is then experimentally demonstrated and the implications for materials discovery and automation are discussed.

36 - MATERIALS SCIENCE↗

SPADES (Scalable Parallel Discrete Events Simulation) [SWR-24-99]

SPADES (Solver for PArallel Discrete Event Simulation) is an open-source parallel discrete event simulation (PDES) package built on the AMReX library. Targeted at solving discrete event systems in parallel, this software package aims to be performance portable and scalable on heterogeneous computing architectures, e.g., graphic processing units (GPU). SPADES implements optimistic synchronization with rollback through an implementation of the Time Warp algorithm. An alternative conservative synchronization approach is also implemented using the Lower Bound on Incoming Time Stamp. In our implementation, logical processes are represented as cells in a grid and event messages are represented as particles. SPADES supports various parallel decomposition strategies, including the use of the Message Passing Interface (MPI) and OpenMP threading. All major GPU architectures (e.g., Intel, AMD, NVIDIA) are supported through the use of performance portability functionalities implemented in AMReX. The SPADES software is released in NREL Software Record SWR-24-99 “SPADES (Scalable Parallel Discrete Events Simulation)”.

Henry de Frahan, Marc [National Renewable Energy L↗

Taking control of compressible modes: bulk viscosity and the turbulent dynamo

Many polyatomic astrophysical plasmas are compressible and out of chemical and thermal equilibrium, introducing a bulk viscosity into the plasma via the internal degrees of freedom of the molecular composition, directly impacting the decay of compressible modes, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$. This is especially important for small-scale, turbulent dynamo processes in the interstellar medium (ISM), which are known to be sensitive to the effects of compression. To control the viscous properties of $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$, we perform trans-sonic, visco-resistive dynamo simulations with additional bulk viscosity $\nu _{\text{bulk}}$, deriving a new $\nu _{\text{bulk}}$ Reynolds number $\text{Re}_{\text{bulk}}$, and viscous Prandtl number $\text{P}\nu \equiv \text{Re}_{\text{bulk}}/ \text{Re}_{\text{shear}}$, where $\text{Re}_{\text{shear}}$ is the shear viscosity Reynolds number. We derive a framework for decomposing $E_{\rm mag}$ growth rates into incompressible and compressible terms via orthogonal tensor decompositions of $\boldsymbol {\nabla }\otimes \mathrm{{\boldsymbol {\mathit {v}}}}$, where $\mathrm{{\boldsymbol {\mathit {v}}}}$ is the fluid velocity. We find that $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ play a dual role, growing and decaying $E_{\rm mag}$, and that field-line stretching is the main driver of growth, even in compressible dynamos. In the absence of $\nu _{\text{bulk}}$ ($\text{P}\nu \rightarrow \infty$), $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ pile up on small-scales, creating a spectral bottleneck, which disappears for $\text{P}\nu \approx 1$. As $\text{P}\nu$ decreases, $\mathrm{{\boldsymbol {\mathit {v}}}}_{\parallel }(\boldsymbol {k})$ are dissipated at increasingly larger scales, in turn suppressing incompressible modes through a coupling between high-k modes. We emphasize the importance of further understanding the role of $\nu _{\text{bulk}}$ in compressible astrophysical plasmas, which we estimate could be as strong as the shear viscosity in the cold ISM, and highlight that compressible direct numerical simulations without bulk viscosity have unresolved compressible mode dissipation scales.

MHD↗

Decomposition and Algorithmic Approaches for Solving Large-Scale Process Family Design Problems

Our most recent work expands the water desalination case study from 76 variants to 10,897 variants using the equation-oriented model built in Pyomo as part of the PARETO project. Using the discretization formulation presented in Stinchfield (2024a), rather than solving for all 10,897 variants simultaneously, we decompose the formulation into subproblems containing subsets of variants from the process family. We solve the overall problem with Progressive Hedging (PH) deployed in parallel on a distributed HPC cluster using the open-source Python package mpi-sppy (Knueven et al., 2023). This approach allowed us to solve this process family design problem to ~1.5% relative optimality gap in about 5 hours; in comparison, Gurobi reached ~50% relative optimality gap in about 6 hours (Stinchfield et al., 2024b). However, this approach still requires discretization of the common unit module design ranges; additionally, PH acts as a heuristic for MILP’s with gap-closing capabilities. Ideally, we would not have to use ML surrogates or discretization to solve this problem, instead solving the process family design problem with the equation-oriented model directly to achieve the most accurate results. However, recall that we did not consider solving the MINLP directly due to complexity and size. In this work, we aim to decompose and solve this large-scale MINLP using a Structured Nonlinear Global Optimization algorithm presented by Cao and Zavala (2019).

Stinchfield, Georgia↗

Scalable Deep-Learning-Accelerated Topology Optimization for Additively Manufactured Materials

Topology optimization (TO) is a popular and powerful computational approach for designing novel structures, materials, and devices. Two computational challenges have limited the applicability of TO to a variety of industrial applications. First, a TO problem often involves a large number of design variables to guarantee sufficient expressive power. Second, many TO problems require a large number of expensive physical model simulations, and those simulations cannot be parallelized. To address these issues, we propose a general scalable deep-learning (DL) based TO framework, referred to as SDL-TO, which utilizes parallel schemes in high performance computing (HPC) to accelerate the TO process for designing additively manufactured (AM) materials. Unlike the existing studies of DL for TO, our framework accelerates TO by learning the iterative history data and simultaneously training on the mapping between the given design and its gradient. The surrogate gradient is learned by utilizing parallel computing on multiple CPUs incorporated with a distributed DL training on multiple GPUs. The learned TO gradient enables a fast online update scheme instead of an expensive update based on the physical simulator or solver. Using a local sampling strategy, we achieve to reduce the intrinsic high dimensionality of the design space and improve the training accuracy and the scalability of the SDL-TO framework. The method is demonstrated by benchmark examples and AM materials design for heat conduction. The proposed SDL-TO framework shows competitive performance compared to the baseline methods but significantly reduces the computational cost by a speed up of around 8.6x over the standard TO implementation.

Bi, Sirui↗

Development of PFLOTRAN Transport Capability for Use in the Waste Isolation Pilot Plant Performance Assessment - 20545

Waste Isolation Pilot Plant (WIPP) performance assessment (PA) calculations estimate the probability of radionuclide release from the repository to the land surface and across the land withdrawal boundary for a regulatory period of 10,000 years after facility closure. Simulations of flow and transport in the repository and the surrounding Salado Formation are foundational to the PA. Because proposed additional waste emplacement panels would result in an asymmetric repository layout, the US Department of Energy (DOE) is preparing to transition to use of a three-dimensional (3-D) model domain for simulation of flow and transport instead of the two-dimensional (2-D) flared grid domain currently used. DOE has charged Sandia National Laboratories with developing the capability necessary to simulate processes affecting flow and transport in the WIPP in PFLOTRAN, an open-source massively parallel multi-phase flow and reactive transport code. The new flow and transport capabilities developed in PFLOTRAN incorporate WIPP-specific process models and will replace the 2-D simulators (BRAGFLO and NUTS) that are currently utilized for Salado flow and transport calculations in WIPP PA. The focus of this paper is on the development of a new Nuclear Waste Transport (NWT) mode in PFLOTRAN that has all of the capabilities necessary for Salado transport simulations, including the ability to handle complete dry-out (100% gas saturation) of arbitrary cells in the model domain, radionuclide mass conservation at step changes in porosity associated with borehole intrusion, and the ability to calculate fluxes on a flared grid. The new PFLOTRAN transport capability and a suite of verification tests were designed around a list of functional requirements for WIPP PA calculations. (authors)

12 MANAGEMENT OF RADIOACTIVE AND NON-RADIOACTIVE W↗

High-fidelity wind farm simulation methodology with experimental validation

The complexity and associated uncertainties involved with atmospheric-turbine-wake interactions produce challenges for accurate wind farm predictions of generator power and other important quantities of interest (QoIs), even with state-of-the-art high-fidelity atmospheric and turbine models. A comprehensive computational study was undertaken with consideration of simulation methodology, parameter selection, and mesh refinement on atmospheric, turbine, and wake QoIs to identify capability gaps in the validation process. For neutral atmospheric boundary layer conditions, the massively parallel large eddy simulation (LES) code Nalu-Wind was used to produce high-fidelity computations for experimental validation using high-quality meteorological, turbine, and wake measurement data collected at the Department of Energy/Sandia National Laboratories Scaled Wind Farm Technology (SWiFT) facility located at Texas Tech University’s National Wind Institute. The wake analysis showed the simulated lidar model implemented in Nalu-Wind was successful at capturing wake profile trends observed in the experimental lidar data.

17 WIND ENERGY↗

Adaptive Spatially Aware I/O for Multiresolution Particle Data Layouts

Large-scale simulations on nonuniform particle distributions that evolve over time are widely used in cosmology, molecular dynamics, and engineering. Such data are often saved in an unstructured format that neither preserves spatial locality nor provides metadata for accelerating spatial or attribute subset queries, leading to poor performance of visualization tasks. Furthermore, the parallel I/O strategy used typically writes a file per process or a single shared file, neither of which is portable or scalable across different HPC systems. We present a portable technique for scalable, spatially aware adaptive aggregation that preserves spatial locality in the output. We evaluate our approach on two supercomputers, Stampede2 and Summit, and demonstrate that it outperforms prior approaches at scale, achieving up to 2.5× faster writes and reads for nonuniform distributions. Furthermore, the layout written by our method is directly suitable for visual analytics, supporting low-latency reads and attribute-based filtering with little overhead.

Usher, Will↗

Rapid Characterization and Statistical Analysis of High-Volume Field-Harvested Photovoltaic Connectors

Photovoltaic (PV) installations heavily depend on connectors for efficient module and string interconnections without requiring skilled labor. Yet this seemingly innocuous component of PV systems is a leading cause of module failures, multiple high-profile fires, and lawsuits in the PV industry. This work aims to answer critical questions regarding why connectors fail and the contributing factors to their failure. The study involves collecting and analyzing more than 17,000 field-harvested connectors from various solar installations across the United States. The vast dataset, which includes connector metadata, visual inspections, and resistance measurements, provides unprecedented insight into the state of health of PV connectors across the US, including the geographic locations, connector types, and installation practices most prone to failures. The work presented here describes a novel rapid characterization method for processing large numbers of connectors and is supported by parallel forensic analysis to discern the root causes of failures as well as a levelized cost of lifetime model to determine the economic ramifications of connector failure. Ultimately, the findings may inform PV developers about the best practices to extend connector longevity and lead to more resilient and reliable PV systems.

connectors↗

A two-level GPU-accelerated incomplete LU preconditioner for general sparse linear systems

This paper presents a parallel preconditioning approach based on incomplete LU (ILU) factorizations in the framework of Domain Decomposition (DD) for general sparse linear systems. We focus on distributed memory parallel architectures, specifically, those that are equipped with graphic processing units (GPUs). In addition to block-Jacobi, we present general purpose two-level ILU Schur complement-based approaches, where different strategies are presented to solve the coarse-level reduced system. These strategies are combined with modified ILU methods in the construction of the coarse-level operator, in order to effectively remove smooth errors by targeting an algebraically smooth vector. We leverage available GPU-based sparse matrix kernels to accelerate the setup and the solve phases of the proposed ILU preconditioner. We evaluate the efficiency of the proposed methods as a smoother for algebraic multigrid (AMG) and as a preconditioner for Krylov subspace methods on challenging anisotropic diffusion problems and a collection of general sparse matrices.

97 MATHEMATICS AND COMPUTING↗

Simulating of magnetic reconnection in solar flares

Magnetic reconnection is an astrophysical process where neighboring magnetic field lines running anti-parallel are reconfigured. Observations of solar flares are what pushed for serious research regarding magnetic reconnection as it seemed to be the underlying mechanism. The reconfiguration of the field lines results in an explosive release of energy as magnetic field energy is converted into plasma kinetic and thermal energies. Our work begins with Athena++, an astrophysical magnetohydrodynamic (MHD) simulation code, and an existing magnetic reconnection setup.

79 ASTRONOMY AND ASTROPHYSICS↗