Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 289 records · Page 16

Simulating ‘Two Ribbon’ Type Solar Flares [Slides]

The underlying mechanism in solar flares in a process known as magnetic reconnection. This astrophysical process is when magnetic field lines, running anti-parallel, break and reconnect. The change in configuration of field lines results in an explosive release of energy as the leftover magnetic energy is converted to kinetic and thermal energies. In this project, we are using an astrophysical magnetohydrodynamic (MHD) simulation code known as Athena++ with a reconnection problem generator file created by Li et al. 2018. Our current research is to successfully implement a radiative cooling term into the MHD equations as it could have important physical effects on the plasma. We use Klimchuk et al. 2008 and SPEX_DM as our chosen cooling functions. The results from the SPEX_DM function show there is a condensation forming that we had predicted but had not seen before with our simulation code. We will continue to analyze the SPEX_DM function and potentially implement thermal conduction.

79 ASTRONOMY AND ASTROPHYSICS↗

Simulating Magnetic Reconnection in ‘Two Ribbon’ Type Solar Flares

Magnetic reconnection is an astrophysical process where neighboring magnetic field lines, facing anti-parallel, are reconfigured. This reconfiguration results in built up magnetic energy being explosively released as it is being converted to plasma kinetic and thermal energies. There are various kinds of simulations used to simulation reconnection; our work begins with Athena++, a magnetohydrodynamic (MHD) simulation code typically used for astrophysical problems, and a reconnection specific code file. The resistive MHD equations are solved with Riemann solvers. There was an initial test run without modifying the code to understand the dynamics of the simulation. We expand on the original reconnection problem file by implementing a radiative cooling term specific to the corona. The radiative cooling is theorized to have an effect on solar coronal plasma and magnetic reconnection dynamics. The cooling term will be tested with various parameters and compared to the case without cooling to study these dynamics. The condensation found in only the with cooling case emphasizes the importance of implementing this feature and will be later tested with a thermal conduction term. We want to determine the parameter regime where non-equilibrium cooling will be important for the reconnection dynamics.

79 ASTRONOMY AND ASTROPHYSICS↗

Applying Time-Parallelization to Turbulent Flows

Parallelization of the temporal domain is explored for the solution of turbulent flows. Multigrid reduction-in-time (MGRIT) is used to advance the large-scale fluid dynamics in time sequentially on the coarsest space-time grid but propagate the information in time parallel on all other levels. The goal of this process is to accurately and efficiently resolve the coarse-scale turbulence structure and use that to drive the fine-scales of the turbulent flow. The extra forcing from nonlinear multigrid facilitates the coupling and interaction between fine and coarse scales, through which the multiscale nonlinear physics is properly captured. Adaptive mesh refinement is employed to finely resolve only the regions with strong gradients, which provides further computational efficiency. The underlying computational fluid dynamics solver is a fourth-order finite-volume scheme with the standard 4-stage Runge-Kutta method. An advanced approach is devised and implemented to enable MGRIT to solve highly turbulent flows successfully. Furthermore, the method is applied to solve a Taylor-Green vortex problem and a doubleshear-layer turbulent mixing flow. Results are promising, validating that MGRIT with the filtering approach has the potential to efficiently solve general turbulent flows.

Computational Fluid Dynamics↗

Porting HEP Parameterized Calorimeter Simulation Code to GPUs

The High Energy Physics (HEP) experiments, such as those at the Large Hadron Collider (LHC), traditionally consume large amounts of CPU cycles for detector simulations and data analysis, but rarely use compute accelerators such as GPUs. As the LHC is upgraded to allow for higher luminosity, resulting in much higher data rates, purely relying on CPUs may not provide enough computing power to support the simulation and data analysis needs. As a proof of concept, we investigate the feasibility of porting a HEP parameterized calorimeter simulation code to GPUs. We have chosen to use FastCaloSim, the ATLAS fast parametrized calorimeter simulation. While FastCaloSim is sufficiently fast such that it does not impose a bottleneck in detector simulations overall, significant speed-ups in the processing of large samples can be achieved from GPU parallelization at both the particle (intra-event) and event levels; this is especially beneficial in conditions expected at the high-luminosity LHC, where extremely high per-event particle multiplicities will result from the many simultaneous proton-proton collisions. We report our experience with porting FastCaloSim to NVIDIA GPUs using CUDA. A preliminary Kokkos implementation of FastCaloSim for portability to other parallel architectures is also described.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Two-Stage Gauss-Seidel Preconditioners and Smoothers for Krylov Solvers on a GPU Cluster: Preprint

Gauss-Seidel (GS) relaxation is often employed as a preconditioner for a Krylov solver or as a smoother for Algebraic Multigrid (AMG). However, the requisite sparse triangular solve is difficult to parallelize on many-core architectures such as graphics processing units (GPUs). In the present study, the performance of the sequential GS relaxation based on a triangular solve is compared with two-stage variants, replacing the direct triangular solve with a fixed number of inner Jacobi-Richardson (JR) iterations. When a small number of inner iterations is sufficient to maintain the Krylov convergence rate, the two-stage GS (GS2) often outperforms the sequential algorithm on many-core architectures. The GS2 algorithm is also compared with JR. When they perform the same number of ops for SpMV (e.g. three JR sweeps compared to two GS sweeps with one inner JR sweep), the GS2 iterations, and the Krylov solver preconditioned with GS2, may converge faster than the JR iterations. Moreover, for some problems (e.g. elasticity), it was found that JR may diverge with a damping factor of one, whereas two-stage GS may improve the convergence with more inner iterations. Finally, to study the performance of the two-stage smoother and preconditioner for a practical problem, these were applied to incompressible uid ow simulations on GPUs.

algebraic multigrid↗

RAPID Manufacturing Institute Final Report

The Rapid Advancement of Process Intensification Deployment (RAPID) Manufacturing Institute, founded in 2017, is a public/private partnership between the U.S. Department of Energy and the American Institute of Chemical Engineers (AIChE). RAPID promotes the development, deployment and commercialization of Process Intensification (PI) and Modular Chemical Process Intensification (MCPI) technologies, enabling U.S. manufacturing to reduce energy consumption, improve process efficiencies and lower investment and operating costs. This mission was carried out through parallel work breakdown structure elements including the establishment of committees to guide the operations and technical direction of RAPID, the establishment of management practices and institute processes, education and workforce development (EWD), and six technical focus areas for the development of technologies to advance PI and MCPI. Throughout the initial six-year cooperative agreement, RAPID worked to meet performance metrics which focused on the operation and sustainment of the institute, education and workforce development and the development of PI and MCPI for the advancement of U.S. manufacturing. All these metrics were successfully met through a total of 43 projects which leveraged $\$$70M Federal with $\$$90M cost share. As a result of these efforts, 84 private and public organizations were brought together by RAPID as members to co-invest in R&D, commercialization and deployment of innovative technologies. In the research portfolio, 82% of the 38 projects achieved > 20% energy efficiency improvement. A RAPID Test Network was developed with 51 testbed facilities to enable access to resources, facilities, tools, and expertise. Eight EWD programs were also developed with over 13,000 impressions. RAPID’s efforts to research, develop, demonstrate, and deploy high-impact PI and modular process technology solutions have enabled reduced energy use, increased sustainability, and improved profitability for U.S. manufacturing.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

TOUGH3-FLAC3D: a modeling approach for parallel computing of fluid flow and geomechanics

The recent development of the TOUGH3 code allows for a faster and more reliable fluid flow simulator. At the same time, new versions of FLAC3D are released periodically, allowing for new features and faster execution. In this paper, we present the first implementation of the coupling between TOUGH3 and FLAC3Dv6/7, maintaining parallel computing capabilities for the coupled fluid flow and geomechanical codes. We compare the newly developed version with analytical solutions and with the previous approach, and provide some performance analysis on different meshes and varying the number of running processors. Finally, we present two case studies related to fault reactivation during CO 2 sequestration and nuclear waste disposal. The use of parallel computing allows for meshes with a larger number of elements, and hence more detailed understanding of thermo-hydro-mechanical processes occurring at depth.

58 GEOSCIENCES↗

Developing an Equity Framework for State Regulatory Decision-Making

The report presents a framework for states that seek to incorporate equity into regulatory decision-making. Berkeley Lab contextualizes approaches and metrics from various states into example processes to exemplify how this may be done, including topics such as the development of equity goals and definitions, intervenor funding, community engagement, performance-based ratemaking, and utility resource planning. The report offers five takeaways and considerations, supported by these examples: 1. Equity comprises multiple tenets and stages, all of which must be considered in parallel. 2. At the start of designing new equity-related processes, it is critical to establish clear and actionable goals, definitions, roles, and responsibilities to ensure progress. 3. Once goals are established, it is critical to align tools and metrics that bridge the gap between what an intervention may do and how it may impact communities and households. 4. Processes should be stakeholder driven. It is important to not only increase education and outreach, but to actively seek out and incorporate feedback from inclusive public processes and build in accountability mechanisms. Processes should be iterative. Feedback loops between evaluations and program design provide the flexibility to better align existing interventions with community priorities and to incorporate equity into future decision-making.

99 GENERAL AND MISCELLANEOUS↗

Diffractive Multiplexing for High-Throughput Roll-to-Roll Laser Patterning of Flexible Organic Photovoltaic Modules (Final Report)

The purpose of this project is to demonstrate a cost-effective, high-throughput roll-to-roll (R2R) process for patterning of flexible, semitransparent organic photovoltaic (OPV) modules by developing diffractive optics-based laser multiplexing (DOL Multiplexing). DOL Multiplexing allows a single, high-powered laser source to perform parallel scribing across the R2R web width in a manner compatible with high process speeds. Such a process could have enormous benefits in terms of increased process speeds and reduced costs, both up-front capital costs, and long-term operational costs, over galvanometer-based step and scan methods or many-laser systems.

14 SOLAR ENERGY↗

Data Mining and Visualization of High-Dimensional ICME Data for Additive Manufacturing

Integrated computational materials engineering (ICME) methods combining CALPHAD with process-based simulations can produce rich, high-dimensional data for alloy and process design. In ICME methods for metallurgical applications, the visualization and interpretation of such high-dimensional data has previously been through heat maps represented in 2 or 3 dimensions. While such an approach is ideal when one variable is varied at a time, in the case of high-dimensional data with multiple variables varied simultaneously, as is the case in additive manufacturing, interpreting the trends through two- or three-dimensional heat maps becomes challenging. Here, we propose a strategy of mixed visual data mining and quantitative analysis for high-dimensional metallurgical and process data using high-throughput thermodynamic calculations. Two case studies show the application of the proposed approach. The first case study investigated the effects of feedstock chemistry on the δ ferrite formation in 316L stainless steel powders used for binder jet additive manufacturing. The second case study linked Scheil–Gulliver calculations to a process model for dissimilar joining of aluminum alloys 5356 and 6111 during laser hot-wire additive manufacturing. Both cases contained thousands of calculated data points, showcasing the utility of visual data analysis through parallel coordinate plotting, Pearson correlation coefficient matrices, and scatter matrices compared to traditional process maps. These visualization techniques can be extended to many additive manufacturing problems to capture process–structure–property relationships for additively manufactured components.

36 MATERIALS SCIENCE↗

h5bench: A unified benchmark suite for evaluating HDF5 I/O performance on pre‐exascale platforms

Summary Parallel I/O is a critical technique for moving data between compute and storage subsystems of supercomputers. With massive amounts of data produced or consumed by compute nodes, high‐performant parallel I/O is essential. I/O benchmarks play an important role in this process; however, there is a scarcity of I/O benchmarks representative of current workloads on HPC systems. Toward creating representative I/O kernels from real‐world applications, we have created h5bench , a set of I/O kernels that exercise hierarchical data format version 5 (HDF5) I/O on parallel file systems in numerous dimensions. Our focus on HDF5 is due to the parallel I/O library's heavy usage in various scientific applications running on supercomputing systems. The various tests benchmarked in the h5bench suite include I/O operations (read and write), data locality (arrays of basic data types and arrays of structures), array dimensionality (one‐dimensional arrays, two‐dimensional meshes, three‐dimensional cubes), I/O modes (synchronous and asynchronous). In this paper, we present the observed performance of h5bench executed along several of these dimensions on existing supercomputers (Cori and Summit) and pre‐exascale platforms (Perlmutter, Theta, and Polaris). h5bench measurements can be used to identify performance bottlenecks and their root causes and evaluate I/O optimizations. As the I/O patterns of h5bench are diverse and capture the I/O behaviors of various HPC applications, this study will be helpful to the broader supercomputing and I/O community.

97 MATHEMATICS AND COMPUTING↗

Process Design and Techno-Economic Analysis of the Modular Staged Pressurized Oxy-Combustion (SPOC) Power Plant for Biomass

This work describes the process design and techno-economic analysis (TEA) of the modular SPOC power plant for biomass firing and coal-biomass co-firing. Two Rankine cycles were considered: a supercritical steam cycle (242 bar, 593°C, 593°C) with 550 MWe net output and a subcritical cycle (166 bar, 566°C, 566°C) with 200 MWe net output. For both cases, 95% carbon capture was modeled, and hybrid poplar biomass was chosen to generate carbon-negative power. In addition, the supercritical 500 MWe case included a 25% biomass co-firing (carbon neutral) case. For both cycles, a 100% Powder River Basin coal firing case was used for comparison purposes. In the SPOC process, oxygen is produced via a cryogenic air separation unit (ASU) and the heat generated from the compression of air is integrated into the steam cycle and utilized for boiler feed water pre-heating. Unique to the SPOC process, the boilers are pressurized and arranged in a series-parallel configuration, with minimized flue gas recirculation. The flue gas is cooled and scrubbed in the direct-contact cooler (DCC) column, and the moisture in the flue gas is condensed, leaving the bottom of the DCC at a sufficiently high temperature such that it can be used for boiler feed water pre-heating, improving plant thermal efficiency. Following drying and purification, CO2 in the flue gas is at the purity required for storage or utilization. The performance data were obtained from process modelling via Aspen Plus®. The stream data from Aspen Plus® were used as an input for the AACE Class 5 cost study. Ultimately, the capital costs, Levelized Cost of Electricity (LCOE), and cost of CO2 captured and avoided were obtained. The HHV efficiency of the carbon negative 550 MWe supercritical SPOC case (34.8%) was clearly above those reported by NETL for the BECCS baseline cases of supercritical pulverized coal with capture (B12B, 31.5%) and the 49% biomass co-firing case with capture (PA3, 29.2%). The HHV efficiency of the carbon-negative subcritical plant is also higher than the subcritical baseline PC plant with capture (case B11B.95) presented by NETL (32% vs 29.7%). The LCOE for the SPOC 100% biomass case was similar to the LCOE for the BECCS 49% biomass with carbon capture case ($147/MWh), and the SPOC carbon neutral case LCOE was lower ($110/MWh) than the cost for the NETL baseline SC coal firing case with 90% carbon capture ($114/MWh).

Magalhaes, Duarte↗

Hybrid Solar System (Final Scientific/Technical Report)

GTI Energy (GTI) teamed with the University of California at Merced (UCM) to scaleup the hybrid solar system (HSS) technology for demonstrating its performance at the US Gypsum (USG) plant in Plaster City, California. The technology integrates two-stage concentrating solar collector with matching particle thermal transport and storage (TSS) system to deliver cost-effective, and on-demand distributed high temperature industrial process heat up to 600°C with solar thermal, in this case to a gypsum kettle, to reduce its fuel use and carbon footprint. Current solar technologies, which reach these temperatures, are not distributable (towers) or cost-effective (dish). The research team developed a conceptual system design for host site retrofit, including preliminary heat balance, process flow diagram, particle to process heat exchanger and equipment placements at the site. Subsequently, parallel efforts were carried out at UCM to design, build and test a 12 m long commercial scale prototype concentrating thermal-only collector system and at GTI to design, build and test a matching 650°C capable particle TTS system. The nominal 50 kWth collector consists of a parabolic trough and three 4 m long two-stage receivers in series. Prior to on-sun testing, a 4 m long receiver was fabricated and successfully tested at 650 °C in a laboratory setting for 100 hrs of continuous operation showing less than 15% radiation loss. A 7 m wide x 17 m long parabolic trough was then installed at UCM for on-sun testing of the 12 m long receiver, and concurrently several 4 m long receivers were built. The optics of the parabolic trough were calibrated, and on-sun test were carried out on 12 m long receivers. During tests, the intense solar radiation (53x) caused the absorber tubes in the receivers to bend, reducing the overall optical efficiency. To address the bending issue, a self-consistent algorithm that includes ray tracing, thermal and deformation models was developed to perform thermal stress analysis on absorbers for parabolic solar collectors. Results obtained with this algorithm showed a dramatic rise in deformation as absorber tube length increases. A combined efficiency parameter that includes the occluded area for the mounts was developed to obtain an optimized tube length obtained. Based on the results, a length of 2.7 m for the absorber + 0.2 m for the coupler was chosen to minimize any bending and optimize optical efficiency while maintaining ease of mounting. The associated particle TTS system was designed, built and successfully tested at GTI. It includes storage, receiving and lock hoppers and piping that simulates the transfer of captured solar energy to an actual industrial furnace. Tests over 77 charge-discharge cycles demonstrated <2% particle degradation, with no problematic particle accumulations and no flow interruptions. The piping pressure drop was about 5 psi. The team also worked with Stanley Consultants (Stanley) to prepare conceptual and preliminary engineering packages to facilitate follow-on development and commercialization efforts. These include process and instrumentation diagram’s (P&ID’s), general arrangements, electrical one-line, project definitions document, equipment data sheets, schedule, and construction cost estimate for 2 MWth system. Updated HSS technology commercialization and customer engagement plans and detailed costs and evaluated market trade-offs and manufacturing.

03 NATURAL GAS↗

Acoustophoretic Additive Manufacturing for Scalable 3D Battery Electrodes

This project focused on investigating two acoustic-based processing methods: a nozzle-based printhead and a chamber that map to two different battery electrode architectures: (1) a line-pattern electrode and (2) a grid-pattern electrode. These two parallel manufacturing and electrode geometry explorations were proposed for the project to understand the process space of acoustic-based manufacturing methods to fabricate patterned battery electrodes. This two-path exploration also allowed us to de-risk the overall project and not rely on a single process for creating patterned electrodes. This project consisted of six high-level tasks aimed at transitioning the concept of acoustic focusing for battery electrodes from a technology readiness level (TRL) of 1 to 3 by project conclusion. Overall, we believe we have developed a practical and high-impact processing method that is chemistry agnostic and suitable for large-area fabrication of both 3D LIBs and other functional material systems where structuring on the scale of tens of microns has the potential to break conventional bulk material trade-offs in performance. In the case of batteries for electric vehicles, structuring 3D LIBs with our acoustic process breaks traditional energy and power trade-offs observed with conventional flat battery packs.

25 ENERGY STORAGE↗

A time-parallel method for scalable heat transfer simulations of additive manufacturing

Here, a major challenge in simulating the thermal behavior in additive manufacturing processes is the disparate length and time scales between transport phenomena occurring in the melt pool and the component. A common simulation approach relies on spatial decomposition for parallel computing, but due to the nature of heat transfer in AM, where most of the computational expenditure is localized near the melt pool, the computational speedup from spatial parallelization saturates quickly. Therefore, additional parallelism by means of time-domain decomposition is needed to fully take advantage of high-performance computing (HPC) resources. This work introduces a time-parallel method to improve the computational scalability of additive manufacturing simulations on HPC systems, while maintaining high temporal resolution of heat transfer near the melt pool. The method, inspired by the nonlinear paraexp formalism, performs an iterative superposition of nonlinear solutions to the initial value problem, integrating the heat equation across overlapping time-parallel intervals. For a single layer of the NIST AMB2018–01 L7 benchmark problem, the method achieves a 38.51x speedup in wall-clock time with a maximum error in the global temperature solution of 0.99%. This reduces the total solution time from 196.72 min to 5.11 min on 128 nodes of the ORNL Frontier supercomputer. The tradeoff between accuracy and total wall-clock time is investigated and recommendations for time-parallel deployment for AM problems are made.

Additive manufacturing↗

Resonant x-ray absorption of strong-field-ionized CF 3 Br

We report on an experimental and theoretical study of strong-field laser ionization of CF 3 Br followed by resonant x-ray absorption at the Br K-edge. Distinct 1s -> 4p, 5p Rydberg transitions of Br q+ (q = 1-4) atomic ions are observed and identified with Hartree-Fock-Slater and relativistic configuration interaction calculations. Time-dependent density functional theory and ab initio molecular dynamics calculations were performed to simulate the dissociative ionization process and the molecular orbitals for the q = 1-4 charge states. Measurements were made with both parallel and perpendicular linear polarizations of the laser and x-rays, but dichroism was not observed, indicating negligible alignment by the laser ionization process. This result is explained by calculations on atomic Br and the molecular simulation.

74 ATOMIC AND MOLECULAR PHYSICS↗

Electromagnetic turbulence simulation of tokamak edge plasma dynamics and divertor heat load during thermal quench

Abstract The edge plasma turbulence and transport dynamics, as well as the divertor power loads during the thermal-quench phase of tokamak disruptions, are numerically investigated with BOUT++’s flux-driven six-field electromagnetic turbulence model. Here, transient yet intense particle and energy sources are applied at the pedestal top to mimic the plasma power drive at the edge induced by a core thermal collapse, which flattens the core temperature profile. Interesting features, such as surging of divertor heat load (up to 50 times) and broadening of heat-flux width (up to four times) on the outer-divertor target plate, are observed in the simulation, in qualitative agreement with experimental observations. The dramatic changes in divertor heat load and width are due to the enhanced plasma turbulence activities inside the separatrix. Two cross-field transport mechanisms, namely, the E × B turbulent convection and the stochastic parallel advection/conduction, are identified to play important roles in this process. First, an elevated edge pressure gradient drives instabilities and subsequent turbulence in the entire pedestal region. The enhanced turbulence not only transports particles and energy radially across the separatrix via the E × B convection, which causes the initial divertor heat-load burst, but it also induces amplified magnetic fluctuation B ˜ . Once themagnetic fluctuation is large enough to break the magnetic flux surface, magnetic flutter effect provides an additional radial transport channel. In the late stage of our simulation, | B ˜ r / B 0 | reaches to 10 −4 level that completely breaks magnetic flux surfaces such that stochastic field lines are directly connecting pedestal top plasma to the divertor target plates or first wall, further contributing to the divertor heat-flux width broadening.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Innovative rail transport of a supersized land-based wind turbine blade

Wind turbine blade logistic providers are being challenged with escalating costs and routing complexities as one-piece blade approach lengths of 75 m in various regions of the U.S. land-based market. New lower cost solutions are needed to enable further reductions in the levelized cost of energy (LCOE) and continued market expansion. In this paper, a novel method of using existing U.S. rail infrastructure to deploy 100-m, one-piece blades to U.S. land-based wind sites is numerically investigated. The study removes the constraint that blades must be kept rigid during transport, and it allows bending to keep blades within a clearance profile while navigating horizontal and vertical curvatures. Novel system optimization and blade design processes consider blade structural constraints and rail logistic constraints in parallel to develop a highly flexible, rail-transportable blade. Results indicate maximum deployment potential in the Interior region of the United States and limited deployment potential in other regions. The study concludes that innovative rail transportation solutions combined with advanced rotor technologies can provide a feasible alternative to segmentation and support continued LCOE reductions in the U.S. land-based wind energy market.

17 WIND ENERGY↗