Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel codes”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 163 records · Page 9

Modeling of small tungsten dust grains in EAST tokamak with NDS-BOUT ++

In order to investigate the transport of small dusts as well as their evolution property along their trajectories, the NDS module is developed under the BOUT++ framework, a highly desirable C++ code package to perform parallel plasma fluid simulations with an arbitrary number of equations in three-dimensional curvilinear coordinates. Due to the severe dust ablation in fusion plasmas, the dust size would decrease from micrometer to nanometer, resulting in impurities. Small dusts in the simulations here are specified as tungsten spheres with the radii on or below the order of submicrometer. The Rayleigh limit is included in the charging process when the dust is ablated to the droplet phase. The simulation results from the NDS module show that a 200 nm radius spherical tungsten dust originated from upper divertor region of EAST Tokamak is ablated completely due to the intense heating from the incoming plasma inside the core region, well consistent with the CCD footage of EAST shot # 81459. Furthermore it is found that the magnetic field dominates the dust transport when the dust radius is below 100 nm during the ablation along the trajectory. Our simulations predict that a 10 nm radius spherical tungsten dust injected from the inner midplane is well constrained by the magnetic field, and it reaches the inner divertor target with a velocity on the order of km/s.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Energetic particle marginal stability profile for HL-2M integrated simulation based on neural network module

Abstract A critical gradient model is employed to develop a module of energetic particle (EP) marginal stability profiles in OMFIT integrated simulations for studying EP transport. Currently, each iteration of transport evolution is approximately 10 min in the integrated simulation, whereas, the EP marginal stability profile, which serves as an input in the integrated simulation could take much longer; the reason being a combination of the TGLFEP and EPtran codes is employed in our previous investigation. To reduce the simulation time, the critical gradient is predicted by a neural network instead of the TGLFEP code, and the EPtran code is revised with parallel computing, so that the running time of this module can be controlled to within 5 min. The predictions are in good agreement with previous approaches. The integrated simulation of HL-2M with Alfven eigenmodes transported by neutral beam EP profiles indicates that EP transport reduces the total pressure and current as expected, but could also under some conditions raise the safety factor in the core, which is favorable for reversed magnetic shear and high-performance plasmas.

Physics↗

miniWeather

The code “miniWeather” solves the 2-D Euler equations with a Finite-Volume method on a regular Cartesian grid with Runge-Kutta time stepping, 4th-order hypervisocisty, and 4th-order spatial accuracy. The code is primarily a parallel programming training mini application in Fortran, C, and C++ with MPI, OpenMP, OpenMP offload, OpenACC, and C++ portability approaches. It can also serve as a mini application for acceptance testing for new machines and compilers for both functionality and performance.

Norman, MatthewR↗

Global Pathway Selection with Zero-RK

Global Pathway Selection (GPS) is an algorithm to effectively generates reduced (skeletal) chemistry mechanisms, which speeds up simulations and can be used as a systematic analytics tool to extract insights from complex reacting system. This release is an extension of the original code to run in parallel and to use LLNL's Zero-RK solver for fast solution of chemical problems.

Whitesides, RussellA↗

Global Pathway Selection with Zero-RK v0.5

Global Pathway Selection (GPS) is an algorithm to effectively generates reduced (skeletal) chemistry mechanisms, which speeds up simulations and can be used as a systematic analytics tool to extract insights from complex reacting system. This release is an extension of the original code to run in parallel and to use LLNL's Zero-RK solver for fast solution of chemical problems.

Whitesides, RussellA [Lawrence Livermore National ↗

Large Eddy Simulations of Turbulence below Antarctic Ice Shelves

Turbulence below ice shelves is of key importance for predicting ice-shelf melt rates and consequently the contribution of ice sheets to sea level rise. In this project we conducted Large-Eddy Simulations (LES) to improve our understanding of sub-ice-shelf ocean turbulence and the relationship between ice-shelf melt rates and ocean conditions. Over the course of the second and final year of this project (FY2020), we accomplished both major code developments for the PArallel Large-eddy simulation Model (PALM; Maronga et al., 2015) and conducted a suite of simulations which form the basis for a publication in preparation.

58 GEOSCIENCES↗

SBIR Phase I Final Report, TACO: Distributed and Heterogeneous Sparse Compiler

Tensor algebra is a powerful tool for computing, but writing optimized codes that operate on sparse tensors can be very complex. This project enables a Tensor Algebra Compiler (TACO) that simplifies this task from man-years to man-days and extends TACO to support complex and large distributed systems. This report details the hypotheses, approaches used, and findings in this project.

97 MATHEMATICS AND COMPUTING↗

Predictive Tools for Customizing Heat Treatment of Additively Manufactured Aerospace Components

Laser-bed powder fusion (LBPF) additive manufacturing is increasingly being used to produce components of complex geometries using the Ni-base superalloy Inconel 718. The composition and the microstructure of the alloy are currently well optimized for wrought components made using conventional manufacturing processes such as rolling, forging, extrusion, etc. The attractive mechanical properties of the alloy result from the underlying austenitic matrix with fine equiaxed grains, and a high density and uniform distribution of the precipitation hardening phase, γ". Heat treatment steps such as homogenization, solutioning and aging are well documented for the wrought alloy. However, when the same wrought alloy compositions are used for the additive manufacturing (AM) processes, the asprocessed microstructure is significantly different, because of the different thermal history associated with LBPF, including rapid solidification and multiple temperature excursions that lead to multiple re-melting and reheating in the solid state. Rapid solidification introduces potential non-equilibrium effects at the moving solid-liquid interfaces that impact the extent of solute segregation, as well as the morphology of the dendritic grains that form. In order to recover the target mechanical properties, AM components have to undergo post-process heat treatments. However, such heat treatments have to be custom designed for the AM process and the component geometry because of the expected vast differences in the microstructure at various locations of a component with complex geometry. The homogenization and precipitation steps should be optimized for the component so that target mechanical properties can be obtained throughout the part. The objective of this research is to utilize High Performance Computing in phase field simulations of microstructure evolution during post-processing of AM components. The physics-based modeling will be beneficial in reducing the experimental effort required for heat treatment process selection, optimization, and certification, thus leading to a significant reduction in energy consumption for AM and post-processing heat treatment. The optimization study will help identify heat treatments steps that are critical for development of a final desired microstructure with the minimum energy input. This combined with shortening of the production cycle (time-to-market) by reducing the number of failed parts (property targets), and reduction in the number of iterations for process optimization, will enable 30-40% savings in the energy costs. Phase field simulations of the degree of homogenization and the effect of local matrix composition on the nucleation and growth of competing precipitating phases were performed using the Microstructure Evolution Using Massively Parallel Phase Field Simulations code developed in-house at the Oak Ridge National Laboratory. The simulations were able to successfully capture the kinetics of nucleation and growth, and morphologies of various precipitating phases as a function of local matrix compositions and composition gradients characteristic of local microstructures arising from location-dependent variations in the thermal conditions. Future work will involve extending the simulations to a length scale consisting of multiple dendrites, so that the effect of homogenization on the coarsening of the dendrites can be simulated and used as an additional input to the optimization of the heat treatment process.

36 MATERIALS SCIENCE↗

xRAGE: A Brief Overview [Slides]

This report provides information about xRAGE, "a multidimensional, multimaterial, massively parallel, Eulerian AMR, multiphysics code." Challenges and future possibilities for xRAGE are also discussed.

97 MATHEMATICS AND COMPUTING↗

Reassessment of Boiling Water Reactor Rod Cusping Capability of VERA-MPACT

This technical report assesses the current modeling capabilities of the Virtual Environment for Reactor Applications (VERA) for minimizing control rod cusping in boiling water reactors. Assessment is performed on a 10 × 10 General Electric (GE)-14 assembly as well as a simplified version of the same assembly. Runs were performed using VERA for three available cusping treatments in the Michigan PArallel Characteristics Transport (MPACT) code: no treatment, subplane treatment, and fitted polynomial correction treatment. Subplane treatment is discussed, but specific results are not provided because of failures in the method of implementation for boiling water reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Reassessment of Boiling Water Reactor Rod Cusping Capability of VERA-MPACT

This technical report assesses the current modeling capabilities of the Virtual Environment for Reactor Applications (VERA) for minimizing control rod cusping in boiling water reactors. Assessment is performed on a 10 × 10 General Electric (GE)-14 assembly as well as a simplified version of the same assembly. Runs were performed using VERA for three available cusping treatments in the Michigan PArallel Characteristics Transport (MPACT) code: no treatment, subplane treatment, and fitted polynomial correction treatment. Subplane treatment is discussed, but specific results are not provided because of failures in the method of implementation for boiling water reactors.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Report on Preliminary Detailed Experimental Plan for Neutron Irradiation of A709 at ATR and HFIR

Advanced nuclear power technologies will use higher temperatures as a means to extract energy at a higher efficiency than current plants and therefore put a larger demand on the structural materials. Improved performance of structural materials could enable greater safety margins, longer plant lifetimes, and reduce maintenance costs. Alloy 709 is championed as the next generation of austenitic alloys for advanced nuclear reactors. In parallel to the ASME code case pursued, the AMMT program is initiating a neutron irradiation campaign to provide first-of-a-kind engineering data to establish operational design parameters and how the mechanical response is modified by environmental factors. This document refines the AMMT neutron irradiation campaign to a 4 year program to support the generation of creep knockdown factors for Alloy 709 and welded Alloy 709. The campaign is divided among two national laboratories, Oak Ridge National Laboratory and Idaho National Laboratory, to use the strengths of each laboratory. Through a cooperative plan, time-independent properties and time-dependent properties will be obtained across a large temperature window, nominally 300°C to 800°C, damage levels up to 10 dpa, and with and without the impacts of transmutation produced helium.

99 GENERAL AND MISCELLANEOUS↗

Report on Preliminary Detailed Experimental Plan for Neutron Irradiation of A709 at ATR and HFIR

Advanced nuclear power technologies will use higher temperatures as a means to extract energy at a higher efficiency than current plants and therefore put a larger demand on the structural materials. Improved performance of structural materials could enable greater safety margins, longer plant lifetimes, and reduce maintenance costs. Alloy 709 is championed as the next generation of austenitic alloys for advanced nuclear reactors. In parallel to the ASME code case pursued, the AMMT program is initiating a neutron irradiation campaign to provide first-of-a-kind engineering data to establish operational design parameters and how the mechanical response is modified by environmental factors. This document refines the AMMT neutron irradiation campaign to a 4 year program to support the generation of creep knockdown factors for Alloy 709 and welded Alloy 709. The campaign is divided among two national laboratories, Oak Ridge National Laboratory and Idaho National Laboratory, to use the strengths of each laboratory. Through a cooperative plan, time-independent properties and time-dependent properties will be obtained across a large temperature window, nominally 300°C to 800°C, damage levels up to 10 dpa, and with and without the impacts of transmutation produced helium.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Extended Physics-Informed Neural Networks (XPINNs): A Generalized Space-Time Domain Decomposition Based Deep Learning Framework for Nonlinear Partial Differential Equations

Here we propose a generalized space-time domain decomposition approach for the physics-informed neural networks (PINNs) to solve nonlinear partial differential equations (PDEs) on arbitrary complex-geometry domains. The proposed framework, named eXtended PINNs ( X P I N N s ), further pushes the boundaries of both PINNs as well as conservative PINNs (cPINNs), which is a recently proposed domain decomposition approach in the PINN framework tailored to conservation laws. Compared to PINN, the XPINN method has large representation and parallelization capacity due to the inherent property of deployment of multiple neural networks in the smaller subdomains. Unlike cPINN, XPINN can be extended to any type of PDEs. Moreover, the domain can be decomposed in any arbitrary way (in space and time), which is not possible in cPINN. Thus, XPINN offers both space and time parallelization, thereby reducing the training cost more effectively. In each subdomain, a separate neural network is employed with optimally selected hyperparameters, e.g., depth/width of the network, number and location of residual points, activation function, optimization method, etc. A deep network can be employed in a subdomain with complex solution, whereas a shallow neural network can be used in a subdomain with relatively simple and smooth solutions. We demonstrate the versatility of XPINN by solving both forward and inverse PDE problems, ranging from one-dimensional to three-dimensional problems, from time-dependent to time-independent problems, and from continuous to discontinuous problems, which clearly shows that the XPINN method is promising in many practical problems. The proposed XPINN method is the generalization of PINN and cPINN methods, both in terms of applicability as well as domain decomposition approach, which efficiently lends itself to parallelized computation. The XPINN code is available on h t t p s : / / g i t h u b . c o m / A m e y a J a g t a p / X P I N N s .

97 MATHEMATICS AND COMPUTING↗

High speed two-dimensional event detection and imaging using an analog interface and a massively parallel processor

A quantitative pulse count (event detection) algorithm with linearity to high count rates is accomplished by combining a high-speed, high frame rate camera with simple logic code run on a massively parallel processor such as a GPU or a FPGA. The parallel processor elements examine frames from the camera pixel by pixel to find and tag events or count pulses. The tagged events are combined to form a combined quantitative event image.

Waugh, Justin↗

Explicit structure-preserving geometric particle-in-cell algorithm in curvilinear orthogonal coordinate systems and its applications to whole-device 6D kinetic simulations of tokamak physics

Explicit structure-preserving geometric particle-in-cell (PIC) algorithm in curvilinear orthogonal coordinate systems is developed. The work reported represents a further development of the structure-preserving geometric PIC algorithm achieving the goal of practical applications in magnetic fusion research. The algorithm is constructed by discretizing the field theory for the system of charged particles and electromagnetic field using Whitney forms, discrete exterior calculus, and explicit non-canonical symplectic integration. In addition to the truncated infinitely dimensional symplectic structure, the algorithm preserves exactly many important physical symmetries and conservation laws, such as local energy conservation, gauge symmetry and the corresponding local charge conservation. As a result, the algorithm possesses the long-term accuracy and fidelity required for first-principles-based simulations of the multiscale tokamak physics. The algorithm has been implemented in the SymPIC code, which is designed for high-efficiency massively-parallel PIC simulations in modern clusters. The code has been applied to carry out whole-device 6D kinetic simulation studies of tokamak physics. A self-consistent kinetic steady state for fusion plasma in the tokamak geometry is numerically found with a predominately diagonal and anisotropic pressure tensor. The state also admits a steady-state sub-sonic ion flow in the range of 10 km s -1 , agreeing with experimental observations and analytical calculations Kinetic ballooning instability in the self-consistent kinetic steady state is simulated. It is shown that high-n ballooning modes have larger growth rates than low-n global modes, and in the nonlinear phase the modes saturate approximately in 5 ion transit times at the 2% level by the E × B flow generated by the instability. These results are consistent with early and recent electromagnetic gyrokinetic simulations.

43 PARTICLE ACCELERATORS↗

Enabling Parallel Execution of System-level Simulations in SAM

This report summarizes the recent code updates related to “element ghosting” in SAM to enable the parallel execution of system-level simulations using multiple processors/cores. Unlike typical MOOSE-based applications, for system-level simulations, SAM mostly deals with a collection of discrete small pieces of meshes, and the connection of physics on these meshes are realized by using “connector” types of components/code structures, such as conjugate heat transfer and flow junctions. The required code implementation is to correctly mark the necessary ghost elements for each type of such components/code structures; thus, the lower-level libraries can correctly perform the necessary data transfer between processors (CPUs) when executed in parallel mode. After the code updates, SAM can now run system-level simulations in the parallel mode. The parallel execution capability was then tested with an ABTR input model with 23k DOFs. Significant speedup was demonstrated when the optimal number of CPUs were used in parallel mode. Future systematic studies on parallelization performance using additional test cases covering different physics/scenarios will be needed to provide additional insights into the scalability of SAM parallelization.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

High-Level Synthesis of Parallel Specifications Coupling Static and Dynamic Controllers

The increased need for efficient ways to implement domain-specific accelerators is driving design methodologies towards the use of abstractions higher than the Register Transfer Level (RTL). In this scenario, High Level Synthesis (HLS) plays a significant role by enabling the automatic generation of custom hardware accelerators starting from high level descriptions (e.g., C code). Conventional HLS tools exploit parallelism mostly at the Instruction Level (ILP). They statically schedule the input specifications, and build centralized Finite State Machine (FSM) controllers. However, aggressive exploitation of ILP in many applications has diminishing returns and, usually, centralized approaches do not efficiently exploit coarser parallelism because FSMs are inherently serial. In this paper we present a HLS framework able to synthesize applications that, beside ILP, also expose Task Level Parallelism (TLP). An application can expose TLP through annotations that identify the parallel functions (i.e., tasks). To generate accelerators that efficiently execute concur- rent tasks, we need to solve several issues: devise a mechanism to support concurrent execution flows, exploit memory parallelism, and manage synchronization. To support concurrent execution flows, we introduce a novel adaptive controller. The adaptive controller is composed of a set of interacting control elements that independently manage the execution of a single operation or function call. These control elements check dependencies and resource constraints at runtime, enabling as soon as possible execution. To support parallel access to shared memories and synchronization, we introduce a novel Hierarchical Memory Interface (HMI). With respect to previous solutions, the proposed interface supports multi-ported memories and atomic memory operations, which commonly occur in parallel programming. Our framework can generate the hardware implementation of C functions by employing two different approaches, depending on its characteristics. If a function exposes TLP, then the framework generates hardware implementations based on the adaptive controller. Otherwise, the framework implements the function by exploiting a more conventional FSM approach, which is optimized for ILP exploitation. We evaluate our framework on a set of parallel applications, and show substantial performance improvements (average speedup of 4.7) with limited area over- heads (average area increase of 5.48 times).

Castellana, Vito G.↗