Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,549 records · Page 86

STORM: Scrape-off layer turbulence in tokamak fusion reactors

The scrape-off layer of a tokamak fusion reactor carries the plasma exhaust from the hot core plasma to the material surfaces of the reactor vessel. The heat loads imposed by the exhaust are a critical limit on the performance of fusion power plants. Turbulent transport of the plasma regulates the width of the scrape-off layer plasma and must be modelled to understand the intensity of these heat loads. STORM is a plasma turbulence code capable of simulating three dimensional turbulence across the full scrape-off layer of a tokamak fusion reactor, using a drift reduced, collisional fluid model. STORM uses mostly finite difference schemes, with a staggered grid in the direction parallel to the magnetic field. We describe the model, geometry and initialisation options used by STORM, as well as the numerical methods, which are implemented using the BOUT++ plasma simulation framework. BOUT++ has been enhanced alongside the development of STORM, providing better support for staggered grid methods. We summarise these enhancements, including a detailed explanation of the parallel derivative methods, which underwent a major update for version 4 of BOUT++.

BOUT++↗

Simulating ‘Two Ribbon’ Type Solar Flares [Slides]

The underlying mechanism in solar flares in a process known as magnetic reconnection. This astrophysical process is when magnetic field lines, running anti-parallel, break and reconnect. The change in configuration of field lines results in an explosive release of energy as the leftover magnetic energy is converted to kinetic and thermal energies. In this project, we are using an astrophysical magnetohydrodynamic (MHD) simulation code known as Athena++ with a reconnection problem generator file created by Li et al. 2018. Our current research is to successfully implement a radiative cooling term into the MHD equations as it could have important physical effects on the plasma. We use Klimchuk et al. 2008 and SPEX_DM as our chosen cooling functions. The results from the SPEX_DM function show there is a condensation forming that we had predicted but had not seen before with our simulation code. We will continue to analyze the SPEX_DM function and potentially implement thermal conduction.

79 ASTRONOMY AND ASTROPHYSICS↗

A GPU ‐Accelerated 3D Unstructured Mesh Based Particle Tracking Code for Multi‐Species Impurity Transport Simulation in Fusion Tokamaks

ABSTRACT This paper presents the multi‐species global impurity transport capability developed in a GPU‐accelerated fully 3D unstructured mesh‐based code, GITRm, to simultaneously track multiple impurity species and handle interactions of these impurities with mixed‐material surfaces. Different computational approaches to model particle‐surface interaction or surface response have been developed and compared. Sheath electric field is taken into account by employing a fast distance‐to‐boundary calculation, which is carried out in parallel on distributed or partitioned meshes on multiple GPUs without the need for any inter‐process communication during the simulation. Several example cases, including two for the DIII‐D tokamak, that is, one with the SAS‐V divertor and the other with the collector probes, are used to demonstrate the utility of the current multi‐species capability. For the DIII‐D probe case, the capability of GITRm to resolve the spatial distribution of particles in localized regions, such as diagnostic probes, within non‐axisymmetric tokamak geometries is demonstrated. These simulations involve up to 320 million particles and utilize up to 48 GPUs.

Nath, Dhyanjyoti D. [Scientific Computation Resear↗

Implementation of a Mesh refinement algorithm into the quasi-static PIC code QuickPIC

Plasma-based acceleration (PBA) has emerged as a promising candidate for the accelerator technology used to build a future linear collider and/or an advanced light source. In PBA, a trailing or witness particle beam is accelerated in the plasma wave wakefield (WF) created by a laser or particle beam driver. The WF is often nonlinear and involves the crossing of plasma particle trajectories in real space and thus particle-in-cell methods are used. The distance over which the drive beam evolves is several orders of magnitude larger than the wake wavelength. This large disparity in length scales is amenable to the quasi-static approach. Three-dimensional (3D), quasi-static (QS), particle-in-cell (PIC) codes, e.g., QuickPIC, have been shown to provide high fidelity simulation capability with 2-4 orders of magnitude speedup over 3D fully explicit PIC codes. In PBA, the witness beam needs to be matched to the focusing forces of the WF to reduce the emittance growth. In some linear collider designs, the matched spot size of the witness beam can be 2 to 3 orders of magnitude smaller than the spot size (and wavelength) of the wakefield. Such an additional disparity in length scales is ideal for mesh refinement where the WF within the witness beam is described on a finer mesh than the rest of the WF. A mesh refinement scheme is described that has been implemented into the 3D QS PIC code, QuickPIC. Very fine (high) resolution is used in a small spatial region that includes the witness beam and progressively coarser resolutions in the rest of the simulation domain. A fast multigrid Poisson solver has been implemented for the field solve on the refined meshes and a Fast Fourier Transform (FFT) based Poisson solver is used for the coarse mesh. The code has been parallelized with both MPI and OpenMP, and the parallel scalability has also been improved by using pipelining. A preliminary adaptive mesh refinement technique is described to optimize the computational time for simulations with an evolving witness beam size. Several test problems are used to verify that the mesh refinement algorithm provides accurate results. Additionally, the results are benchmarked against highly resolved simulations exhibiting near-azimuthal symmetry, performed using QPAD—a novel hybrid QS PIC code that uses a PIC description in the coordinates (r, ct – z) and a gridless description in the azimuthal angle, Φ.

Linear collider↗

Scalable parallel communications

Coarse-grain parallelism in networking (that is, the use of multiple protocol processors running replicated software sending over several physical channels) can be used to provide gigabit communications for a single application. Since parallel network performance is highly dependent on real issues such as hardware properties (e.g., memory speeds and cache hit rates), operating system overhead (e.g., interrupt handling), and protocol performance (e.g., effect of timeouts), we have performed detailed simulations studies of both a bus-based multiprocessor workstation node (based on the Sun Galaxy MP multiprocessor) and a distributed-memory parallel computer node (based on the Touchstone DELTA) to evaluate the behavior of coarse-grain parallelism. Our results indicate: (1) coarse-grain parallelism can deliver multiple 100 Mbps with currently available hardware platforms and existing networking protocols (such as Transmission Control Protocol/Internet Protocol (TCP/IP) and parallel Fiber Distributed Data Interface (FDDI) rings); (2) scale-up is near linear in n, the number of protocol processors, and channels (for small n and up to a few hundred Mbps); and (3) since these results are based on existing hardware without specialized devices (except perhaps for some simple modifications of the FDDI boards), this is a low cost solution to providing multiple 100 Mbps on current machines. In addition, from both the performance analysis and the properties of these architectures, we conclude: (1) multiple processors providing identical services and the use of space division multiplexing for the physical channels can provide better reliability than monolithic approaches (it also provides graceful degradation and low-cost load balancing); (2) coarse-grain parallelism supports running several transport protocols in parallel to provide different types of service (for example, one TCP handles small messages for many users, other TCP's running in parallel provide high bandwidth service to a single application); and (3) coarse grain parallelism will be able to incorporate many future improvements from related work (e.g., reduced data movement, fast TCP, fine-grain parallelism) also with near linear speed-ups.

Maly, K.↗

Group implicit concurrent algorithms in nonlinear structural dynamics

During the 70's and 80's, considerable effort was devoted to developing efficient and reliable time stepping procedures for transient structural analysis. Mathematically, the equations governing this type of problems are generally stiff, i.e., they exhibit a wide spectrum in the linear range. The algorithms best suited to this type of applications are those which accurately integrate the low frequency content of the response without necessitating the resolution of the high frequency modes. This means that the algorithms must be unconditionally stable, which in turn rules out explicit integration. The most exciting possibility in the algorithms development area in recent years has been the advent of parallel computers with multiprocessing capabilities. So, this work is mainly concerned with the development of parallel algorithms in the area of structural dynamics. A primary objective is to devise unconditionally stable and accurate time stepping procedures which lend themselves to an efficient implementation in concurrent machines. Some features of the new computer architecture are summarized. A brief survey of current efforts in the area is presented. A new class of concurrent procedures, or Group Implicit algorithms is introduced and analyzed. The numerical simulation shows that GI algorithms hold considerable promise for application in coarse grain as well as medium grain parallel computers.

Ortiz, M.↗

Accelerating massively parallel hemodynamic models of coarctation of the aorta using neural networks

Comorbidities such as anemia or hypertension and physiological factors related to exertion can influence a patient’s hemodynamics and increase the severity of many cardiovascular diseases. Observing and quantifying associations between these factors and hemodynamics can be difficult due to the multitude of co-existing conditions and blood flow parameters in real patient data. Machine learning-driven, physics-based simulations provide a means to understand how potentially correlated conditions may affect a particular patient. Here, we use a combination of machine learning and massively parallel computing to predict the effects of physiological factors on hemodynamics in patients with coarctation of the aorta. We first validated blood flow simulations against in vitro measurements in 3D-printed phantoms representing the patient’s vasculature. We then investigated the effects of varying the degree of stenosis, blood flow rate, and viscosity on two diagnostic metrics – pressure gradient across the stenosis (ΔP) and wall shear stress (WSS) - by performing the largest simulation study to date of coarctation of the aorta (over 70 million compute hours). Using machine learning models trained on data from the simulations and validated on two independent datasets, we developed a framework to identify the minimal training set required to build a predictive model on a per-patient basis. We then used this model to accurately predict ΔP (mean absolute error within 1.18 mmHg) and WSS (mean absolute error within 0.99 Pa) for patients with this disease.

59 BASIC BIOLOGICAL SCIENCES↗

Digital system for structural dynamics simulation

State-of-the-art digital hardware and software for the simulation of complex structural dynamic interactions, such as those which occur in rotating structures (engine systems). System were incorporated in a designed to use an array of processors in which the computation for each physical subelement or functional subsystem would be assigned to a single specific processor in the simulator. These node processors are microprogrammed bit-slice microcomputers which function autonomously and can communicate with each other and a central control minicomputer over parallel digital lines. Inter-processor nearest neighbor communications busses pass the constants which represent physical constraints and boundary conditions. The node processors are connected to the six nearest neighbor node processors to simulate the actual physical interface of real substructures. Computer generated finite element mesh and force models can be developed with the aid of the central control minicomputer. The control computer also oversees the animation of a graphics display system, disk-based mass storage along with the individual processing elements.

Krauter, A. I.↗

Assessing Confidence in Pliocene Sea Surface Temperatures to Evaluate Predictive Models

In light of mounting empirical evidence that planetary warming is well underway, the climate research community looks to palaeoclimate research for a ground-truthing measure with which to test the accuracy of future climate simulations. Model experiments that attempt to simulate climates of the past serve to identify both similarities and differences between two climate states and, when compared with simulations run by other models and with geological data, to identify model-specific biases. Uncertainties associated with both the data and the models must be considered in such an exercise. The most recent period of sustained global warmth similar to what is projected for the near future occurred about 3.33.0 million years ago, during the Pliocene epoch. Here, we present Pliocene sea surface temperature data, newly characterized in terms of level of confidence, along with initial experimental results from four climate models. We conclude that, in terms of sea surface temperature, models are in good agreement with estimates of Pliocene sea surface temperature in most regions except the North Atlantic. Our analysis indicates that the discrepancy between the Pliocene proxy data and model simulations in the mid-latitudes of the North Atlantic, where models underestimate warming shown by our highest-confidence data, may provide a new perspective and insight into the predictive abilities of these models in simulating a past warm interval in Earth history.This is important because the Pliocene has a number of parallels to present predictions of late twenty-first century climate.

simulation↗

Computer simulation and design of a three degree-of-freedom shoulder module

An in-depth kinematic analysis of a three degree of freedom fully-parallel robotic shoulder module is presented. The major goal of the analysis is to determine appropriate link dimensions which will provide a maximized workspace along with desirable input to output velocity and torque amplification. First order kinematic influence coefficients which describe the output velocity properties in terms of actuator motions provide a means to determine suitable geometric dimensions for the device. Through the use of computer simulation, optimal or near optimal link dimensions based on predetermined design criteria are provided for two different structural designs of the mechanism. The first uses three rotational inputs to control the output motion. The second design involves the use of four inputs, actuating any three inputs for a given position of the output link. Alternative actuator placements are examined to determine the most effective approach to control the output motion.

Marco, David↗

Exploring model complexity in machine learned potentials for simulated properties

Abstract Machine learning (ML) enables the development of interatomic potentials with the accuracy of first principles methods while retaining the speed and parallel efficiency of empirical potentials. While ML potentials traditionally use atom-centered descriptors as inputs, different models such as linear regression and neural networks map descriptors to atomic energies and forces. This begs the question: what is the improvement in accuracy due to model complexity irrespective of descriptors? We curate three datasets to investigate this question in terms of ab initio energy and force errors: (1) solid and liquid silicon, (2) gallium nitride, and (3) the superionic conductor Li $$_{10}$$ 10 Ge(PS $$_{6}$$ 6 ) $$_{2}$$ 2 (LGPS). We further investigate how these errors affect simulated properties and verify if the improvement in fitting errors corresponds to measurable improvement in property prediction. By assessing different models, we observe correlations between fitting quantity (e.g. atomic force) error and simulated property error with respect to ab initio values. Graphical abstract

Rohskopf, A. (ORCID:0000000227128296)↗

Production of electron conics by stochastic acceleration parallel to the magnetic field

Electron conics are enhancements in the electron flux at the edges of the electron loss cone. Such enhancements are a common feature in the electron distribution in the auroral zone. In analogy with ion conics, it has been suggested that electron conics are produced by waves which accelerate electrons perpendicular to the magnetic field. However, using a test particle simulation of the electron distribution it is shown that electron conics can be produced purely by stochastic acceleration of the electrons parallel to a dipole magnetic field. A possible wave mode that can produce parallel acceleration is the Alfven-ion cyclotron mode that has recently been shown to modulate the high energy part of the inverted-V electron distribution.

Temerin, Michael A.↗

Fluid Dynamics Effects on Microstructure Prediction in Single-Laser Tracks for Additive Manufacturing of IN625

Single-track laser fusion were simulated using a heat-transfer-solidification-only (HTS) model and its extension with fluid dynamics (HTS_FD) model using a parallel open-source code, which included laminar fluid dynamics, flat-free surface of the molten alloy, heat transfer, phase-change, evaporation, and surface tension phenomena. The results illustrate that the fluid dynamics affects the solidification and ensuing microstructure. For the HTS_FD simulations, thermal gradient, G was found to exhibit a maximum at the extremity of the solidified pool ( i.e. , at the free surface), while for HTS simulations, G exhibited a maximum around the entire edge of the solidified pool. HTS_FD simulations predicted a wider range of cooling rates than the HTS simulations, exhibited an increased spread in the solidification speed, V variation within the melt-pool with respect to the HTS model results. Primary dendrite arm spacing (PDAS) were evaluated based on power law correlations and marginal stability theory models using the ( G , V ) from HTS and HTS_FD simulations to quantify the effect of the fluid dynamics on the microstructure. At low-laser powers and low-scan speeds, the PDAS obtained with the fluid dynamics model (HTS_FD) was larger by more than 30 pct with respect to the PDAS calculated with the simple HTS model. A new PDAS correlation, i.e. , \( \lambda_{1} \left[ {\mu {\text{m}}} \right] = 832\;G\left[ {\text{K/m}} \right]^{ - 0.5} V\left[ {\text{m/s}} \right]^{ - 0.25} \) , which uses the ( G , V ) results from the HTS_FD model was developed and validated against experimental results.

36 MATERIALS SCIENCE↗

Operational performance of a low cost, air mass 2 solar simulator

The present work describes briefly the design, construction, and operation of a low cost air mass 2 solar simulator, and then presents the performance characteristics of a modified version in terms of total irradiance, uniformity of irradiance, spectral distribution, and beam subtense angle. The simulator consists of an array of 143 tungsten halogen lamps and a corresponding array of 143 Fresnel lenses parallel and in front of the lenses, so that a 1.2 m by 1.2 m area is irradiated with uniform collimated irradiance. The performance of this 143-lamp version is compared with that of the 12-lamp prototype. It was found that the larger design required a lower lamp voltage for an equivalent average total irradiance in the test plane. However, distribution of total irradiance was not as good for the large simulator.

Yass, K.↗

Simulation of lithium transport using the BOUT++ framework

A numerical model that calculates the collisional interactions between the lithium atoms from a lithium pellet and the background plasmas has been upgraded. The ion density ($N_t$), electron temperature ($T_e$), ion temperature ($T_i$) and parallel ion velocity ($V_{∥, i}$) are used to characterize the background plasmas. The lithium atom density ($N^{a}_{Li}$) and parallel velocity ($V_{∥,a}$) of lithium atoms evolve with time. For each lithium ion, the density ($N_{Li^{n+}}$), temperature ($T_{Li^{n+}}$) and parallel velocity ($V_{∥, Li^{n+}}$) are self-consistently calculated. A C-mod lower single null equilibrium is used to generate the grid for the BOUT++ simulation. The lithium atoms can be fully ionized to $Li^{3+}$ in ~2 μs. The rapid radial and poloidal expansion of the lithium ions are found in the simulation. After the collision interaction process, the electron temperature rapidly decreases at the pellet location; then, it rapidly poloidally expands, and the temperature at the pellet location starts to recover. The electron pressure increases at the pellet location despite the decrease in electron temperature because of the extra electrons from the lithium ionization. The ion pressure profile decreases in the pellet location due to the decrease in ion temperature.

74 ATOMIC AND MOLECULAR PHYSICS↗

Performance and Accuracy Implications of Parallel Split Physics-Dynamics Coupling in the Energy Exascale Earth System Atmosphere Model

Simultaneous calculation of atmospheric processes is faster than calculating processes one at a time. This type of parallelism is beneficial or perhaps even necessary to provide good performance on modern supercomputers, which achieve faster performance through increased processor count rather than improved clock speed. The scalability of the Energy Exascale Earth System Model (E3SM) Atmosphere Model (EAM) is limited by the fluid dynamics which scales up to the number of mesh cells in the global mesh. In contrast, the suite of physics parameterizations in EAM is scalable up to the total number of physics columns, which is an order of magnitude greater than the number of mesh cells. A proposed solution to unlocking the greater potential performance from the physics suite is to solve the physics and dynamics in parallel. This work represents a first attempt at parallel splitting of the grid-scale fluid dynamics model and the subgrid-scale physics parameterizations in a global atmosphere model. We will demonstrate that switching to parallel physics-dynamics coupling extends the scalability of the EAM to up to 3 times the previous peak scalability limit and is up to 20% faster than the sequentially split coupling at the highest core counts and the same time step. Decadal simulations of both coupling approaches show very little impact to the model climate. This improved performance does not come without drawbacks, however. Parallel splitting requires a shorter time step and other modifications which largely offset performance gains. A mass fixer is required for conservation. Techniques for mitigating these issues are also discussed.

97 MATHEMATICS AND COMPUTING↗

Potential Flow Interactions With Directional Solidification

The effect of convective melt motion on the growth of morphological instabilities in crystal growth has been the focus of many studies in the past decade. While most of the efforts have been directed towards investigating the linear stability aspects, relatively little attention has been devoted to experimental and numerical studies. In a pure morphological case, when there is no flow, morphological changes in the solid-liquid interface are governed by heat conduction and solute distribution. Under the influence of a convective motion, both heat and solute are redistributed, thereby affecting the intrinsic morphological phenomenon. The overall effect of the convective motion could be either stabilizing or destabilizing. Recent investigations have predicted stabilization by a flow parallel to the interface. In the case of non-parallel flows, e.g., stagnation point flow, Brattkus and Davis have found a new flow-induced morphological instability that occurs at long wavelengths and also consists of waves propagating against the flow. Other studies have addressed the nonlinear aspects (Konstantinos and Brown, Wollkind and Segel)). In contrast to the earlier studies, our present investigation focuses on the effects of the potential flow fields typically encountered in Hele-Shaw cells. Such a Hele-Shaw cell can simulate a gravity-free environment in the sense that buoyancy-driven convection is largely suppressed, and hence negligible. Our interest lies both in analyzing the linear stability of the solidification process in the presence of potential flow fields, as well as in performing high-accuracy nonlinear simulations. Linear stability analysis can be performed for the flow configuration mentioned above. It is observed that a parallel potential flow is stabilizing and gives rise to waves traveling downstream. We have built a highly accurate numerical scheme which is validated at small amplitudes by comparing with the analytically predicted results for the pure morphological case. We have been able to observe nonlinear effects at larger times. Preliminary results for the case when flow is imposed also provide good validation at small amplitudes.

Buddhavarapu, Sudhir S.↗

A parallel p ‐adaptive discontinuous Galerkin method for the Euler equations with dynamic load‐balancing on tetrahedral grids

Abstract A novel p ‐adaptive discontinuous Galerkin (DG) method has been developed to solve the Euler equations on three‐dimensional tetrahedral grids. Hierarchical orthogonal basis functions are adopted for the DG spatial discretization while a third order TVD Runge‐Kutta method is used for the time integration. A vertex‐based limiter is applied to the numerical solution in order to eliminate oscillations in the high order method. An error indicator constructed from the solution of order and is used to adapt degrees of freedom in each computational element, which remarkably reduces the computational cost while still maintaining an accurate solution. The developed method is implemented with under the Charm++ parallel computing framework. Charm++ is a parallel computing framework that includes various load‐balancing strategies. Implementing the numerical solver under Charm++ system provides us with access to a suite of dynamic load balancing strategies. This can be efficiently used to alleviate the load imbalances created by p ‐adaptation. A number of numerical experiments are performed to demonstrate both the numerical accuracy and parallel performance of the developed p ‐adaptive DG method. It is observed that the unbalanced load distribution caused by the parallel p ‐adaptive DG method can be alleviated by the dynamic load balancing from Charm++ system. Due to this, high performance gain can be achieved. For the testcases studied in the current work, the parallel performance gain ranged from 1.5× to 3.7×. Therefore, the developed p ‐adaptive DG method can significantly reduce the total simulation time in comparison to the standard DG method without p ‐adaptation.

97 MATHEMATICS AND COMPUTING↗