Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 649 records · Page 36

Massively parallel modeling and inversion of electrical resistivity tomography data using PFLOTRAN

Abstract. Electrical resistivity tomography (ERT) is a broadly accepted geophysical method for subsurface investigations. Interpretation of field ERT data usually requires the application of computationally intensive forward modeling and inversion algorithms. For large-scale ERT data, the efficiency of these algorithms depends on the robustness, accuracy, and scalability on high-performance computing resources. In this regard, we present a robust and highly scalable implementation of forward modeling and inversion algorithms for ERT data. The implementation is publicly available and developed within the framework of PFLOTRAN, an open-source, state-of-the-art massively parallel subsurface flow and transport simulation code. The forward modeling is based on a finite-volume discretization of the governing differential equations, and the inversion uses a Gauss–Newton optimization scheme. To evaluate the accuracy of the forward modeling, two examples are first presented by considering layered (1D) and 3D earth conductivity models. The computed numerical results show good agreement with the analytical solutions for the layered earth model and results from a well-established code for the 3D model. Inversion of ERT data, simulated for a 3D model, is then performed to demonstrate the inversion capability by recovering the conductivity of the model. To demonstrate the parallel performance of PFLOTRAN's ERT process model and inversion capabilities, large-scale scalability tests are performed by using up to 131 072 processes on a leadership class supercomputer. These tests are performed for the two most computationally intensive steps of the ERT inversion: forward modeling and Jacobian computation. For the forward modeling, we consider models with up to 122 ×106 degrees of freedom (DOFs) in the resulting system of linear equations and demonstrate that the code exhibits almost linear scalability on up to 10 000 DOFs per process. On the other hand, the code shows superlinear scalability for the Jacobian computation, mainly because all computations are fairly evenly distributed over each process with no parallel communication.

58 GEOSCIENCES↗

Scaling theory of three-dimensional magnetic reconnection spreading

We develop a first-principles scaling theory of the spreading of three-dimensional (3D) magnetic reconnection of finite extent in the out of plane direction. This theory addresses systems with or without an out of plane (guide) magnetic field, and with or without Hall physics. The theory reproduces known spreading speeds and directions with and without guide fields, unifying previous knowledge in a single theory. New results include: (1) Reconnection spreads in a particular direction if an x-line is induced at the interface between reconnecting and non-reconnecting regions, which is controlled by the out of plane gradient of the electric field in the outflow direction. (2) The spreading mechanism for anti-parallel collisionless reconnection is convection, as is known, but for guide field reconnection it is magnetic field bending. We confirm the theory using 3D two-fluid and resistive-magnetohydrodynamics simulations. (3) The theory explains why anti-parallel reconnection in resistive-magnetohydrodynamics does not spread. (4) The simulation domain aspect ratio, associated with the free magnetic energy, influences whether reconnection spreads or convects with a fixed x-line length. (5) We perform a simulation initiating anti-parallel collisionless reconnection with a pressure pulse instead of a magnetic perturbation, finding spreading is unchanged rather than spreading at the magnetosonic speed as previously suggested. The results provide a theoretical framework for understanding spreading beyond systems studied here, and are important for applications including two-ribbon solar flares and reconnection in Earth’s magnetosphere.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Ion-beam-driven electrostatic ion cyclotron instabilities

Results are presented of a particle simulation study of the electrostatic ion-cyclotron (EIC) instability driven by a parallel ion beam. The results of this simulation study demonstrate the nonlinear consequences of nonresonant EIC waves destabilized by an ion beam parallel to the magnetic field for the case of a large beam velocity. As a consequence of the instability, it is shown that the beam ions are heated strongly in the perpendicular direction and suffer a strong anomalous friction via EIC waves which leads to the beam slowing down. Simulation results indicate that the anomalous slowing down of beam ions by EIC waves is much larger than that from the classical electron-drag, and the perpendicular collision frequency measured from perpendicular beam heating is as large as that from Bohm diffusion. It is concluded that the ion beam driven EIC wave is a viable mechanism for the transfer of ion parallel beam energy to the ion perpendicular energy.

Okuda, H.↗

Asynchronous-many-task systems: Challenges and opportunities - Scaling an AMR astrophysics code on exascale machines using Kokkos and HPX

Dynamic and adaptive mesh refinement is pivotal in high-resolution, multi-physics, multi-model simulations, necessitating precise physics resolution in localized areas across expansive domains. Today’s supercomputers’ extreme heterogeneity presents a significant challenge for dynamically adaptive codes, highlighting the importance of achieving performance portability at scale. Our research focuses on astrophysical simulations, particularly stellar mergers, to elucidate early universe dynamics. Here, we present Octo-Tiger, leveraging Kokkos, HPX, and SIMD for portable performance at scale in complex, massively parallel adaptive multi-physics simulations. Octo-Tiger supports diverse processors, accelerators, and network backends. Experiments demonstrate exceptional scalability across several heterogeneous supercomputers including Perlmutter, Frontier, and Fugaku, encompassing major GPU architectures and x86, ARM, and RISC-V CPUs. Parallel efficiency of 47.59% (110,080 cores and 6880 hybrid A100 GPUs) on a full-system run on Perlmutter (26% HPCG peak performance) and 51.37% (using 32,768 cores and 2048 MI250X) on Frontier are achieved.

97 MATHEMATICS AND COMPUTING↗

Cyclic behavior at quasi-parallel collisionless shocks

Large scale one-dimensional hybrid simulations with resistive electrons have been carried out of a quasi-parallel high-Mach-number collisionless shock. The shock initially appears stable, but then exhibits cyclic behavior. For the magnetic field, the cycle consists of a period when the transition from upstream to downstream is steep and well defined, followed by a period when the shock transition is extended and perturbed. This cyclic shock solution results from upstream perturbations caused by backstreaming gyrating ions convecting into the shock. The cyclic reformation of a sharp shock transition can allow ions, at one time upstream because of reflection or leakage, to contribute to the shock thermalization.

Burgess, D.↗

Classic and Quantum Task-Based Intelligent Runtime for QIRs Running on Multiple QPUs

High-performance computing systems are rapidly evolving into heterogeneous platforms that fuse quantum accelerators with traditional classical processing units (CPUs) and graphical processing units (GPUs). This convergence calls for runtimes capable of managing both classical and quantum workloads in a unified manner. We introduce an intelligent, task-based runtime that marries the Intelligent RuntIme System (IRIS) asynchronous scheduler with a quantum programming stack through the Quantum Intermediate Representation Execution Engine (QIR-EE). Our design allows programs written in the quantum intermediate representation (QIR) to be dispatched concurrently to a variety of back-ends, including multiple quantum simulators and nascent quantum processors, enabling genuine hybrid execution on a single node. To illustrate its practicality, we partition a 4-qubit and 20-qubit circuit into three sub-circuits using quantum circuit cutting via the QCut library. Each sub-circuit is simulated independently by the QIR-EE driver within IRIS, after which a classical post-processing step merges the simulation results to recover the outcome of the original full-circuit computation. This case study demonstrates how finer task granularity can enable the parallel execution and lower the simulation burden per quantum task while preserving overall accuracy, highlighting the feasibility of our hybrid approach.

Miniskar, Narasinga Rao [ORNL] (ORCID:000000018259↗

Less can be more: Insights on the role of electrode microstructure in redox flow batteries from two-dimensional direct numerical simulations

Understanding how to structure a porous electrode to facilitate fluid, mass, and charge transport is key to enhancing the performance of electrochemical devices, such as fuel cells, electrolyzers, and redox flow batteries (RFBs). Here, using a parallel computational framework, direct numerical simulations are carried out on idealized porous electrode microstructures for RFBs. Strategies to improve an electrode design starting from a regular lattice are explored. By introducing vacancies in the ordered arrangement, it is possible to achieve higher voltage efficiency at a given current density, thanks to improved mixing of reactive species, despite reducing the total reactive surface. Careful engineering of the location of vacancies, resulting in a density gradient, outperforms disordered configurations. Our simulation framework is a new tool to explore transport phenomena in RFBs, and our findings suggest new ways to design performant electrodes.

25 ENERGY STORAGE↗

A Versatile Simulated Data Transport Layer for in Situ Workflows Performance Evaluation

In situ processing does not only allow scientific applications to face the explosion in data volume and velocity but also to address the time constraints of many simulation-analysis workflows by providing scientists with early insights about their applications at runtime. Multiple frameworks implement the concept of a data transport layer (DTL) to enable such in situ workflows. These tools are very versatile, directly or indirectly access the data generated on the same node, another node of the same compute cluster, or a completely distinct node, and allow data publishers and subscribers to run on the same computing resources or not. This versatility puts on researchers the onus of taking key decisions related to resource allocation and how to transport data to ensure the most efficient execution of their in situ workflows. However, domain scientists and workflow practitioners lack the appropriate tools to assess the respective performance of particular design and deployment options. In this paper we introduce a versatile simulated DTL designed to provide researchers with insights on the respective performance of different execution scenarios of in situ workflows. This open-source, standalone library builds on the SimGrid toolkit and can be linked to any SimGrid-based simulator. It facilitates the evaluation of the performance behavior, at scale, of different data transport configurations and the study of the effects of resource allocation strategies. We demonstrate the scalability, versatility, and accuracy of this simulated DTL by reproducing the execution of two synthetic benchmarks and of a real-world in situ workflow composed of an MPI application and a parallel data analysis. Results of simulations run on a single core show that the proposed library can simulate the interactions of tens of thousands of simulated processes deployed on two interconnected commodity clusters in a few seconds, and the execution by a thousand simulated processes of an in situ workflow in less than three minutes.

Suter, Fred [ORNL] (ORCID:0000000319021955)↗

A multiarchitecture parallel-processing development environment

A description is given of the hardware and software of a multiprocessor test bed - the second generation Hypercluster system. The Hypercluster architecture consists of a standard hypercube distributed-memory topology, with multiprocessor shared-memory nodes. By using standard, off-the-shelf hardware, the system can be upgraded to use rapidly improving computer technology. The Hypercluster's multiarchitecture nature makes it suitable for researching parallel algorithms in computational field simulation applications (e.g., computational fluid dynamics). The dedicated test-bed environment of the Hypercluster and its custom-built software allows experiments with various parallel-processing concepts such as message passing algorithms, debugging tools, and computational 'steering'. Such research would be difficult, if not impossible, to achieve on shared, commercial systems.

Townsend, Scott↗

Near-specular reflection of ions at quasi-parallel shocks

One-dimensional hybrid simulations and a semianalytical model of the shock front are employed to investigate the source regions in the incident ion phase space of reflected and transmitted ions, the relative importance of the electric and magnetic forces in the reflection of incident ions, and how these characteristics change in time. The phase space origin of reflected particles and the fraction of incident ions that reflect are found to depend on the electromagnetic field structure of the shock front at the time the ions encounter it. The reflection fraction is maximized when the electric field along the shock normal (Ex) and the noncoplanar magnetic field (By) are at their maximum values. When Ex and By are large, the reflection process produces a beam that is cooler, more dense, and closer to specular than when they are small.

Mckean, M. E.↗

Interagency Agreement No. DE-SC0006988 (Final Technical Report)

The overarching goal of this project was to advance understanding of deep convection updraft microphysics—the source of long-lived stratiform ice—by combining detailed mining of in situ and remote sensing data from the MC3E field campaign with detailed 3D simulations. The project began with parallel work, first strictly on the remote-sensing observation side via dedicated analysis of polarimetric radar signatures, which is a relatively new area, on the one hand. On the other hand, a more traditional but well-formulated preparation and comparison of detailed aerosol-aware simulations with in situ measurements was prepared. In the final phase of this work, these parallel elements were brought together into an integrated analysis of in situ and remote-sensing observations with model results. The project also supported team member involvement in collaborative activities.

54 ENVIRONMENTAL SCIENCES↗

Research in Computational Astrobiology

We present results from several projects in the new field of computational astrobiology, which is devoted to advancing our understanding of the origin, evolution and distribution of life in the Universe using theoretical and computational tools. We have developed a procedure for calculating long-range effects in molecular dynamics using a plane wave expansion of the electrostatic potential. This method is expected to be highly efficient for simulating biological systems on massively parallel supercomputers. We have perform genomics analysis on a family of actin binding proteins. We have performed quantum mechanical calculations on carbon nanotubes and nucleic acids, which simulations will allow us to investigate possible sources of organic material on the early earth. Finally, we have developed a model of protobiological chemistry using neural networks.

Chaban, Galina↗

Kinetic study of shock formation and particle acceleration in laser-driven quasi-parallel magnetized collisionless shocks

Quasi-parallel magnetized collisionless shocks are believed to be one of the most efficient accelerators in the universe. Compared to quasi-perpendicular shocks, quasi-parallel shocks are more difficult to form in the laboratory and to simulate because of their large spatial scales and long formation times. Our two-dimensional particle-in-cell simulations show that the early stages of quasi-parallel shock formation are achievable in experiments planned for the National Ignition Facility and that particles accelerated by diffusive shock acceleration (DSA) are expected to be observable in the experiment. Repetitive ion acceleration by crossings of the shock front, a key feature of DSA, is seen in the simulations. Other characteristic features of quasi-parallel shocks such as upstream wave excitation by energetic ions are also observed, and energy partition between the ions and the electrons in the downstream of the shock is briefly discussed.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Particle acceleration magnetic field generation, and emission in Relativistic pair jets

Plasma waves and their associated instabilities (e.g., the Buneman instability, two-streaming instability, and the Weibel instability) are responsible for particle acceleration in relativistic pair jets. Using a 3-D relativistic electromagnetic particle (REMP) code, we have investigated particle acceleration associated with a relativistic pair jet propagating through a pair plasma. Simulations show that the Weibel instability created in the collisionless shock accelerates particles perpendicular and parallel to the jet propagation direction. Simulation results show that this instability generates and amplifies highly nonuniform, small-scale magnetic fields, which contribute to the electron's transverse deflection behind the jet head. The "jitter' I radiation from deflected electrons can have different properties than synchrotron radiation which is calculated in a uniform magnetic field. This jitter radiation may be important to understanding the complex time evolution and/or spectral structure in gamma-ray bursts, relativistic jets, and supernova remnants. The growth rate of the Weibel instability and the resulting particle acceleration depend on the magnetic field strength and orientation, and on the initial particle distribution function. In this presentation we explore some of the dependencies of the Weibel instability and resulting particle acceleration on the magnetic field strength and orientation, and the particle distribution function.

Nishikawa, K.-I.↗

Cyber-Power Co-Simulation for End-to-End Synchrophasor Network Analysis and Applications

The resiliency, reliability and security of the next generation cyber-power smart grid depend upon efficiently leveraging advanced communication and computing technologies. Also, developing real-time data-driven applications is critical to enable wide-area monitoring and control of the cyber-power grid given high-resolution data from Phasor Measurement Units (PMUs). North American Synchrophasor Initiative Network (NASPlnet) provides guidance for PMU data exchanges. With the advancement in networking and grid operation, it is necessary to evaluate the performance of different data flow architectures suggested by NASPInet and analyze the impact on applications. Therefore, we need a cyber-power co-simulation framework that supports very large-scale co-simulation capable of running in parallel, high-performance computing platforms and capturing real-life network behavior. This work presents an end-to-end automated and user-driven cyber-power co-simulation using NS3 to model communication networks, GridPACK to model the power grid, and HELICS as a co-simulation engine. Comparative analysis of latency in synchrophasor networks and a performance evaluation of a power system stabilizer application utilizing PMU data in an IEEE 39 bus test system is presented using this cosimulation testbed.

Mustafa, Hussain M.↗

To Exascale and Beyond—The Simple Cloud-Resolving E3SM Atmosphere Model (SCREAM), a Performance Portable Global Atmosphere Model for Cloud-Resolving Scales

The new generation of heterogeneous CPU/GPU computer systems offer much greater computational performance but are not yet widely used for climate modeling. One reason for this is that traditional climate models were written before GPUs were available and would require an extensive overhaul to run on these new machines. In addition, even conventional “high–resolution” simulations don't currently provide enough parallel work to keep GPUs busy, so the benefits of such overhaul would be limited for the types of simulations climate scientists are accustomed to. The vision of the Simple Cloud-Resolving Energy Exascale Earth System (E3SM) Atmosphere Model (SCREAM) project is to create a global atmospheric model with the architecture to efficiently use GPUs and horizontal resolution sufficient to fully take advantage of GPU parallelism. After 5 years of model development, SCREAM is finally ready for use. In this paper, we describe the design of this new code, its performance on both CPU and heterogeneous machines, and its ability to simulate real-world climate via a set of four 40 day simulations covering all 4 seasons of the year.

54 ENVIRONMENTAL SCIENCES↗

Parallel Programming in MCNP6

Monte Carlo N-Particle (MCNP)1 is a general-purpose Monte Carlo particle transport code developed by Los Alamos National Laboratory (LANL). To efficiently handle long simulations, MCNP version 6 (MCNP6) supports parallel execution using two primary programming models: • Shared-memory task-based threading using OpenMP (Open Multi-Processing), and • Distributed-memory calculations using MPI (Message Passing Interface). The OpenMP and MPI programming models enable MCNP6 to scale from desktop systems to high-performance computing (HPC) clusters, allowing users to run MCNP in one of three parallel modes: • OpenMP-only, • MPI-only, and • Hybrid (MPI + OpenMP). The choice of parallelization mode depends on the underlying computer architecture and the characteristics of the simulation problem.

97 MATHEMATICS AND COMPUTING↗