Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel simulation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 703 records · Page 39

Parallel hybrid quantum-classical machine learning for kernelized time-series classification

Supervised time-series classification garners widespread interest because of its applicability throughout a broad application domain including finance, astronomy, biosensors, and many others. Here, in this work, we tackle this problem with hybrid quantum-classical machine learning, deducing pairwise temporal relationships between time-series instances using a timeseries Hamiltonian kernel (TSHK). A TSHK is constructed with a sum of inner products generated by quantum states evolved using a parameterized time evolution operator. This sum is then optimally weighted using techniques derived from multiple kernel learning. Because we treat the kernel weighting step as a differentiable convex optimization problem, our method can be regarded as an end-to-end learnable hybrid quantum-classical-convex neural network, or QCC-net, whose output is a data set-generalized kernel function suitable for use in any kernelized machine learning technique such as the support vector machine (SVM). Using our TSHK as input to a SVM, we classify univariate and multivariate time-series using quantum circuit simulators and demonstrate the efficient parallel deployment of the algorithm to 127-qubit superconducting quantum processors using quantum multi-programming.

97 MATHEMATICS AND COMPUTING↗

First-Principles Grand-Canonical Simulations of Water Adsorption in Proton-Exchanged Zeolites

Water appears by design or as impurities in many important reactive systems. For those catalyzed by porous solid acids, such as widely used zeolites, experimentally quantifying the amount or elucidating the structure of adsorbed water clusters at reaction conditions is challenging, while computational studies (e.g., first-principles molecular dynamics simulations) often need to assume the loading to examine solvation effects. Furthermore, we perform first-principles grand-canonical simulations to predict water adsorption to H-ZSM-5 zeolites under specified experimental conditions. Presampling with inexpensive force fields and a pool-based parallelization algorithm are used to improve simulation efficiency, while molecular dynamics is used to sample configurations involving hydronium species. We observe that H + exchange dramatically increases the hydrophilicity of zeolite MFI and an appreciable amount of water is present at very low relative humidities. At all conditions examined, the zeolitic protons are found to dissociate readily from the surface basic sites and become mobile by participating in the hydrogen-bonded chains of adsorbed water molecules.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Particle Acceleration in Relativistic Jets due to Weibel Instability

Shock acceleration is an ubiquitous phenomenon in astrophysical plasmas. Plasma waves and their associated instabilities (e.g., the Buneman instability, two-streaming instability, and the Weibel instability) created in the shocks are responsible for particle (electron, positron, and ion) acceleration. Using a 3-D relativistic electromagnetic particle (REMP) code, we have investigated particle acceleration associated with a relativistic jet front propagating through an ambient plasma with and without initial magnetic fields. We find only small differences in the results between no ambient and weak ambient magnetic fields. Simulations show that the Weibel instability created in the collisionless shock front accelerates particles perpendicular and parallel to the jet propagation direction. The simulation results show that this instability is responsible for generating and amplifying highly nonuniform, small-scale magnetic fields, which contribute to the electron s transverse deflection behind the jet head. The jitter radiation (Medvedev 2000) from deflected electrons has different properties than synchrotron radiation which is calculated in a uniform magnetic field. This jitter radiation may be important to understanding the complex time evolution and/or spectral structure in gamma-ray bursts, relativistic jets, and supernova remnants.

Nishikawa, K.↗

Implementation of Active Sites in DSMC to Capture Pitting of Oxidizing Carbon Materials

In this work we demonstrate a newly developed capability to capture pitting of carbon fibers in DSMC simulations, specifically using the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code. State-of-the-art reactive surface models in DSMC compute collision dependent carbon consumption rates (usually through desorption of CO) based on a set of surface reactions that has been derived from molecular beam experiments. The reactivity on each carbon surface element is constant in those models, such that the carbon surface recedes uniformly as a result of ablation. However, it is well known that in reality the carbon surface has locally different reaction rates due to the presence of defects at the atomic scale. These defective sites have a much higher reactivity than the average sites (2-3 orders of magnitude) and are first to react during ablation leading to its removal. This causes all the neighboring atoms to be defective and increase their reactivity, thus leading to the localized carbon removal around these ”active” sites. In this manner, these highly reactive defective sites serve as nucleation sites for the formation and growth of etch pits with potentially detrimental effects on the structural integrity. Recently a detailed surface chemistry framework was developed in SPARTA, capable of incorporating various reaction mechanisms such as adsorption, desorption, Eley-Rideal (ER) and Langmuir-Hinshelwood (LH) mechanisms. Within this framework, we have implemented the capability of a single surface having multiple site sets with different reactivities. Using this feature, we can simulate the presence of active sites on carbon surfaces, whose reactivity is much greater than an average site as a result of defects. We have implemented the active site fraction as a property of surface elements within SPARTA, which is directly proportional to the local reactivity of each surface element. By introducing an initial distribution of the active site fraction across the carbon surface, and propagating it in a manner that mimics the evolution of real reacting carbon surfaces, we are able to capture the formation and growth of etch pits as a result of surface consumption reactions such as oxidation.

DSMC↗

Implementation of active sites to capture pitting of oxidizing carbon materials in DSMC.

In this work we demonstrate a newly developed capability to capture pitting of carbon fibers in DSMC simulations, specifically using the Stochastic PArallel Rarefied-gas Time-accurate Analyzer (SPARTA) code [1]. State-of-the-art reactive surface models in DSMC compute collision dependent carbon consumption rates (usually through desorption of CO) based on a set of surface reactions that has been derived from molecular beam experiments [2]. The reactivity on each carbon surface element is constant in those models, such that the carbon surface recedes uniformly as a result of ablation. However, it is well known that in reality the carbon surface has locally different reaction rates due to the presence of defects at the atomic scale [3]. These defective sites have a much higher reactivity than the average sites (2-3 orders of magnitude) and are first to react during ablation leading to its removal. This causes all the neighboring atoms to be defective and increase their reactivity, thus leading to the localized carbon removal around these ”active” sites (as shown in Fig. 1). In this manner, these highly reactive defective sites serve as nucleation sites for the formation and growth of etch pits with potentially detrimental effects on the structural integrity. Recently a detailed surface chemistry framework was developed in SPARTA, capable of incorporating various reaction mechanisms such as adsorption, desorption, Eley-Rideal (ER) and Langmuir-Hinshelwood (LH) mechanisms [4]. Within this framework, we have implemented the capability of a single surface having multiple site sets with different reactivities. Using this feature, we can simulate the presence of active sites on carbon surfaces, whose reactivity is much greater than an average site as a result of defects. We have implemented the active site fraction as a property of surface elements within SPARTA, which is directly proportional to the local reactivity of each surface element. By introducing an initial distribution of the active site fraction across the carbon surface, and propagating it in a manner that mimics the evolution of real reacting carbon surfaces, we are able to capture the formation and growth of etch pits as a result of surface consumption reactions such as oxidation.

DSMC↗

Implementation and Characterization of Three-Dimensional Particle-in-Cell Codes on Multiple-Instruction-Multiple-Data Massively Parallel Supercomputers

A three-dimensional electrostatic particle-in-cell (PIC) plasma simulation code has been developed on coarse-grain distributed-memory massively parallel computers with message passing communications. Our implementation is the generalization to three-dimensions of the general concurrent particle-in-cell (GCPIC) algorithm. In the GCPIC algorithm, the particle computation is divided among the processors using a domain decomposition of the simulation domain. In a three-dimensional simulation, the domain can be partitioned into one-, two-, or three-dimensional subdomains ("slabs," "rods," or "cubes") and we investigate the efficiency of the parallel implementation of the push for all three choices. The present implementation runs on the Intel Touchstone Delta machine at Caltech; a multiple-instruction-multiple-data (MIMD) parallel computer with 512 nodes. We find that the parallel efficiency of the push is very high, with the ratio of communication to computation time in the range 0.3%-10.0%. The highest efficiency (> 99%) occurs for a large, scaled problem with 64(sup 3) particles per processing node (approximately 134 million particles of 512 nodes) which has a push time of about 250 ns per particle per time step. We have also developed expressions for the timing of the code which are a function of both code parameters (number of grid points, particles, etc.) and machine-dependent parameters (effective FLOP rate, and the effective interprocessor bandwidths for the communication of particles and grid points). These expressions can be used to estimate the performance of scaled problems--including those with inhomogeneous plasmas--to other parallel machines once the machine-dependent parameters are known.

Lyster, P. M.↗

Parallel machine architecture and compiler design facilities

The objective is to provide an integrated simulation environment for studying and evaluating various issues in designing parallel systems, including machine architectures, parallelizing compiler techniques, and parallel algorithms. The status of Delta project (which objective is to provide a facility to allow rapid prototyping of parallelized compilers that can target toward different machine architectures) is summarized. Included are the surveys of the program manipulation tools developed, the environmental software supporting Delta, and the compiler research projects in which Delta has played a role.

Kuck, David J.↗

Advancements in real-time engine simulation technology

The approaches used to develop real-time engine simulations are reviewed. Both digital and hybrid (analog and digital) techniques are discussed and specific examples of each are cited. These approaches are assessed from the standpoint of their usefulness for digital engine control development. A number of NASA-sponsored simulation research activities, aimed at exploring real-time simulation techniques, are described. These include the development of a microcomputer-based, parallel processor system for real-time engine simulation.

Szuch, J. R.↗

Comparison of Procedures for Dual and Triple Closely Spaced Parallel Runways

A human-in-the-loop high fidelity flight simulation experiment was conducted, which investigated and compared breakout procedures for Very Closely Spaced Parallel Approaches (VCSPA) with two and three runways. To understand the feasibility, usability and human factors of two and three runway VCSPA, data were collected and analyzed on the dependent variables of breakout cross track error and pilot workload. Independent variables included number of runways, cause of breakout and location of breakout. Results indicated larger cross track error and higher workload using three runways as compared to 2-runway operations. Significant interaction effects involving breakout cause and breakout location were also observed. Across all conditions, cross track error values showed high levels of breakout trajectory accuracy and pilot workload remained manageable. Results suggest possible avenues of future adaptation for adopting these procedures (e.g., pilot training), while also showing potential promise of the concept.

triple approaches↗

Integrated hydrogeophysical modelling and data assimilation for geoelectrical leak detection

Time-lapse electrical resistivity tomography (ERT) measurements provide indirect observations of hydrological processes in the Earth's shallow subsurface at high spatial and temporal resolution. ERT has been used in the past decades to detect leaks and monitor the evolution of associated contaminant plumes. Specifically, inverted resistivity images allow visualization of the dynamic changes in the structure of the plume. However, existing methods do not allow the direct estimation of leak parameters (e.g. leak rate, location, etc.) and their uncertainties. We propose an ensemble-based data assimilation framework that evaluates proposed hydrological models against observed time-lapse ERT measurements without directly inverting for the resistivities. Each proposed hydrological model is run through the parallel coupled hydro-geophysical simulation code PFLOTRAN-E4D to obtain simulated ERT measurements. The ensemble of model proposals is then updated using an iterative ensemble smoother. In this paper, we demonstrate the proposed framework on synthetic and field ERT data from controlled tracer injection experiments. Our results show that the approach allows joint identification of contaminant source location, initial release time, and solute loading from the cross-borehole time-lapse ERT data, alongside with an assessment of uncertainties in these estimates. We demonstrate a reduction in site-wide uncertainty by comparing the prior and posterior plume mass discharges at a selected image plane. This framework is particularly attractive to sites that have previously undergone extensive geological investigation (e.g., nuclear sites). It is well suited to complement ERT imaging and we discuss practical issues in its application to field problems.

58 GEOSCIENCES↗

Diagonal Shadows Could Cause Arcs in Thin-Film Modules With P4 Scribes

Thin-film photovoltaic (PV) modules are often made using monolithic integration (MLI), regardless of absorber technology. MLI modules sometimes use a fourth pattern of scribe lines, P4, to divide modules into parallel substrings of cells. We simulated diagonal shadows in such modules and show that they cause a voltage difference across P4. This voltage can be enough to cause an arc across P4. An arc inside a PV module can cause burned polymers and broken glass. Such packaging failures create a risk of fire or electric shock in any PV module. These hazards go beyond the permanent loss of efficiency that vertical shadows can cause. Here, we propose several solutions to this potential problem.

14 SOLAR ENERGY↗

Pervasive shifts in forest dynamics in a changing world

Forest dynamics arise from the interplay of environmental drivers and disturbances with the demographic processes of recruitment, growth, and mortality, subsequently driving biomass and species composition. However, forest disturbances and subsequent recovery are shifting with global changes in climate and land use, altering these dynamics. Changes in environmental drivers, land use, and disturbance regimes are forcing forests toward younger, shorter stands. Rising carbon dioxide, acclimation, adaptation, and migration can influence these impacts. Recent developments in Earth system models support increasingly realistic simulations of vegetation dynamics. In parallel, emerging remote sensing datasets promise qualitatively new and more abundant data on the underlying processes and consequences for vegetation structure. In combination, these advances hold promise for improving the scientific understanding of changes in vegetation demographics and disturbances.

54 ENVIRONMENTAL SCIENCES↗

Fierro Version 2.x

FIERRO is a parallel C++ code designed to simulate fluid mechanics, heat transfer, and solid mechanics in two- and three-dimensional space. FIERRO is written to run on homogeneous (CPU) and heterogeneous (CPU+GPU) high performance computing machines. Fierro can aid a) modeling and design efforts that have historically relied on commercial implicit and explicit finite element codes, b) numerical methods research, c) manufacturing research, and d) computer science research. The code contains diverse numerical methods to solve the governing physics equations for both quasi-static and dynamic problems. Mathematical optimization solvers are coupled to the numerical methods to research topology and shape optimization that has application to additive manufacturing, and to create novel numerical approaches. Phase-field methods with micromechanical solvers are provided to simulate microstructure formation and evolution in manufacturing processes. The micromechanical solvers can also help research efforts create continuum-scale constitutive models for solids, as a function of the microstructure, in situ in a calculation or in a stand-alone manner. No physical data exists within the code.

Morgan, Nathaniel↗

PFLOTRAN 5

PFLOTRAN leverages massively parallel, high performance computing to simulate large-scale non-isothermal multiphase flow, multicomponent reactive transport and electrical resistivity tomography (ERT) problems in the subsurface environment. Researchers have employed PFLOTRAN to simulate these Earth system processes on leadership class supercomputers for over two decades. The code is designed to predict the future estate of environmental systems and better inform stakeholders in the regulatory decision making process (e.g., fate of contaminants, long-term stewardship for nuclear waste, impact of climate change, etc.). A diverse team of scientists oversees PFLOTRAN development and maintenance under an open-source licensing agreement and manages contributions from an international community of researchers.

Hammond, Glenn↗

Fierro

FIERRO is a parallel C++ code designed to simulate fluid mechanics, heat transfer, and solid mechanics in two- and three dimensional space. FIERRO is written to run on homogeneous (CPU) and heterogeneous (CPU+GPU) high performance computing machines. Fierro can aid a) modeling and design efforts that have historically relied on commercial implicit and explicit finite element codes, b) numerical methods research, c) manufacturing research, and d) computer science research. The code contains diverse numerical methods to solve the governing physics equations for both quasi-static and dynamic problems. Mathematical optimization solvers are coupled to the numerical methods to research topology and shape optimization that has application to additive manufacturing, and to create novel numerical approaches. Phase-field methods with micromechanical solvers are provided to simulate microstructure formation and evolution in manufacturing processes. The micromechanical solvers can also help research efforts create continuum-scale constitutive models for solids, as a function of the microstructure, in situ in a calculation or in a stand-alone manner. No physical data exists within the code.

Morgan, Nathaniel↗

Tracer Gas Model Development and Verification in PFLOTRAN

Tracer gases, whether they are chemical or isotopic in nature, are useful tools in examining the flow and transport of gaseous or volatile species in the underground. One application is using detection of short-lived argon and xenon radionuclides to monitor for underground nuclear explosions. However, even chemically inert species, such as the noble gases, have bene observed to exhibit non-conservative behavior when flowing through porous media containing certain materials, such as zeolites, due to gas adsorption processes. This report details the model developed, implemented, and tested in the open source and massively parallel subsurface flow and transport simulator PFLOTRAN for future use in modeling the transport of adsorbing tracer gases.

07 ISOTOPE AND RADIATION SOURCES↗

DGaaS: GPU as a Service on Distributed Computing System

In the rapidly evolving landscape of scientific computing, Graphics Processing Units (GPUs) have become indispensable for their unparalleled ability to handle parallel tasks in complex calculations, simulations, and data analysis. Their utility is further magnified in machine learning and AI applications, where they significantly accelerate model training and predictive analytics. Within this context, the Triton Inference Server emerges as a pivotal open-source tool, specializing in AI inferencing and optimizing GPU utilization across various platforms and frameworks. This paper presents an in-depth study on distributed High Throughput Computing (HTC), specifically focusing on the HTCondor framework and its resource provisioning tools, GlideinWMS and HEPCloud. These systems enable large-scale scientific experiments like CMS and DUNE to efficiently access and utilize vast computational resources. The paper explores the core architectural components of GlideinWMS, including jobs, user pools, and worker nodes, and discusses their integration with GPUs and the Triton server. The primary aim of this research is to develop a solution that optimizes GPU utilization by leveraging Glideins and containers. This approach allows computational jobs, particularly those involving AI models, to use GPUs only when essential, thereby facilitating efficient sharing of limited GPU resources. To validate this architecture, the study conducted three key tests involving custom scripts, container-based servers, and Triton server deployments. However, the study faces challenges, notably in locating the Triton server and ensuring secure remote access. To address these issues, future work will focus on developing a proxy mechanism and enhancing security protocols. In conclusion, this study offers a comprehensive roadmap for effective and efficient GPU utilization in distributed High Throughput Computing. It aims to contribute significantly to the scientific community by solving pressing problems and implementing robust solutions in collaboration with the GlideinWMS and HEPCloud teams. The research sets the stage for a more efficient, scalable, and cost-effective paradigm in scientific computing.

97 MATHEMATICS AND COMPUTING↗