Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PARALLEL PROCESSING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Parallel simulation via SPPARKS of on-lattice kinetic and Metropolis Monte Carlo models for materials processing

Abstract SPPARKS is an open-source parallel simulation code for developing and running various kinds of on-lattice Monte Carlo models at the atomic or meso scales. It can be used to study the properties of solid-state materials as well as model their dynamic evolution during processing. The modular nature of the code allows new models and diagnostic computations to be added without modification to its core functionality, including its parallel algorithms. A variety of models for microstructural evolution (grain growth), solid-state diffusion, thin film deposition, and additive manufacturing (AM) processes are included in the code. SPPARKS can also be used to implement grid-based algorithms such as phase field or cellular automata models, to run either in tandem with a Monte Carlo method or independently. For very large systems such as AM applications, the Stitch I/O library is included, which enables only a small portion of a huge system to be resident in memory. In this paper we describe SPPARKS and its parallel algorithms and performance, explain how new Monte Carlo models can be added, and highlight a variety of applications which have been developed within the code.

36 MATERIALS SCIENCE↗

Bulk viscosity from Urca processes: n p e μ matter in the neutrino-transparent regime

We study the bulk viscosity of moderately hot and dense, neutrino-transparent relativistic npeμ matter arising from weak-interaction direct Urca processes. This work parallels our recent study of the bulk viscosity of npeμ matter with a trapped neutrino component. The nuclear matter is modeled in a relativistic density functional approach with two different parametrizations—DDME2 (which does not allow for the low-temperature direct-Urca process at any density) and NL3 (which allows for low-temperature direct-Urca process above a low-density threshold). Here, we compute the equilibration rates of Urca processes of neutron decay and lepton capture, as well as the rate of the muon decay, and find that the muon decay process is subdominant to the Urca processes at temperatures T ≥ 3 MeV in the case of DDME2 model and T ≥ 1 MeV in the case of NL3 model. Thus, the Urca-process-driven bulk viscosity is computed with the assumption that pure leptonic reactions are frozen. As a result the electronic and muonic Urca channels contribute to the bulk viscosity independently and at certain densities the bulk viscosity of npeμ matter shows a double-peak structure as a function of temperature instead of the standard one-peak (resonant) form. In the final step, we estimate the damping timescales of density oscillations by the bulk viscosity. We find that, e.g., at a typical oscillation frequency f = 1~kHz, the damping of oscillation is most efficient at temperatures 3 ≤ T ≤ 5~MeV and densities n B ≤ 2n 0 where they can affect the evolution of the post-merger object.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

High-performance data management for whole slide image analysis in digital pathology

When dealing with giga-pixel digital pathology in whole-slide imaging, a notable proportion of data records holds relevance during each analysis operation. For instance, when deploying an image analysis algorithm on whole-slide images (WSI), the computational bottleneck often lies in the input-output (I/O) system. This is particularly notable as patch-level processing introduces a considerable I/O load onto the computer system. However, this data management process could be further paralleled, given the typical independence of patch-level image processes across different patches. This paper details our endeavors in tackling this data access challenge by implementing the Adaptable IO System version 2 (ADIOS2). Our focus has been constructing and releasing a digital pathology-centric pipeline using ADIOS2, which facilitates streamlined data management across WSIs. Additionally, we’ve developed strategies aimed at curtailing data retrieval times. The performance evaluation encompasses two key scenarios: (1) a pure CPU-based image analysis scenario (“CPU scenario”), and (2) a GPU-based deep learning framework scenario (“GPU scenario”). Our findings reveal noteworthy outcomes. Under the CPU scenario, ADIOS2 showcases an impressive two-fold speed-up compared to the brute-force approach. In the GPU scenario, its performance stands on par with the cutting-edge GPU I/O acceleration framework, NVIDIA Magnum IO GPU Direct Storage (GDS). From what we know, this appears to be among the initial instances, if any, of utilizing ADIOS2 within the field of digital pathology. The source code has been made publicly available at https://github.com/hrlblab/adios.

Wang, Xiao↗

Role of surface diffusion in formation of unique reactivity for graphite oxidation: Time-resolved measurements in a pulsed diffusion reactor

Quantification of oxidation kinetics is essential to develop graphitic materials for diverse applications: from refractories found in gas-cooled nuclear reactors to catalysts needed for chemical manufacturing. In this work, using well-defined highly oriented pyrolytic graphite, low-pressure isotopic transient experiments combined with controlled annealing periods, we resolve the role of surface diffusion and quantify oxidation kinetics with nanomole-precision. We observe an unexpected increase in reactivity following annealing which is explained by the role of surface diffusion increasing the probability for trapping mobile oxygen at more reactive edge sites. Here, the locus of adsorption and spillover to the basal plane is distinct from the trapping location creating a more active oxygen species. Isotopic products reflect the population dynamics of oxygen added at the edge and surface diffusion that relocates basal plane oxygen to more reactive edge sites. Since this process proceeds in parallel with direct oxidation reactions, it is not likely to be observed using steady-state or conventional ‘bulk’ characterization techniques. Our unique time-resolved non-equilibrium measurement in a well-defined transport regime, enables observation of three distinct behaviors: short-term deactivation due to the balance of rates in oxygen supply/product formation, reactivity increases due to surface diffusion and longer-term reactivity increase with oxygen accumulation.

36 MATERIALS SCIENCE↗

Development of a Reactive Force Field for Simulating Photoinitiated Acrylate Polymerization

Light-driven and photo-curable polymer based additive manufacturing (AM) has enormous potential due to its excellent resolution and precision. Acrylated radical chain-growth polymerized resins are widely used in photopolymer AM due to their fast kinetics, and often serve as a departure point for developing other resin materials for photopolymer-based AM technologies. For successful control of the photopolymer resins, the molecular basis of the acrylate free-radical polymerization has to be understood in detail. We present an optimized reactive force field (ReaxFF) for molecular dynamics (MD) simulations of acrylate polymer resins that captures radical polymerization thermodynamics and kinetics. The force field is trained against an extensive training set including density functional theory (DFT) calculations of reaction pathways along the radical polymerization from methyl acrylate to methyl butyrate, bond dissociation energies, and structures and partial charges of several molecules and radicals. We also found that it was critical to train the force field against an incorrect, nonphysical reaction pathway observed in simulations that used parameters not optimized for acrylate polymerization. As a result, the parameterization process utilizes a parallelized search algorithm, and the resulting model can describe polymer resin formation, crosslinking density, conversion rate, and residual monomers of the complex acrylate mixtures.

36 MATERIALS SCIENCE↗

In Situ Imaging of Faujasite Surface Growth Reveals Unique Pathways of Zeolite Crystallization

Zeolite crystallization occurs by complex processes involving a variety of possible mechanisms. The sol gel media used to prepare zeolites leads to heterogeneous mixtures of solution and solid states with diverse solute species. At later stages of zeolite synthesis when growth occurs predominantly from solution, classical two-dimensional nucleation and spreading of layers on crystal surfaces via the addition of soluble species is the dominant pathway. At earlier stages, these processes occur in parallel with nonclassical pathways involving crystallization by particle attachment (CPA). The relative roles of solution- and solid-state species in zeolite crystallization have been a subject of debate. Here, in this work, we investigate the growth mechanism of a commercially relevant zeolite, faujasite (FAU). In situ atomic force microscopy (AFM) measurements reveal that supernatant solutions extracted from a conventional FAU synthesis at various times do not result in growth, indicating that FAU growth predominantly occurs from the solid state through a disorder-to-order transition of amorphous precursors. Elemental analysis shows that supernatant solutions are significantly more siliceous than both the original growth mixture and the FAU zeolite product; however, in situ AFM studies using a dilute clear solution with a lower Si/Al ratio revealed three-dimensional growth of surfaces that is distinct from layer-by-layer and CPA pathways. This unique mechanism of growth differs from those observed in studies of other zeolites. Given that relatively few zeolite frameworks have been the subject of mechanistic investigation by in situ techniques, these observations of FAU crystallization raise the question whether its growth pathway is characteristic of other zeolite structures.

36 MATERIALS SCIENCE↗

Adaptive Grid Redistribution for a 1D Model of Turbulence and Clouds

In global atmospheric models, resolving stratocumulus (Sc) in the vertical is computationally expensive. However, Sc appear only under special meteorological conditions. Therefore, there is motivation to refine the vertical grid levels adaptively. In order to facilitate the possibility of parallelization on graphical processing units, our grid adaptation method prescribes the number of vertical levels a priori. Then grid levels are relocated toward altitude ranges in need of refinement. Because the method relocates existing grid levels, rather than adding extra levels, there is a risk of creating regions with overly coarse grid spacing, that is, voids in the grid mesh. To prevent such voids from forming, a simple method is developed to impose a maximum grid spacing. To decide where to place enhanced resolution, the authors develop an empirical mesh refinement criterion. It refines grid spacing near the ground, near strong temperature gradients, and within clouds. Our grid adaptation method is implemented in a single-column model and evaluated on four test cases: decaying stratocumulus, developing shallow cumulus, a quasi-stationary stratocumulus deck, and the diurnal cycle of a dry boundary layer. In the stratocumulus cases, mesh refinement leads to improvements in both the time evolution of fields and their time averages. The other two cases show smaller differences.

Carstensen, Steffen [Univ. of Wisconsin, Milwaukee↗

Angularly resolved photoionization dynamics in atoms and molecules combining temporally and spectrally resolved experiments at ATTOLab and Synchrotron SOLEIL

We report results for XUV-IR two-photon ionization of Ar, Ne, NO, and O2, where an XUV attosecond pulse train is superimposed with a synchronized IR pulse, obtained at the ATTOLab laser facility using electron–ion coincidence 3D momentum spectroscopy. Temporally resolved photoelectron angular distributions providing angle-resolved time-delays for np ionization of Ar and Ne, achieved by reconstruction of attosecond beating by interference of two-photon transitions through a unified formalism (Joseph et al. in J Phys B At Mol Opt Phys 53:184007, 2020), are summarized. For inner valence XUV-IR dissociative photoionization of NO and O2 molecules, we report electron–ion kinetic energy correlation diagrams and disentangle the dissociative photoionization processes relying on parallel XUV experiments at Synchrotron SOLEIL. For ionization into the NO+(c3Π) ionic state, extending the formalism developed for single-photon ionization, we focus on photoelectron angular distributions averaged on the delay between the XUV and the IR field in the field frame, molecular frame, and electron frame of reference.

Joseph, J↗

Interactive Web Application for Traffic Simulation Data Management and Visualization

As traffic simulation software becomes more effective for realistically simulating and analyzing traffic dynamics and vehicle interactions on the mesoscopic and microscopic level, the management, dissemination, and collaborative visualization of traffic simulation results produced by individual transportation planners presents a significant challenge. Existing online content management systems have a very limited capability in allowing users to query specific traffic simulation scenarios and geospatially visualize simulation results through shareable and interactive web interfaces. This paper presents a web-based application for promoting the archiving, sharing, and visualization of large-scale traffic simulation outputs. The application is developed to enhance cyber-physical controls, communications, and public education for collaborative transportation planning. Unique features of the web application include: (a) allowing users to upload their new traffic simulation scenarios (parameters and outputs), as well as search existing scenarios using easily accessible interfaces; (b) optimizing simulation output files with heterogeneous data formats and projected coordinate systems for web-based storage and management using a scalable and searchable data/metadata standard; (c) standardizing user-uploaded simulation outputs using web interfaces and data processing libraries with parallel computing capacity; and (d) providing shareable web visual interfaces for visualizing the traffic flow and signal information stored in simulation outputs (e.g., regional traffic patterns and individual vehicle interactions) and visually comparing multiple simulation outputs both spatially and temporally. Furthermore, the paper presents the conceptual design and implementation of this application, and demonstrates the application’s performance for sharing, comparing, and visualizing simulation outputs from VISSIM and SUMO, two commonly used traffic simulation software programs.

96 KNOWLEDGE MANAGEMENT AND PRESERVATION↗

Towards a more general understanding of the algorithmic utility of recurrent connections

Lateral and recurrent connections are ubiquitous in biological neural circuits. Yet while the strong computational abilities of feedforward networks have been extensively studied, our understanding of the role and advantages of recurrent computations that might explain their prevalence remains an important open challenge. Foundational studies by Minsky and Roelfsema argued that computations that require propagation of global information for local computation to take place would particularly benefit from the sequential, parallel nature of processing in recurrent networks. Such “tag propagation” algorithms perform repeated, local propagation of information and were originally introduced in the context of detecting connectedness, a task that is challenging for feedforward networks. Here, we advance the understanding of the utility of lateral and recurrent computation by first performing a large-scale empirical study of neural architectures for the computation of connectedness to explore feedforward solutions more fully and establish robustly the importance of recurrent architectures. In addition, we highlight a tradeoff between computation time and performance and construct hybrid feedforward/recurrent models that perform well even in the presence of varying computational time limitations. We then generalize tag propagation architectures to propagating multiple interacting tags and demonstrate that these are efficient computational substrates for more general computations of connectedness by introducing and solving an abstracted biologically inspired decision-making task. Our work thus clarifies and expands the set of computational tasks that can be solved efficiently by recurrent computation, yielding hypotheses for structure in population activity that may be present in such tasks.

59 BASIC BIOLOGICAL SCIENCES↗

GIGA-Lens: Fast Bayesian Inference for Strong Gravitational Lens Modeling

We present GIGA-Lens: a gradient-informed, GPU-accelerated Bayesian framework for modeling strong gravitational lensing systems, implemented in TensorFlow and JAX. The three components, optimization using multistart gradient descent, posterior covariance estimation with variational inference, and sampling via Hamiltonian Monte Carlo, all take advantage of gradient information through automatic differentiation and massive parallelization on graphics processing units (GPUs). We test our pipeline on a large set of simulated systems and demonstrate in detail its high level of performance. The average time to model a single system on four Nvidia A100 GPUs is 105 s. The robustness, speed, and scalability offered by this framework make it possible to model the large number of strong lenses found in current surveys and present a very promising prospect for the modeling of ${ \mathcal O }({10}^{5})$ lensing systems expected to be discovered in the era of the Vera C. Rubin Observatory, Euclid, and the Nancy Grace Roman Space Telescope.

79 ASTRONOMY AND ASTROPHYSICS↗

Parallel Simulation of Quantum Networks with Distributed Quantum State Management

Quantum network simulators offer the opportunity to cost-efficiently investigate potential avenues for building networks that scale with the number of users, communication distance, and application demands by simulating alternative hardware designs and control protocols. Several quantum network simulators have been recently developed with these goals in mind. As the size of the simulated networks increases, however, sequential execution becomes time-consuming. Parallel execution presents a suitable method for scalable simulations of large-scale quantum networks, but the unique attributes of quantum information create unexpected challenges. In this work, we identify requirements for parallel simulation of quantum networks and develop the first parallel discrete-event quantum network simulator by modifying the existing serial simulator SeQUeNCe. Our contributions include the design and development of a quantum state manager (QSM) that maintains shared quantum information distributed across multiple processes. We also optimize our parallel code by minimizing the overhead of the QSM and decreasing the amount of synchronization needed among processes. Using these techniques, we observe a speedup of 2 to 25 times when simulating a 1,024-node linear network topology using 2 to 128 processes. We also observe an efficiency greater than 0.5 for up to 32 processes in a linear network topology of the same size and with the same workload. We repeat this evaluation with a randomized workload on a caveman network. We also introduce several methods for partitioning networks by mapping them to different parallel simulation processes. We have released the parallel SeQUeNCe simulator as an open source tool alongside the existing sequential version.

97 MATHEMATICS AND COMPUTING↗

An Accelerated Clip Algorithm for Unstructured Meshes: A Batch-Driven Approach

The clip technique is a popular method for visualizing complex structures and phenomena within 3D unstructured meshes. Meshes can be clipped by specifying a scalar isovalue to produce an output unstructured mesh with its external surface as the isovalue. Similar to isocontouring, the clipping process relies on scalar data associated with the mesh points, including scalar data generated by implicit functions such as planes, boxes, and spheres, which facilitates the visualization of results interior to the grid. In this paper, we introduce a novel batch-driven parallel algorithm based on a sequential clip algorithm designed for high-quality results in partial volume extraction. Our algorithm comprises five passes, each progressively processing data to generate the resulting clipped unstructured mesh. The novelty lies in the use of fixed-size batches of points and cells, which enable rapid workload trimming and parallel processing, leading to a significantly improved memory footprint and run-time performance compared to the original version. On a 32-core CPU, the proposed batch-driven parallel algorithm demonstrates a run-time speed-up of up to 32.6x and a memory footprint reduction of up to 4.37x compared to the existing sequential algorithm. The software is currently available under an open-source license in the VTK visualization system.

Tsalikis, Spiros↗

Geometric GNNs for charged particle tracking at GlueX

Nuclear physics experiments are aimed at uncovering the fundamental building blocks of matter. The experiments involve high-energy collisions that produce complex events with many particle trajectories. Tracking charged particles resulting from collisions in the presence of a strong magnetic field is critical to enable the reconstruction of particle trajectories and precise determination of interactions. It is traditionally achieved through combinatorial approaches that scale worse than linearly as the number of hits grows. Since particle hit data naturally form a point cloud and can be structured as graphs, graph neural networks (GNNs) emerge as an intuitive and effective choice for this task. In this study, we evaluate the GNN model for track finding on the data from the GlueX experiment at Jefferson Lab. We use simulation data to train the model and test on both simulation and real GlueX measurements. We demonstrate that GNN-based track finding outperforms the currently used traditional method at GlueX in terms of segment-based efficiency at a fixed purity while providing faster inferences. We show that the GNN model can achieve significant speedup by processing multiple events in batches, which exploits the parallel computation capability of graphical processing units (GPUs). Finally, we compare the GNN implementation on GPU and field-programmable gate array and describe the trade-off.

batched GNN pipeline↗

Computational Complexity of Neuromorphic Algorithms

Neuromorphic computing has several characteristics that make it an extremely compelling computing paradigm for post Moore computation. Some of these characteristics include intrinsic parallelism, inherent scalability, collocated processing and memory, and event-driven computation. While these characteristics impart energy efficiency to neuromorphic systems, they do come with their own set of challenges. One of the biggest challenges in neuromorphic computing is to establish the theoretical underpinnings of the computational complexity of neuromorphic algorithms. In this paper, we take the first steps towards defining the space and time complexity of neuromorphic algorithms. Specifically, we describe a model of neuromorphic computation and state the assumptions that govern the computational complexity of neuromorphic algorithms. Next, we present a theoretical framework to define the computational complexity of a neuromorphic algorithm. We explicitly define what space and time complexities mean in the context of neuromorphic algorithms based on our model of neuromorphic computation. Finally, we leverage our approach and define the computational complexities of six neuromorphic algorithms: constant function, successor function, predecessor function, projection function, neuromorphic sorting algorithm and neighborhood subgraph extraction algorithm.

Date, Prasanna↗

Comparing the Performance of Julia on CPUs versus GPUs and Julia-MPI versus Fortran-MPI: a case study with MPAS-Ocean (Version 7.1)

Abstract. Some programming languages are easy to develop at the cost of slow execution, while others are fast at runtime but much more difficult to write. Julia is a programming language that aims to be the best of both worlds – a development and production language at the same time. To test Julia's utility in scientific high-performance computing (HPC), we built an unstructured-mesh shallow water model in Julia and compared it against an established Fortran-MPI ocean model, the Model for Prediction Across Scales–Ocean (MPAS-Ocean), as well as a Python shallow water code. Three versions of the Julia shallow water code were created: for single-core CPU, graphics processing unit (GPU), and Message Passing Interface (MPI) CPU clusters. Comparing identical simulations revealed that our first version of the Julia model was 13 times faster than Python using NumPy, where both used an unthreaded single-core CPU. Further Julia optimizations, including static typing and removing implicit memory allocations, provided an additional 10–20× speed-up of the single-core CPU Julia model. The GPU-accelerated Julia code was almost identical in terms of performance to the MPI parallelized code on 64 processes, an unexpected result for such different architectures. Parallelized Julia-MPI performance was identical to Fortran-MPI MPAS-Ocean for low processor counts and ranges from 2× faster to 2× slower for higher processor counts. Our experience is that Julia development is fast and convenient for prototyping but that Julia requires further investment and expertise to be competitive with compiled codes. We provide advice on Julia code optimization for HPC systems.

54 ENVIRONMENTAL SCIENCES↗

Optical neural engine for solving scientific partial differential equations

Abstract Solving partial differential equations (PDEs) is the cornerstone of scientific research and development. Data-driven machine learning (ML) approaches are emerging to accelerate time-consuming and computation-intensive numerical simulations of PDEs. Although optical systems offer high-throughput and energy-efficient ML hardware, their demonstration for solving PDEs is limited. Here, we present an optical neural engine (ONE) architecture combining diffractive optical neural networks for Fourier space processing and optical crossbar structures for real space processing to solve time-dependent and time-independent PDEs in diverse disciplines, including Darcy flow equation, the magnetostatic Poisson’s equation in demagnetization, the Navier-Stokes equation in incompressible fluid, Maxwell’s equations in nanophotonic metasurfaces, and coupled PDEs in a multiphysics system. We numerically and experimentally demonstrate the capability of the ONE architecture, which not only leverages the advantages of high-performance dual-space processing for outperforming traditional PDE solvers and being comparable with state-of-the-art ML models but also can be implemented using optical computing hardware with unique features of low-energy and highly parallel constant-time processing irrespective of model scales and real-time reconfigurability for tackling multiple tasks with the same architecture. The demonstrated architecture offers a versatile and powerful platform for large-scale scientific and engineering computations.

Tang, Yingheng (ORCID:0009000153622546)↗

Mechanistic insights into the pyrolysis of poly (vinyl chloride)

The accumulation of unmanaged plastic waste in the environment has a devastating impact upon marine life and human health. Catalytic and thermal pyrolysis are promising technologies toward the efficient utilization of plastic waste. Yet, the processing of polymers, such as polyvinyl chloride (PVC), that decompose into corrosive compounds remains a major challenge. In this work, we employ density functional theory (DFT) and thermogravimetric analysis (TGA) to explore the elementary chemical steps that underpin the thermal decomposition of PVC. We determine that the dehydrochlorination reaction (i.e., 1 st stage of thermal decomposition) begins in tertiary chloride defects and propagates via HCl–mediated autocatalysis of internal allylic (IA) chloride groups. The latter groups, when in the vicinity of π–conjugated polymer segments, release HCl in a facile manner. We predict that other compounds, including hydrogen halides and H 2 O could also catalyze PVC’s dehydrochlorination. We suggest that hydrogen halides are the most efficient catalysts for this process, while other compounds like H 2 O may slow down the dehydrochlorination compared to purely HCl-catalyzed process, because of the dilution of the produced HCl. This result is corroborated by TGA experiments. Additionally, we study the thermochemistry and kinetics of polyene chain crosslinking and the formation of aromatics. The former reaction proceeds in parallel with the dehydrochlorination process (i.e., during the 1st stage of thermal decomposition), whilst the latter may occur at high temperatures (i.e., during the 2 nd stage of thermal decomposition). Furthermore, this work contributes to the fundamental understanding of molecular scale phenomena that take place during PVC pyrolysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗