Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “PARALLEL PROCESSING”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Geometric GNNs for charged particle tracking at GlueX

Nuclear physics experiments are aimed at uncovering the fundamental building blocks of matter. The experiments involve high-energy collisions that produce complex events with many particle trajectories. Tracking charged particles resulting from collisions in the presence of a strong magnetic field is critical to enable the reconstruction of particle trajectories and precise determination of interactions. It is traditionally achieved through combinatorial approaches that scale worse than linearly as the number of hits grows. Since particle hit data naturally form a point cloud and can be structured as graphs, graph neural networks (GNNs) emerge as an intuitive and effective choice for this task. In this study, we evaluate the GNN model for track finding on the data from the GlueX experiment at Jefferson Lab. We use simulation data to train the model and test on both simulation and real GlueX measurements. We demonstrate that GNN-based track finding outperforms the currently used traditional method at GlueX in terms of segment-based efficiency at a fixed purity while providing faster inferences. We show that the GNN model can achieve significant speedup by processing multiple events in batches, which exploits the parallel computation capability of graphical processing units (GPUs). Finally, we compare the GNN implementation on GPU and field-programmable gate array and describe the trade-off.

batched GNN pipeline↗

Computational Complexity of Neuromorphic Algorithms

Neuromorphic computing has several characteristics that make it an extremely compelling computing paradigm for post Moore computation. Some of these characteristics include intrinsic parallelism, inherent scalability, collocated processing and memory, and event-driven computation. While these characteristics impart energy efficiency to neuromorphic systems, they do come with their own set of challenges. One of the biggest challenges in neuromorphic computing is to establish the theoretical underpinnings of the computational complexity of neuromorphic algorithms. In this paper, we take the first steps towards defining the space and time complexity of neuromorphic algorithms. Specifically, we describe a model of neuromorphic computation and state the assumptions that govern the computational complexity of neuromorphic algorithms. Next, we present a theoretical framework to define the computational complexity of a neuromorphic algorithm. We explicitly define what space and time complexities mean in the context of neuromorphic algorithms based on our model of neuromorphic computation. Finally, we leverage our approach and define the computational complexities of six neuromorphic algorithms: constant function, successor function, predecessor function, projection function, neuromorphic sorting algorithm and neighborhood subgraph extraction algorithm.

Date, Prasanna↗

High-performance computing in water resources hydrodynamics

In this work, we present a vision of future water resources hydrodynamics codes that can fully utilize the strengths of modern high-performance computing. The advances to computing power, formerly driven by the improvement of central processing unit processors, now focus on parallel computing and, in particular, the use of graphics processing units (GPUs). However, this shift to a parallel framework requires refactoring the code to make efficient use of the data as well as changing even the nature of the algorithm that solves the system of equations. These concepts along with other features such as the precision for the computations, dry regions management, and input/output data are analyzed in this paper. A 2D multi-GPU flood code applied to a large-scale test case is used to corroborate our statements and ascertain the new challenges for the next-generation parallel water resources codes.

54 ENVIRONMENTAL SCIENCES↗

Comparing the Performance of Julia on CPUs versus GPUs and Julia-MPI versus Fortran-MPI: a case study with MPAS-Ocean (Version 7.1)

Abstract. Some programming languages are easy to develop at the cost of slow execution, while others are fast at runtime but much more difficult to write. Julia is a programming language that aims to be the best of both worlds – a development and production language at the same time. To test Julia's utility in scientific high-performance computing (HPC), we built an unstructured-mesh shallow water model in Julia and compared it against an established Fortran-MPI ocean model, the Model for Prediction Across Scales–Ocean (MPAS-Ocean), as well as a Python shallow water code. Three versions of the Julia shallow water code were created: for single-core CPU, graphics processing unit (GPU), and Message Passing Interface (MPI) CPU clusters. Comparing identical simulations revealed that our first version of the Julia model was 13 times faster than Python using NumPy, where both used an unthreaded single-core CPU. Further Julia optimizations, including static typing and removing implicit memory allocations, provided an additional 10–20× speed-up of the single-core CPU Julia model. The GPU-accelerated Julia code was almost identical in terms of performance to the MPI parallelized code on 64 processes, an unexpected result for such different architectures. Parallelized Julia-MPI performance was identical to Fortran-MPI MPAS-Ocean for low processor counts and ranges from 2× faster to 2× slower for higher processor counts. Our experience is that Julia development is fast and convenient for prototyping but that Julia requires further investment and expertise to be competitive with compiled codes. We provide advice on Julia code optimization for HPC systems.

54 ENVIRONMENTAL SCIENCES↗

Optical neural engine for solving scientific partial differential equations

Abstract Solving partial differential equations (PDEs) is the cornerstone of scientific research and development. Data-driven machine learning (ML) approaches are emerging to accelerate time-consuming and computation-intensive numerical simulations of PDEs. Although optical systems offer high-throughput and energy-efficient ML hardware, their demonstration for solving PDEs is limited. Here, we present an optical neural engine (ONE) architecture combining diffractive optical neural networks for Fourier space processing and optical crossbar structures for real space processing to solve time-dependent and time-independent PDEs in diverse disciplines, including Darcy flow equation, the magnetostatic Poisson’s equation in demagnetization, the Navier-Stokes equation in incompressible fluid, Maxwell’s equations in nanophotonic metasurfaces, and coupled PDEs in a multiphysics system. We numerically and experimentally demonstrate the capability of the ONE architecture, which not only leverages the advantages of high-performance dual-space processing for outperforming traditional PDE solvers and being comparable with state-of-the-art ML models but also can be implemented using optical computing hardware with unique features of low-energy and highly parallel constant-time processing irrespective of model scales and real-time reconfigurability for tackling multiple tasks with the same architecture. The demonstrated architecture offers a versatile and powerful platform for large-scale scientific and engineering computations.

Tang, Yingheng (ORCID:0009000153622546)↗

Mechanistic insights into the pyrolysis of poly (vinyl chloride)

The accumulation of unmanaged plastic waste in the environment has a devastating impact upon marine life and human health. Catalytic and thermal pyrolysis are promising technologies toward the efficient utilization of plastic waste. Yet, the processing of polymers, such as polyvinyl chloride (PVC), that decompose into corrosive compounds remains a major challenge. In this work, we employ density functional theory (DFT) and thermogravimetric analysis (TGA) to explore the elementary chemical steps that underpin the thermal decomposition of PVC. We determine that the dehydrochlorination reaction (i.e., 1 st stage of thermal decomposition) begins in tertiary chloride defects and propagates via HCl–mediated autocatalysis of internal allylic (IA) chloride groups. The latter groups, when in the vicinity of π–conjugated polymer segments, release HCl in a facile manner. We predict that other compounds, including hydrogen halides and H 2 O could also catalyze PVC’s dehydrochlorination. We suggest that hydrogen halides are the most efficient catalysts for this process, while other compounds like H 2 O may slow down the dehydrochlorination compared to purely HCl-catalyzed process, because of the dilution of the produced HCl. This result is corroborated by TGA experiments. Additionally, we study the thermochemistry and kinetics of polyene chain crosslinking and the formation of aromatics. The former reaction proceeds in parallel with the dehydrochlorination process (i.e., during the 1st stage of thermal decomposition), whilst the latter may occur at high temperatures (i.e., during the 2 nd stage of thermal decomposition). Furthermore, this work contributes to the fundamental understanding of molecular scale phenomena that take place during PVC pyrolysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

IDAES-PSE 2.7.0 Release

The Institute for the Design of Advanced Energy Systems (IDAES) Integrated Platform is a versatile computational environment offering extensive process systems engineering (PSE) capabilities for optimizing the design and operation of complex, interacting technologies and systems. IDAES enables users to efficiently search vast, complex design spaces to discover the lowest cost solutions while supporting the full process modeling lifecycle, from conceptual design to dynamic optimization and control. The extensible, open platform empowers users to create models of novel processes and rapidly develop custom analyses, workflows, and end-user applications. IDAES-PSE 2.7.0 Release Highlights New features: AutoScaler and CustomScalerBase classes: Such tools are the core of the new scaling framework being implemented in IDAES. Wider adoption of scaling tools among users will result in quicker and more robust model solutions. Scaler for equilibrium reactor and saponification properties: These scaler models are examples to follow for how to use the new scaling tools. ONNX Surrogate support from Optimization & Machine Learning Toolkit (OMLT): ONNX is an open standard format to save and load ML/AI models that is widely supported by all major frameworks. This capability makes it easier for IDAES users to create surrogate models and use them without having to support each framework individually. 1D Membrane Model for CO2 Capture and Utilization: Supports ongoing efforts for modeling and optimizing polymer membrane processes for CO2 capture and conversion into formic acid. StreamScaler unit model: Unrelated to the CustomScalerBase, this unit model allows a stream’s extensive variables to be scaled by a fixed factor. This allows streams being processed by multiple units in parallel to be scaled down to unit scale and scaled back up to process scale. Bug fixes or improvements: Scaling, EoS, Diagnostics tool, Modular Properties, tests & documentation Deprecations: Old Cubic EoS

AS↗

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Supercomputing Pipelines Search for Therapeutics Against COVID-19

The urgent search for drugs to combat SARS-CoV-2 has included the use of supercomputers. The use of general-purpose graphical processing units (GPUs), massive parallelism, and new software for high-performance computing (HPC) has allowed researchers to search the vast chemical space of potential drugs faster than ever before. We developed a new drug discovery pipeline using the Summit supercomputer at Oak Ridge National Laboratory to help pioneer this effort, with new platforms that incorporate GPU-accelerated simulation and allow for the virtual screening of billions of potential drug compounds in days compared to weeks or months for their ability to inhibit SARS-COV-2 proteins. Here, this effort will accelerate the process of developing drugs to combat the current COVID-19 pandemic and other diseases.

60 APPLIED LIFE SCIENCES↗

Characterization and Rationalization of Microstructural Evolution in GRCop-84 Processed by Laser-Powder Bed Fusion (L-PBF)

In this study, prismatic geometries of GRCop-84 [Cu-8Cr-4Nb (at. pct)] were built with laser-powder bed fusion (L-PBF) process. The samples were sectioned parallel or perpendicular to the build direction and characterized in the as-built and after post-processing with a hot-isostatically pressing (HIP) treatment. The microstructure and phase evolutions were evaluated with optical microscopy, scanning electron microscopy (SEM), electron backscattered diffraction (EBSD), and high-temperature X-ray diffraction (HTXRD) up to 1223 K. The samples in the as-built conditions exhibited predominantly columnar epitaxial and misoriented Cu-FCC grains. The microstructure evolutions are discussed based on locations within the overall build geometry, the dynamics of small melt pool shape and sectioning effects. The above grain structure did not change significantly during post-process HIP treatment. The stability of this FCC grain structure is attributed to the formation of primary stable Cr 2 Nb (Laves phase) during L-PBF, even before the emergence of FCC grains from liquid. The stability of Cr 2 Nb in both as-built and HIPed samples were evaluated using high-temperature X-ray diffraction measurements and compared with that of gas-atomized powder. The significance of these results is discussed with reference to aerospace applications.

36 MATERIALS SCIENCE↗

An open source fast fluid dynamics model for data center thermal management

Although computational fluid dynamics (CFD) has been widely adopted to improve data center thermal management, the high computational demand limits its applications, such as multivariate optimal design and operation. Fast fluid dynamics (FFD), which has been applied for fast airflow simulation, shows great potential. However, few research applied FFD for optimal design and operation of data center thermal management. This research improves the FFD model for data centers and conducts a comprehensive evaluation and demonstration. First, the FFD model is improved by solving the advection and diffusion equations together using an upwind scheme instead of a semi-Lagrangian advection solver in the conventional FFD model. Second, new features for data centers are added, such as a pressure correction method to simulate plenum airflow and dynamic boundary conditions for IT racks. The new FFD model is first validated with two indoor environment cases and the results show that the new FFD model has slightly better overall prediction accuracy and faster speed compared to the conventional FFD model. It is also observed that both FFD models achieve acceptable accuracy, except for a few localized disparities with experimental data, which might be due to simplified handling of turbulence viscosity near the boundaries. Furthermore, validation with a real data center shows that the FFD model achieves a similar level of accuracy as CFD when compared to the experimental measurements with some level of uncertainties. It is then demonstrated for data center optimal design and operation, which saves 53.4–58.8% of annual energy while still meeting the thermal requirements. In conclusion, with a much faster speed and comparable accuracy compared to CFD, the FFD model parallelized on a graphics processing unit is promising for practical model-based data center early design and operation.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Nanoscale Probing of Electrical Memory Effects in van der Waals Layered PdSe 2

Tunable electronic materials that can be switched between different impedance states are fundamental to the hardware elements for neuromorphic computing architectures. This “brain-like” computing paradigm uses highly paralleled and colocated data processing, leading to greatly improved energy efficiency and performance compared to traditional architectures in which data have to be frequently transferred between processor and memory. In this work, we use scanning microwave impedance microscopy for nanoscale electrical and electronic characterization of two-dimensional layered semiconductor PdSe 2 to probe neuromorphic properties. The local resolution of tens of nanometers reveals significant differences in electronic behavior between and within PdSe 2 nanosheets (NSs). In particular, we detected both n-type and p-type behaviors, although previous reports only point to ambipolar n-type dominating characteristics. Nanoscale capacitance–voltage curves and subsequent calculation of characteristic maps revealed a hysteretic behavior originating from the creation and erasure of Se vacancies as well as the switching of defect charge states. In addition, stacks consisting of two NSs show enhanced resistive and capacitive switching, which is attributed to trapped charge carriers at the interfaces between the stacked NSs. Stacking n- and p-type NSs results in a combined behavior that allows one to tune electrical characteristics. In conclusion, as local inhomogeneities of electrical and electronic behavior can have a significant impact on the overall device performance, the demonstrated nanoscale characterization and analysis will be applicable to a wide range of semiconducting materials.

2D material↗

Formation of late-generation atmospheric compounds inhibited by rapid deposition

Reactive organic carbon species are important fuel for atmospheric chemical reactions, including the formation of secondary organic aerosol. However, in parallel to atmospheric oxidation processes, deposition can remove compounds from the atmosphere and impact downstream environments. To understand the impact of deposition on atmospheric oxidation, we present a framework for predicting and visualizing the fate of a molecule on the basis of the physicochemical properties of compounds (Henry’s law constant, vapour pressure and reaction rate constants), which are used to estimate timescales for oxidation and deposition. Further, by implementing our deposition rates in chemical models, we show that deposition substantially suppresses atmospheric reactivity and aerosol formation by removing early-generation products and preventing the formation of large fractions (up to 90%) of downstream, late-generation compounds. Deposition is frequently missing in the laboratory experiments and detailed chemical modelling, which probably biases our understanding of atmospheric composition.

54 ENVIRONMENTAL SCIENCES↗

Laboratory study of the PFRC-2's initial plasma densification stages

Initial plasma densification by odd-parity rotating magnetic fields (RMF o ) applied to the linear magnetized Princeton field-reversed configuration (PFRC-2) device with fill gases at pressures near 1 mTorr proceeds through two phases: a slow one, characterized by a rise time $τ_{slow}$ ~ 100 $μ$s, followed by a fast one, characterized by $τ_{fast}$ ~ 10 $μ$s. The transition from slow to fast occurs at a line-integral-averaged electron density, t n e , near 2$\times$ 10 11 cm –3 , independent of magnetic field. Here, over most of the range of experimental parameters investigated, as the PFRC-2 axial magnetic field strength was increased, RMF o power decreased, gas fill pressure lowered, or lower atomic mass unit (AMU) fill gas used, the duration of the slow phase lengthened from 50 $μ$s to longer than 10 ms after the RMF o power began. The post-fast-phase maximum n e increases with the fill-gas AMU, exceeding 5 × 10 13 cm –3 for Ar. The slow phase is consistent with atomic physics processes and field-parallel sound-speed losses. The fast phase may be explained by improved axial confinement, possibly augmented by radial or axial contraction of the plasma. Another possible explanation, a large increase in electron temperature, is inconsistent with x-ray emission. The n e behavior is discussed in relation to the E to H transition.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Understanding detachment of the W7-X island divertor

The fundamental behavior of the W7-X island divertor under detached conditions, which has been theoretically predicted with the EMC3-Eirene code, is re-examined here under the experimental conditions achieved so far and compared with the first experimental results. Both simulations and experiments cover a range of divertor configurations and plasma parameters, and show the following common trends: (1) with rising impurity radiation, the target heat load decreases 'uniformly' over the entire target surface in the sense that both the peak and average heat loads can drop by an order of magnitude. Impurity radiation (mainly from intrinsic carbon) occurs primarily at the plasma edge and the resulting negative impact on the stored energy is less than 10%. (2) When the total radiation exceeds a critical level, the target particle flux (the recycling flux Γrecy) begins to fall and can drop by a factor of 3–5 at high radiation levels without an obvious indication of significant volume recombination. (3) While Γ recy decreases, the divertor neutral pressure continues to build up and reaches a maximum, at which point Γ recy has declined significantly. (4) During detachment, the electron temperature at the last closed flux surface falls in a way that is not quantitatively understandable from parallel classical heat conduction processes. This paper presents a physical explanation of the numerical/experimental results described above. Furthermore, using the EMC3-Eirene code as a diagnostic tool, we are able, apparently for the first time, to provide a full quantitative analysis of each transport channel in the island divertor, aiming to clarify how the island divertor plasma self-regulates to maintain particle, energy, and momentum balance under detached conditions.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Solar Cell Metallization Wear is Sensitive to Loading Frequency

Here, we performed cyclic loading of photovoltaic laminates with precracked silicon cells to explore if and how loading frequency and contact pressure influence the ensuing gridline wear-out process. A measurement of parallel resistance across cracked gridlines on a laminated cell coupon was used as the metric for gridline electrical contact degradation. A statistical analysis of variance (ANOVA) analysis of the experimental results discerned that loading frequency is a more significant factor than contact pressure for gridline degradation.

14 SOLAR ENERGY↗

A Two-Stage Decomposition Approach for AC Optimal Power Flow

The alternating current optimal power flow (AC-OPF) problem is critical to power system operations and planning, but it is generally hard to solve due to its nonconvex and large-scale nature. Furthermore, this paper proposes a scalable decomposition approach in which the power network is decomposed into a master network and a number of subnetworks, where each network has its own AC-OPF subproblem. This formulates a two-stage optimization problem and requires only a small amount of communication between the master and subnetworks. The key contribution is a smoothing technique that renders the response of a subnetwork differentiable with respect to the input from the master problem, utilizing properties of the barrier problem formulation that naturally arises when subproblems are solved by a primal-dual interior-point algorithm. Consequently, existing efficient nonlinear programming solvers can be used for both the master problem and the subproblems. The advantage of this framework is that speedup can be obtained by processing the subnetworks in parallel, and it has convergence guarantees under reasonable assumptions. The formulation is readily extended to instances with stochastic subnetwork loads. Numerical results show favorable performance and illustrate the scalability of the algorithm which is able to solve instances with more than 11 million buses.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High energy density picoliter-scale zinc-air microbatteries for colloidal robotics

The recent interest in microscopic autonomous systems, including microrobots, colloidal state machines, and smart dust, has created a need for microscale energy storage and harvesting. However, macroscopic materials for energy storage have noted incompatibilities with microfabrication techniques, creating substantial challenges to realizing microscale energy systems. Here, we photolithographically patterned a microscale zinc/platinum/SU-8 system to generate the highest energy density microbattery at the picoliter (10 −12 liter) scale. The device scavenges ambient or solution-dissolved oxygen for a zinc oxidation reaction, achieving an energy density ranging from 760 to 1070 watt-hours per liter at scales below 100 micrometers lateral and 2 micrometers thickness in size. The parallel nature of photolithography processes allows 10,000 devices per wafer to be released into solution as colloids with energy stored on board. Within a volume of only 2 picoliters each, these primary microbatteries can deliver open circuit voltages of 1.05 ± 0.12 volts, with total energies ranging from 5.5 ± 0.3 to 7.7 ± 1.0 microjoules and a maximum power near 2.7 nanowatts. We demonstrated that such systems can reliably power a micrometer-sized memristor circuit, providing access to nonvolatile memory. We also cycled power to drive the reversible bending of microscale bimorph actuators at 0.05 hertz for mechanical functions of colloidal robots. Additional capabilities, such as powering two distinct nanosensor types and a clock circuit, were also demonstrated. The high energy density, low volume, and simple configuration promise the mass fabrication and adoption of such picoliter zinc-air batteries for micrometer-scale, colloidal robotics with autonomous functions.

Robotics↗