Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Parallel processing”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

Deep Reinforcement Learning-Based Control of Energy Storage for Interarea Oscillation Damping

With the increasing electricity consumption and lack of transmission investment, today's power systems are operated much closer to their limits, raising concerns of inter-area oscillations that deteriorate the system stability. Here, this article presents a novel energy storage placement and control approach for enhanced damping of interarea oscillations. Combining the residual analysis and dominant mode analysis, we are able to identify the advantageous locations for placing energy storage that achieve improved damping performance. To overcome the challenges, such as fixed control parameters and insufficient damping, we propose to use a deep reinforcement learning-based approach for energy storage control. A state-of-the-art guided surrogate-gradient-based evolutionary strategy is used to train a learning agent in a robust, efficient, and reproducible manner. Parallel computing is also adopted to speed up the training process. The proposed strategy has been tested on both medium and large-scale systems. The proposed methods have demonstrated their effectiveness in mitigating various interarea oscillations within a timeframe of 20 s, thereby averting system collapse and enhancing power grid stability effectively.

25 ENERGY STORAGE↗

High-entropy 1D halide perovskite piezoelectrics found by megalibrary synthesis and rapid nonlinear optical screening

Piezoelectric molecular crystals offer excellent compositional and structural tunability and sustainable processability. However, their discovery is slow, primarily due to the serial synthesis and screening processes used. Here, we report an approach that combines massively parallel megalibrary synthesis with scanning second harmonic generation (SHG) microscopy for rapid screening of piezoelectric molecular crystals. Megalibraries consisting of more than 1,000,000 compositionally distinct but positionally encoded TMCM x TMA (1–x) Cd y Pb (1–y) ClzBr (3–z) (TMCM: trimethylchloromethylammonium, TMA: tetramethylammonium; 0 ≤ x ≤ 1, 0 ≤ y ≤ 1, 0 ≤ z ≤ 3) nanocrystals were synthesized. The megalibraries were rapidly screened by SHG microscopy to identify notable noncentrosymmetric structures, which were then tested for piezoelectricity, facilitating discovery of a high-entropy noncentrosymmetric material with a large d 33 (TMCM 0.75 TMA 0.25 Cd 0.75 Pb 0.25 Cl 1.5 Br 1.5 , 42.8 picocoulombs per newton). Furthermore, this approach enabled systematic investigation of the Curie temperature (T C )–composition relationship in the TMCMCdCl z Br (3–z) system, facilitating reverse design of materials with targeted T C . Our work establishes a powerful approach to accelerate the discovery and design of unusual piezoelectrics for next-generation electronics and optics.

Li, Jun [Northwestern University, Evanston, IL (Un↗

Integral Kernel Methods for Nonlinear Parabolic-Elliptic Systems

Nonlinear parabolic-elliptic systems arise in many physical, biological, and chemical phenomena such as chemotaxis, ion transport, self-gravitating particles, and Brownian vortices. Existing methods struggle with the strong coupling and high nonlinearity and nonlocality of some of these systems, especially the ill-conditioned, convection-dominated problems. To overcome numerical difficulties, current approaches rely on initial guesses, preconditioning, or iterative techniques with no convergence guarantees. They might suffer from poor scalability, large memory usage, and difficulty to parallelize. Inspired by the connection of parabolic-elliptic systems to stochastic processes, we introduce a novel meshless, monolithic, and fully explicit method that naturally encapsulates the elliptic and parabolic operators into a single step which updates each node deterministically with global information. By being fully quadrature-based, it avoids solving systems of discretized equations and does not utilize initial guesses or preconditioning, while requiring little memory and being easy to parallelize. We first derive the method in an integral kernel formulation with quadratic complexity in the number of integration nodes and then leverage kernel-independent fast multipole methods (FMM) to present a scalable algorithm with linear complexity. We provide numerical examples for the Poisson-Nernst-Planck equations in one, two, and three dimensions, together with the derivation of the integral kernel for each case. Furthermore, the examples demonstrate the fast convergence and scalability of the FMM-accelerated algorithm, as well as its suitability for convection-dominated problems, making it competitive against traditional PDE solvers.

PDE systems↗

Time reversal-odd effects in QCD and beyond

In this paper we will describe a parallel between phenomenological and symmetry-based descriptions of the process dependence of the time-reversal-odd phenomena in QCD, such as the Sivers effect, with the goal of defining the essential elements that lead to such process dependence in QCD. This will then serve as a starting point to explore possible generalizations in a different gauge theory.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

The Kokkos Ecosystem [Brief]

In 2016/2017, the field of High-Performance Computing (HPC) entered a new era driven by fundamental physics challenges to produce ever more energy and cost-efficient processors. Since the convergence on the Message-Passing Interface (MPI) standard in the mid-1990s, application developers enjoyed a seemingly static view of the underlying machine — that of a distributed collection of homogeneous nodes executing in collaboration. However, after almost two decades of dominance, the sole use of MPI to derive parallelism acted as a limiter to improved future performance. While MPI is widely expected to continue to function as the basic mechanism for communication between compute nodes for the immediate future, additional parallelism is required on the computing node itself if high performance and efficiency goals are to be realized. When reviewing the architectures of the top HPC systems today, the change in paradigm is clear: the compute nodes of the leading machines in the world are either powered by many-core chips with a few dozen cores each, or use heterogeneous designs, where traditional CPUs marshal work to massively parallel compute accelerators which has as many as 200,000 processing threads in flight simultaneously. Complicating matters further for application developers, each processor vendor has its own preferred way of writing code for their architecture.The Kokkos EcoSystem was released by Sandia in 2017 to address this new era in HPC system design by providing a vendor independent performance portable programming system for scientific, engineering, and mathematical software applications written in the C++ programming language. Using Kokkos, application developers can be more productive because they will not have to create and maintain separate versions of their software for each architecture, nor will they have to be experts in each architecture's peculiar requirements. Instead, they will have a single method of programming for the diverse set of modern HPC architectures. While Kokkos started in 2011 as a programming model only, it soon became clear that complex applications needed more. It is also critical to have a portable mathematical functions and developers need tools to debug their applications, gain insight into the performance characteristics of their codes and tune algorithm performance parameters through automated processes. The Kokkos EcoSystem addresses those needs through its three main components: the Kokkos Core programming model, the Kokkos Kernels math library, and the Kokkos Tools project.

97 MATHEMATICS AND COMPUTING↗

Cable-Driven Parallel Robot (CDPR) for Panelized Envelope Retrofits: Feasible Workspace Analysis

Recent decades have seen remarkable progress in the field of robotic-assisted construction. Cable-driven parallel robots (CDPRs) emerge as promising tools for automating construction processes, due to their advantageous features such as scalability, reconfigurability, compact design, and high payload-to-weight ratio. This paper uses a simple static model to determine the feasibility of a CDPR for overclad panel installation in building envelope retrofits. Given that the building facade needs to be a subset of the CDPR’s wrench-feasible workspace, we focus on the sensitivity of the workspace concerning various cable arrangements and CDPR frame sizes (e.g., height and width extensions). Our analysis indicates that no cable arrangement satisfies the requirement of complete facade coverage and avoids cable-to-panel collisions. Thus, frame extension is needed to enhance coverage. However, in densely populated areas where width extension is limited by space constraints, height extension alone is insufficient to guarantee full facade coverage. This paper pioneers the investigation of CDPRs for panelized envelope retrofits, showcasing their advantages and limitations and paving the way for further research and development.

Liu, Yifang↗

Hot Rolling of ZK60 Magnesium Alloy with Isotropic Tensile Properties from Tubing Made by Shear Assisted Processing and Extrusion (ShAPE)

In the present work, we utilized Shear Assisted Processing and Extrusion (ShAPE), a solid-phase processing technique, to extrude hollow tubes of ZK60 Mg alloy. Hot rolling was performed on these as-extruded tubes (after slitting them longitudinally) to thickness reductions of 37%, 68%, and 93% to investigate their viability as rolling feedstock material. EBSD analysis showed the formation of twinned grains in the ShAPE processed material and a gradual re-orientation of the basal texture parallel to the extrusion direction with each rolling step. Moreover, an equiaxed grain size of 5.15 ± 3.39 μm was obtained in the ShAPE extruded material, and the microstructure was retained even after 93% rolling reduction. The rolled sheets also showed excellent tensile strengths and no mechanical anisotropy, a critical characteristic for formability. The unique microstructures developed and their excellent mechanical properties, combined with the ease of scalability of the process, make ShAPE a promising alternative to existing methods for producing rolling feedstock material.

36 MATERIALS SCIENCE↗

Massively parallel and universal approximation of nonlinear functions using diffractive processors

Nonlinear computation is essential for a wide range of information processing tasks, yet implementing nonlinear functions using optical systems remains a challenge due to the weak and power-intensive nature of optical nonlinearities. Overcoming this limitation without relying on nonlinear optical materials could unlock unprecedented opportunities for ultrafast and parallel optical computing systems. Here, we demonstrate that large-scale nonlinear computation can be performed using linear optics through optimized diffractive processors composed of passive phase-only surfaces. In this framework, the input variables of nonlinear functions are encoded into the phase of an optical wavefront—e.g., via a spatial light modulator (SLM)—and transformed by an optimized diffractive structure with spatially varying point-spread functions to yield output intensities that approximate a large set of unique nonlinear functions–all in parallel. We provide proof establishing that this architecture serves as a universal function approximator for an arbitrary set of bandlimited nonlinear functions, also covering wavelength-multiplexed nonlinear functions as well as multi-variate and complex-valued functions that are all-optically cascadable. Our analysis also indicates the successful approximation of typical nonlinear activation functions commonly used in neural networks, including the sigmoid, tanh, ReLU (rectified linear unit), and softplus. We numerically demonstrate the parallel computation of one million distinct nonlinear functions, accurately executed at wavelength-scale spatial density at the output of a diffractive optical processor. Furthermore, we experimentally validated this framework using in situ optical learning and approximated 35 unique nonlinear functions in a single shot using a compact setup consisting of an SLM and an image sensor. These results establish diffractive optical processors as a scalable platform for massively parallel universal nonlinear function approximation, paving the way for new capabilities in analog optical computing based on linear materials.

Rahman, Md Sadman Sakib [University of California,↗

Addressing Load Imbalance in Bioinformatics and Biomedical Applications: Efficient Scheduling across Multiple GPUs

Computational bioinformatics and biomedical applications frequently contain heterogeneously sized units of work or tasks, for instance due to variability in the sizes of biological sequences and molecules. Variable-sized workloads lead to load imbalances in parallel implementations which detract from efficiency and performance. Many modern computing resources now have multiple graphics processing units(GPUs) per computer for acceleration. These multiple GPU resources need to be used efficiently through balancing of workloads across the GPUs. OpenMP is a portable directive-based parallel programming API used ubiquitously in bioscience applications to program CPUs; recently, the use of OpenMP directives for GPU acceleration has become possible. Here, motivated by experiences with imbalanced loads in GPU-accelerated bioinformatics applications, we address the load balancing problem using OpenMP task-to-GPU scheduling combined with OpenMP GPU offloading for multiply heterogeneous workloads – loads with both variable input sizes, and simultaneously, variable convergence rates for algorithms with a stochastic component – scheduled across multiple GPUs. We aim to develop strategies which are both easy to use and have lower overheads, and may be incorporated incrementally in existing programs which already make use of OpenMP for CPU-based threading in order to make use of multi-GPU computers. We test different combinations of input size variability and convergence rate variability, and characterize the effects of these different scenarios on the performance of scheduling strategies across multiple GPUs with OpenMP. We present several dynamic scheduling solutions for different parallel patterns, explore optimizations, and provide publicly available example computational kernels to make these strategies easy to use in programs. This work will enable application developers to efficiently and easily use multiple GPUs for imbalanced workloads found in bioinformatics and biomedical applications.

Thavappiragasam, Mathialakan↗

Roles of Carburization and Oxidation on Transition-Metal Catalyzed Surface Reactions

Many industrial chemical transformation processes are based on reactions catalyzed by transition metals. In some cases the nano-size particles used as catalysts are synthesized simultaneously or right before the initiation of the main catalytic reaction. Synthesis treatments involve complex processes such as reductions followed by calcinations in oxygenated environments coupled to the presence of gases used for the main reaction. Synthesis details are crucial because they have important consequences on the catalyst shape, size, and chemical composition, and may also affect the structure and composition of the catalyst support. All these effects may alter the main catalyzed reaction in unexpected ways. This work focuses on the synthesis of the overall catalytic system that includes catalyst and support taking place in parallel with the initial stages of the main catalyzed reaction. Two important processes are characterized: carburization and oxidation of the transition metal catalyst during synthesis or initial stages of the main reaction. The objective is to understand how the coking mechanism coupled to surface oxidation and changes in the substrate influence the dry reforming of methane taken as a model reaction. The approach uses first principles computational methods to investigate the catalyst synthesis and characterization and catalytic reaction mechanisms. These studies are complemented with experimental characterization done by the group of Dr. Renu Sharma at the National Institute of Standards and Technology (unfunded collaborator) using high resolution transmission electron microscopy. Specific goals are to identify the role of the synthesis process on the initial nanocatalyst structure and composition, and to determine how that structure and composition may evolve after being exposed to the reactants atmosphere. These studies facilitate a systematic analysis of conditions where the nanoparticle composition can be optimized regarding activity and stability. Moreover, the fundamental study of catalyst/support interactions yields valuable insights for other catalytic processes and is a first step toward controlled nanocatalyst synthesis.

36 MATERIALS SCIENCE↗

Mesoscopic Modeling and Rapid Simulation of Incremental Changes in Epidemic Scenarios on GPUs

In simulation-based studies and analyses of epidemics, a major challenge lies in resolving the conflict between fidelity of models and the speed of their simulation. Another related challenge arises in dealing with the large number of what–if scenarios that need to be explored. Here, we describe new computational methods that together provide an approach to dealing with both challenges. A mesoscopic modeling approach is described that strikes a middle ground between macroscopic models based on coupled differential equations and microscopic models built on fine-grained behaviors at the individual entity level. The mesoscopic approach offers the ability to incorporate complex compositions of multiple layers of dynamics even while retaining the potential for aggregate behaviors at varying levels. It also is an excellent match to the accelerator-based architectures of modern computing platforms in which graphical processing units (GPUs) can be exploited for fast simulation via the parallel execution mode of single instruction multiple thread (SIMT). The challenge of simulating a large number of scenarios is addressed via a method of sharing model state and computation across a tree of what–if scenarios that are localized, incremental changes to a large base simulation. A combination of the mesoscopic modeling approach and the incremental what–if scenario tree evaluation has been implemented in the software on modern GPUs. Synthetic simulation scenarios are presented to demonstrate the computational characteristics of our approach. Results from the experiments with large population data, including USA, UK, and India, illustrate the modeling methodology and computational performance on thousands of synthetically generated what–if scenarios. Execution of our implementation scaled to 8192 GPUs of supercomputing platforms demonstrates the ability to rapidly evaluate what–if scenarios several orders of magnitude faster than the conventional methods.

97 MATHEMATICS AND COMPUTING↗

Optimizing the hit finding algorithm for liquid argon TPC neutrino detectors using parallel architectures

Neutrinos are particles that interact rarely, so identifying them requires large detectors which produce lots of data. Processing this data with the computing power available is becoming even more difficult as the detectors increase in size to reach their physics goals. Liquid argon time projection chamber (LArTPC) neutrino experiments are expected to grow in the next decade to have 100 times more wires than in currently operating experiments, and modernization of LArTPC reconstruction code, including parallelization both at data- and instruction-level, will help to mitigate this challenge. The LArTPC hit finding algorithm is used across multiple experiments through a common software framework. In this paper we discuss a parallel implementation of this algorithm. Using a standalone setup we find speedup factors of two times from vectorization and 30–100 times from multi-threading on Intel architectures. The new version has been incorporated back into the framework so that it can be used by experiments. On a serial execution, the integrated version is about 10 times faster than the previous one and, once parallelization is enabled, more speedups comparable to the standalone program are achieved.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Density and Magnetic Field Asymmetric Kelvin‐Helmholtz Instability

Abstract The Kelvin‐Helmholtz (KH) instability can transport mass, momentum, magnetic flux, and energy between the magnetosheath and magnetosphere, which plays an important role in the solar‐wind‐magnetosphere coupling process for different planets. Meanwhile, strong density and magnetic field asymmetry are often present between the magnetosheath (MSH) and magnetosphere (MSP), which could affect the transport processes driven by the KH instability. Our magnetohydrodynamics simulation shows that the KH growth rate is insensitive to the density ratio between the MSP and the MSH in the compressible regime, which is different than the prediction from linear incompressible theory. When the interplanetary magnetic field (IMF) is parallel to the planet's magnetic field, the nonlinear KH instability can drive a double mid‐latitude reconnection (DMLR) process. The total double reconnected flux depends on the KH wavelength and the strength of the lower magnetic field. When the IMF is anti‐parallel to the planet's magnetic field, the nonlinear interaction between magnetic reconnection and the KH instability leads to fast reconnection (i.e., close to Petschek reconnection even without including kinetic physics). However, the peak value of the reconnection rate still follows the asymmetric reconnection scaling laws. We also demonstrate that the DMLR process driven by the KH instability mixes the plasma from different regions and consequently generates different types of velocity distribution functions. We show that the counter‐streaming beams can be simply generated via the change of the flux tube connection and do not require parallel electric fields.

Astronomy & Astrophysics↗

Science of the Van Allen Probes Science Operations Centers

The Van Allen Probes mission operations materialized through a distributed model in which operational responsibility was divided between the Mission Operations Center (MOC) and separate instrument specific SOCs. The sole MOC handled all aspects of telemetering and receiving tasks as well as certain scientifically relevant ancillary tasks. Each instrument science team developed individual instrument specific SOCs proficient in unique capabilities in support of science data acquisition, data processing, instrument performance, and tools for the instrument team scientists. In parallel activities, project scientists took on the task of providing a significant modeling tool base usable by the instrument science teams and the larger scientific community. With a mission as complex as Van Allen Probes, scientific inquiry occurred due to constant and significant collaboration between the SOCs and in concert with the project science team. Planned cross-instrument coordinated observations resulted in critical discoveries during the seven-year mission. Instrument cross-calibration activities elucidated a more seamless set of data products. Specific topics include post-launch changes and enhancements to the SOCs, discussion of coordination activities between the SOCs, SOC specific analysis software, modeling software provided by the Van Allen Probes project, and a section on lessons learned. One of the most significant lessons learned was the importance of the original decision to implement individual team SOCs providing timely and well-documented instrument data for the NASA Van Allen Probes Mission scientists and the larger magnetospheric and radiation belt scientific community.

47 OTHER INSTRUMENTATION↗

Revealing Deformation Mechanisms in Polymer-Grafted Thermoplastic Elastomers via In Situ Small-Angle X-ray Scattering

The tunable properties of thermoplastic elastomers (TPEs), through polymer chemistry manipulations, enable these technologically critical materials to be employed in a broad range of applications. The need to “dial-in” the mechanical properties and responses of TPEs generally requires the design and synthesis of new macromolecules. In these designs, TPEs with nonlinear macromolecular architectures outperform the mechanical properties of their linear copolymer counterparts, but the differences in the deformation mechanism providing enhanced performance are unknown. Here, in situ small-angle X-ray scattering (SAXS) measurements during uniaxial extension reveal distinct deformation mechanisms between a commercially available linear poly(styrene)–poly(butadiene)–poly(styrene) (SBS) triblock copolymer and the grafted SBS version containing grafted poly(styrene) (PS) chains from the poly(butadiene) (PBD) midblock. The neat SBS (φ SBS = 100%) sample deforms congruently with the macroscopic dimensions, with the domain spacing between spheres increasing and decreasing along and transverse to the stretch direction, respectively. At high extensions, end segment pullout from the PS-rich domains is detected, which is indicated by a disordering of SBS. Conversely, the PS-grafted SBS that is 30 vol % SBS and 70% styrene (φ SBS = 30%) exhibits a lamellar morphology, and in situ SAXS measurements reveal an unexpected deformation mechanism. During deformation, there are two simultaneous processes: significant lamellar domain rearrangement to preferentially orient the lamellae planes parallel to the stretch direction and crazing. The samples whiten at high strains as expected for crazing, which corresponds with the emergence of features in the 2D SAXS pattern during stretching consistent with fibril-like structures that bridge the voids in crazes. The significant domain rearrangement in the grafted copolymers is attributed to the new junctions formed across multiple PS domains by the grafting of a single chain. In conclusion, the in situ SAXS measurements provide insights into the enhanced mechanical properties of grafted copolymers that arise through improved physical cross-linking that leads to nanostructure domain reorientation for self-reinforcement and craze formation where fibrils help to strengthen the polymer.

36 MATERIALS SCIENCE↗

Generalized large optics fabrication multiplexing

High precision astronomical optics are manufactured through deterministic computer controlled optical surfacing processes, such as subaperture small tool polishing, magnetorheological finishing, bonnet tool polishing, and ion beam figuring. Due to the small tool size and the corresponding tool influence function, large optics fabrication is a highly time-consuming process. The framework of multiplexed figuring runs for the simultaneous use of two or more tools is presented. This multiplexing process increases the manufacturing efficiency and reduces the overall cost using parallelized subaperture tools.

36 MATERIALS SCIENCE↗

Multiphysics analysis system for heat pipe cooled micro-reactors employing PRAGMA-OpenFOAM-ANLHTP

A multiphysics analysis system for neutronics/thermo-mechanical/heat-pipe analysis of heat pipe cooled micro-reactors was developed using the PRAGMA code as the neutronics engine. PRAGMA, which used to be a GPU-based continuous-energy MC code for power reactor applications, now has an extended geometry package to handle geometries with unstructured meshes generated by the ANSYS Design-Modeler and Meshing. The NVIDIA ray tracing engine OptiX was exploited for efficient neutron transport on unstructured geometry. On the multiphysics side, the open-source CFD tool OpenFOAM and one-dimensional heat pipe analysis code ANLHTP were adopted. The manager-worker system based on the MPI dynamic process management (DPM) model enables efficient coupling of codes employing different parallelization schemes. With all the features, the multiphysics analysis of the one-sixth symmetrical MegaPower 2D core was performed and it demonstrated the benefit of the tight integration of the three-way coupling system and one-to-one geometry coupling strategy. (authors)

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Nonresonant particle acceleration in strong turbulence: Comparison to kinetic and MHD simulations

Collisionless, magnetized turbulence offers a promising framework for the generation of nonthermal high-energy particles in various astrophysical sites. Yet, the detailed mechanism that governs particle acceleration has remained subject to debate. By means of 2D and 3D particle-in-cell, as well as 3D (incompressible) magnetohydrodynamic (MHD) simulations, we test here a recent model of nonresonant particle acceleration in strongly magnetized turbulence, which ascribes the energization of particles to their continuous interaction with the random velocity flow of the turbulence, in the spirit of the original Fermi model. To do so, we compare, for a large number of particles that were tracked in the simulations, the predicted and the observed histories of particles momenta. The predicted history is that derived from the model, after extracting from the simulations, at each point along the particle trajectory, the three force terms that control acceleration: the acceleration of the field line velocity projected along the field line direction, its shear projected along the same direction, and its transverse compressive part. Overall, we find a clear correlation between the model predictions and the numerical experiments, indicating that this nonresonant model can successfully account for the bulk of particle energization through Fermi-type processes in strongly magnetized turbulence. Additionally we also observe that the parallel shear contribution tends to dominate the physics of energization in the particle-in-cell simulations, while in the magnetohydrodynamic incompressible simulation, both the parallel shear and the transverse compressive term provide about equal contributions.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗