Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 271 records · Page 15

Chromatin Changes in Phytochrome Interacting Factor-Regulated Genes Parallel Their Rapid Transcriptional Response to Light

As sessile organisms, plants must adapt to a changing environment, sensing variations in resource availability and modifying their development in response. Light is one of the most important resources for plants, and its perception by sensory photoreceptors (e.g., phytochromes) and subsequent transduction into long-term transcriptional reprogramming have been well characterized. Chromatin changes have been shown to be involved in photomorphogenesis. However, the initial short-term transcriptional changes produced by light and what factors enable these rapid changes are not well studied. Here, we define rapidly light-responsive, Phytochrome Interacting Factor (PIF) direct-target genes (LRP-DTGs). We found that a majority of these genes also show rapid changes in Histone 3 Lysine-9 acetylation (H3K9ac) in response to the light signal. Detailed time-course analysis of transcript and chromatin changes showed that, for light-repressed genes, H3K9 deacetylation parallels light-triggered transcriptional repression, while for light-induced genes, H3K9 acetylation appeared to somewhat precede light-activated transcript accumulation. However, direct, real-time imaging of transcript elongation in the nucleus revealed that, in fact, transcriptional induction actually parallels H3K9 acetylation. Collectively, the data raise the possibility that light-induced transcriptional and chromatin-remodeling processes are mechanistically intertwined. Histone modifying proteins involved in long term light responses do not seem to have a role in this fast response, indicating that different factors might act at different stages of the light response. This work not only advances our understanding of plant responses to light, but also unveils a system in which rapid chromatin changes in reaction to an external signal can be studied under natural conditions.

59 BASIC BIOLOGICAL SCIENCES↗

Oblique instability of quasi-parallel whistler waves in the presence of cold and warm electron populations

Whistler waves propagating nearly parallel to the ambient magnetic field experience a nonlinear instability due to transverse currents when the background plasma has a population of sufficiently low energy electrons. Intriguingly, this nonlinear process may generate oblique electrostatic waves, including whistlers near the resonance cone with properties resembling oblique chorus waves in the Earth’s magnetosphere. Focusing on the generation of oblique whistlers, earlier analysis of the instability is extended here to the case where low-energy background plasma consists of both a “cold” population with energy of a few eV and a “warm” electron component with energy of the order of 100 eV. This is motivated by spacecraft observations in the Earth’s magnetosphere where oblique chorus waves were shown to interact resonantly with the warm electrons. The main new results are: 1) the instability producing oblique electrostatic waves is sensitive to the shape of the electron distribution at low energies. In the whistler range of frequencies, two distinct peaks in the growth rate are typically present for the model considered: a peak associated with the warm electron population at relatively low wavenumbers and a peak associated with the cold electron population at relatively high wavenumbers; 2) overall, the instability producing oblique whistler waves near the resonance cone persists (with a reduced growth rate) even in the cases where the temperature of the cold population is relatively high, including cases where cold population is absent and only the warm population is included; 3) particle-in-cell simulations show that the instability leads to heating of the background plasma and formation of characteristic plateau and beam features in the parallel electron distribution function in the range of energies resonant with the instability. The plateau/beam features have been previously detected in spacecraft observations of oblique chorus waves. However, they have been attributed to external sources and have been proposed to be the mechanism generating oblique chorus. In the present scenario, the causality link is reversed and the instability generating oblique whistler waves is shown to be a possible mechanism for formation of the plateau and beam features.

79 ASTRONOMY AND ASTROPHYSICS↗

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)↗

Parallel Diffusion Coefficient of Energetic Charged Particles in the Inner Heliosphere from the Turbulent Magnetic Fields Measured by Parker Solar Probe

Diffusion coefficients of energetic charged particles in turbulent magnetic fields are a fundamental aspect of diffusive transport theory but remain incompletely understood. In this work, we use quasi-linear theory to evaluate the spatial variation of the parallel diffusion coefficient κ ∥ from the measured magnetic turbulence power spectra in the inner heliosphere. We consider the magnetic field and plasma velocity measurements from Parker Solar Probe made during Orbits 5–13. The parallel diffusion coefficient is calculated as a function of radial distance from 0.062 to 0.8 au, and the particle energy from 100 keV to 1 GeV. We find that κ ∥ increases exponentially with both heliocentric distance and energy of particles. The fluctuations in κ ∥ are related to the episodes of large-scale magnetic structures in the solar wind. By fitting the results, we also provide an empirical formula of κ ∥ = (5.16 ± 1.22) × 10 18 r 1.17 ± 0.08 E 0.71 ± 0.02 (cm 2 s -1 ) in the inner heliosphere, which can be used as a reference in studying the transport and acceleration of solar energetic particles as well as the modulation of cosmic rays.

79 ASTRONOMY AND ASTROPHYSICS↗

Electron Influence on the Parallel Proton Firehose Instability in 10-moment, Multifluid Simulations

Instabilities driven by pressure anisotropy play a critical role in modulating the energy transfer in space and astrophysical plasmas. For the first time, we simulate the evolution and saturation of the parallel proton firehose instability using a multifluid model without adding artificial viscosity. These simulations are performed using a 10-moment, multifluid model with local and gradient relaxation heat-flux closures in high-β proton–electron plasmas. When these higher-order moments are included and pressure anisotropy is permitted to develop in all species, we find that the electrons have a significant impact on the saturation of the parallel proton firehose instability, modulating the proton pressure anisotropy as the instability saturates. Even for lower β's more relevant to heliospheric plasmas, we observe a pronounced electron energization in simulations using the gradient relaxation closure. Our results indicate that resolving the electron pressure anisotropy is important to correctly describe the behavior of multispecies plasma systems.

79 ASTRONOMY AND ASTROPHYSICS↗

A Scalable Parallel Hypergraph Generator (HyGen)

Graphs are extensively used to model real-world complex systems. An edge in a graph can model pairwise relationships. However, multiway relationships (connections between three or more vertices) are common in many complex systems such as cellular process, image segmentation, and circuit design. A graph edge cannot model multiway relationships. A hypergraph, which can connect more than two vertices, is thus a better option to model multiway relationships. A large-scale hypergraph analysis has the potential to find useful insights from a complex system and assist in knowledge discovery. Currently a limited number of hypergraphs exists that are representative of real-world datasets. Moreover, real-world hypergraph datasets are small in size and inadequate to incorporate future needs. A graph generator that can produce large-scale synthetic hypergraphs can solve the above mentioned problems. In this paper, we present a scalable parallel hypergraph generator (HyGen) based on the Message Passing Interface (MPI) standard. To generate hypergraphs, HyGen takes the following parameter values as inputs: i) number of vertices, ii) number of hyperedges, iii) number of clusters, iv) vertex distribution, v) hyperedge distribution, vi) local cluster cardinality, and vii) global cluster cardinality. We have demonstrated that HyGen can generate hypergraphs of various sizes in a scalable fashion. HyGen takes approximately four minutes to generate a hypergraph with 4.8 million vertices, 1.6 million hyperedges, and 800 clusters using 1,024 processes on a leadership class computing platform. Our strong and weak scaling experiments on supercomputers demonstrate that HyGen can quickly create large-scale hypergraphs in a parallel manner, thus providing a useful capability for hypergraph analysis.

Hasan, S M Shamimul↗

Accelerating Collective Communication in Data Parallel Training across Deep Learning Frameworks

This work develops new techniques within Horovod, a generic communication library supporting data parallel training across deep learning frameworks. In particular, we improve the Horovod control plane by implementing a new coordination scheme that takes advantage of the characteristics of the typical data parallel training paradigm, namely the repeated execution of collectives on the gradients of a fixed set of tensors. Using a caching strategy, we execute Horovod’s existing coordinator-worker logic only once during a typical training run, replacing it with a more efficient decentralized orchestration strategy using the cached data and a global intersection of a bitvector for the remaining training duration. Next, we introduce a feature for end users to explicitly group collective operations, enabling finer grained control over the communication buffer sizes. To evaluate our proposed strategies, we conduct experiments on a world-class supercomputer — Summit. We compare our proposals to Horovod’s original design and observe 2x performance improvement at a scale of 6000 GPUs; we also compare them against tf.distribute and torch.DDP and achieve 12% better and comparable performance, respectively, using up to 1536 GPUs; we compare our solution against BytePS in typical HPC settings and achieve about 20% better performance on a scale of 768 GPUs. Finally, we test our strategies on a scientific application (STEMDL) using up to 27,600 GPUs (the entire Summit) and show that we achieve a near-linear scaling of 0.93 with a sustained performance of 1.54 exaflops (with standard error +- 0.02) in FP16 precision.

Romero, Joshua↗

Drivers for paralleled semiconductor switches

An apparatus includes a plurality of parallel-connected semiconductor switches (e.g., silicon carbide (SiC) metal oxide semiconductor field effect transistors (MOSFETs) or other wide-bandgap semiconductor switches) and a plurality of driver circuits having outputs configured to be coupled to control terminals of respective ones of the plurality of semiconductor switches and configured to drive the parallel-connected semiconductor switches responsive to a common switch state control signal. The driver circuits may have respective different power supplies, which may be adjustable. Respective output resistors may couple respective ones of the driver circuits to respective ones of the semiconductor switches. The output resistors may be adjustable.

Nojima, Geraldo↗

High speed two-dimensional event detection and imaging using an analog interface and a massively parallel processor

A quantitative pulse count (event detection) algorithm with linearity to high count rates is accomplished by combining a high-speed, high frame rate camera with simple logic code run on a massively parallel processor such as a GPU or a FPGA. The parallel processor elements examine frames from the camera pixel by pixel to find and tag events or count pulses. The tagged events are combined to form a combined quantitative event image.

Waugh, Justin↗

Evaluating Portable Parallelization Strategies for Heterogeneous Architectures in High Energy Physics

High-energy physics (HEP) experiments have developed millions of lines of code over decades that are optimized to run on traditional x86 CPU systems. However, we are seeing a rapidly increasing fraction of floating point computing power in leadership-class computing facilities and traditional data centers coming from new accelerator architectures, such as GPUs. HEP experiments are now faced with the untenable prospect of rewriting millions of lines of x86 CPU code, for the increasingly dominant architectures found in these computational accelerators. This task is made more challenging by the architecture-specific languages and APIs promoted by manufacturers such as NVIDIA, Intel and AMD. Producing multiple, architecture-specific implementations is not a viable scenario, given the available person power and code maintenance issues. The Portable Parallelization Strategies team of the HEP Center for Computational Excellence is investigating the use of Kokkos, SYCL, OpenMP, std::execution::parallel and alpaka as potential portability solutions that promise to execute on multiple architectures from the same source code, using representative use cases from major HEP experiments, including the DUNE experiment of the Long Baseline Neutrino Facility, and the ATLAS and CMS experiments of the Large Hadron Collider. This cross-cutting evaluation of portability solutions using real applications will help inform and guide the HEP community when choosing their software and hardware suites for the next generation of experimental frameworks. We present the outcomes of our studies, including performance metrics, porting challenges, API evaluations, and build system integration.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Plastic Parallel Pathways Platform - 4P Model

Global momentum is building towards a circular economy capable of keeping plastics in use and out of waste streams. Given that 79% of all plastic produced since 1950 has accumulated in landfills or the natural environment,rapid implementation of various end-of-life (EoL) management technologies will be needed to reach this target. However, it can be challenging to develop an effective plastic EoL strategy when the available options - chemical or molecular recycling, energy recovery, upcycling, downcycling, closed-loop (plastic-to-plastic) or open-loop (plastic-to-x) recycling, among others - can generate products ranging from low-grade to virgin-quality plastic and from fuels to value-added chemicals. We present a flexible material flow model capable of analyzing the effects of both plastic-to-plastic and plastic-to-x EoL management strategies on the U.S. PET economy. This Plastic Parallel Pathways Platform (4P) assesses the environmental impacts, costs, and circularity of a PET system in which waste is managed through six potential EoL pathways: landfill, incineration with energy recovery, pyrolysis to fuel oil, upcycling to glass fiber reinforced plastic (GFRP), mechanical recycling to low-grade PET, and chemical recycling (glycolysis) to bottle-grade PET. We compare the pathways across multiple metrics using multi-criteria decision analysis (MCDA) and then use a brute force algorithm to predict an optimal combination of EoL pathways to minimize greenhouse gas (GHG) emissions and costs and maximize circularity. This work highlights the need to implement a diverse portfolio of EoL strategies in parallel to enable a PET economy that meets environmental, economic, and circularity requirements simultaneously.

downcycling↗

Parallel performance of algebraic multigrid domain decomposition

Algebraic multigrid (AMG) is a widely used scalable solver and preconditioner for large-scale linear systems resulting from the discretization of a wide class of elliptic PDEs. While AMG has optimal computational complexity, the cost of communication has become a significant bottleneck that limits its scalability as processor counts continue to grow on modern machines. This article examines the design, implementation, and parallel performance of a novel algorithm, algebraic multigrid domain decomposition (AMG-DD), designed specifically to limit communication. The goal of AMG-DD is to provide a low-communication alternative to standard AMG V-cycles by trading some additional computational overhead for a significant reduction in communication cost. Numerical results show that AMG-DD achieves superior accuracy per communication cost compared with AMG, and speedup over AMG is demonstrated on a large GPU cluster.

97 MATHEMATICS AND COMPUTING↗

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING↗

Solving larger maximum clique problems using parallel quantum annealing

Quantum annealing has the potential to find low energy solutions of NP-hard problems that can be expressed as quadratic unconstrained binary optimization problems. However, the hardware of the quantum annealer manufactured by D-Wave Systems, which we consider in this work, is sparsely connected and moderately sized (on the order of thousands of qubits), thus necessitating a minor-embedding of a logical problem onto the physical qubit hardware. The combination of relatively small hardware sizes and the necessity of a minor-embedding can mean that solving large optimization problems is not possible on current quantum annealers. In this research, we show that a hybrid approach combining parallel quantum annealing with graph decomposition allows one to solve larger optimization problem accurately. We apply the approach to the Maximum Clique problem on graphs with up to 120 nodes and 6395 edges.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Parallelized POD-based suboptimal economic model predictive control of a state-constrained Boussinesq approximation

Motivated by an energy efficient building application, we want to optimize a quadratic cost functional subject to the Boussinesq approximation of the Navier-Stokes equations and to bilateral state and control constraints. Since the computation of such an optimal solution is numerically costly, we design an efficient strategy to compute a sub-optimal (but applicationally acceptable) solution with significantly reduced computational effort. We employ an economic Model Predictive Control (MPC) strategy to obtain a feedback control. The MPC sub-problems are based on a linear-quadratic optimal control problem subjected to mixed control and state constraints and a convection-diffusion equation, reduced with proper orthogonal decomposition. Finally, to solve each sub-problem, we apply a primal-dual active set strategy. The method can be fully parallelized, which enables the solution of large problems with real-world parameters.

97 MATHEMATICS AND COMPUTING↗

Spatiotemporal parallelization of an analytical heat conduction model for additive manufacturing via a hybrid OpenMP + MPI approach

The ability to do thermal simulations for entire additive manufacturing builds is a key computational problem facing the additive manufacturing community; however, complex numerical models considering multiple physical phenomena currently do not have the capacity for simulations at this scale. To this end, conduction only analytic models offer a viable approach due to the massive drop in computational expense. In this work, we extend an existing implementation which uses a governing equation which can be evaluated at any point in space and time. This implementation already utilizes OpenMP with a spatial decompositions scheme stemming from a melt pool tracking algorithm. Furthermore, we then combine this with a parallel in time (PinT) approach to make the problem highly parallelizable. The new scheme, which uses MPI for internode communication and OpenMP for intranode communication, is shown to scale very well across multiple computational nodes. This approach results in the ability to simulate the 3D solidification conditions for entire layers of additively manufactured parts in minutes making part scale thermal simulations more practical.

36 MATERIALS SCIENCE↗

QRCODE: Massively parallelized real-time time-dependent density functional theory for periodic systems

We present a new software module, QRCODE (Quantum Research for Calculating Optically Driven Excitations), for massively parallelized real-time time-dependent density functional theory (RT-TDDFT) calculations of periodic systems in the open-source Qbox software package. Our approach utilizes a custom implementation of a fast Fourier transformation scheme that significantly reduces inter-node message passing interface (MPI) communication of the major computational kernel and shows impressive scaling up to 16,344 CPU cores. In addition to improving computational performance, QRCODE contains a suite of various time propagators for accurate RT-TDDFT calculations. As benchmark applications of QRCODE, we calculate the current density and optical absorption spectra of hexagonal boron nitride (h-BN) and photo-driven reaction dynamics of the ozone-oxygen reaction. We also calculate the second and higher harmonic generation of monolayer and multi-layer boron nitride structures as examples of large material systems. Our optimized implementation of RT-TDDFT in QRCODE enables large-scale calculations of real-time electron dynamics of chemical and material systems with enhanced computational performance and impressive scaling across several thousand CPU cores.

97 MATHEMATICS AND COMPUTING↗

Parallel computing for power system climate resiliency: Solving a large-scale stochastic capacity expansion problem with mpi-sppy

Here we propose a nodal stochastic generation and transmission expansion planning model that incorporates the output from high-resolution global climate models through load and generation availability scenarios. We implement our model in Pyomo and perform computational studies on a realistically-sized test case of the California electric grid in a high performance computing environment. We propose model reformulations and algorithm tuning to efficiently solve this large problem using a variant of the Progressive Hedging Algorithm. We utilize the parallelization capabilities and overall versatility of mpi-sppy, exploiting its hub-and-spoke architecture to concurrently obtain inner and outer bounds on an optimal expansion plan. Initial results show that instances with 360 representative days on a system with over 8,000 buses can be solved to within 5% of optimality in under 4 h of wall clock time, a first step towards solving a large-scale power system expansion planning problem across a wide range of climate-informed operational scenarios.

24 POWER TRANSMISSION AND DISTRIBUTION↗