Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallelization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 1,369 records · Page 76

Flow regime and Reynolds number variation effects on the mixing behavior of parallel flows

The hydraulic single-phase mixing of three parallel rectangular channels is experimentally investigated at various Reynolds numbers (Re) and flow regime combinations. Particle Image Velocimetry results for seven mixing cases are presented and discussed with varying Re combinations ranging from 1,824 to 20,844. While all cases result in the same Re ratio of ~0.69 between the inner and outer flows, two cases represent multi-regime mixing with the inner-outer regime pair of laminar-transitional and transitional-turbulent, while the other 5 cases are all characteristic of turbulent mixing with varying levels of turbulence. The outer channels initially share characteristics with a backward facing step. The center channel is found to initially behave like a slot jet, but then sees a significant increase in velocity decay. This inner flow velocity decay increased dramatically in the laminar-transitional mixing case, whose centerline velocity decay was ~6 times larger than the decay in the turbulent mixing cases. Second order statistics revealed a consistent mixing layer thickness of ~0.1 hydraulic diameters for all the cases but showed more intense shearing in the multi-regime mixing cases. The combined point and thereby the mixing layer length is determined using centerline velocity decay profiles, which show a much more aggressive mixing in multi-regime flows. Multi-regime mixing demonstrated superior characteristics relative to turbulent mixing due to a more dramatic velocity decay in the inner flow and a shorter mixing length. The contributions of this work include communicating the benefits of multi-regime mixing and providing detailed characterization efforts that can serve future efforts for validating computational models. Here this research also lays the groundwork for future studies aimed at achieving high levels of mixing without a severe penalty in pressure drop.

42 ENGINEERING↗

Tusas: A fully implicit parallel approach for coupled phase-field equations

In this study, we develop a fully-coupled, fully-implicit approach for phase-field modeling of solidification in metals and alloys. Predictive simulation of solidification in pure metals and metal alloys remains a significant challenge in the field of materials science, as microstructure formation during the solidification process plays a critical role in the properties and performance of the solid material. Our simulation approach consists of a finite element spatial discretization of the fully-coupled nonlinear system of partial differential equations at the microscale, which is treated implicitly in time with a preconditioned Jacobian-free Newton-Krylov method. The approach is algorithmically scalable as well as efficient due to an effective preconditioning strategy based on algebraic multigrid and block factorization. We implement this approach in the open-source Tusas framework, which is a general, flexible tool developed in C++ for solving coupled systems of nonlinear partial differential equations. The performance of our approach is analyzed in terms of algorithmic scalability and efficiency, while the computational performance of Tusas is presented in terms of parallel scalability and efficiency on emerging heterogeneous architectures. We demonstrate that modern algorithms, discretizations, and computational science, and heterogeneous hardware provide a robust route for predictive phase-field simulation of microstructure evolution during additive manufacturing.

97 MATHEMATICS AND COMPUTING↗

A Fine-grained Asynchronous Bulk Synchronous parallelism model for PGAS applications

The Partitioned Global Address Space (PGAS) model is well suited for executing irregular applications on cluster-based systems, due to its efficient support for short, one-sided messages. Separately, the actor model has been gaining popularity as a productive asynchronous message-passing approach for distributed objects in enterprise and cloud computing platforms, typically implemented in languages such as Erlang, Scala or Rust. To the best of our knowledge, there has been no past work on using the actor model to deliver both productivity and scalability to irregular PGAS applications with large number of small messages. In this paper, we introduce a new programming system for PGAS applications, in which point-to-point remote operations can be expressed as fine-grained asynchronous actor messages. In our approach, the programmer does not need to worry about programming complexities related to message aggregation and termination detection. Our approach can be viewed as extending the classical Bulk Synchronous Parallelism model with fine-grained asynchronous communications within a phase or superstep. Here, we believe that our approach offers a desirable point in the productivity-performance space for PGAS applications, with more scalable performance and higher productivity relative to past approaches. Specifically, for seven irregular mini-applications from the Bale Kernels and three graph kernels executed using 2048 cores in the NERSC Cori system, our approach shows geometric mean performance improvements of ≥ 20X relative to standard PGAS versions (UPC and OpenSHMEM) while maintaining comparable productivity to those versions.

97 MATHEMATICS AND COMPUTING↗

A parallel variable population multi-objective optimizer for accelerator beam dynamics optimization

The simultaneous optimization of multiple objective functions is needed in many particle accelerator applications. In this paper, we present a parallel evolution based multi-objective optimizer that uses a variable population from generation to generation and an external storage to save good solutions. Two heuristic optimization methods, one uses the unified differential evolution and the other uses the real-coded genetic algorithm, are included in the optimizer to generate next generation candidate solutions, and are compared in the test examples. Finally, as an application, we applied this optimizer to the beam dynamics design optimization of a photoinjector and attained the optimal front solutions after 200 generations with the unified differential evolution offspring production scheme.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

A parallel strategy for density functional theory computations on accelerated nodes

Using the Löwdin orthonormalization of tall-skinny matrices as a proxy-app for wavefunction-based Density Functional Theory solvers, we investigate a distributed memory parallel strategy focusing on Graphics Processing Unit (GPU)-accelerated nodes as available on some of the top ranked supercomputers at the present time. Here we present numerical results in the strong limit regime, as it is particularly relevant for First-Principles Molecular Dynamics. We also examine how matrix product-based iterative solvers provide a competitive alternative to dense eigensolvers on GPUs, allowing to push the strong scaling limit of these computations to a larger number of distributed tasks. Our strategy, which relies on replicated Gram matrices and efficient collective communications using the NCCL library, leads to a time-to-solution under 0.5 s for the Löwdin orthonormalization of a tall-skinny matrix of 3000 columns on Summit at Oak Ridge Leadership Facility (OLCF). Given the similarity in computational operations between one iteration of a DFT solver and this proxy-app, this shows the possibility of solving accurately the DFT equations well under a minute for 3000 electronic wave functions, and thus perform First-Principles molecular dynamics of physical systems much larger than traditionally solved on CPU systems.

97 MATHEMATICS AND COMPUTING↗

The heat transfer coefficient associated with a moving packed bed of silica particles flowing through parallel plates

Concentrating Solar Power (CSP) with thermal energy storage has the potential to be a renewable energy technology with long duration, inexpensive energy storage. Higher temperature operation increases the conversion efficiency and reduces the cost of energy storage. Several emerging CSP designs utilize particles as the solar receiver due to their high temperature stability and low cost. In some designs, the particles are also used as the thermal energy storage (TES) media. In either configuration, an energy transfer is required between the hot particles and the working fluid in the power cycle in a Particle-to-Fluid Heat Exchanger (PtFHX). Understanding the heat transfer between a moving packed bed and a stationary surface is critical to the successful design of a PtFHX. In this paper, a test facility is described in which a moving packed bed of silica sand with particle size 100–600 μm is introduced into the channel formed by two parallel plates, one of which is heated. The effective static thermal conductivity of the particles used for the test are separately measured over the entire range of test temperatures. The inlet and outlet bulk temperatures of the particle flow are measured as are the surface temperatures at several axial locations along the centerline of the plate. The result is the measurement of heat transfer coefficient as a function of temperature for several velocities. The uncertainty of the measurements is presented and the results are compared to model results found in the literature.

14 SOLAR ENERGY↗

Extreme-scale EV charging infrastructure planning for last-mile delivery using high-performance parallel computing

Here, this paper addresses stochastic charger location and allocation problems under queue congestion for last-mile delivery using electric vehicles (EVs). The objective is to decide where to open charging stations and how many chargers of each type to install, subject to budgetary and waiting-time constraints. We formulate the problem as a mixed-integer non-linear program, where each station-charger pair is modeled as a multiserver queue with stochastic arrivals and service times to capture the notion of waiting in fleet operations. The model is extremely large, with billions of variables and constraints for a typical metropolitan area; even loading the model in solver memory is difficult, let alone solving it. To address this challenge, we develop a Lagrangian-based dual decomposition framework that decomposes the problem by station and leverages parallelization on high-performance computing systems, where the subproblems are solved by using a cutting plane method and their solutions are collected at the master level. We also develop a three-step rounding heuristic to transform the fractional subproblem solutions into feasible integral solutions. Computational experiments on data from the Chicago metropolitan area with hundreds of thousands of households and thousands of candidate stations show that our approach produces high-quality solutions in cases where existing exact methods cannot even load the model in memory. We also analyze various policy scenarios, demonstrating that combining existing depots with newly built stations under multiagency collaboration substantially reduces costs and congestion. These findings offer a scalable and efficient framework for developing sustainable large-scale EV charging networks.

Capacity allocation↗

Internal Standard Triggered-Parallel Reaction Monitoring Mass Spectrometry Enables Multiplexed Quantification of Candidate Biomarkers in Plasma

Despite advances in proteomic technologies, clinical translation of plasma biomarkers remains low, partly due to a major bottleneck between the discovery of candidate biomarkers and costly clinical validation studies. Due to a dearth of multiplexable assays, generally only a few candidate biomarkers are tested, and the validation success rate is accordingly low. Previously, mass spectrometry-based approaches have been used to fill this gap but feature poor quantitative performance and were generally limited to hundreds of proteins. Here, we demonstrate the capability of an internal standard triggered-parallel reaction monitoring (IS-PRM) assay to greatly expand the numbers of candidates that can be tested with improved quantitative performance. The assay couples immunodepletion and fractionation with IS-PRM and was developed and implemented in human plasma to quantify 5176 peptides representing 1314 breast cancer biomarker candidates. Characterization of the IS-PRM assay demonstrated the precision (median % CV of 7.7%), linearity (median R 2 > 0.999 over 4 orders of magnitude), and sensitivity (median LLOQ < 1 fmol, approximately) to enable rank-ordering of candidate biomarkers for validation studies. Using three plasma pools from breast cancer patients and three control pools, 893 proteins were quantified, of which 162 candidate biomarkers were verified in at least one of the cancer pools and 22 were verified in all three cancer pools. The assay greatly expands capabilities for quantification of large numbers of proteins and is well suited for prioritization of viable candidate biomarkers.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Modified Energy Span Analysis of Catalytic Parallel Pathways and Selectivity

Mechanistic modeling provides vital insights into catalytic reactions. To analyze complex reaction networks with parallel pathways, we leverage the graph theory approach of the Energy Span Model (ESM) to develop a modified energy span analysis (MESA). A new method of cycle plots is proposed to perform reaction pathways analysis visually. We demonstrate this method on two published models: one describing carbon monoxide oxidation and the other simulating ethylene conversion to propanal via hydroformylation or ethane via hydrogenation. Fundamental insights explain kinetic observables, such as a reactant’s negative reaction order. General principles are revealed, such as rate-determining surface species being outside the primary reaction flux cycle and pathway selectivity being a purely kinetic property when reaction conditions are not near equilibrium. Lastly, we demonstrate MESA’s consistency with published microkinetic modeling results, highlighting this technique’s extension of the ESM to heterogeneous catalysts using collision theory to describe adsorption steps and concentration effects.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Parallel Implementation of Nonadditive Gaussian Process Potentials for Monte Carlo Simulations

A strategy is presented to implement Gaussian process potentials in molecular simulations through parallel programming. Attention is focused on the three-body nonadditive energy, though all algorithms extend straightforwardly to the additive energy. The method to distribute pairs and triplets between processes is general to all potentials. Results are presented for a simulation box of argon, including full box and atom displacement calculations, which are relevant to Monte Carlo simulation. Data on speed-up are presented for up to 120 processes across four nodes. A 4-fold speed-up is observed over five processes, extending to 20-fold over 40 processes and 30-fold over 120 processes.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Solitary Magnetic Structures at Quasi–Parallel Collisionless Shocks: Formation

Solitary magnetic structures known as SLAMS (short large-amplitude magnetic structures) have been considered as essential elements of collisionless shocks with quasi-parallel geometries. Yet the physics underlying their formation remains an open question. In this paper, we use measurements from the magnetospheric multiscale mission combined with fully kinetic simulations to study the formation of SLAMS. We find that gyro-resonance between solar wind ions and right-hand circularly polarized electromagnetic waves results in magnetic field amplification. Gyro-trapping by the growing magnetic field builds up the plasma density that further enhances the current and field. The solitary nature of SLAMS stems from a beat-like magnetic field envelope where the maximum sets the initial location for nonlinear growth. Our results present a conceptual advance on SLAMS, and may shed new light on the open question of magnetic field amplification at astrophysical shocks.

58 GEOSCIENCES↗

Global Hybrid Simulations of Interaction Between Interplanetary Rotational Discontinuity and Bow Shock/Magnetosphere: Can Ion-Scale Magnetic Reconnection be Driven by Rotational Discontinuity Downstream of Quasi-Parallel Shock?

Ion-scale magnetic reconnection has been observed downstream of the terrestrial quasi-parallel (Q-∥) shock. Whether it is driven by interplanetary discontinuities or turbulent Q-∥ shock, however, is unclear. Using three-dimensional global hybrid simulation, we investigate the generation of magnetic reconnection downstream of the Q-∥ shock, while an interplanetary rotational discontinuity (RD) is launched to the bow shock. Cases with various solar wind Alfvén Mach numbers, M A = 3.0 to 8, and propagation directions of the RD are presented. The propagation direction n is assumed to be in the GSE xz plane and pointing earthward, with n = (-sin(θ 12 /2),0,-cos(θ 12 /2)), where θ 12 is the angle between the upstream (B1) and downstream (B2) magnetic fields across the transmitted RD. It is found that magnetic reconnection occurs inside the RD downstream of the Q-∥ shock, forming flux ropes extending along the dawn-dusk direction about tens of ion inertial lengths. Large-amplitude low-frequency waves originated from the Q-∥ shock lead to the bending and squeezing of the field lines around the RD, which play an important role in triggering reconnection inside the RD. As the RD impacts the dayside magnetopause, magnetopause reconnection takes place between the field lines behind the RD and geomagnetic field lines. Nevertheless, no reconnection is found downstream of the Q-∥ shock itself or outside the RD in the magnetosheath. The existent and structure of reconnection in the magnetosheath are found to strongly depend on the parameters M A and n. Our simulation shows that ion-scale magnetic reconnection is driven by an external driver in the form of the compression of an RD around the bow shock and in the magnetosheath, rather than caused by the turbulent Q-∥ shock alone.

79 ASTRONOMY AND ASTROPHYSICS↗

Predicting synthetic mRNA stability using massively parallel kinetic measurements, biophysical modeling, and machine learning

Abstract mRNA degradation is a central process that affects all gene expression levels, though it remains challenging to predict the stability of a mRNA from its sequence, due to the many coupled interactions that control degradation rate. Here, we carried out massively parallel kinetic decay measurements on over 50,000 bacterial mRNAs, using a learn-by-design approach to develop and validate a predictive sequence-to-function model of mRNA stability. mRNAs were designed to systematically vary translation rates, secondary structures, sequence compositions, G-quadruplexes, i-motifs, and RppH activity, resulting in mRNA half-lives from about 20 seconds to 20 minutes. We combined biophysical models and machine learning to develop steady-state and kinetic decay models of mRNA stability with high accuracy and generalizability, utilizing transcription rate models to identify mRNA isoforms and translation rate models to calculate ribosome protection. Overall, the developed model quantifies the key interactions that collectively control mRNA stability in bacterial operons and predicts how changing mRNA sequence alters mRNA stability, which is important when studying and engineering bacterial genetic systems.

Cetnar, Daniel P.↗

Massively parallel reporter assays and mouse transgenic assays provide correlated and complementary information about neuronal enhancer activity

High-throughput massively parallel reporter assays (MPRAs) and phenotype-rich in vivo transgenic mouse assays are two potentially complementary ways to study the impact of noncoding variants associated with psychiatric diseases. Here, we investigate the utility of combining these assays. Specifically, we carry out an MPRA in induced human neurons on over 50,000 sequences derived from fetal neuronal ATAC-seq datasets and enhancers validated in mouse assays. We also test the impact of over 20,000 variants, including synthetic mutations and 167 common variants associated with psychiatric disorders. We find a strong and specific correlation between MPRA and mouse neuronal enhancer activity. Four out of five tested variants with significant MPRA effects affected neuronal enhancer activity in mouse embryos. Mouse assays also reveal pleiotropic variant effects that could not be observed in MPRA. Our work provides a catalog of functional neuronal enhancers and variant effects and highlights the effectiveness of combining MPRAs and mouse transgenic assays.

Kosicki, Michael↗

On-chip parallel processing of quantum frequency comb

Abstract The frequency degree of freedom of optical photons has been recently explored for efficient quantum information processing. Significant reduction in hardware resources and enhancement of quantum functions can be expected by leveraging the large number of frequency modes. Here, we develope an integrated photonic platform for the generation and parallel processing of quantum frequency combs (QFCs). Cavity-enhanced parametric down-conversion with Sagnac configuration is implemented to generate QFCs with identical spectral distributions. On-chip quantum interference of different frequency modes is simultaneously realized with the same photonic circuit. High interference visibility is maintained across all frequency modes with the identical circuit setting. This enables the on-chip reconfiguration of QFCs. By deterministically separating QFCs without spectral filtering, we further demonstrate high-dimensional Hong-Ou-Mandel effect. Our work provides the critical step for the efficient implementation of quantum information processing with integrated photonics using the frequency degree of freedom.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Ion and Electron Acoustic Bursts during Anti-Parallel Reconnection Driven by Lasers

Magnetic reconnection converts magnetic energy into thermal and kinetic energy in plasma. Among the numerous candidate mechanisms, ion acoustic instabilities driven by the relative drift between ions and electrons (or equivalently, electric current) have been suggested to play a critical role in dissipating magnetic energy in collisionless plasmas. However, their existence and effectiveness during reconnection have not been well understood due to ion Landau damping and difficulties in resolving the Debye length scale in the laboratory. We report a sudden onset of ion acoustic bursts measured by collective Thomson scattering in the exhaust of anti-parallel magnetically driven reconnection using high-power lasers. The ion acoustic bursts are followed by electron acoustic bursts with electron heating and bulk acceleration. We reproduce these observations with one- and two-dimensional particle-in-cell simulations in which an electron outflow jet drives ion acoustic instabilities, forming double layers. These layers induce electron two-stream instabilities that generate electron acoustic bursts and energize electrons. Our results demonstrate the importance of ion and electron acoustic dynamics during reconnection when ion Landau damping is ineffective, a condition applicable to a range of astrophysical plasmas including near-Earth space, stellar flares and black hole accretion engines.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Parallelized telecom quantum networking with an ytterbium-171 atom array

The integration of quantum computers and sensors into a quantum network enables new capabilities in quantum information science. Most networks with atom-like qubits operate at visible or near-ultraviolet wavelengths and require conversion to the telecom band for long-distance communication, which reduces efficiency and potentially introduces noise. In this article we report high-fidelity entanglement between ytterbium-171 atoms and optical photons generated directly in the telecommunication band, where fibre loss is low. The nuclear spin of the atom is entangled with a single photon in the time-bin basis, yielding a high atom-measurement-corrected atom–photon Bell state fidelity. This can be further improved by addressing photon measurement errors. By imaging the atom array onto an optical fibre array, we also implement a parallelized networking protocol that can increase the remote entanglement rate proportionately with the number of channels. We also preserve coherence on a memory qubit during operations on communication qubits. These results support the integration of atomic systems into scalable quantum networks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗