Engineering PapersSearch

SEARCH · Engineering Papers

Results for “parallelism”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Role of perturbed parallel magnetic field effects in predicting turbulent transport in NSTX

This study presents analysis of gyrokinetic simulations on the National Spherical Torus Experiment (NSTX) to investigate the effects of electromagnetic fields on plasma turbulence and transport. The simulations, performed with varying levels of fidelity using the gyrokinetic CGYRO code, include electrostatic (ES), single-field electromagnetic (EM1), and two-field electromagnetic (EM2) models. A detailed comparison across the simulation database reveals that electromagnetic effects increase both predicted growth rates and quasilinear fluxes, with EM2 simulations producing stronger turbulence than ES and EM1 cases. Quasilinear modeling using QLGYRO demonstrates that while the perturbed parallel magnetic field (δB ∥ ) does not drastically affect the total flux at experimental gradients, it leads to a shift in the dominant instability, altering mode structures from microtearing to kinetic ballooning modes (KBMs). The proximity of the plasma profiles to the KBM threshold is explored, with the experimental conditions being near the onset of KBM-driven transport. The KBM, with its large growth rates, is identified as a potential driver of electron temperature flattening, as it can rapidly transport heat across flux surfaces. Performing stability analysis shows core-localized unstable a low- mode that could contribute to the flattening at the early times of the discharge. TGYRO predictive modeling, incorporating both TGLF and QLGYRO, indicates that the inclusion of δB ∥ significantly improves the accuracy of temperature profile predictions in NSTX high-beta plasmas, although challenges remain in modeling the sharp flux discontinuities caused by KBM-driven instabilities.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

VerifyIO: Ensuring Correctness of Consistency Semantics in Parallel I/O

Abstract—High-performance computing (HPC) applications generate and consume substantial amounts of data, typically managed by parallel file systems. These applications access file systems either through the POSIX interface or by using highlevel I/O libraries. While the POSIX consistency model remains dominant in HPC, emerging file systems and popular I/O libraries increasingly adopt alternative consistency models that relax semantics in various ways, creating significant challenges for correctness and portability. This paper addresses these challenges by proposing a trace-driven I/O consistency verification workflow, implemented in our open-source tool, VerifyIO, which collects execution traces, detects data conflicts, and verifies proper synchronization against specified consistency models. Our extensive evaluation of 91 test case executions across three widely used I/O libraries with four I/O consistency models reveals critical consistency issues at both application and implementation levels.

Consistency Semantics

Revealing Parallel Inter‐ and Intra‐Ligand Charge Transfer Dynamics in [Ru(L) 2 (dppz)] 2+ Molecular Lightswitch with N K‐Edge X‐Ray Absorption Spectroscopy

In photoactive metal complexes the localization of photoexcited charges dictates the site of chemical reactivity, but few studies measure the charge redistribution in these systems with spatial precision. Herein, we track the inter- and intra-ligand charge transfer processes that underpin light-driven charge separation in the well-studied “molecular lightswitch” [Ru(bpy) 2 dppz] 2+ (aqueous [Ruthenium II (2,2′-bipyridine)2(dipyrido[3,2-a:2′,3′-c]phenazine)] 2+ [Cl − ] 2 ) by probing the electronic structure of ligand nitrogen atoms in real-time using ultrafast X-ray absorption spectroscopy and first principles calculations. We confirm the localization of excited electron density on the phenazine N atoms of dppz and we newly identify two parallel electron transfer pathways to populate this state. Sub-70 fs electron transfer to the phenazine portion of dppz is observed and attributed to intra-ligand electron transfer following Ru-to-dppz metal-to-ligand charge transfer (MLCT) excitation. This fast charge transfer was not reported in prior ultrafast studies. The slower (ca. 2 ps) charge transfer reported extensively in time-resolved optical absorption and emission studies is reassigned here to inter-ligand electron “hopping” between nearly isoenergetic ligand moieties following Ru-to-bpy MLCT excitation. In conclusion, the results demonstrate much faster charge separation than previously identified in this well-studied system, highlighting how extended azaacene ligand motifs promote the competitive charge transfer processes needed to drive light-driven electron transfer chemistry.

Donor-acceptor systems

Multi-task Parallelism for Robust Pre-training of Graph Foundation Models on Multi-source, Multi-fidelity Atomistic Modeling Data

Graph foundation models using graph neural networks promise sustainable, efficient atomistic modeling. To tackle challenges of processing multi-source, multi-fidelity data during pre-training, recent studies employ multi-task learning, in which shared message passing layers initially process input atomistic structures regardless of source, then route them to multiple decoding heads that predict data-specific outputs. This approach stabilizes pre-training and enhances a model’s transferability to unexplored chemical regions. Preliminary results on approximately four million structures are encouraging, yet questions remain about generalizability to larger, more diverse datasets and scalability on supercomputers. We propose a multi-task parallelism method that distributes each head across computing resources with GPU acceleration. Implemented in the open-source HydraGNN architecture, our method was trained on over 24 million structures from five datasets and tested on the Perlmutter, Aurora, and Frontier supercomputers, demonstrating efficient scaling on all three highly heterogeneous super-computing architectures.

Lupo Pasini, Massimiliano [ORNL] (ORCID:0000000249

Dynamic analysis of fully constrained Cable-Driven Parallel Robots for automated prefabricated component installation

This paper presents a dynamic analysis and validation framework to assess a fully constrained six-anchor Cable-Driven Parallel Robot (CDPR) for automated installation of prefabricated facade components. Compared with conventional eight-anchor systems, the six-anchor configuration simplifies setup and reduces cost, but it also reduces control authority, shrinks the wrench-feasible workspace, and tightens orientation limits. Consequently, it is unclear a priori whether dynamically feasible trajectories exist to move the end effector from pickup to the facade. A constrained trajectory optimization is formulated to enforce the system dynamics, cable-tension bounds, and pose/velocity limits, and the framework is evaluated in simulation at three levels: (i) an idealized reference model, (ii) a lab-scale prototype model incorporating measured anchor misalignments and identified damping, and (iii) a full-scale three-story building model with load decomposition for structural feasibility checks. Across these scenarios, the analysis shows that optimal, constraint-satisfying trajectories exist that move the end effector from pickup to installation while maintaining a near-plumb, level orientation at the final pose. Collectively, this multi-scale dynamic analysis and validation framework supports the deployment readiness of the six-anchor CDPR and provides a prototype-based sensitivity case study of how measured anchor placement deviations affect feasibility.

CDPR

A novel xylosylated fucoglucuronan in Penium reveals structural parallels to rhamnogalacturonan-I and its broad evolutionary footprint in lower plants

Green algae inhabit aquatic environments across the planet and play a crucial role in sustaining the global ecosystem. Ancestors of some Charophytes adapted to terrestrial conditions and eventually evolved into land plants. Extant green algae have inherited traits from their ancestors and evolved into their current morphological and chemical forms, as reflected by their cell walls with distinct shapes and compositions. To illuminate the evolution of plant cell walls and bridge the gap between green algae and land plants, we investigated the charophyte Penium margaritaceum, a close relative of terrestrial plants. We discovered a previously unknown polysaccharide in both its culture medium and cell wall. This polysaccharide, termed xylosylated fucoglucuronan (XFG), possesses a rhamnogalacturonan-I (RG-I)-like backbone composed of repeating [-3-α-Fucp-(1,4)-α-GlcpA-] disaccharides that are extensively xylosylated and acetylated. Surveying approximately 20 non-vascular plants revealed that XFG and RG-I (or related structures) first emerge in certain Chlorophyceae and subsequently co-occur throughout lineages along the evolutionary trajectory to bryophytes, thereby bridging aquatic green algae to early land plants. The striking structural parallels between XFG, RG-I, and ulvan suggest a shared evolutionary origin, offering new insight into how plant cell walls adapted during the transition from marine to freshwater environments and ultimately to land.

Algae

Benchmarking of massively parallel phase-field codes for directional solidification

We present a detailed benchmark comparing two state-of-the-art phase-field implementations for simulating alloy solidification under experimentally relevant conditions. The study investigates the directional solidification of Al-3wt%Cu under high-velocity solidification conditions and SCN-0.46wt% camphor under microgravity conditions from National Aeronautics and Space Administration (NASA) DECLIC-DSI-R experiments. Both codes, one employing finite-difference discretization with uniform mesh and GPU-acceleration (GPU-PF) and the other one employing finite-element discretization with adaptive-mesh and CPU-parallelization (PRISMS-PF), solve the same quantitative phase-field formulation that incorporates an anti-trapping current for the solidification of dilute alloys. We evaluate the predictions of each code for dendritic morphology, primary spacing, and tip dynamics in both 2D and 3D, as well as their numerical convergence and computational performance. While existing benchmark problems have primarily focused on simplified or small-scale simulations, they do not reflect the computational and modeling challenges posed by employing experimentally relevant time and length scales. Our results provide a practical framework for assessing phase-field code performance as well as validating and facilitating their application in integrated computational materials engineering (ICME) workflows that require integration with realistic experimental data.

36 MATERIALS SCIENCE

Massively parallel axisymmetric fluid model for streamer discharges

A highly parallelizable fluid plasma simulation tool based upon the first-order drift-diffusion equations is discussed. Atmospheric pressure plasmas have densities and gradients that require small element sizes in order to accurately simulate the plasm resulting in computational meshes on the order of millions to tens of millions of elements for realistic size plasma reactors. To enable simulations of this nature, parallel computing is required and must be optimized for the particular problem. Here, a finite-volume, electrostatic drift-diffusion implementation for low-temperature plasma is discussed. The implementation is built upon the Message Passing Interface (MPI) library in C++ using Object Oriented Programming. The underlying numerical method is outlined in detail and benchmarked against simple streamer formation from other streamer codes. Electron densities, electric field, and propagation speeds are compared with the reference case and show good agreement. Convergence studies are also performed showing a minimal space step of approximately 4 μm required to reduce relative error to below 1% during early streamer simulation times and even finer space steps are required for longer times. Additionally, strong and weak scaling of the implementation are studied and demonstrate the excellent performance behavior of the implementation up to 100 million elements on 1024 processors. Lastly, different advection schemes are compared for the simple streamer problem to analyze the influence of numerical diffusion on the resulting quantities of interest.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Extreme-scale EV charging infrastructure planning for last-mile delivery using high-performance parallel computing

Here, this paper addresses stochastic charger location and allocation problems under queue congestion for last-mile delivery using electric vehicles (EVs). The objective is to decide where to open charging stations and how many chargers of each type to install, subject to budgetary and waiting-time constraints. We formulate the problem as a mixed-integer non-linear program, where each station-charger pair is modeled as a multiserver queue with stochastic arrivals and service times to capture the notion of waiting in fleet operations. The model is extremely large, with billions of variables and constraints for a typical metropolitan area; even loading the model in solver memory is difficult, let alone solving it. To address this challenge, we develop a Lagrangian-based dual decomposition framework that decomposes the problem by station and leverages parallelization on high-performance computing systems, where the subproblems are solved by using a cutting plane method and their solutions are collected at the master level. We also develop a three-step rounding heuristic to transform the fractional subproblem solutions into feasible integral solutions. Computational experiments on data from the Chicago metropolitan area with hundreds of thousands of households and thousands of candidate stations show that our approach produces high-quality solutions in cases where existing exact methods cannot even load the model in memory. We also analyze various policy scenarios, demonstrating that combining existing depots with newly built stations under multiagency collaboration substantially reduces costs and congestion. These findings offer a scalable and efficient framework for developing sustainable large-scale EV charging networks.

Capacity allocation

Predicting synthetic mRNA stability using massively parallel kinetic measurements, biophysical modeling, and machine learning

Abstract mRNA degradation is a central process that affects all gene expression levels, though it remains challenging to predict the stability of a mRNA from its sequence, due to the many coupled interactions that control degradation rate. Here, we carried out massively parallel kinetic decay measurements on over 50,000 bacterial mRNAs, using a learn-by-design approach to develop and validate a predictive sequence-to-function model of mRNA stability. mRNAs were designed to systematically vary translation rates, secondary structures, sequence compositions, G-quadruplexes, i-motifs, and RppH activity, resulting in mRNA half-lives from about 20 seconds to 20 minutes. We combined biophysical models and machine learning to develop steady-state and kinetic decay models of mRNA stability with high accuracy and generalizability, utilizing transcription rate models to identify mRNA isoforms and translation rate models to calculate ribosome protection. Overall, the developed model quantifies the key interactions that collectively control mRNA stability in bacterial operons and predicts how changing mRNA sequence alters mRNA stability, which is important when studying and engineering bacterial genetic systems.

Cetnar, Daniel P.

Massively parallel reporter assays and mouse transgenic assays provide correlated and complementary information about neuronal enhancer activity

High-throughput massively parallel reporter assays (MPRAs) and phenotype-rich in vivo transgenic mouse assays are two potentially complementary ways to study the impact of noncoding variants associated with psychiatric diseases. Here, we investigate the utility of combining these assays. Specifically, we carry out an MPRA in induced human neurons on over 50,000 sequences derived from fetal neuronal ATAC-seq datasets and enhancers validated in mouse assays. We also test the impact of over 20,000 variants, including synthetic mutations and 167 common variants associated with psychiatric disorders. We find a strong and specific correlation between MPRA and mouse neuronal enhancer activity. Four out of five tested variants with significant MPRA effects affected neuronal enhancer activity in mouse embryos. Mouse assays also reveal pleiotropic variant effects that could not be observed in MPRA. Our work provides a catalog of functional neuronal enhancers and variant effects and highlights the effectiveness of combining MPRAs and mouse transgenic assays.

Kosicki, Michael

Parallelized telecom quantum networking with an ytterbium-171 atom array

The integration of quantum computers and sensors into a quantum network enables new capabilities in quantum information science. Most networks with atom-like qubits operate at visible or near-ultraviolet wavelengths and require conversion to the telecom band for long-distance communication, which reduces efficiency and potentially introduces noise. In this article we report high-fidelity entanglement between ytterbium-171 atoms and optical photons generated directly in the telecommunication band, where fibre loss is low. The nuclear spin of the atom is entangled with a single photon in the time-bin basis, yielding a high atom-measurement-corrected atom–photon Bell state fidelity. This can be further improved by addressing photon measurement errors. By imaging the atom array onto an optical fibre array, we also implement a parallelized networking protocol that can increase the remote entanglement rate proportionately with the number of channels. We also preserve coherence on a memory qubit during operations on communication qubits. These results support the integration of atomic systems into scalable quantum networks.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS

Lab-Scale Cable-Driven Parallel Robot Prototype for Automated Prefabricated Component Manipulation

This paper presents the design and evaluation of a lab-scale cable-driven parallel robot (CDPR) developed as a flexible platform for automated installation of prefabricated components onto exterior building envelopes. Traditional manual installation methods for prefabricated components, which depend on scaffolding, cranes, cherry pickers, and verbal coordination, are not only labor-intensive and error-prone but also face significant limitations in dense urban environments due to site access constraints. To address these challenges, we developed a lab-scale CDPR platform capable of autonomously transporting building envelope components from a designated pickup zone to their target installation location, minimizing the need for human intervention. This study describes the system’s mechanical design, actuation architecture, real-time feedback system, and control strategy of the CDPR, and evaluates its performance in a laboratory environment. The robot’s actuation system uses torque control for end-effector manipulation. The robot’s real-time pose feedback comes from a construction-grade total station and a wireless inertial measurement unit (IMU), which together support precise end-effector control. Experimental results demonstrate the successful integration of the hardware, sensing, state estimation, and control subsystems. Preliminary tests showed that our lab-scale prototype can position the end effector with an error of less than 3 mm, which is a level of precision not previously achieved by existing CDPRs in construction applications. The key findings are twofold: (1) torque-only control is necessary but not sufficient for minimizing final pose error, and (2) incorporating real-time pose feedback can achieve the desired placement accuracy.

Liu, Yifang [Oak Ridge National Laboratory (ORNL),

LibERI—A portable and performant multi-GPU accelerated library for electron repulsion integrals via OpenMP offloading and standard language parallelism

A portable and performant graphics processing unit (GPU)-accelerated library for electron repulsion integral (ERI) evaluation, named LibERI, has been developed and implemented via directive-based (e.g., OpenMP and OpenACC) and standard language parallelism (e.g., Fortran DO CONCURRENT). Offloaded ERIs consist of integrals over low and high contraction s, p, and d functions using the rotated-axis and Rys quadrature methods. GPU codes are factorized based on previous developments with two layers of integral screening and quartet presorting. In this work, the density screening is moved to the GPU to enhance the computational efficacy for large molecular systems. Here, the L-shells in the Pople basis set are also separated into pure S and P shells to increase the ERI homogeneity and reduce atomic operations and the memory footprint. LibERI is compatible with any quantum chemistry drivers supporting the MolSSI Driver Interface. Benchmark calculations of LibERI interfaced with the GAMESS software package were carried out on various GPU architectures and molecular systems. The results show that the LibERI performance is comparable to other state-of-the-art GPU-accelerated codes (e.g., TeraChem and GMSHPC) and, in some cases, outperforms conventionally developed ERI CUDA kernels (e.g., QUICK) while fully maintaining portability.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Nonlinear Poisson–Boltzmann solutions for charged parallel plates: When opposite charges repel

I present an exact solution of the Poisson–Boltzmann equation for two parallel plates and discuss the solution properties. I discuss in more detail plates with opposite charges: In this case, there are two critical separations, L c,1 < L c,2 . For separations less than L c,1 , the force between plates is repulsive. It switches to attractive at L c,1 , but with the electric potential having the same sign on both plates. For L > L c,2 , the force remains attractive, and the potential at the plates has the same sign as the charge on each plate. I also describe charge regulation, determined by pK a , and provide formulas for both the critical distance where oppositely charged plates repel and their charging process. Finally, the implications of these results for the nanoparticle assembly, as driven by electrostatic interactions, are also discussed.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH

Mechanical design of a parallel flexure-based RADSI instrument for curved x-ray mirror metrology

Modern synchrotron x-ray beamlines demand reflective optics with higher surface profile accuracy to achieve diffraction-limited focusing. This necessitates advanced metrology instruments capable of delivering repeatable measurements in the nanometer to sub-nanometer range. Slope ranges exceeding 15 mrad (0.86°) and greater pose significant challenges for mirror metrology using conventional interferometric methods. Here, to address this, we present a new relative angle determinable stitching interferometry instrument featuring a parallel flexure-based mechanical design. This approach enhances vibration and thermal stability while maintaining a compact and lightweight system. Initial measurements of a cylindrical mirror with a 16 m radius of curvature and a slope range of 5 mrad demonstrate nanometer-level repeatability. Comprehensive system characterization suggests the potential for achieving sub-nanometer repeatability with further refinement to the instrument.

36 MATERIALS SCIENCE

Saturation of the kinetic ballooning instability due to the electron parallel nonlinearity

The electron parallel nonlinearity (EPN) is implemented in the gyrokinetic particle-in-cell turbulence code GEM [Y. Chen and S. E. Parker, J. Comp. Phys. 220, 839 (2007)]. Application to the Cyclone Base Case reveals a strong effect of EPN on the saturated heat transport above the kinetic ballooning mode (KBM) threshold. Evidence is provided to show that the strong effect is associated with the electron radial motion due to magnetic fluttering, which turns fine structures of the KBM eigenmode in radius into fine structures in velocity and increases the magnitude of the EPN term in the kinetic equation.

Gyrokinetic simulations