Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Nonlinear simulations of GAEs in NSTX-U

A set of nonlinear simulations has been performed in order to study the nonlinear evolution of unstable global Alfvén eigenmodes in the National Spherical Torus Experiment-Upgrade (NSTX-U). Results of the single toroidal mode number, n, simulations are compared with a full nonlinear simulation (all toroidal harmonics included). In single-n simulations, the conservation of two integrals of motion of a particle in a cyclotron resonance with a monochromatic wave is demonstrated, resulting in a one-dimensional evolution of the particle distribution in (E,μ,pϕ) phase-space. Nonlinear simulations (both single-n and full nonlinear) show a significant redistribution of the resonant fast ions, especially in the pitch parameter. Thus, the changes in the resonant particle's parallel and perpendicular energies can be several times larger than the total particle energy change, with only a small fraction transferring into the excitation of the mode itself. This implies that even a relatively small amplitude mode can significantly modify the beam distribution in the resonant region. For the NSTX-U case considered, the single-n simulation results are close to full nonlinear simulation only for the most unstable mode, in which case the saturation amplitudes and changes in the fast ion distribution are comparable. In contrast, peak amplitudes of subdominant modes in all-n simulations are smaller by a factor of 3–10 compared to single-n runs due to the flattening of the beam ion distribution by the fastest growing mode.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

First observations of solar halo gamma rays over a full solar cycle

We analyze 15 years of Fermi-LAT data and produce a detailed model of the Sun’s inverse-Compton scattering emission (solar halo), which is powered by interactions between ambient cosmic-ray electrons and positrons with sunlight. By developing a novel analysis method to analyze moving sources, we robustly detect the solar halo at energies between 31.6 MeV and 100 GeV, and angular extensions up to 45° from the Sun, providing new insight into spatial regions where there are no direct measurements of the Galactic cosmic-ray flux. The large statistical significance of our signal allows us to subdivide the data and provide the first 𝛾-ray probes into the time variation and azimuthal asymmetry of the solar modulation potential, finding time-dependent changes in solar modulation both parallel and perpendicular to the ecliptic plane. Our results are consistent with (but with independent uncertainties from) local cosmic-ray measurements, unlocking new probes into astrophysical processes near the solar surface.

79 ASTRONOMY AND ASTROPHYSICS

A Practical Framework for Simulating Time-Resolved Spectroscopy Based on a Real-Time Dyson Expansion

Time-resolved spectroscopy is a powerful tool for probing electron dynamics in molecules and solids, revealing transient phenomena on subfemtosecond time scales. The interpretation of experimental results is often enhanced by parallel numerical studies, which can provide insight and validation for experimental hypotheses. However, developing a theoretical framework for simulating time-resolved spectra remains a significant challenge. The most suitable approach involves the many-body nonequilibrium Green's function formalism, which accounts for crucial dynamical many-body correlations during time evolution. While these dynamical correlations are essential for observing emergent behavior in time-resolved spectra, they also render the formalism prohibitively expensive for large-scale simulations. Substantial effort has been devoted to reducing this computational cost─through approximations and numerical techniques─while preserving the key dynamical correlations. The ultimate goal is to enable first-principles simulations of time-dependent systems ranging from small molecules to large, periodic, multidimensional solids. Here, in this perspective, we outline key challenges in developing practical simulations for time-resolved spectroscopy, with a particular focus on Green's function methodologies. We highlight a recent advancement toward a scalable framework: the real-time Dyson expansion (RT-DE) [Phys. Rev. Lett. 2024, 133, 226902]. We introduce the theoretical foundation of RT-DE and discuss strategies for improving scalability, which have already enabled simulations of system sizes beyond the reach of previous fully dynamical approaches. We conclude with an outlook on future directions for extending RT-DE to first-principles studies of dynamically correlated, nonequilibrium systems.

Reeves, Cian C. [Univ. of California, Santa Barbar

Rotational coherence dominates early-time dynamics and produces long-time revivals in the S 2 state of azulene

Here, the ultrafast dynamics of azulene have been debated for decades, with reported picosecond decay constants variously attributed to intramolecular vibrational redistribution (IVR), internal conversion, or rotational dephasing. Using polarization- and femtosecond time-resolved resonance-enhanced multiphoton ionization spectroscopy with a nanosecond delay window, we disentangle this long-standing inconsistency and show that the early 2–5 ps decay component arises entirely from the rotational dephasing of an excited-state wavepacket. Identical time constants extracted from the decay of the parallel signal and the rise of the perpendicular signal across multiple vibronic origins provide an unambiguous rotational anisotropy signature, eliminating the need for IVR-based interpretations. Extending the measurement window to 1.3 ns reveals well-structured J-type and C-type rotational coherence revivals in S 2 azulene on top of the well-documented fluorescence decay, demonstrating that both the short- and long-time dynamics contain information about the coherent rotational dynamics. These results show that azulene, and by extension polycyclic aromatic hydrocarbons, can sustain structured rotational coherence deep into the nanosecond regime, positioning PAHs as model systems for quantum-coherent wavepacket dynamics and providing a framework for probing coherence, decoherence, and rotational control in electronically rich molecular systems.

Zhan, Jie [University of Georgia, Athens, GA (Unit

Relative molecular orientation can impact the onset of plasticity in molecular crystals

Abstract Creating or moving dislocations is the first step to dissipating mechanical energy via plastic deformation under contact loading. In molecular crystals there is both a lattice that defines crystal orientation and a relative orientation of the basis of the molecules. We define a normalization parameter which relates strain at yield, the hardness of the bulk crystal, and a distance parameter analogous to a Burgers vector that nominally predicts the relative ease of initiating plasticity in this broad class of materials. Analyzing the yield behavior of 10 different molecular crystals of varying space groups shows the inter-molecular orientation predicts the experimentally observed applied stress needed to nucleate dislocations. When molecules are oriented ‘parallel’ relative to one another the normalized maximum shear stress at the onset of plasticity is on the order of 3–5 times lower than when molecules within the crystal are ‘anti-parallel’, and molecules with a more equiaxed shape fall in between these bounds. This provides an initial indication of a structural feature which predicts the relative ease of initiating plasticity during contact loading in molecular crystals.

36 MATERIALS SCIENCE

Virtual Time III, Part 3: Throttling and Message Cancellation

This is Part 3 of a trio of papers that unify in a natural way the two historically distinct parallel discrete event synchronization paradigms, optimistic and conservative, combining the best properties of both into a single framework called Unified Virtual Time (UVT). In this part, we survey the synchronization effects that can be achieved by restricting to corner cases the relationships permitted among the control variables, GVT, CVT, TVT, and LVT, which were defined in Part 1. Here we also survey various throttling policies from the literature and describe how they can be implemented in UVT by controlling the value of TVT, including policies that can take advantage of rollback in addition to LP blocking. A significant result is a new category of efficient and higher precision throttling algorithms for optimistic execution that are based on optimistic lookahead, defined in a way that is symmetric to what we now call the conservative lookahead information that is traditionally used for conservative synchronization. Finally, we present a novel algorithm allowing the choice between lazy and aggressive cancellation to be made on a message-by-message basis using either external logic expressed in the model code, or policy code internal to the simulator, or a mixture of both.

throttling

Rate-induced aging effects on Parallel-Plate Avalanche Counter (PPAC) caused by heavy ion beams

The Facility for Rare Isotope Beams (FRIB) is one of the premier scientific user facilities for nuclear science with radioactive beams, capable of producing most (approximately 80%) of the isotopes expected to exist, from oxygen to uranium, at energies up to 200 MeV/u. With the increase in beam power from the present 10 kW to the planned 400 kW, FRIB experiments are about to enter a new era. An unprecedented rate capability as well as stable performance of all the planned instrumentation intended for beam diagnostics and beam tuning is required at the expected high beam intensities (> 1 MHz). A summary of aging phenomena at high heavy-ion beam rates observed in the Advanced Rare Isotope Separator (ARIS) detectors for beam diagnostics, including Parallel Plate Avalanche Counters (PPAC) and plastic scintillation for time-of-flight measurements, is discussed. Current research and development project to mitigate rate-induced aging are presented.

Aging effects

Tula: Optimizing Time, Cost, and Generalization in Distributed Large-Batch Training

Distributed training increases the number of batches processed per iteration either by scaling-out (adding more nodes) or scaling-up (increasing the batch-size). However, the largest configuration does not necessarily yield the best performance. Horizontal scaling introduces additional communication overhead, while vertical scaling is constrained by computation cost and device memory limits. Thus, simply increasing the batch-size leads to diminishing returns: training time and cost decrease initially but eventually plateaus, creating a knee-point in the time/cost vs. batch-size pareto curve. The optimal batch-size therefore depends on the underlying model, data and available compute resources. Large batches also suffer from worse model quality due to the well-known “generalization gap”. In this paper, we present Tula, an online service that automatically optimizes time, cost, and convergence quality for large-batch training of convolutional models. It combines parallel-systems modeling with statistical performance prediction to identify the optimal batchsize. Tula predicts training time and cost within 7.5−14% error across multiple models, and achieves up to 20× overall speedup and improves test accuracy by ≈9% on average over standard large-batch training on various vision tasks, thus successfully mitigating the generalization gap and accelerating training at the same time.

Tyagi, Sahil [ORNL] (ORCID:0009000783144745)

High-throughput synthesis of high-entropy alloys via parallelized electric field assisted sintering

Materials discovery and design is an expensive and time-consuming process, though necessary to advance many engineering fields. In this work, a novel tooling design is utilized in conjunction with electric field assisted sintering (EFAS) to effectively create a new high-throughput synthesis technique: parallelized EFAS. Through this technique, a wide range of material compositions and geometries can be synthesized in parallel as isolated samples or as part of contiguous arrays. Multiple tooling designs are explored to examine both the flexibility and limitations of the technique. A series of increasing complex alloys is produced simultaneously using in situ alloying, beginning with pure Ni and adding equimolar constituents up to the septenary high-entropy alloy AlCoCrCuFeMnNi. Microstructural characterization reveals each sample is effectively fully dense and chemically homogenous while exhibiting phases in agreement with CALPHAD predictions. Scalability of parallelized EFAS is then experimentally demonstrated and the implications for materials discovery and automation are discussed.

36 - MATERIALS SCIENCE

1D model of SOL with self consistent calculation of non-coronalradiation and applications to shallow field line angles

A 1D model of the scrape off layer (SOL) is created with a physics based calculation of the impurity radiation and the inclusion of the perpendicular heat flux q ⊥ compared to previous 1D models. The calculation of the impurity radiation includes the addition of a model using parallel and perpendicular motion to calculate the impurity confinement time τ and the inclusion of charge exchange between neutral hydrogen and impurity ions. Reasonable agreement is found between the model and the 2D fluid code SOLPS-ITER. Additionally, an improvement is shown over the constant τ used in previous works The model is used to examine the effects of shallow field line angles at ITER levels of parallel heat flux (q || ). A large increase in radiation results in an increase in the upstream parallel heat flux which can still be in detachment at shallow field line angles for both q ⊥ and the calculated τ for traditional impurities (argon and nitrogen) as well as boron. The physics behind this increase in radiation at shallow field line angles is shown to be a decrease in τ and a change in the temperature profile with the inclusion of q ⊥ .

ITER

ZERNIPAX: A fast and accurate Zernike polynomial calculator in Python

Zernike polynomials serve as an orthogonal basis on the unit disc, and have proven to be effective in optics simulations, astrophysics, and more recently in plasma simulations. Unlike Bessel functions, Zernike polynomials are inherently finite and smooth at the disc center (r=0), ensuring continuous differentiability along the axis. This property makes them particularly suitable for simulations, requiring no additional handling at the origin. We developed ZERNIPAX, an open-source Python package capable of utilizing CPU/GPUs, leveraging Google's JAX package and available on GitHub as well as the Python software repository PyPI. Furthermore, our implementation of the recursion relation between Jacobi polynomials significantly improves computation time compared to alternative methods by use of parallel computing while still performing more accurately for high-mode numbers.

Astrophysics

Two-level overlapping additive Schwarz preconditioner for training scientific machine learning applications

In this work we introduce a novel two-level overlapping additive Schwarz preconditioner for accelerating the training of scientific machine learning applications. The design of the proposed preconditioner is motivated by the nonlinear two-level overlapping additive Schwarz preconditioner. The neural network parameters are decomposed into groups (subdomains) with overlapping regions. In addition, the network’s feed-forward structure is indirectly imposed through a novel subdomain-wise synchronization strategy and a coarse-level training step. Through a series of numerical experiments, which consider physicsinformed neural networks and operator learning approaches, we demonstrate that the proposed two-level preconditioner significantly speeds up the convergence of the standard (LBFGS) optimizer while also yielding more accurate machine learning models. Moreover, the devised preconditioner is designed to take advantage of model-parallel computations, which can further reduce the training time.

97 MATHEMATICS AND COMPUTING

Distributed Stochastic Optimization of a Neural Representation Network for Time-Space Tomography Reconstruction

4D time-space reconstruction of dynamic events or deforming objects using X-ray computed tomography (CT) is an important inverse problem in non-destructive evaluation. Conventional back-projection based reconstruction methods assume that the object remains static for the duration of several tens or hundreds of X-ray projection measurement images (reconstruction of consecutive limited-angle CT scans). However, this is an unrealistic assumption for many in-situ experiments that causes spurious artifacts and inaccurate morphological reconstructions of the object. To solve this problem, we propose to perform a 4D time-space reconstruction using a distributed implicit neural representation (DINR) network that is trained using a novel distributed stochastic training algorithm. Our DINR network learns to reconstruct the object at its output by iterative optimization of its network parameters such that the measured projection images best match the output of the CT forward measurement model. Here, we use a forward measurement model that is a function of the DINR outputs at a sparsely sampled set of continuous valued 4D object coordinates. Unlike previous neural representation architectures that forward and back propagate through dense voxel grids that sample the object's entire time-space coordinates, we only propagate through the DINR at a small subset of object coordinates in each iteration resulting in an order-of-magnitude reduction in memory and compute for training. DINR leverages distributed computation across several compute nodes and GPUs to produce high-fidelity 4D time-space reconstructions. We use both simulated parallel-beam and experimental cone-beam X-ray CT datasets to demonstrate the superior performance of our approach.

36 MATERIALS SCIENCE

Optimizing inference of segmentation on high-resolution images in MLExchange

MLExchange is a machine learning (ML) operations platform providing web user-interfaces (UIs) for data visualization and analysis pipelines at synchrotron facilities. Among these UIs is the segmentation app which helps synchrotron users utilize ML algorithms to automatically segment high-resolution scientific images with minimal manual annotation effort. In this work, we share code optimizations that significantly speed up the segmentation inference workflow of large data in short time. By optimizing the sequence of CPU-GPU data transfers and introducing CPU parallelization to key operations, we improve the per-device, per-image frame computational efficiency and observe close to 3×$$\times$$ speedup over the original segmentation inference workflow run time when utilizing a single GPU. Further adaptations enabling multi-GPU inference yield more than 40×$$\times$$ speedup with 100 GPUs compared to the optimized single GPU inference workflow. This acceleration of the segmentation inference workflow will provide MLExchange users with easy access to segmentation results with little wait time.

Lu, Shizhao

Core inductive electric field during sawtooth crashes on DIII-D

Sawtooth crashes on tokamak plasmas exhibit relaxation much faster than resistive time scales via a mechanism not fully understood. Using core magnetic measurements from the Radial Interferometer-Polarimeter (RIP) diagnostic on the DIII-D tokamak, Grad–Shafranov equilibria constrained by internal magnetic measurements that have high time resolution (<1μs) can be computed, allowing analysis of how equilibrium parameters such as safety factor q, current density J, and parallel electric field E ∥ , particularly on-axis, evolve. At the sawtooth crash, on-axis safety factor q 0 is observed to rise by 5% but remain below 1 throughout the cycle, and on-axis current density J 0 is observed to drop by 5%. On-axis parallel electric field E ∥ (0) is found to be balanced by ηJ 0 (resistivity times on-axis current density) except during the 200 µs crash period, where E ∥ (0) reaches 22 V m –1 , exceeding ηJ 0 by a factor of more than 2000. These first measurements in tokamak plasmas verify that generalized Ohm's law is not balanced during the crash by resistive effects alone; this is a finding expected due to the relaxation being much faster than resistive timescales. As a result, measurement of the electric field during the tokamak sawtooth serves to illuminate the physical mechanisms at work.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Progressive Hedging Decomposition for Solutions of Large-Scale Process Family Design Problems

Rapid, wide-scale deployment of green process systems, such as carbon capture or water desalination systems, is essential for combatting climate change. Methods relying on traditional design or modularity fail to capture the benefits of both economies of numbers and economies of scale. We have proposed process family design, which designs a family of processes simultaneously exploiting opportunities for common elements. In previous work, we explored different optimization formulations to solve this problem. In this work, we develop a decomposition approach to tackle larger problems efficiently. We solve a water desalination case study, which is too large to solve within a reasonable timeframe with the discretization formulation. We exploit the block angular structure of the discretization problem to decompose and solve using Progressive Hedging (PH). We use the open-source Python package mpi-sppy to execute PH which allows us to leverage parallelization and a HPC cluster to further improve solution time.

Stinchfield, Georgia

Observations of Subduction, Downward Heat Flux and Dense Filament Collapse in the Northern Gulf of Mexico

Submesoscale processes are important contributors to the global heat budget and generally support upward heat transport through restratification. However, in salinity‐stratified regions, such as the northern Gulf of Mexico with its influx of freshwater from the Mississippi‐Atchafalaya river system, temperature can act like a passive tracer and submesoscale processes can contribute to downward heat transport. Oceanic heat content is a factor in many environmental risks the region faces, for example, hurricane intensification, and marine heatwaves. During the 2022 field campaign of the Submesoscales Under Near‐Resonant Inertial Shear Experiment, a sampling plan was developed to study such submesoscale processes in high resolution. Over 31 hr, four assets (two research ships and two remotely controlled boats) drove in parallel across a dense filament, capturing its evolution in time and space. The observations show that surface waters, warmed by daytime solar radiation, were subducted and that the associated overturning circulation transported heat below the surface layer where it was later irreversibly mixed away. The estimated downward heat flux was as strong as the concurrent net air‐sea heat flux into the ocean. The filament was then observed to rapidly collapse which we attribute to boundary layer turbulence and the breakdown of geostrophic balance. The collapsing fronts display behaviors indicative of gravity currents. These observations highlight how in salinity‐stratified regions, frontal dynamics can be associated with downward heat flux and how the submesoscale can play an important role in the oceanic heat budget.

58 GEOSCIENCES

Skipper-in-CMOS: Nondestructive Readout With Subelectron Noise Performance for Pixel Detectors

The Skipper-in-CMOS image sensor integrates the nondestructive readout capability of skipper charge coupled devices (Skipper-CCDs) with the high conversion gain of a pinned photodiode (PPD) in a CMOS imaging process while taking advantage of in-pixel signal processing. This allows both single photon counting as well as high frame rate readout through highly parallel processing. The first results obtained from a ${15} \times {15}~\mu $ m2 pixel cell of a Skipper-in-CMOS sensor fabricated in Tower Semiconductor’s commercial 180-nm CMOS image sensor process are presented. Measurements confirm the expected reduction of the readout noise with the number of samples down to deep subelectron noise of $0.15\text {e}^ - $ , demonstrating the charge transfer operation from the PPD and the single photon counting operation when the sensor is exposed to light. This article also discusses new testing strategies employed for its operation and characterization.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND