Engineering PapersSearch

SEARCH · Engineering Papers

Results for “Parallel in time”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

On how avalanches penetrate the SOL and broaden heat loads

Recent experiments reported a correlation between power law core temperature spectra and Dα emission, suggesting that heat avalanches penetrate the SOL. This paper derives a threshold criterion for avalanche penetration using a reduced model. Avalanches with ($\bigtriangledown$$\tilde{T}$) rms > $\bigtriangledown$$\tilde{T}$ crit at the separatrix are predicted to penetrate, and so broaden the SOL and heat load distribution. $\bigtriangledown$$\tilde{T}$ crit is ~ 1/$τ$ ∥ , where $τ$ ∥ is the parallel heat flow time through the SOL. Penetration occurs when avalanches are strong enough to steepen sufficiently to shock at the separatrix. A positive correlation is found between the nonlinear drive for steepening and the penetration depth. In particular, penetration depth exceeds that of the heuristic drift limit when shocks form. Implications for numerical and physical experiments are also discussed.

SOL heat load

A Portfolio Approach to Massively Parallel Bayesian Optimization

One way to reduce the time of conducting optimization studies is to evaluate designs in parallel rather than just one-at-a-time. For expensive-to-evaluate black-boxes, batch versions of Bayesian optimization have been proposed. They work by building a surrogate model of the black-box to simultaneously select multiple designs via an infill criterion. Still, despite the increased availability of computing resources that enable large-scale parallelism, the strategies that work for selecting a few tens of parallel designs for evaluations become limiting due to the complexity of selecting more designs. It is even more crucial when the black-box is noisy, necessitating more evaluations as well as repeating experiments. Here we propose a scalable strategy that can keep up with massive batching natively, focused on the exploration/exploitation trade-off and a portfolio allocation. We compare the approach with related methods on noisy functions, for mono and multi-objective optimization tasks. These experiments show orders of magnitude speed improvements over existing methods with similar or better performance.

97 MATHEMATICS AND COMPUTING

Toucan: A performance portable, scalable implementation of the DECA algorithm

In the field of additive manufacturing (AM), cellular automata (CA) is extensively used to simulate microstructural evolution during solidification. However, while traditional CA approaches are relatively fast, they still require a substantial number of time steps, are limited to moderate volumes, and are relatively difficult to improve through parallelism due to the highly localized nature of the solidification front. Here, to address these issues of time to solution and load balancing, we introduce Toucan, a parallel, performance-portable, and scalable code written in C++ with the Kokkos library that leverages the discrete event inspired cellular automata (DECA) algorithm to perform parallel-in-time (PinT) grain growth simulations. Toucan effectively mitigates load balancing issues by distributing the computational workload more evenly across processors, enhancing scalability and efficiency. We conduct both strong and weak scaling studies on up to 64 GPUs on the Frontier supercomputer, demonstrating that Toucan significantly outperforms the current state-of-the-art, time-stepped CA code, ExaCA, on both single and multi-GPU simulations. Even in AM-specific weak scaling scenarios, Toucan maintains near-ideal scaling, in contrast to the linear increase observed with ExaCA due to the moving laser raster pattern. This study highlights Toucan’s potential to transform microstructural simulations in AM by radically improving both efficiency and scalability over existing methods.

36 MATERIALS SCIENCE

Speedup of UEDGE Parameter Scans Using Machine-Learning Optimized OpenMP Parallelization and a Continuation Solver

This article presents the OpenMP parallelization of the preconditioning Jacobian assembly and right‐hand side residual evaluation in UEDGE. A continuation algorithm, utilizing the internal NKSOL implicit Jacobian‐Free Newton‐Krylov solver to efficiently scan physical parameters, is also presented. The implemented parallelization reduces the computational time for a benchmark scan run on 32 threads by compared to the serial version when using trained random forest regression models to identify the optimal decomposition of the system of equations. Random forest regression models applied to the UEDGE time‐dependent and continuation solver algorithms did not yield meaningful improvement in computational performance. A benchmark DIII‐D gas injection rate scan in the 0.35–0.75 kA interval, performed on a test cluster using the parallelized code and continuation solver, produced 1066 steady‐state solutions with a 22 s average wall‐clock computational time per steady‐state solution.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY

Randomized Preconditioned Solvers for Strong Constraint 4D-Var Data Assimilation

The Strong Constraint 4D Variational (SC-4DVAR) data assimilation method is widely used in climate and weather applications. SC-4DVAR involves solving a minimization problem to compute the maximum a posteriori estimate, which we tackle using the Gauss-Newton method. The computation of the descent direction is expensive since it involves the solution of a large-scale and potentially ill-conditioned linear system, solved using the preconditioned conjugate gradient (PCG) method. Here, to address this cost, we efficiently construct scalable preconditioners using three different randomization techniques, which all rely on a certain low-rank structure involving the Gauss-Newton Hessian. The proposed techniques come with theoretical guarantees on the condition number, and at the same time, are amenable to parallelization. We also develop an adaptive approach to estimate the sketch size and choose between the reuse or recomputation of the preconditioner. We demonstrate the performance and effectiveness of our methodology on two representative model problems—the Burgers and barotropic vorticity equation—showing a drastic reduction in both the number of PCG iterations and the number of Gauss-Newton Hessian products after including the preconditioner construction cost.

Gauss-Newton

Advancing specialized biofoundries via automated adaptive laboratory evolution

Adaptive laboratory evolution (ALE) is a powerful strategy for improving microbial phenotypes by harnessing natural selection under defined environmental conditions. Through applying selection regimes, beneficial mutations accumulate, enabling the generation of strains with enhanced properties. However, conventional ALE is labor-intensive and difficult to scale, limiting reproducibility and broader discovery of evolutionary principles. Recent advances in robotics, automation, and computational infrastructure are transforming ALE into a scalable, data-rich experimental paradigm. Automated platforms enable standardized and complex protocols, real-time monitoring, and highly parallel evolution campaigns, improving consistency while generating longitudinal datasets that reveal convergent adaptive mechanisms. Here, we discuss the role of specialized biofoundries in advancing automated ALE and enabling large-scale evolutionary engineering. We review major automated ALE formats and outline key design principles for effective ALE biofoundries, highlighting how automated ALE can support autonomous experimentation and AI-guided strain engineering.

59 BASIC BIOLOGICAL SCIENCES

Reversible to irreversible transitions for ac driven skyrmions on periodic substrates

Abstract Using atomistic simulations, we investigate the dynamical behavior of magnetic skyrmions in dimer and trimer molecular crystal arrangements, as well as bipartite lattices at 3/2 and 5/2 fillings, under ac driving over a square array of anisotropy defects. For low ac amplitudes, at all fillings reversible motion appears in which the skyrmions return to their original positions at the end of each ac drive cycle and the diffusion is zero. We also identify two distinct irreversible regimes. The first is a translating regime in which the skyrmions form channels of flow in opposing directions and translate by one substrate lattice constant per ac drive cycle. The translating state appears in the dimer and trimer arrangements, and produces pronounced peaks in the diffusivity in the direction perpendicular to the external drive. For larger ac amplitudes, we find chaotic irreversible motion in which the skyrmions can randomly exchange places with each other over time, producing long-time diffusive behavior both parallel and perpendicular to the ac driving direction.

36 MATERIALS SCIENCE

A Comparison of GPU-Accelerated Multiphase CFD Solvers on the Polaris Supercomputer: Part 1

This report is in support of the Innovative and Novel Computational Impact on Theory and Experiment (INCITE) program sponsored by the U.S. Department of Energy (USDOE). With INCITE-level resources, one project, titled BubblyFlow, was granted computational resources for the 2025 calendar year on the Polaris supercomputer at the Argonne Leadership Computing Facility (ALCF). The project aims to conduct simulations to understand the fundamental characteristics of turbulent bubbly flow phenomena in nature. Staff at the ALCF and Argonne’s Computational Science division, along with collaborators at the City College of New York and University of Illinois at Chicago, helped a summer student to assess the accuracy and performance of two high performance computing (HPC) codes. Both codes, ImExLBM and FluTAS, are fundamentally different in their mathematical and numerical modeling. However, both may be used to solve the same physical problem. The collaboration sought to better understand the differences between both codes in terms of accuracy and efficiency. This would ultimately help the BubblyFlow project better utilize resources and establish a knowledge-base of code capabilities in future simulation campaigns. We compare ImExLBM and FluTAS, two high-performance multiphase computational fluid dynamics (CFD) solvers, in terms of physical fidelity, time-to-solution, and parallel efficiency. We validate ImExLBM (Implicit-Explicit Lattice Boltzmann Method) against a canonical benchmark and assess it’s performance relative to FluTAS (Fluid Transport Accelerated Solver), a well-established open-source CFD code.

97 MATHEMATICS AND COMPUTING

Correcting beam space charge effects in Active-Target Time Projection Chamber

By providing a large gaseous volume for nuclear interactions while simultaneously recording the tracks of resulting reaction products, an active target serves as both a thick target and a detector. Once a reaction occurs, the emitted charged fragments strip electrons from the target gas along their path as they transverse the detector. Collection of these stripped electrons allow for detection of the product tracks. As beam intensity increases, the resulting ionization in the active target can significantly distort this collection of electrons. If left uncorrected, the resulting measurements could be wrong. In this paper, we investigate the impact of the space charge produced by heavy radioactive beams within the Active Target - Time Projection Chamber at Michigan State University. The beams are injected parallel to the electric field of the time projection chamber which is operated without a magnetic field for this experiment. Furthermore, we analyze the rate dependence of the space charge effects and demonstrate that they can be modeled and effectively corrected.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND

Parallelizing autotuning for HPC applications: Unveiling the potential of the speculation strategy in Bayesian optimization

In the exascale computing era, tuning High-Performance Computing (HPC) applications has become a significant computational challenge. Although Bayesian optimization (BO) has emerged as a promising tool for HPC performance tuning, the BO workflow is inherently sequential (i.e., one function evaluation at a time) and cannot leverage the huge amount of parallel resources present in modern supercomputers, resulting in a considerable underutilization of their computational capabilities. This paper explores the trade-off between search quality and parallelism in BO, investigating a diverse set of methods. Building upon both previous approaches from the literature and novel methodologies introduced in this work, our study provides a deep analysis to accelerate BO performance tuning. By examining a set of synthetic functions and practical HPC applications, our exploration analyzes the interaction among various BO methods for parallelization, the quantity of parallel resources, the runtime distribution of target HPC applications, and the costs associated with different search orchestration mechanisms that have been overlooked in previous studies. Compared to sequential BO, our novel methodology achieves comparable quality while demonstrating robust scalability in search time as the amount of parallel resources increases; it also outperforms a state-of-the-art tuner, which supports parallelization, achieving up to 3.67x faster search time. We provide high-value insights for practitioners seeking to leverage the power of parallel computing for efficient HPC application tuning. Additionally, to further assist researchers in accelerating the performance tuning of their HPC applications, we provide an extension of an existing open-source tuning framework that incorporates our methods.

Bayesian optimization

Hydrotreatment of Nylon 66 and Amide Model Compounds Over Sulfided NiMo Catalysts

Molybdenum sulfide-based catalysts, such as nickel–molybdenum on alumina (NiMoS x /Al 2 O 3 ), are widely used in hydrotreating and have potential for catalyzing waste plastic conversion via hydrogenolysis, yet their performance, such as reaction kinetics and network, for amide-rich polymer feeds is poorly defined. Here we combine Nylon 66 with the amide model compound, N,N-dibutylhexanediamide (DBDAD), to quantify hydrodeoxygenation (HDO) and hydrodenitrogenation (HDN) chemistry in a stirred batch reactor (53 bar H 2 , 280–320°C). DBDAD conversion is near-linear with time, indicating strong adsorption of the substrates on the active sites. Time-resolved product identification indicates parallel C─O first-cleavagedeoxygenation (DO) and C─N first-cleavagedenitrogenation (DN) sequences proceeding through amine and diol intermediates, respectively, to C 4 ─C 6 alkanes. Increasing temperature shifts selectivity toward DN, decreasing the initial r(DO)/r(DN) from 1.38 (280°C) to 0.69 (320°C), with an apparent activation energy of 173 kJ mol −1 for DBDAD conversion. At 300°C, nylon 66 converts faster than DBDAD, producing a complex mixture of oxygen- and nitrogen-containing species and an initial rate ratio r(DO)/r(DN) of 1.6. No heteroaromatic nitrogen products are detected by the method used. These results provide reaction pathways and product signatures relevant to hydro-processing catalysts exposed to polyamide-derived streams.

Nylon 66

Interfacial Inversion of Stealth Surfactants

Amphiphilic macromolecular surfactants segregate to liquid–liquid interfaces, thereby reducing the interfacial tension and free energy. Here, we investigated “stealth surfactants” in the form of core–shell bottlebrush polymers comprised of pH-responsive diblock copolymer side chains forming a hydrophilic core and a hydrophobic shell, enabling solubility in oil. At liquid–liquid interfaces, these polymers undergo a structural “inversion”, with hydrophilic blocks segregating into the aqueous phase and hydrophobic blocks residing in the oil phase. The reconfiguration kinetics and surfactant properties are influenced by multiple factors, including the molecular weights of the backbone and side chain components, the hydrophilic-to-hydrophobic balance of the side chains, and the pH of the aqueous phase. An observed nonmonotonic dependence of interfacial tension with time is attributed to a progressive structural inversion, where the projected area of the macromolecule onto the interface decreases. To validate this inversion hypothesis, interfacial properties were characterized by sum-frequency generation vibrational spectroscopy, which revealed configurational changes of the core–shell bottlebrush polymers at the fluid interface and revealed a pH-dependent interfacial coverage. Coarse-grained molecular dynamics simulations supported these experimental findings, showing that the pH-responsive core and hydrophobic shell assume a time-averaged configuration with orientations parallel and perpendicular to the plane of the interface, respectively. These findings open routes to design multistimuli-responsive polymeric surfactants and compatibilizers, expanding their potential applications in advanced interfacial systems.

Stealth surfactants

Low-Temperature Plasma-Based Metrology of Lithium-Ion Battery Electrode Materials (CRADA Final Report)

As part of the Cyclotron Road program, SirenOpt Inc. evaluated its low-temperature plasma-based metrology sensor prototype for measuring multiple critical properties of lithium-ion battery electrode materials in parallel and in real-time. Cost-effective, minimal-waste manufacturing of high-performance battery electrode materials will be vital for achieving society’s net-zero carbon emission goals. Because existing electrode metrology sensors cannot operate within most sections of manufacturing lines, manufacturers often complete hundreds of processing steps before they can test their products and detect problems. When manufacturers perform these offline tests, they typically only test a small portion of the manufactured products. Current electrode manufacturing thus often yields many low-quality products, or off-spec products that must be thrown away all together. For example, at least 6% of the total lithium-ion battery manufacturing cost (i.e., over $250 million/year for the average gigafactory) is devoted to processing defective electrodes that are not scrapped until performance tests are failed during late-stage quality control checks. Electrode variability also leads manufacturers to build extra cells into battery packs to reduce the risk of poor performance. For example, many electric vehicle (EV) manufacturers include up to 10% more cells than needed, which substantially increases the cost and weight of the final EV product. The SirenOpt sensor can potentially enable early detection of poorly manufactured electrodes and allow them to be removed earlier from manufacturing lines, which can save battery manufacturers (hundreds of) millions of dollars per year. The sensor can further be used to improve product quality by accelerating R&D and process optimization, improving quality control, and enabling real-time process control. Overall, a real-time, in-situ metrology strategy can create unprecedented opportunities for implementation of smart manufacturing practices and advanced quality and process control solutions to realize higher battery electrode throughput and performance.

25 ENERGY STORAGE

Comparative Analysis of Radial and Random Microstructures of Mesophase Pitch Carbon Fibers

Carbon fibers (CF) with radial and random microstructures are produced. Here, these fibers are subjected to identical treatment before being mechanically tested and analyzed with Weibull analysis, with the results revealing a statistically significant difference in tensile strengths of 2.23 GPa for random CF and 1.69 GPa for radial CF. Raman mapping probed the crystalline structure perpendicular to the fiber axis and found a uniform structure, while wide‐angle X‐ray diffraction showed a significant difference of 7.5 Å in the crystallites’ basal lengths parallel to the fiber. Small‐angle X‐ray scattering is completed parallel to the fiber for the first time. A cross‐section Guinier plot of the 1D azimuthal integration is generated assuming symmetric scattering, and the parallel scatterers are found to have a similar length scale to the crystallite's length, validating the testing method. Finally, transmission electron microscopy is completed on the longitudinal cross‐section of each fiber. The radial carbon fiber is found to have a core–shell structure, as evidenced further by fast Fourier transform images. Through all studies, it is shown that the structure developed during mesophase pitch spinning altered the microstructure, thus impacting the mechanical properties, confirming a direct relationship between processing, structure, and properties.

Scherschel, Alexander [Univ. of Virginia, Charlott

Using Hardware-In-The-Loop Methodology to Develop Test Systems

Hardware in the Loop (HIL) testing methodologies have become widespread in industry. Typically, they focus on developing control algorithms for systems such as autonomous vehicles or aircraft. An oft overlooked aspect of product development is the design and fabrication of a test system for validating that the product meets requirements. Abstractly, a test system differs little from a control system—testers provide signals to the unit, monitor feedback, and base decisions on the results. While the time scales may differ, the functionalities are conceptually similar. Viewed in this light, it becomes natural to extend HIL approaches to tester development. By replacing a physical unit with a proxy model deployed to a real-time or pseudo real-time target, test systems can be developed in parallel with the design and fabrication of a first production unit. This saves considerable time in the life cycle from conceptual design to realized product. This manuscript demonstrates the process flow using a capacitive discharge unit as an exemplar.

42 ENGINEERING

Real-Time Bayesian Inference at Extreme Scale: A Digital Twin for Tsunami Early Warning Applied to the Cascadia Subduction Zone

We present a Bayesian inversion-based digital twin that employs acoustic pressure data from seafloor sensors, along with 3D coupled acoustic–gravity wave equations, to infer earthquake-induced spatiotemporal seafloor motion in real time and forecast tsunami propagation toward coastlines for early warning with quantified uncertainties. Our target is the Cascadia subduction zone, with one billion parameters. Computing the posterior mean alone would require 50 years on a 512 GPU machine. Instead, exploiting the shift invariance of the parameter-to-observable map and devising novel parallel algorithms, we induce a fast offline–online decomposition. The offline component requires just one adjoint wave propagation per sensor; using MFEM, we scale this part of the computation to the full El Capitan system (43,520 GPUs) with 92% weak parallel efficiency. Moreover, given real-time data, the online component exactly solves the Bayesian inverse and forecasting problems in 0.2 seconds on a modest GPU system, a ten-billion-fold speedup.

97 MATHEMATICS AND COMPUTING

Enabling Parallel Performance and Portability of Solid Mechanics Simulations Across CPU and GPU Architectures

Efficiently simulating solid mechanics is vital across various engineering applications. As constitutive models grow more complex and simulations scale up in size, harnessing the capabilities of modern computer architectures has become essential for achieving timely results. This paper presents advancements in running parallel simulations of solid mechanics on multi-core CPUs and GPUs using a single-code implementation. This portability is made possible by the C++ matrix and array (MATAR) library, which interfaces with the C++ Kokkos library, enabling the selection of fine-grained parallelism backends (e.g., CUDA, HIP, OpenMP, pthreads, etc.) at compile time. MATAR simplifies the transition from Fortran to C++ and Kokkos, making it easier to modernize legacy solid mechanics codes. We applied this approach to modernize a suite of constitutive models and to demonstrate substantial performance improvements across different computer architectures. This paper includes comparative performance studies using multi-core CPUs along with AMD and NVIDIA GPUs. Results are presented using a hypoelastic–plastic model, a crystal plasticity model, and the viscoplastic self-consistent generalized material model (VPSC-GMM). The results underscore the potential of using the MATAR library and modern computer architectures to accelerate solid mechanics simulations.

Morgan, Nathaniel (ORCID:0000000276118449)

Cardinal: Seismic and Geoacoustic Array Processing

Data collected via seismic and infrasound array deployments are leveraged in the geosciences to detect and characterize a myriad of natural and anthropogenic sources. These deployments consist of numerous sensors placed in a predetermined configuration to amplify signal strength and improve the efficacy of array processing techniques used to measure signal directionality and waveform coherence. High‐fidelity feature extraction is often predicated on interstation distance as well as the frequency content and wavelength of an incident signal. Numerous array processing softwares analyze data in sequential frequency bands to obtain a more detailed characterization of a signal. However, current algorithms are limited in their ability to determine optimal array configuration for each band. We introduce an open‐source Python code, called Cardinal, to process seismic and infrasound array data in discretized time–frequency space with the option of applying an adaptive array design to determine optimal subarray configuration for each frequency band. To reduce computational time, the array processing step can be run in parallel using multithreading. Furthermore, the software has the capability to aggregate array processing results from different time–frequency pixels to produce separate sets of detections, or families, with added utility via the application of an adaptive semblance threshold, which aids in isolating signals‐of‐interest from coherent background noise. Upon appropriate configuration, Cardinal exhibits the potential to combine distinct seismic and infrasound phases into separate families.

Adaptive Array