Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “parallel time integration”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Energetic particle marginal stability profile for HL-2M integrated simulation based on neural network module

Abstract A critical gradient model is employed to develop a module of energetic particle (EP) marginal stability profiles in OMFIT integrated simulations for studying EP transport. Currently, each iteration of transport evolution is approximately 10 min in the integrated simulation, whereas, the EP marginal stability profile, which serves as an input in the integrated simulation could take much longer; the reason being a combination of the TGLFEP and EPtran codes is employed in our previous investigation. To reduce the simulation time, the critical gradient is predicted by a neural network instead of the TGLFEP code, and the EPtran code is revised with parallel computing, so that the running time of this module can be controlled to within 5 min. The predictions are in good agreement with previous approaches. The integrated simulation of HL-2M with Alfven eigenmodes transported by neutral beam EP profiles indicates that EP transport reduces the total pressure and current as expected, but could also under some conditions raise the safety factor in the core, which is favorable for reversed magnetic shear and high-performance plasmas.

Physics↗

Phased plan for the implementation of the time-resolving magnetic recoil spectrometer on the National Ignition Facility (NIF)

The time-resolving magnetic recoil spectrometer (MRSt) is a transformative diagnostic that will be used to measure the time-resolved neutron spectrum from an inertial confinement fusion implosion at the National Ignition Facility (NIF). It uses a CD foil on the outside of the hohlraum to convert fusion neutrons to recoil deuterons. An ion-optical system positioned outside the NIF target chamber energy-disperses and focuses forward-scattered deuterons. A pulse-dilation drift tube (PDDT) subsequently dilates, un-skews, and detects the signal. While the foil and ion-optical system have been designed, the PDDT requires more development before it can be implemented. Therefore, a phased plan is presented that first uses the foil and ion-optical systems with detectors that can be implemented immediately—namely CR-39 and hDISC streak cameras. These detectors will allow the MRSt to be commissioned in an intermediate stage and begin collecting data on a reduced timescale, while the PDDT is developed in parallel. A CR-39 detector will be used in phase 1 for the measurement of the time-integrated neutron spectra with excellent energy-resolution, necessary for the energy calibration of the system. Streak cameras will be used in phase 2 for measurement of the time-resolved spectrum with limited spectral coverage, which is sufficient to diagnose the time-resolved ion temperature. Simulations are presented that predict the performance of the streak camera detector, indicating that it will achieve excellent burn history measurements at current yields, and good time-resolved ion-temperature measurements at yields above 3 × 10 17 . The PDDT will be used for optimal efficiency and resolution in phase 3.

47 OTHER INSTRUMENTATION↗

High Yield Xray Imager Final Design Review

The High Yield Xray Imager (HYXI) is a new NIF target diagnostic system currently under development. The goal of HYXI is to provide high-fidelity, high temporal resolution x-ray imaging capability on high yield NIF implosions at 10MJ and above. The HYXI instrument design concept is based on the combination of two technologies that have been successfully utilized at the NIF on previous instruments, electron pulse-dilation and hybrid-CMOS sensor imaging. The combination of these two techniques will give HYXI sufficient data quality to ascertain differences in hot spot formation dynamics between high and low yield implosions. This information will highlight the critical hot spot conditions needed for ignition and burn. The HYXI design leverages the successful operation of the PDIXI x-ray imager at the NIF on multi MJ yield shots. A new radiation tolerant CMOS imaging array (HYPERION) is being developed to eliminate the significant background noise which limits the data quality of PDIXI. We successfully placed the contract with Advanced hCMOS Systems (AHS) to develop the HYPERION sensor, which fulfils our criteria to place long lead time item procurements by end of FY24. The HYXI Final Design Review was completed at the end of Q4 FY24 (Sep 24 th and Sep 30 th ). The HYXI project is a multi-year effort with a phased approach to be bring up system functionality over time in parallel with the development and fabrication effort of the HYPERION CMOS imaging array. In Phase 1, time-integrated x-ray images on NIF DT experiments will be collected starting in Q3 FY25. In Phase 2 of the project, time-resolved imaging with HYXI utilizing a spare microchannel plate detector back-end will begin in Q3 FY26. Phase 3 concludes the project with the installation of the HYPERION sensor array and the final performance qualification of the HYXI instrument which is scheduled for Q3 FY27 as discussed in the PDR and MRT report on this project in FY23.

42 ENGINEERING↗

Efficient derivative computation for unsteady fatigue-constrained nonlinear aero-structural wind turbine blade optimization

Gradient-based optimization offers significant efficiency advantages for wind turbine blade design, but its application has often been limited by the cost and accuracy of finite-difference derivative calculations, especially when fatigue constraints are considered. In this work, we systematically compare and evaluate four differentiation techniques, namely algorithmic differentiation, implicit differentiation, sparsity exploitation, and parallelization, to determine their effectiveness in computing accurate gradients through time-domain aero-structural simulations. By integrating these techniques with unsteady nonlinear aerodynamic and structural models, we develop software designed for accurate gradient computation. We show that combining these techniques addresses memory and runtime challenges associated with long simulations required by design load cases. Specifically, the most effective combination reduces derivative computation wall time by over an order of magnitude compared to finite differencing while maintaining superior accuracy. We demonstrate this approach in a proof-of-concept aero-structural optimization of a wind turbine blade that improves the cost of energy by 12.78 %. This comparative study establishes a viable approach for fatigue-aware blade design that balances computational efficiency with modeling accuracy.

17 WIND ENERGY↗

Computational Algorithms for Unit Commitment with AC Power Flows (Final Report)

Security-constrained unit commitment (SCUC) is a key component in power system operations. When AC power flow constraints are considered in the SCUC model (AC-SCUC), the problem becomes extremely difficult due to its discrete and non-convex nature, as described in “Grid Optimization Competition Challenge 3 Problem Formulation (GOCC)”. There are four main challenges: (i) Discrete decisions regarding unit online/offline status and start-up/shut-down procedures for every single unit. The number of discrete decision variables increases considerably when a system integrates multiple generators; (ii) Configuration-based combined-cycle formulations, and multi-commodity models that include ramping products, spin/non-spin products, and regulation up/down products. The combined-cycle units introduce additional discrete decision variables and auxiliary service products further complicate the model by connecting multi-commodity products’ continuous and discrete variables; (iii) SCUC models with AC power flow constraints are far more complex due to massive bilinear terms in the large-scale nonlinear power balance equations. The nonlinear power balance equations are further complicated by the discrete step control variables of shunts; (iv) N − 1 contingency analysis. The size of the model increases linearly with the number of contingencies considered, greatly increasing the size of the optimization model. Accordingly, there is an emergent need to develop a robust algorithm capable of deriving a high-quality solution in a short time and passing through contingency tests simultaneously. In this project, we explore innovative techniques to address this challenging problem by integrating advanced polyhedral theory, approximation methods, relaxation strategies, decomposition techniques, and parallel computing. Each technique approaches the problem from a different perspective, leveraging its specific strengths to tackle distinct challenges. Each individual method has demonstrated its effectiveness in the PI’s previous research. Their integration is expected to significantly reduce the computational time required to solve the proposed complex problem. Successful completion of this project has the potential to transform the industry by enhancing optimization solvers capable of handling large-scale day-ahead energy market clearing models within strict time constraints, while incorporating AC power flow constraints. This advancement will lead to reduced overall generation costs and, consequently, increased social welfare.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Genetic algorithm optimization of nuclear criticality experiment for reduction of intermediate-energy 239 Pu nuclear data uncertainties

Nuclear criticality experiments are conducted to investigate specific nuclear data important for safe handling and storage of fissile materials, reactor design and operation, and the validation of radiation transport codes. Incorrect or uncertain nuclear data can prohibitively impact operational safety limits, reactor licensing, and predictive simulation capability; therefore, integral measurements from criticality experiments are necessary and should be performed frequently. To maximize the impact of the integral measurements, it is important to consider experiment geometry, material selection, and component dimensions. When taking these considerations into account, the experiment design process becomes iterative and very time intensive. This work utilizes a genetic algorithm to efficiently explore potential nuclear criticality experiment designs for the Laboratory Directed Research & Development project PARADIGM (PARallel Approach of Differential and InteGral Measurements) at Los Alamos National Laboratory. In this paper, the building blocks of the genetic algorithm are discussed in detail, the genetic algorithm methodology is verified, and the genetic algorithm is used to produce three candidate experiment models for the final PARADIGM design. The three candidate models produced by the genetic algorithm consist of copper-reflected assemblies containing 14 repeating units of alumina, graphite, boron, and plutonium plates. Furthermore, in addition to the optimization results, final design considerations are also discussed for designs with a height and/or weight very close to or slightly above assembly machine operational limits.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

BeeSwarm: Enabling Parallel Scaling Performance Measurement in Continuous Integration for HPC Applications

Testing is one of the most important steps in software development–it ensures the quality of software. Continuous Integration (CI) is a widely used testing standard that can report software quality to the developer in a timely manner during development progress. Performance, especially scalability, is another key factor for High Performance Computing (HPC) applications. There are many existing profiling and performance tools for HPC applications, but none of these are integrated into CI tools. In this work, we propose BeeSwarm, an HPC container based parallel scaling performance system that can be easily applied to the current CI test environments. BeeSwarm is mainly designed for HPC application developers who need to monitor how their applications can scale on different compute resources. We demonstrate BeeSwarm using a multi-physics HPC application with Travis CI, GitLab CI and GitHub Actions while using ChameleonCloud and Google Compute Engine as the compute backends. Finally, our results show that BeeSwarm can be used for scalability and performance testing of HPC applications.

97 MATHEMATICS AND COMPUTING↗

Battery asset management with cycle life prognosis

We report Battery Asset Management problem determines the minimum cost replacement schedules for each individual asset in a group of battery assets that operate in parallel. Battery cycle life varies under different operating conditions including temperature, depth of discharge (DOD), charge rate, etc., and a battery deteriorates due to usage, which cannot be handled by current asset management models. This paper presents a new battery asset management methodology where battery cycle life prognosis is integrated with parallel asset management to reduce lifecycle cost of the Battery Energy Storage Systems (BESS). For the battery failure time prognosis, a nonlinear physics-based battery capacity fade model is developed and incorporated in parallel asset management model to update battery capacity over time. Experiment results have shown that the developed battery asset management methodology can be conveniently used to facilitate BESS asset management decision making thereby decreasing asset lifecycle costs.

25 ENERGY STORAGE↗

Tokamak ITG-KBM transition benchmarking with the mixed variables/pullback transformation electromagnetic gyrokinetic scheme

Electromagnetic gyrokinetic simulation of high temperature plasma is required to predict confinement in magnetic fusion devices and has posed challenges for existing codes. In this paper, we demonstrate successful global gyrokinetic simulation of the ion temperature gradient-driven mode-kinetic ballooning mode transition in a toroidal fusion plasma test case using the mixed variables/pullback transformation (MV/PT) scheme with the particle-in-cell codes XGC and ORB5, and compare to results from a conventional continuum code from the literature. Furthermore, the MV/PT scheme combines explicit time integration with mitigation of the well-known electromagnetic gyrokinetic “cancelation problem.” We calculate eigenmodes in the electrostatic and parallel vector potentials, and find good agreement in growth rate, real frequency, and the normalized plasma pressure of mode transition.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Attribute-Aware RBFs: Interactive Visualization of Time Series Particle Volumes Using RT Core Range Queries

Smoothed-particle hydrodynamics (SPH) is a mesh-free method used to simulate volumetric media in fluids, astrophysics, and solid mechanics. Visualizing these simulations is problematic because these datasets often contain millions, if not billions of particles carrying physical attributes and moving over time. Radial basis functions (RBFs) are used to model particles, and overlapping particles are interpolated to reconstruct a high-quality volumetric field; however, this interpolation process is expensive and makes interactive visualization difficult. Existing RBF interpolation schemes do not account for color-mapped attributes and are instead constrained to visualizing just the density field. To address these challenges, we exploit ray tracing cores in modern GPU architectures to accelerate scalar field reconstruction. We use a novel RBF interpolation scheme to integrate per-particle colors and densities, and leverage GPU-parallel tree construction and refitting to quickly update the tree as the simulation animates over time or when the user manipulates particle radii. We also propose a Hilbert reordering scheme to cluster particles together at the leaves of the tree to reduce tree memory consumption. Finally, we reduce the noise of volumetric shadows by adopting a spatially temporal blue noise sampling scheme. Our method can provide a more detailed and interactive view of these large, volumetric, time-series particle datasets than traditional methods, leading to new insights into these physics simulations.

Particle Volumes↗

A time-parallel multiple-shooting method for large-scale quantum optimal control

Quantum optimal control plays a crucial role in quantum computing by providing the interface between compiler and hardware. Solving the optimal control problem is particularly challenging for multi-qubit gates, due to the exponential growth in computational complexity with the system's dimensionality and the deterioration of optimization convergence. To ameliorate the computational complexity of time-integration, this paper introduces a multiple-shooting approach in which the time domain is divided into multiple windows and the intermediate states at window boundaries are treated as additional optimization variables. Further, this enables parallel computation of state evolution across time-windows, significantly accelerating objective function and gradient evaluations. Since the initial state matrix in each window is only guaranteed to be unitary upon convergence of the optimization algorithm, the conventional gate trace infidelity is replaced by a generalized infidelity that is convex for non-unitary state matrices. Continuity of the state across window boundaries is enforced by equality constraints. A quadratic penalty optimization method is used to solve the constrained optimal control problem, and an efficient adjoint technique is employed to calculate the gradients in each iteration. We demonstrate the effectiveness of the proposed method through numerical experiments on quantum Fourier transform gates in systems with 2, 3, and 4 qubits, noting a speedup of 80x for evaluating the gradient in the 4-qubit case, highlighting the method's potential for optimizing control pulses in multi-qubit quantum systems.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

MFC 5.0: An exascale many-physics flow solver

Many problems of interest in engineering, medicine, and the fundamental sciences rely on high-fidelity flow simulation, making performant computational fluid dynamics solvers a mainstay of the open-source software community. Previous work MFC 3.0 was made a published, documented, and open-source solver via Bryngelson et al. Comp. Phys. Comm. (2021) with numerous physical features, numerical methods, and scalable infrastructure. MFC 5.0 is a significant update to MFC 3.0, featuring a broad set of well-established and novel physical models and numerical methods, as well as the introduction of GPU and APU (or superchip) acceleration. Here, we exhibit state-of-the-art performance and ideal scaling on the first two exascale supercomputers, OLCF Frontier and LLNL El Capitan. Combined with MFC’s single-accelerator performance, MFC achieves exascale computation in practice, and achieved the largest-to-date public CFD simulation at 200 trillion grid points as a 2025 ACM Gordon Bell Prize finalist. New physical features include the immersed boundary method, N-fluid phase change, Euler–Euler and Euler–Lagrange sub-grid bubble models, fluid-structure interaction, hypo- and hyper-elastic materials, chemically reacting flow, two-material surface tension, magnetohydrodynamics (MHD), and more. Numerical techniques now represent the current state-of-the-art, including general relaxation characteristic boundary conditions, WENO variants, Strang splitting for stiff sub-grid flow features, and low Mach number treatments. Weak scaling to tens of thousands of GPUs on OLCF Summit and Frontier and LLNL El Capitan achieves efficiencies within 5% of ideal to over 90% of their respective system sizes. Strong scaling results for a 16-times increase in device count show parallel efficiencies over 90% on OLCF Frontier. MFC’s software stack has undergone further improvements, including continuous integration, which ensures code resilience and correctness through over 300 regression tests; metaprogramming, which reduces code length while maintaining performance portability; and code generation for computing chemical reactions

Computational fluid dynamics↗

RAPID

Parallel computer code for the simulator for dynamics of power systems which has the capability to initiate the system and create different faults for the dynamic analysis. The code is based on time-parallel method (Parareal) with Adaptive Method Reduction (AMR). The coarse solvers for the Parareal algorithm include several Semi Analytical Solution methods. Also, Integrated simulation of coupled transmission and distribution systems can be studied.

Simunovic, Srdjan [Oak Ridge National Lab. (ORNL),↗

Seeing in with X-rays: 4D Strain and Thermometry Measurements for Thermal-Mechanical Testing

Understanding temperature-dependent material decomposition and structural deformation induced by combined thermal-mechanical environments is critical for safety qualification of hardware under accident scenarios. Seeing in with X-rays elucidated the physics necessary to develop X-ray strain and thermometry diagnostics for use in optically opaque environments. Two parallel thermometry schemes were explored: X-ray fluorescence and X-ray diffraction of inorganic doped ceramics– colloquially known as thermographic phosphors. Two parallel surface strain techniques–Path-Integrated Digital Image Correlation and Frequency Multiplexed Digital Image Correlation–were demonstrated. Finally, preliminary demonstration of time-resolved digital volume correlation was performed by taking advantage of limited view reconstruction techniques. Additionally, research into blended ceramic-metal coatings was critical to generating intrinsic thermographic patterns for the future combination of X-ray strain and thermometry measurements.

36 MATERIALS SCIENCE↗

A Cryogenic Readout IC with 100 KSPS in-Pixel ADC for Skipper CCD-in-CMOS Sensors

The Skipper CCD-in-CMOS Parallel Read-Out Circuit (SPROCKET) is a mixed-signal front-end design for the readout of Skipper CCD-in-CMOS image sensors. SPROCKET is fabricated in a 65 nm CMOS process and each pixel occupies a 50$\mu$m $\times$ 50$\mu$m footprint. SPROCKET is intended to be heterogeneously integrated with a Skipper-in-CMOS sensor array, such that one readout pixel is connected to a multiplexed array of nine Skipper-in-CMOS pixels to enable massively parallel readout. The front-end includes a variable gain preamplifier, a correlated double sampling circuit, and a 10-bit serial successive approximation register (SAR) ADC. The circuit achieves a sample rate of 100 ksps with 0.48 $\mathrm{e^-_{rms}}$ equivalent noise at the input to the ADC. SPROCKET achieves a maximum dynamic range of 9,000 $e^-$ at the lowest gain setting (or 900 $e^-$ at the lowest noise setting). The circuit operates at 100 Kelvin with a power consumption of 40 $\mu W$ per pixel. A SPROCKET test chip was submitted in September 2022, and test results will be presented at the conference.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Towards fast, accurate predictions of RF simulations via data-driven modeling: Forward and lateral models

Three machine learning techniques (multilayer perceptron, random forest, and Gaussian process) provide fast surrogate models for lower hybrid current drive (LHCD) simulations. A single GENRAY/CQL3D simulation without radial diffusion of fast electrons requires several minutes of wall-clock time to complete, which is acceptable for many purposes, but too slow for integrated modeling and real-time control applications. More accurate simulations with fast electron diffusion are even slower, requiring multiple hours of run time with parallel processing. The machine learning models use a database of 16,000+ GEN-RAY/CQL3D simulations for training, validation, and testing. Latin hypercube sampling methods implemented in πScope ensure that the database covers the range of 9 input parameters (n e0 , T e0 , I p , B t , R 0 , n ∥︀ , Z e f f , V loop , P LHCD ) with sufficient density in all regions of parameter space. The surrogate models reduce the computation time from minutes-hours to ms with high accuracy across the input parameter space. Data-driven surrogate models also allow for solving inverse and “lateral” problems. A surrogate model for the inverse problem maps from a desired current drive or power deposition profile to a set of input parameters that would result in such a profile, while a surrogate model for the lateral problem maps from a measured experimental quantity such as hard x-ray emission to a current drive or power deposition profile. In conclusion, the πScope database creation workflow is flexible and applicable to other RF simulation codes such as TORIC.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Modeling of streamflow in a 30 km long reach spanning 5 years using OpenFOAM 5.x

Abstract. Developing accurate and efficient modeling techniques for streamflow at the tens-of-kilometers spatial scale and multi-year temporal scale is critical for evaluating and predicting the impact of climate- and human-induced discharge variations on river hydrodynamics. However, achieving such a goal is challenging because of limited surveys of streambed hydraulic roughness, uncertain boundary condition specifications, and high computational costs. We demonstrate that accurate and efficient three-dimensional (3-D) hydrodynamic modeling of natural rivers at 30 km and 5-year scales is feasible using the following three techniques within OpenFOAM, an open-source computational fluid dynamics platform: (1) generating a distributed hydraulic roughness field for the streambed by integrating water-stage observation data, a rough wall theory, and a local roughness optimization and adjustment strategy; (2) prescribing the boundary condition for the inflow and outflow by integrating precomputed results of a one-dimensional (1-D) hydraulic model with the 3-D model; and (3) reducing computational time using multiple parallel runs constrained by 1-D inflow and outflow boundary conditions. Streamflow modeling for a 30 km long reach in the Columbia River (CR) over 58 months can be achieved in less than 6 d using 1.1 million CPU hours. The mean error between the modeled and the observed water stages for our simulated CR reach ranges from −16 to 9 cm (equivalent to approximately ±7 % relative to the average water depth) at seven locations during most of the years between 2011 and 2019. We can reproduce the velocity distribution measured by the acoustic Doppler current profiler (ADCP). The correlation coefficients of the depth-averaged velocity between the model and ADCP measurements are in the range between 0.71 and 0.83 at 75 % of the survey cross sections. With the validated model, we further show that the relative importance of dynamic pressure versus hydrostatic pressure varies with discharge variations and topography heterogeneity. Given the model's high accuracy and computational efficiency, the model framework provides a generic approach to evaluate and predict the impacts of climate- and human-induced discharge variations on river hydrodynamics at tens-of-kilometers and decadal scales.

58 GEOSCIENCES↗

An Open-Source Parallel EMT Simulation Framework

As the integration level of inverter-based resources (IBRs) increases, ensuring the reliable operation of the bulk power systems requires the use of electromagnetic transient (EMT) simulation tools to identify and mitigate system-wide stability risks. Conducting EMT studies for large-scale, IBR-rich grids, however, is challenging due to the inherent computational bottleneck caused by the underlying high-fidelity models and required small time steps. This paper introduces ParaEMT: an open-source, generic EMT simulation framework designed to accelerate simulations by leveraging advanced parallel computational technologies, such as high-performance computers. This paper presents a comprehensive exposition of ParaEMT, covering its modeling library, simulation strategy, framework structure, operational procedures, and auxiliary features, alongside its extensible parallel computational architecture. Notably, ParaEMT is a publicly accessible and modularized framework written in Python, thereby facilitating future development and the integration of new models and algorithms. The accuracy and efficiency of ParaEMT are demonstrated by rigorous validations via multiple case studies.

electromagnetic transient simulation↗