Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “algorithmic differentiation”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

GEAR-MC and Differential-Operator Methods Applied to Electron-Photon Transport in the Integrated TIGER Series

The sensitivity analysis algorithms that have been developed by the radiation transport community in multiple neutron transport codes, such as MCNP and SCALE, are extensively used by fields such as the nuclear criticality community. However, these techniques have seldom been considered for electron transport applications. In the past, the differential-operator method with the single scatter capability has been implemented in Sandia National Laboratories’ Integrated TIGER Series (ITS) coupled electron-photon transport code. This work is meant to extend the available sensitivity estimation techniques in ITS by implementing an adjoint-based sensitivity method, GEAR-MC, to strengthen its sensitivity analysis capabilities. To ensure the accuracy of this method being extended to coupled electron-photon transport, it is compared against the central-difference and differential-operator methodologies to estimate sensitivity coefficients for an experiment performed by McLaughlin and Hussman. Energy deposition sensitivities were calculated using all three methods, and the comparison between them has provided confidence in the accuracy of the newly implemented method. Unlike the current implementation of the differential-operator method in ITS, the GEAR-MC method was implemented with the option to calculate the energy-dependent energy deposition sensitivities, which are the sensitivity coefficients for energy deposition tallies to energy-dependent cross sections. The energy-dependent cross sections could be the cross sections for the material, elements in the material, or reactions of interest for the element. Further, these sensitivities were compared to the energy-integrated sensitivity coefficients and exhibited a maximum percentage difference of 2.15%.

42 ENGINEERING↗

Verification of MOOSE/Bison's Heat Conduction Solver Using Combined Spatiotemporal Convergence Analysis

Bison is a computational physics code that uses the finite element method to model the thermo-mechanical response of nuclear fuel. Since Bison is used to inform high-consequence decisions, it is important that its computational results are reliable and predictive. One important step in assessing the reliability and predictive capabilities of a simulation tool is the verification process, which quantifies numerical errors in a discrete solution relative to the exact solution of the mathematical model. One step in the verification process—called code verification—ensures that the implemented numerical algorithm is a faithful representation of the underlying mathematical model, including partial differential or integral equations, initial and boundary conditions, and auxiliary relationships. In this paper, the code verification process is applied to spatiotemporal heat conduction problems in Bison. Simultaneous refinement of the discretization in space and time is employed to reveal any potential mistakes in the numerical algorithms for the interactions between the spatial and temporal components of the solution. For each verification problem, the correct spatial and temporal order of accuracy is demonstrated for both first- and second-order accurate finite elements and a variety of time-integration schemes. Furthermore, these results provide strong evidence that the Bison numerical algorithm for solving spatiotemporal problems reliably represents the underlying mathematical model in MOOSE. The selected test problems can also be used in other simulation tools that numerically solve for conduction or diffusion.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

A Comprehensive Review of Latent Space Dynamics Identification Algorithms for Intrusive and Non-Intrusive Reduced-Order-Modeling

Numerical solvers of partial differential equations (PDEs) have been widely employed for simulating physical systems. However, the computational cost remains a major bottleneck in various scientific and engineering applications, which has motivated the development of reduced-order models (ROMs). Recently, machine-learning-based ROMs have gained significant popularity and are promising for addressing some limitations of traditional ROM methods, especially for advection dominated systems. In this chapter, we focus on a particular framework known as Latent Space Dynamics Identification (LaSDI), which transforms the high-fidelity data, governed by a PDE, to simpler and low-dimensional latent-space data, governed by ordinary differential equations (ODEs). These ODEs can be learned and subsequently interpolated to make ROM predictions. Each building block of LaSDI can be easily modulated depending on the application, which makes the LaSDI framework highly flexible. In particular, we present strategies to enforce the laws of thermodynamics into LaSDI models (tLaSDI), enhance robustness in the presence of noise through the weak form (WLaSDI), select high-fidelity training data efficiently through active learning (gLaSDI, GPLaSDI), and quantify the ROM prediction uncertainty through Gaussian processes (GPLaSDI). We demonstrate the performance of different LaSDI approaches on Burgers equation, a non-linear heat conduction problem, and a plasma physics problem, showing that LaSDI algorithms can achieve relative errors of less than a few percent and up to thousands of times speed-ups.

Computational Engineering, Finance, and Science (c↗

Three dimensions, two microscopes, one code: Automatic differentiation for x-ray nanotomography beyond the depth of focus limit

Conventional tomographic reconstruction algorithms assume that one has obtained pure projection images, involving no within-specimen diffraction effects nor multiple scattering. Advances in x-ray nanotomography are leading toward the violation of these assumptions, by combining the high penetration power of x-rays, which enables thick specimens to be imaged, with improved spatial resolution that decreases the depth of focus of the imaging system. We describe a reconstruction method where multiple scattering and diffraction effects in thick samples are modeled by multislice propagation and the 3D object function is retrieved through iterative optimization. We show that the same proposed method works for both full-field microscopy and for coherent scanning techniques like ptychography. Our implementation uses the optimization toolbox and the automatic differentiation capability of the open-source deep learning package TensorFlow, demonstrating a straightforward way to solve optimization problems in computational imaging with flexibility and portability.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Electromagnetic Transient (EMT) Simulation Algorithm for Evaluation of Photovoltaic (PV) Generation Systems

Use of inverter-based resources facilitating renewable energy resources such as photovoltaic (PV) generation is increasing rapidly with decreasing costs and reduced emissions associated. To accommodate such rapid growth of inverter-based resources like PV systems, electromagnetic transient (EMT) simulation models of both PV systems and grids are required to analyze the interaction of PVs in the grid (like the post-event analysis). In addition, the EMT simulation would help with the planning of future power grid with a large number of PVs as well as other inverter-based distributed generation systems. In this paper, the EMT simulation models of PV systems and grids are developed based on the differential algebraic equations (DAEs) representing their EMT dynamics. Furthermore, advanced simulation algorithms including numerical stiffness-based hybrid discretization, DAE clustering and aggregation, multi-order integration, and matrix splitting approaches are applied to accelerate the EMT simulation. The proposed algorithm was applied to 125 PV inverters within 52-bus medium-voltage (MV) distribution grid.

Choi, Jongchan↗

A modified model parametrization algorithm for solving a special type of heat and mass transfer systems

A new method for solving nonlinear heat and mass transfer design tasks was considered. Systems using the Number of Transfer Units (NTU) method are a special type of mathematical model of heat and mass exchangers. It was observed, that the NTU models in a form of differential-algebraic equations (DAEs) cannot be directly solved with higher values of NTU. The requirements for consistent initial conditions, as well as numerical limitations of DAEs solvers, result, that the solution to the considered design problems that cannot be obtained by a classical direct shooting procedure. To overcome the presented difficulties, the αDAE model optimization algorithm was adjusted for solving NTU-based models. The new approach consists of 3 main steps: 1) task discretization by a multiple-shooting approach, 2) design an appropriate function $f_{NTU}$(α) to effectively influence the variability of the state variables described by dynamical relations, 3) the iterative numerical optimization algorithm for the new parametrized system. Moreover, computations can be performed by a chosen numerical optimization approach, which can be communicated with an available outer procedure for solving differential-algebraic equations. The presented algorithm was implemented and applied to solve the design task with the NTU model of a counter-flow exchanger. Here, the new approach was used to modify the system dynamics to influence the difficulty of the considered problem. Finally, the presented method enabled failure-free numerical computations for the higher values of the NTU parameter.

97 MATHEMATICS AND COMPUTING↗

Parallel projection—An improved return mapping algorithm for finite element modeling of shape memory alloys

Here, we present a novel finite element analysis of inelastic structures containing Shape Memory Alloys (SMAs). Phenomenological constitutive models for SMAs lead to material nonlinearities, that require substantial computational effort to resolve. Finite element analysis methods, which rely on Gauss quadrature integration schemes, must solve two sets of coupled differential equations: one at the global level and the other at the local, i.e. Gauss point level. In contrast to the conventional return mapping algorithm, which solves these two sets of coupled differential equations separately using a nested Newton procedure, we propose a scheme to solve the local and global differential equations simultaneously. In the process we also derive closed-form expressions used to update the internal/constitutive state variables, and unify the popular closest-point and cutting plane methods with our formulas. Numerical testing indicates that our method allows for larger thermomechanical loading steps and provides increased computational efficiency, over the standard return mapping algorithm.

42 ENGINEERING↗

Deliberate Motion Analytics Fused Radar and Video Test Results Deployed Beyond the Perimeter Fence in a High Noise Environment

Security systems that protect the nation’s critical facilities must be capable of detecting physical intrusions in all weather conditions. Intrusion detection sensors in a perimeter with a high nuisance alarm rate (NAR) significantly undermine detection performance and degrade security system effectiveness. This research demonstrated a fused sensor system that can differentiate foliage and weather-induced nuisance alarms from those caused by intruders, providing reliable detection within a two-fence perimeter or beyond the fence. A key element of this work is the creation and application of a “deliberate motion algorithm” that fuses alarm data from radar and video analytics to create video motion detection fused radar system. The two-layer architecture of the algorithm uses machine learning, multi-hypothesis tracking, and Dynamic Bayes Nets to differentiate intruder alarms from weather induced alarms.

47 OTHER INSTRUMENTATION↗

Discovering causal structure with reproducing-kernel Hilbert space ε -machines

We merge computational mechanics’ definition of causal states (predictively equivalent histories) with reproducing-kernel Hilbert space (RKHS) representation inference. The result is a widely applicable method that infers causal structure directly from observations of a system’s behaviors whether they are over discrete or continuous events or time. A structural representation—a finite- or infinite-state kernel ϵ-machine—is extracted by a reduced-dimension transform that gives an efficient representation of causal states and their topology. In this way, the system dynamics are represented by a stochastic (ordinary or partial) differential equation that acts on causal states. We introduce an algorithm to estimate the associated evolution operator. Paralleling the Fokker–Planck equation, it efficiently evolves causal-state distributions and makes predictions in the original data space via an RKHS functional mapping. We demonstrate these techniques, together with their predictive abilities, on discrete-time, discrete-value infinite Markov-order processes generated by finite-state hidden Markov models with (i) finite or (ii) uncountably infinite causal states and (iii) continuous-time, continuous-value processes generated by thermally driven chaotic flows. The method robustly estimates causal structure in the presence of varying external and measurement noise levels and for very high-dimensional data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Hyporheic-zone Processes and Stream Oxygen Dynamics: Insights from a Multiscale Reactive Transport Model: Modeling Archive

This archive contains the data and Python scripts required to reproduce the analyses and figures in the study: Gomez-Velez, J. D., Rathore, S. S., Cohen, M. J., & Painter, S. L. (2025). Hyporheic-zone Processes and Stream Oxygen Dynamics: Insights from a Multiscale Reactive Transport Model. Submitted to Water Resources Research. The analysis utilizes the subgrid model Advection Dispersion Equation with Lagrangian Subgrids (ADELS) implemented in the Advanced Terrestrial Simulator (ATS; https://amanzi.github.io/ats/stable/). In this case, the ATS and Amanzi versions are (1) ATS version 1.5.1_f5ba18f8 and (2) Amanzi version 1.6-dev_53444cca4. The repository includes a Jupyter Notebook and the necessary data (Pandas DataFrames stored as pickle files) to generate the figures for the manuscript. Additionally, it contains Python scripts to create ATS input files, run the ATS simulations, and post-process the results. Finally, it provides routines for parameter estimation using the Single-Station Metabolism (SSM) model with the Differential Evolution Adaptive Metropolis (DREAM) Markov Chain Monte Carlo (MCMC) algorithm with ZS enhancements (DREAM-ZS).

54 ENVIRONMENTAL SCIENCES↗

JAXtronomy: A JAX port of lenstronomy

Gravitational lensing is a phenomenon where light bends around massive objects, resulting in distorted images seen by an observer. Studying gravitationally lensed systems provides insights into cosmology and astrophysics, including constraints of the expansion rate of the Universe and the distribution of dark matter. Thus, we introduce JAXtronomy, a re-implementation of the gravitational lensing software package lenstronomy (Birrer, 2021; Birrer & Amara, 2018) using JAX (Bradbury et al., 2018). JAX is a Python library that uses an accelerated linear algebra (XLA) compiler to improve the performance of computing software. Our core design principle of JAXtronomy is to maintain an identical API to that of lenstronomy. The main JAX features utilized in JAXtronomy are just-in-time compilation, which can lead to significant reductions in execution time, and automatic differentiation, which allows for the implementation of gradient-based algorithms that were previously impossible. Additionally, JAX allows code to be run on GPUs or parallelized across CPU cores, further boosting the performance of JAXtronomy.

astronomy↗

Binary Quantum Control Optimization with Uncertain Hamiltonians

Optimizing the controls of quantum systems plays a crucial role in advancing quantum technologies. The time-varying noises in quantum systems and the widespread use of inhomogeneous quantum ensembles raise the need for high-quality quantum controls under uncertainties. In this paper, we consider a stochastic discrete optimization formulation of a discretized binary optimal quantum control problem involving Hamiltonians with predictable uncertainties. We propose a sample-based reformulation that optimizes both risk-neutral and risk-averse measurements of control policies, and solve these with two gradient-based algorithms using sum-up-rounding approaches. Furthermore, we discuss the differentiability of the objective function and prove upper bounds of the gaps between the optimal solutions to binary control problems and their continuous relaxations. We conduct numerical simulations on various sized problem instances based on two applications of quantum pulse optimization; we evaluate different strategies to mitigate the impact of uncertainties in quantum systems. In conclusion, we demonstrate that the controls of our stochastic optimization model achieve significantly higher quality and robustness compared with the controls of a deterministic model.

conditional value-at-risk (CVaR)↗

GPU-enabled extreme-scale turbulence simulations: Fourier pseudo-spectral algorithms at the exascale using OpenMP offloading

Fourier pseudo-spectral methods for nonlinear partial differential equations are of wide interest in many areas of advanced computational science, including direct numerical simulation of three-dimensional (3-D) turbulence governed by the Navier-Stokes equations in fluid dynamics. This paper presents a new capability for simulating turbulence at a new record resolution up to 35 trillion grid points, on the world's first exascale computer, Frontier, comprising AMD MI250x GPUs with HPE's Slingshot interconnect and operated by the US Department of Energy's Oak Ridge Leadership Computing Facility (OLCF). Key programming strategies designed to take maximum advantage of the machine architecture involve performing almost all computations on the GPU which has the same memory capacity as the CPU, performing all-to-all communication among sets of parallel processes directly on the GPU, and targeting GPUs efficiently using OpenMP offloading for intensive number-crunching including 1-D Fast Fourier Transforms (FFT) performed using AMD ROCm library calls. With 99% of computing power on Frontier being on the GPU, leaving the CPU idle leads to a net performance gain via avoiding the overhead of data movement between host and device except when needed for some I/O purposes. Memory footprint including the size of communication buffers for MPI_ALLTOALL is managed carefully to maximize the largest problem size possible for a given node count. Detailed performance data including separate contributions from different categories of operations to the elapsed wall time per step are reported for five grid resolutions, from 2048 3 on a single node to 32768 3 on 4096 or 8192 nodes out of 9408 on the system. Both 1D and 2D domain decompositions which divide a 3D periodic domain into slabs and pencils respectively are implemented. The present code suite (labeled by the acronym GESTS, GPUs for Extreme Scale Turbulence Simulations) achieves a figure of merit (in grid points per second) exceeding goals set in the Center for Accelerated Application Readiness (CAAR) program for Frontier. The performance attained is highly favorable in both weak scaling and strong scaling, with notable departures only for 2048 3 where communication is entirely intra-node, and for 32768 3 , where a challenge due to small message sizes does arise. Communication performance is addressed further using a lightweight test code that performs all-to-all communication in a manner matching the full turbulence simulation code. Performance at large problem sizes is affected by both small message size due to high node counts as well as dragonfly network topology features on the machine, but is consistent with official expectations of sustained performance on Frontier. Overall, although not perfect, the scalability achieved at the extreme problem size of 32768 3 (and up to 8192 nodes — which corresponds to hardware rated at just under 1 exaflop/sec of theoretical peak computational performance) is arguably better than the scalability observed using prior state-of-the-art algorithms on Frontier's predecessor machine (Summit) at OLCF. New science results for the study of intermittency in turbulence enabled by this code and its extensions are to be reported separately in the near future.

3D fast Fourier transform↗

ZEUS: An Efficient GPU Optimization Method Integrating PSO, BFGS, and Automatic Differentiation

We introduce a novel, efficient computational method, ZEUS, for numerical optimization, and provide an open-source implementation. It has four key ingredients: (1) particle swarm optimization (PSO), (2) the use of the Broyden-Fletcher-Goldfarb-Shanno (BFGS) method, (3) automatic differentiation (AD), and (4) GPUs. Our approach addresses the computational challenges inherent in high-dimensional, non-convex optimization problems. In the first phase of the algorithm, we get a potentially good set of starting points using PSO. Thereafter, we run BFGS independently in parallel from these starting points. BFGS is one of the best-performing algorithms for numerical optimization. However, it requires the gradient of the function being optimized. ZEUS integrates automatic differentiation into BFGS thus avoiding the need for the user to calculate derivatives explicitly. The use of GPUs allows ZEUS to speed up the calculations substantially. We carry out systematic studies to explore the trade-offs between the number of PSO iterations taken, starting points, and BFGS iteration depth. We show that a handful of iterations of PSO can improve global convergence when combined with BFGS. We also present performance studies using common test functions. The source code can be found at https://github.com/fnal-numerics/global-optimizer-gpu.

Soos, Dominik [Old Dominion U.]↗

Learning Stochastic Parametric Differentiable Predictive Control Policies

We present a scalable unsupervised learning-based method for obtaining explicit control policies for model predictive control problems for stochastic linear systems with additive uncertainties subject to nonlinear chance constraints. We call the proposed method stochastic parametric differentiable predictive control (SP-DPC), which extends the recently proposed deterministic DPC policy optimization algorithm. We formulate the SP-DPC as a deterministic approximation to the stochastic parametric constrained optimal control problem via independent sampling of the problem's parameters and uncertainties. This formulation allows us to directly compute the policy gradients via automatic differentiation of the problem's value function, evaluated over sampled parameters and uncertainties. In particular, the computed expectation of the problem's value function is backpropagated through the finite-time closed-loop system rollouts parametrized by a known nominal system dynamics model and neural control policy. We also provide theoretical probabilistic guarantees on closed-loop stability and chance constraints satisfaction for systems controlled by learned neural policies. We demonstrate the computational efficiency and scalability of the proposed policy optimization algorithm in three numerical examples, including systems with a large number of states or subject to nonlinear constraints.

Drgona, Jan↗

Verification of Bison fission product species conservation under TRISO reactor conditions

When assessing the reliability and predictive capabilities of a simulation tool, code verification is used to ensure that the implemented numerical algorithm is a faithful representation of its underlying mathematical model, including partial differential or integral equations, initial and boundary conditions, and auxiliary relationships. During this process, numerical results in a discrete solution are compared to the analytical solution of the mathematical model. Here, in this paper, the code verification process is applied to one-dimensional spatiotemporal problems that exercise partial differential equation governing the conservation of fission product species (or mass diffusion). Numerical experiments were performed in the Bison fuel performance code to evaluate its predictive capability under various TRISO reactor conditions such as base irradiation and safety heating test conditions for either short- or long-lived fission product species, as well as a case concerning evaporation from the outer surface of a particle. The code predictions were compared with the expected exact results obtained from the analytical expressions, and the fact that they demonstrate the correct analytical behavior provides strong evidence of proper numerical algorithm implementation.

07 ISOTOPE AND RADIATION SOURCES↗

Solving the Bernstein-Vazirani problem using Majorana-based topological quantum algorithms

Executing quantum algorithms using Majorana zero modes—a major milestone for the field of topological quantum computing—requires a platform that can be scaled to large quantum registers, can be controlled in real time and space, and a braiding protocol that uses the unique properties of these exotic particles. Here, we demonstrate the first successful simulation of a Majorana-based, fault-tolerant quantum algorithm to solve the Bernstein-Vazirani problem in two-dimensional magnet-superconductor hybrid structures from initialization to read-out of the final many-body state. Utilizing the Majorana zero modes’ topological properties, we introduce an optimized braiding protocol for the algorithm and a scalable architecture for its implementation with an arbitrary number of qubits. We visualize the algorithm protocol in real time and space by computing the non-equilibrium density of states, which is proportional to the time-dependent differential conductance, and the non-equilibrium charge density, which assigns a unique signature to each final state of the algorithm.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗