Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “linear equations”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

Milestone 49 Report: Batched Sparse LA Phase 5 Implementation

Batched sparse linear algebra operations in general, and solvers in particular, have become the major algorithmic development activity and foremost performance engineering effort in the numerical software libraries work on modern hardware with accelerators such as GPUs. Many applications, ECP and non-ECP alike, require simultaneous solutions of many small linear systems of equations that are structurally sparse in one form or another. In order to move towards high hardware utilization levels, it is important to provide these applications with appropriate interface designs to be both functionally efficient and performance portable and give full access to the appropriate batched sparse solvers running on modern hardware accelerators prevalent across DOE supercomputing sites since the inception of ECP. To this end, we present here a summary of recent advances on the interface designs in use by HPC software libraries supporting batched sparse linear algebra and the development of sparse batched kernel codes for solvers and preconditioners. We also address the potential interoperability opportunities to keep the corresponding software portable between the major hardware accelerators from AMD, Intel, and NVIDIA, while maintaining the appropriate disclosure levels conforming to the active NDA agreements. The presented interface specifications include a mix of batched band, sparse iterative, and sparse direct solvers with their accompanying functionality that is already required by the application codes or we anticipated to be needed in the near future. This report summarizes progress in Kokkos Kernels and the xSDK libraries MAGMA, Ginkgo, hypre, PETSc, and SuperLU.

97 MATHEMATICS AND COMPUTING↗

Juqbox.jl

The Juqbox.jl package implements functionality for solving the quantum optimal control problem for realizing logical gates in closed quantum systems. The dynamics of the quantum system is modeled by Schroedinger's equation, which takes to form of a linear system of ordinary differential equations (ODE). Juqbox.jl solves this ODE by numerical time stepping and applies a gradient-based optimization technique to determine control pulses for driving an initial state to a final state, according to the desired logical gate transformation. To evaluate the gradient, Juqbox.jl applies the ``first discretize, then optimize'' approach based on a discrete adjoint time stepping technique. The actual optimization is performed by the open source Ipopt libaray. Juqbox.jl is written in the Julia programming language which, among many other features, provides a convenient interface to the Ipopt library.

PETERSSON, NILSA.↗

A Quantum Approach for Implementing Fixed-Point Arithmetic in Solving Ordinary Differential Equations

Differential equations (DEs) serve as fundamental tools in mathematical modeling across scientific disciplines, yet classical numerical solvers face limitations with large-scale or computationally intensive problems. This study explores a quantum-inspired approach to solving DEs, combining quantum- inspired techniques with classical methods. It focuses on fixed- point arithmetic on quantum circuits, utilizing basic quantum gates to manipulate DE solutions. We expand upon the techniques introduced by Zanger et al. [Quantum, 5, 502 (2021)] by offering a precise computation for a fixed-point signed multiplication scheme, while also presenting a quantum circuit capable of executing the fixed-point division algorithm. We demonstrate the feasibility of our approach through the simulation of a linear Ordinary Differential Equation (ODE), where initial conditions and parameters are encoded into quantum circuits using fixed- point representation. By executing sequences of quantum gates mimicking numerical integration steps, we obtain approximate solutions to the ODE with specified fixed-point precision.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

A Quantum Approach for Implementing Fixed-Point Arithmetic in Solving Ordinary Differential Equations

Differential equations (DEs) serve as fundamental tools in mathematical modeling across scientific disciplines, yet classical numerical solvers face limitations with large-scale or computationally intensive problems. This study explores a quantum-inspired approach to solving DEs, combining quantum-inspired techniques with classical methods. It focuses on fixed-point arithmetic on quantum circuits, utilizing basic quantum gates to manipulate DE solutions. We expand upon the techniques introduced by Zanger et al. [Quantum, 5, 502 (2021)] by offering a precise computation for a fixed-point signed multiplication scheme, while also presenting a quantum circuit capable of executing the fixed-point division algorithm. We demonstrate the feasibility of our approach through the simulation of a linear Ordinary Differential Equation (ODE), where initial conditions and parameters are encoded into quantum circuits using fixed-point representation. By executing sequences of quantum gates mimicking numerical integration steps, we obtain approximate solutions to the ODE with specified fixed-point precision.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Learning in Modal Space: Solving Time-Dependent Stochastic PDEs Using Physics-Informed Neural Networks

One of the open problems in scientific computing is the long-time integration of nonlinear stochastic partial differential equations (SPDEs), especially with arbitrary initial data. We address this problem by taking advantage of recent advances in scientific machine learning and the spectral dynamically orthogonal (DO) and borthogonal (BO) methods for representing stochastic processes. The recently introduced DO/BO methods reduce the SPDE to solving a system of deterministic PDEs and a system of stochastic ordinary differential equations. Specifically, we propose two new physics-informed neural networks (PINNs) for solving time-dependent SPDEs, namely the neural network (NN)-DO/BO methods. The proposed methods incorporate the DO/BO constraints into the loss function (along with the modal decomposition of the SPDE) with an implicit form instead of generating explicit expressions for the temporal derivatives of the DO/BO modes. Hence, the NN-DO/BO methods can overcome some of the drawbacks of the original DO/BO methods. For example, we do not need the assumption that the covariance matrix of the random coefficients is invertible as in the original DO method, and we can remove the assumption of no eigenvalue crossing as in the original BO method. Moreover, the NN-DO/BO methods can be used to solve time-dependent stochastic inverse problems with the same formulation and same computational complexity as for forward problems. Furthermore, we demonstrate the capability of the proposed methods via several numerical examples, namely: (1) A linear stochastic advection equation with deterministic initial condition: we obtain good results with the proposed methods, while the original DO/BO methods cannot be applied directly in this case. (2) Long-time integration of the stochastic Burgers' equation: we show the good performance of NN-DO/BO methods, especially the effectiveness of the NN-BO approach for such problems with many eigenvalue crossings during the whole time evolution, while the original BO method fails. (3) Nonlinear reaction diffusion equation: we consider both the forward problem and the inverse problems, including very noisy initial point values, to investigate the flexibility of the NN-DO/BO methods in handling inverse and mixed type problems. Taken together, these simulation results demonstrate that the NN-DO/BO methods can be employed to effectively quantify uncertainty propagation in a wide range of physical problems, but future work should address the efficiency issue of PINNs for forward problems.

97 MATHEMATICS AND COMPUTING↗

Local reduced-order modeling for electrostatic plasmas by physics-informed solution manifold decomposition

Despite advancements in high-performance computing and modern numerical algorithms, computational cost remains prohibitive for multi-query kinetic plasma simulations. Here, in this work, we develop data-driven reduced-order models (ROMs) for collisionless electrostatic plasma dynamics, based on the kinetic Vlasov-Poisson equation. Our ROM approach projects the equation onto a linear subspace defined by the proper orthogonal decomposition (POD) modes. We introduce an efficient tensorial method to update the nonlinear term using a precomputed third-order tensor. We capture multiscale behavior with a minimal number of POD modes by decomposing the solution manifold into multiple time windows and creating temporally local ROMs. We consider two strategies for decomposition: one based on the physical time and the other based on the electric field energy. Applied to the 1D1V Vlasov–Poisson simulations, that is, prescribed E-field, Landau damping, and two-stream instability, we demonstrate that our ROMs accurately capture the total energy of the system both for parametric and time extrapolation cases. The temporally local ROMs are more efficient and accurate than the single ROM. In addition, in the two-stream instability case, we show that the energy-windowing reduced-order model (EW-ROM) is more efficient and accurate than the time-windowing reduced-order model (TW-ROM). With the tensorial approach, EW-ROM solves the equation approximately 90 times faster than Eulerian simulations while maintaining a maximum relative error of 7.5% for the training data and 11% for the testing data.

Electrostatic plasmas↗

A Characteristics Approach to the Finite Element Method

Herein, we present a new method for solving the linear Boltzmann transport equation. Two commonly used and well-understood methods for solving partial differential equations are the method of characteristics (MOC) and the finite element method (FEM). We propose a new method that combines the fundamental concept of the FEM with the analytic solution from the MOC to obtain coefficients for the FEM basis function expansion. Traditionally, coefficients for the FEM basis function expansion are obtained via matrix inversion. Instead, we solve for the coefficients with the MOC and represent the underlying fields with the basis function expansion using these coefficients. We provide a convergence study for our method with results from two sets of FEM basis functions: Gauss-Legendre and Gauss-Lobatto sets. We also compare two different variations of our method categorized as short characteristics and intermediate characteristics.

42 ENGINEERING↗

A smeared crack modeling framework accommodating multi-directional fracture at finite strains

A generic smeared crack modeling framework predicated on the deformation gradient decomposition (DGD) approach is proposed for use in dynamic fracture problems at finite strains, accommodating failure along multiple mutually orthogonal fracture planes embedded within an independently defined bulk material model. Within this constitutive framework, the traction equilibrium conditions imposed at each failure surface are used to determine the associated crack displacements stored as internal state variables. In general, the enforcement of interfacial equilibrium entails the implicit solution of a non-linear system of equations within the constitutive update procedure. However, if inertial effects arising due to the relative motion of the fractured material are incorporated within the model, the traction equilibrium conditions are shown to give rise to corresponding dynamic equations of motion governing the time-evolution of the crack opening displacements. For dynamic problems, an explicit time-integration procedure is devised to efficiently update the material state, subject to a set of internal frictionless contact constraints to prevent material inversion. Finally, the efficacy of the proposed modeling framework is investigated through several benchmark dynamic fracture problems run within the explicit finite element code DYNA3D.

42 ENGINEERING↗

Asynchronous Richardson iterations: theory and practice

We consider asynchronous versions of the first- and second-order Richardson methods for solving linear systems of equations. These methods depend on parameters whose values are chosen a priori. We explore the parameter values that can be proven to give convergence of the asynchronous methods. This is the first such analysis for asynchronous second-order methods. We find that for the first-order method, the optimal parameter value for the synchronous case also gives an asynchronously convergent method. For the second-order method, the parameter ranges for which we can prove asynchronous convergence do not contain the optimal parameter values for the synchronous iteration. In practice, however, the asynchronous second-order iterations may still converge using the optimal parameter values, or parameter values close to the optimal ones, despite this result. We explore this behavior with a multithreaded parallel implementation of the asynchronous methods.

97 MATHEMATICS AND COMPUTING↗

Leveraging operator learning to accelerate convergence of the preconditioned conjugate gradient method

We propose a new deflation strategy to accelerate the convergence of the preconditioned conjugate gradient (PCG) method for solving parametric large-scale linear systems of equations. Unlike traditional deflation techniques that rely on eigenvector approximations or recycled Krylov subspaces, we generate the deflation subspaces using operator learning, specifically the Deep Operator Network (DeepONet). To this aim, we introduce two complementary approaches for assembling the deflation operators. The first approach approximates near-null space vectors of the discrete PDE operator using the basis functions learned by the DeepONet. The second approach directly leverages solutions predicted by the DeepONet. To further enhance convergence, we also propose several strategies for prescribing the sparsity pattern of the deflation operator. Here, a comprehensive set of numerical experiments encompassing steady-state, time-dependent, scalar, and vector-valued problems posed on both structured and unstructured geometries is presented and demonstrates the effectiveness of the proposed DeepONet-based deflated PCG method, as well as its generalization across a wide range of model parameters and problem resolutions.

Deflation↗

Simulation toolkit for digital material characterization of large image-based microstructures

In this paper, an efficient image-based simulation toolkit for material characterization is presented, which is scalable to work from personal computers to workstations. The effective thermal conductivity, elasticity, and permeability are evaluated employing a computational homogenization framework based on the Finite Element Method (FEM). Two complementary open-source packages are presented: one developed in Python, which can convert digital images into voxel meshes (pyTomoviewer); the other developed in Julia, that can run numerical simulations to compute effective material properties (chpack). Also, a CUDA C version of chpack is provided (chfem_gpu). They were designed to deal with large multi-phase models, so strategies were devised to minimize their memory footprint, while avoiding a high toll on execution time. The voxel-based approach significantly simplifies the FEM meshes and allows efficient matrix-free implementations. In that sense, to handle large linear systems of equations, the element-by-element (EBE) technique is adopted, in conjunction with a low-memory implementation of the Preconditioned Conjugate Gradient (PCG) method. Finally, the code was thoroughly tested on an artificial geometry made of a square array of cylinders, for which analytical solutions exist, as well as on a real micro-tomographic reconstruction of FiberForm TM , a carbon preform commonly used in thermal protection systems.

36 MATERIALS SCIENCE↗

QUIC-URB and QUIC-fire extension to complex terrain: Development of a terrain-following coordinate system

Ensemble-based approaches to prescribed fire planning cannot be supported by CFD-based models like FIRETEC and WFDS because they are too computationally expensive and cannot leverage LES approaches like CAWFE and WRF-SFIRE because too coarse of resolution. QUIC-Fire was developed to fill this gap but it cannot currently address complex terrain, typical for instance of the Western United States. In this paper, we describe the extension of the diagnostic wind model QUIC-URB, the wind engine of QUIC-Fire, to a terrain-following coordinate system. In particular, the paper presents the mathematical derivation of the wind solver leading to a linear system of equations that are solved through the successive over-relaxation method. The model is validated against a standard test used in previous works (the Askervein Hill) and against a new dataset from measurements in the Socorro Mountains, New Mexico. The terrain-following implementation captured the correct phenomenology for the isolated Askervein Hill, with a wind speed up at the top of the hill. We report the model agreed well with measurements on the upwind side of the peak, but overestimated speed-up on the downwind side of the hill. This is due to the inability of the model to generate flow separation and wake-eddy dynamics. On a common laptop, the divergence-free wind field was obtained in 6 s, making the solver appealing for coupled fire–atmosphere simulations. The Socorro Mountain was highly complex, with many cliff faces, peaks, and valleys. Although the model captures the magnitude and direction of inlet and outlet areas of the domain, it performs rather poorly in the valley region and in the regions near the steep cliffs. Hence, the model shows good agreement with data in areas of open sloped terrain but lacks in areas where flow separation and thermally driven effects may be present (neither effect was addressed in this work). Results highlight that future work should focus on the implementation of parameterizations of wake-eddies, similar to QUIC-URB’s building parameterizations, and on thermodynamic-driven flow.

54 ENVIRONMENTAL SCIENCES↗

Analysis and prediction of reaction kinetics using the degree of rate control

“Degree of rate control” (DRC) analysis provides a quantitative approach for analysing the kinetics of multi-step reaction mechanisms that has been widely applied to both heterogeneous and homogeneous catalysis research, as well as electrocatalysis. The DRC of any given transition state or intermediate is defined as a partial derivative such that it approximately equals the fractional increase in net rate to the product of interest per differential decrease in its standard-state free energy for that species (÷RT), holding constant the standard-state free energies of all other transition states and intermediates. Even very complex mechanisms usually have only a few species with non-zero DRCs and are thus the species whose interactions with the catalyst most strongly affect the net rate. These key DRC values thus offer a simple and intuitive route to optimize catalyst materials, especially with the assistance of computational methods like density functional theory (DFT). These high-DRC species are also the species whose energetics must be most accurately measured or calculated to achieve an accurate kinetic model for any reaction mechanism. In simple cases with a single “rate-determining step”, the DRC for its transition state (TS) is + 1. Catalyst-bound intermediates, on the other hand, often have negative DRCs equal to a small integer times their fractional occupancy of catalyst sites. The apparent activation energy equals a weighted average of the standard-state enthalpies (relative to reactants) of all the species (intermediates, transition states and products) in the reaction mechanism, each weighted by its DRC (+RT). It has been shown that the apparent transfer coefficient in electrocatalysis, an inverted form of the Tafel slope, is a weighted average of the number of electrons transferred to generate each intermediate or product species in the mechanism, weighted again by the DRC. Quantitative analysis of kinetic isotope effects (KIEs, or the ratio of net rates for different reactant isotopes) in complex mechanisms has shown that the logarithm of the KIE equals the weighted average over all species in the mechanism of the difference between the two isotopes in their standard-state free energies (÷RT), again weighted by the DRC. The reaction orders with respect to fluid-phase concentrations of reactants, products and intermediates have also been proven to be directly related to DRCs. Thus, there are numerous experimental observables which equate to short linear combinations of DRCs, so that combinations of experimental measurements might provide access to DRC values. Since its invention for transition states in 1994 and its generalization to include intermediates in 2009, the DRC has thus far mainly been calculated for microkinetic models based either entirely on DFT or on DFT with the key energies (i.e., those for high-DRC species) being fine-tuned to match experiments. The relationships summarized in this work provide new opportunities for using experiments earlier in the development and optimization of microkinetic models that require input from computational catalysis.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Site heterogeneity and broad surface-binding isotherms in modern catalysis: Building intuition beyond the Sabatier principle

Learning the science of heterogeneous catalysis and electrocatalysis always starts with the simple case of a flat, uniform surface with an ideal adsorbate. It has of course been recognized for a century that real catalysts are more complicated. For the increasingly complex catalysts of the 21st century, this Perspective argues that surface heterogeneity and non-ideal binding isotherms are central features, and their implications need to be incorporated in current thinking. A variety of systems are described herein where catalyst complexity leads to broad, non-Langmuirian surface isotherms for the binding of hydrogen atoms – and this occurs even for ideal, flat Pt(111) surfaces. Modern catalysis employs nanoscale materials whose surfaces have substantial step, edge, corner, impurity, and other defect sites, and they increasingly have both metallic and non-metallic elements M n X m , including metal oxides, chalcogenides, pnictides, carbides, doped carbons, etc. The surfaces of such catalysts are often not crystal facets of the bulk phase underneath, and they typically have a variety of potential active sites. Catalytic surfaces in operando are often non-stoichiometric, amorphous, dynamic, and impure, and often vary from one part of the surface to another. Understanding of the issues that arise at such nanoscale, multi-element catalysts is just beginning to emerge. Yet these catalysts are widely discussed using Brønsted/Bell-Evans-Polanyi (BEP) relations, volcano plots, Tafel slopes, the Butler-Volmer equation, and other linear free energy relations (LFERs), which all depend on the implicit assumption that the active sites are “similar” and that surface adsorption is close to ideal. These assumptions underly the ubiquitous intuition based on the Sabatier Principle, that the fastest catalysis will occur when key intermediates have free energies of adsorption that are not too strong nor too weak. Current catalysis research often aims to minimize the complexity of non-ideal isotherms through experimental and computational design (e.g., the use of single crystal surfaces), and these studies are the foundation of the field. In contrast, this Perspective argues that the heterogeneity of binding sites and binding energies is an inherent strength of these catalysts. Here, this diversity makes many nanoscale catalysts inherently a high-throughput screen wrapped in a tiny package. Only by making the heterogeneity part of the foundation of catalysis models, sorting the types of active sites and dissecting non-ideal binding isotherms, will modern catalysis learn to harness the inherent diversity of real catalysts. Controlling and exploiting diversity rather than avoiding it will help to optimize complex modern catalysts and catalytic conditions.

Mayer, James M.↗

Accelerating eigenvalue computation for nuclear structure calculations via perturbative corrections

Subspace projection methods utilizing perturbative corrections have been proposed for computing the lowest few eigenvalues and corresponding eigenvectors of large Hamiltonian matrices. In this paper, we build upon these methods and introduce the term Subspace Projection with Perturbative Corrections (SPPC) method to refer to this approach. We tailor the SPPC for nuclear many-body Hamiltonians represented in a truncated configuration interaction subspace, i.e., the no-core shell model (NCSM). We use the hierarchical structure of the NCSM Hamiltonian to partition the Hamiltonian as the sum of two matrices. The first matrix corresponds to the Hamiltonian represented in a small configuration space, whereas the second is viewed as the perturbation to the first matrix. Eigenvalues and eigenvectors of the first matrix can be computed efficiently. Because of the split, perturbative corrections to the eigenvectors of the first matrix can be obtained efficiently from the solutions of a sequence of linear systems of equations defined in the small configuration space. These correction vectors can be combined with the approximate eigenvectors of the first matrix to construct a subspace from which more accurate approximations of the desired eigenpairs can be obtained. We show by numerical examples that the SPPC method can be more efficient than conventional iterative methods for solving large-scale eigenvalue problems such as the Lanczos, block Lanczos and the locally optimal block preconditioned conjugate gradient (LOBPCG) method. The method can also be combined with other methods to avoid convergence stagnation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

A hybrid Monte Carlo-deterministic second moment method with efficient variance reduction

In this work, we present a hybrid method that combines Monte Carlo with deterministic finite element methods to solve a linear Boltzmann transport equation. Our hybrid method runs orders of magnitude faster than Monte Carlo, without sacrificing accuracy, for a proxy problem from radiative transfer that contains both optically-thick and optically-thin material. We believe that this is the first demonstration of a hybrid Second Moment Method in more than one spatial dimension, the first to consider more than one material, and the first to use variance reduction. Our variance reduction approach arises from an asymptotic analysis in which we show that the magnitude of the scattering source grows without bound. We transform the problem to compute the deviation of the radiation intensity from isotropy. The magnitude of the source in the transformed problem is bounded, and the quality of the hybrid method solution is dramatically improved by a substantial reduction in the variance.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Weak-Bonding Elements Lead to High Thermoelectric Performance in BaSnS 3 and SrSnS 3 : A First-Principles Study

SnS2, an earth-abundant and ecofriendly material, is limited as a thermoelectric material because of the high lattice thermal conductivity κ L and low carrier mobility μ. By introducing weak-bonding elements Ba or Sr into the SnS 2 framework, we discovered two SnS 2 -based materials BaSnS3 and SrSnS3 with the calculated low κL values of 0.15 and 0.17 W m -1 K -1 , respectively, along the a-axis. The low group velocity and high lattice anharmonicity originating from the weakened and distorted Sn–S bonding network are found in both systems. Moreover, the vibrations of Ba and Sr induce low-lying optical phonons, which strongly couple with the acoustic phonons and strengthen the phonon scattering rates. Compared to SnS 2 , both compounds present lower single-band effective masses, smaller deformation potential constants, and better band convergence, which enhance μ with an insignificantly reduced effective mass. By solving the linearized Boltzmann transport equation with a nonempirical carrier lifetime, we predict excellent ZT values of 2.89 and 2.77 along the a-axis at 900 K in BaSnS 3 and SrSnS 3 , respectively. Further phase diagram calculations of Ba 1–x Sr x SnS 3 solid solutions propose a new compound, Ba 0.5 Sr 0.5 SnS 3 , with an even higher ZT of 3.0. Our work analyzes explicitly how weak-bonding elements enhance μ and suppress κL simultaneously in SnS 2 -analogous systems with a series of compounds nominated as potential high-performance thermoelectric materials.

36 MATERIALS SCIENCE↗