Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “direct solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

jaxhps: An elliptic PDE solver built with machine learning in mind

Elliptic partial differential equations (PDEs) can model many physical phenomena, such as electrostatics, acoustics, wave propagation, and diffusion. In scientific machine learning settings, a high-throughput PDE solver may be required to generate a training dataset, run in the inner loop of an iterative algorithm, or interface directly with a deep neural network. To provide value to machine learning users, such a PDE solver must be compatible with standard automatic differentiation frameworks, scale efficiently when run on graphics processing units (GPUs), and maintain high accuracy for a large range of input parameters. We have designed the jaxhps package with these use-cases in mind by implementing a highly efficient and accurate solver for elliptic problems with native hardware acceleration and automatic differentiation support.

97 MATHEMATICS AND COMPUTING↗

Micro-continuum approach for mineral precipitation

Abstract Rates and extents of mineral precipitation in porous media are difficult to predict, in part because laboratory experiments are problematic. It is similarly challenging to implement numerical methods that model this process due to the need to dynamically evolve the interface of solid material. We developed a multiphase solver that implements a micro-continuum simulation approach based on the Darcy–Brinkman–Stokes equation to study mineral precipitation. We used the volume-of-fluid technique in sharp interface implementation to capture the propagation of the solid mineral surface. Additionally, we utilize an adaptive mesh refinement method to improve the resolution of near interface simulation domain dynamically. The developed solver was validated against both analytical solution and Arbitrary Lagrangian–Eulerian approach to ensure its accuracy on simulating the propagation of the solid interface. The precipitation of barite (BaSO 4 ) was chosen as a model system to test the solver using variety of simulation parameters: different geometrical constraints, flow conditions, reaction rate and ion diffusion. The growth of a single barite crystal was simulated to demonstrate the solver’s capability to capture the crystal face specific directional growth.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Shadow molecular dynamics for flexible multipole models

Shadow molecular dynamics provide an efficient and stable atomistic simulation framework for flexible charge models with long-range electrostatic interactions. Shadow molecular dynamics simulations are driven by approximate “shadow” Born–Oppenheimer potentials for which the exact charges and forces are directly accessible without relying on costly (and approximate) iterative solvers. While previous implementations have been limited to atomic monopole charge distributions, we extend this approach to flexible multipole models. We derive detailed expressions for the shadow energy functions, potentials, and force terms, explicitly incorporating monopole–monopole, dipole–monopole, and dipole–dipole interactions. In our formulation, both atomic monopoles and atomic dipoles are treated as extended dynamical variables alongside the propagation of the nuclear degrees of freedom. We demonstrate that introducing the additional dipole degrees of freedom preserves the stability and accuracy previously seen in monopole-only shadow molecular dynamics simulations. In addition, we present a shadow molecular dynamics scheme where the monopole charges are held fixed while the dipoles remain flexible. Our extended shadow dynamics provide a framework for stable, computationally efficient, and versatile molecular dynamics simulations involving long-range interactions between flexible multipoles. This is of particular current interest in combination with machine-learned interatomic potentials, including long-range electrostatic interactions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fourier Neural Networks as Function Approximators and Differential Equation Solvers

We present a Fourier neural network (FNN) that can be mapped directly to the Fourier decomposition. The choice of activation and loss function yields results that replicate a Fourier series expansion closely while preserving a straightforward architecture with a single hidden layer. The simplicity of this network architecture facilitates the integration with any other higher-complexity networks, at a data pre- or postprocessing stage. We validate this FNN on naturally periodic smooth functions and on piecewise continuous periodic functions. We showcase the use of this FNN for modeling or solving partial differential equations with periodic boundary conditions. The main advantages of the current approach are the validity of the solution outside the training region, interpretability of the trained model, and simplicity of use.

Fourier decomposition↗

A modified model parametrization algorithm for solving a special type of heat and mass transfer systems

A new method for solving nonlinear heat and mass transfer design tasks was considered. Systems using the Number of Transfer Units (NTU) method are a special type of mathematical model of heat and mass exchangers. It was observed, that the NTU models in a form of differential-algebraic equations (DAEs) cannot be directly solved with higher values of NTU. The requirements for consistent initial conditions, as well as numerical limitations of DAEs solvers, result, that the solution to the considered design problems that cannot be obtained by a classical direct shooting procedure. To overcome the presented difficulties, the αDAE model optimization algorithm was adjusted for solving NTU-based models. The new approach consists of 3 main steps: 1) task discretization by a multiple-shooting approach, 2) design an appropriate function $f_{NTU}$(α) to effectively influence the variability of the state variables described by dynamical relations, 3) the iterative numerical optimization algorithm for the new parametrized system. Moreover, computations can be performed by a chosen numerical optimization approach, which can be communicated with an available outer procedure for solving differential-algebraic equations. The presented algorithm was implemented and applied to solve the design task with the NTU model of a counter-flow exchanger. Here, the new approach was used to modify the system dynamics to influence the difficulty of the considered problem. Finally, the presented method enabled failure-free numerical computations for the higher values of the NTU parameter.

97 MATHEMATICS AND COMPUTING↗

Mesoflow: An Open-Source Reacting Flow Solver for Catalysis at Mesoscale

We present the capabilities and software performance metrics of our open-source continuum solver for catalysis, Mesoflow, developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex catalyst surface morphologies directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous catalyst interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPUs) compared to a single processor for representative problem sizes (2 million cell mesh). We will also present a brief introduction on how to build and use this software for application problems pertaining to catalytic upgrading and gas transport within porous catalyst particles.

adaptive meshing↗

HTR-1.3 solver: Predicting electrified combustion using the hypersonic task-based research solver

Here this manuscript presents an updated open-source version of the Hypersonics Task-based Research (HTR) solver. The solver, whose main features are presented in Di Renzo et al. (2020) and Di Renzo & Pirozzoli (2021), is designed for direct numerical simulation of reacting flows at high Reynolds numbers. This new version extends the applications of the HTR solver to turbulent combustion in the presence of external electric fields. In particular, a new distributed Poisson solver compatible with heterogeneous architectures has been incorporated in the algorithm to compute the electric potential distribution in bi-periodic configurations. The drift fluxes of the electrically charged species are now included in the transport equations using a targeted essentially non-oscillatory scheme. A verification of these new features of the solver is provided using one-dimensional burner stabilized flames, whereas a three dimensional turbulent flame is utilized to discuss the scalability of the proposed numerical tool.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

mesoflow [SWR-22-56]

Mesoflow is a continuum scale simulation tool developed specifically for modeling transport and chemistry at the mesoscale. Our solver utilizes Cartesian block-structured adaptive mesh refinement to resolve complex surface morphologies (of catalysts/biomass particles among others) directly obtained from X-ray tomography data. An immersed boundary based formulation enables rapid representation of complex geometries prevalent in most mesoporous interfaces. The solver is developed on top of open-source performance portable library, AMReX, providing parallel execution capabilities on current and upcoming high-performance-computing (HPC) architectures. Our flexible software framework enables integration of complex chemical mechanisms at heterogenous interfaces and time-split algorithms for circumventing highly disparate reaction and flow time-scales. Our current studies indicate a ten-fold performance gain by using graphics-processing-units (GPU) compared to a single processor for representative problem sizes (2 million cell mesh).

Sitaraman, Hariswaran↗

Enabling attractive-repulsive potentials in binary-collision-approximation monte-carlo codes for ion-surface interactions

Abstract Binary Collision Approximation (BCA) codes for ion-material interactions, such as SRIM, Tridyn, F-TRIDYN, and SDtrimSP, have historically been limited to screened Coulomb potentials even at low energies due to the difficulty in numerically solving the Distance of Closest Approach (DOCA) problem for attractive-repulsive potentials. Techniques such as direct n-body simulation or modifications to Newton’s method are either prohibitively costly or not guaranteed to work for all potentials. Advanced rootfinding techniques, such as companion matrix solvers, offer a solution. For many attractive-repulsive potentials, however, a companion matrix cannot be used directly, because there is no way to put the associated functions into a monomial basis form. A complementary technique is proxy rootfinding—by finding the best-fit polynomial approximant of a function, the zeros of the approximant can be guaranteed to be close to the zeros of the function. Using the Chebyshev basis and grid offers additional guarantees with regards to the quality of the approximation, the speed of convergence, and the avoidance of Runge’s phenomenon. By finding Chebyshev interpolants and using the Chebyshev-Frobenius companion matrix, the zeros of any real function on a bounded domain can be found. Here we show that using an Adaptive Chebyshev Proxy Rootfinder with Automatic Subdivision (ACPRAS) with appropriate scaling functions, numerical issues presented by attractive-repulsive potentials, including those of scale, can be handled. Using these techniques, we show that it is possible to include any physically reasonable interatomic potential in a BCA code, and to guarantee correctness of the resulting scattering angle calculations.

Materials Science↗

Application of linear prolongation to coarse mesh finite difference acceleration in CASMO5

The Coarse-Mesh Finite Difference (CMFD) method has been used for over a decade to accelerate the convergence of the Method of Characteristics (MOC) solution to the two- dimensional particle transport equation in CASMO5. Numerical testing, along with widespread use in production-level calculations, have shown that the current CMFD implementation provides stability and robustness for a wide range of realistic reactor physics problems. However, the recent development of linear prolongation has attracted attention from the community as a way to further improve the performance and stability of CMFD. Two interpolation methods for linear prolongation are presented in this work and implemented into a test version of CASMO5. The performance of the proposed interpolations, referred to as the bilinear and linear directional schemes, is evaluated in terms of runtime relative to the default constant or uniform scaling update. Numerical results indicate that the use of linear prolongation can reduce the transport solver runtime on average by approximately 10% when tested with two hundred randomly selected cases. The new directional linear interpolation, combined with default constant boundary updates, is found to provide the highest reduction in runtime for the cases analyzed. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Magnetic tunnel junction random number generators applied to dynamically tuned probability trees driven by spin orbit torque

Abstract Perpendicular magnetic tunnel junction (pMTJ)-based true-random number generators (RNGs) can consume orders of magnitude less energy per bit than CMOS pseudo-RNGs. Here, we numerically investigate with a macrospin Landau–Lifshitz-Gilbert equation solver the use of pMTJs driven by spin–orbit torque to directly sample numbers from arbitrary probability distributions with the help of a tunable probability tree. The tree operates by dynamically biasing sequences of pMTJ relaxation events, called ‘coinflips’, via an additional applied spin-transfer-torque current. Specifically, using a single, ideal pMTJ device we successfully draw integer samples on the interval [0, 255] from an exponential distribution based on p -value distribution analysis. In order to investigate device-to-device variations, the thermal stability of the pMTJs are varied based on manufactured device data. It is found that while repeatedly using a varied device inhibits ability to recover the probability distribution, the device variations average out when considering the entire set of devices as a ‘bucket’ to agnostically draw random numbers from. Further, it is noted that the device variations most significantly impact the highest level of the probability tree, with diminishing errors at lower levels. The devices are then used to draw both uniformly and exponentially distributed numbers for the Monte Carlo computation of a problem from particle transport, showing excellent data fit with the analytical solution. Finally, the devices are benchmarked against CMOS and memristor RNGs, showing faster bit generation and significantly lower energy use.

77 NANOSCIENCE AND NANOTECHNOLOGY↗

Fast inversion, preconditioned quantum linear system solvers, fast Green's-function computation, and fast evaluation of matrix functions

Preconditioning is the most widely used and effective way for treating ill-conditioned linear systems in the context of classical iterative linear system solvers. We introduce a quantum primitive called fast inversion, which can be used as a preconditioner for solving quantum linear systems. The key idea of fast inversion is to directly block encode a matrix inverse through a quantum circuit implementing the inversion of eigenvalues via classical arithmetics. We demonstrate the application of preconditioned linear system solvers for computing single-particle Green's functions of quantum many-body systems, which are widely used in quantum physics, chemistry, and materials science. We analyze the complexities in three scenarios: the Hubbard model, the quantum many-body Hamiltonian in the plane-wave-dual basis, and the Schwinger model. We also provide a method for performing Green's function calculation in second quantization within a fixed-particle manifold and note that this approach may be valuable for simulation more broadly. Aside from solving linear systems, fast inversion also allows us to develop fast algorithms for computing matrix functions, such as the efficient preparation of Gibbs states. Furthermore, we introduce two efficient approaches for such a task, based on the contour-integral formulation and the inverse transform, respectively.

97 MATHEMATICS AND COMPUTING↗

A Perspective on Quantum Computing Applications in Quantum Chemistry Using 25-100 Logical Qubits

The intersection of quantum computing and quantum chemistry represents a promising frontier for achieving quantum utility in domains of both scientific and societal relevance. Owing to the exponential growth of classical resource requirements for simulating quantum systems, quantum chemistry has long been recognized as a natural candidate for quantum computation. This perspective focuses on identifying scientifically meaningful use cases where early fault-tolerant quantum computers, which are considered to be equipped with approximately 25-100 logical qubits, could deliver tangible impact. While recent advances in classical computing have pushed the boundaries of tractable simulations to unprecedented scales, this logical-qubit regime represents the first window where quantum devices can pursue qualitatively distinct strategies, such as polynomial-scaling phase estimation, direct simulation of quantum dynamics, and active-space embedding, that remain challenging for classical solvers, such as multireference charge-transfer and conical-intersection states central to photochemistry and materials design. We highlight near-term opportunities in algorithm and software design, discuss representative chemical problems suited for quantum acceleration, and propose strategic roadmaps and collaborative pathways for advancing practical quantum utility in quantum chemistry.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Comparison of Some RANS Solvers

We will take a look at solving the Reynolds-averaged Navier-Stokes (RANS) equations that are encountered in the context of wind farm performance simulations and optimizations. We will compare some of the more popular ways to solve these equations with a focus on using iterative solvers for the linear solve. We will compare their performance and reliability to a direct solve as we scale the problem both by adding more and parallel resources and by increasing the size of the domain, both in two and three dimensions. There are many strategies that can be applied to solving the RANS equations, some are very efficient, while others are very insensitive to, for example, the Reynolds number. The first contender we will consider is the Pressure-Convection-Diffusion (PCD) preconditioner. Early results suggest that PCD is indeed a very efficient solver, in particular in two dimensions, as long as the Reynold's number remains small. Next we will try to reorder our degrees of freedom such that we can use GMRES with ILU for our linear solve. Another popular choice we will consider for solving the RANS equations is SIMPLE (and its derivatives). For all of our implementations we make use of either FEniCS or Firedrake, basing our work on both existing implementations of some of these solvers while also writing new extensions for others.

CFD↗

Multi-variance replica exchange SGMCMC for inverse and forward problems via Bayesian PINN

Physics-informed neural network (PINN) has been successfully applied in solving a variety of nonlinear non-convex forward and inverse problems. However, the training is challenging because of the non-convex loss functions and the multiple optima in the Bayesian inverse problem. In this work, we propose a multi-variance replica exchange stochastic gradient Langevin dynamics method to tackle the challenge of the multiple local optima in the optimization and the challenge of the multiple modal posterior distribution in the inverse problem. Replica exchange methods are capable of escaping from the local traps and accelerating the convergence; two chains with different temperatures are designed where the low temperature chain aims for the local convergence, and the target of the high temperature chain is to travel globally and explore the whole loss function entropy landscape. However, it may not be efficient to solve mathematical inversion problems by using the vanilla replica method directly since the method doubles the computational cost in evaluating the forward solvers (likelihood functions) in the two chains. To address this issue, we propose to make different assumptions on the energy function estimation and this facilities one to use solvers of different fidelities in the likelihood function evaluation. More precisely, one can use a solver with low fidelity in the high temperature chain while using a solver with high fidelity in the low temperature chain. Our proposed method significantly lowers the computational cost in the high temperature chain, meanwhile preserving the accuracy and converging very fast. Here we give an unbiased estimate of the swapping rate and give an estimation of the discretization error of the scheme. To verify our idea, we design and solve four inverse problems which have multiple modes. The proposed method is also employed to train the Bayesian PINN to solve the forward and inverse problems; faster and more accurate convergence has been observed when compared to the stochastic gradient Langevin dynamics (SGLD) method and vanilla replica exchange methods.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO Users Manual (Rev. 3.19)

This manual describes the use of the Portable, Extensible Toolkit for Scientific Computation (PETSc) and the Toolkit for Advanced Optimization (TAO) for the numerical solution of partial differential equations and related problems on high-performance computers. PETSc/TAO is a suite of data structures and routines that provide the building blocks for the implementation of large-scale application codes on parallel (and serial) computers. PETSc uses the MPI standard for all distributed memory communication. PETSc/TAO includes a large suite of parallel linear solvers, nonlinear solvers, time integrators, and opti mization that may be used in application codes written in Fortran, C, C++, and Python (via petsc4py; see Getting Started). PETSc provides many of the mechanisms needed within parallel application codes, such as parallel matrix and vector assembly routines. The library is organized hierarchically, enabling users to employ the level of abstraction that is most appropriate for a particular problem. By using techniques of object-oriented programming, PETSc provides enormous flexibility for users. PETSc is a sophisticated set of software tools; as such, for some users it initially has a much steeper learning curve than packages such as MATLAB or a simple subroutine library. In particular, for individuals without some computer science background, experience programming in C, C++, python, or Fortran and experience using a debugger such as gdb or lldb, it may require a significant amount of time to take full advantage of the features that enable efficient software use. However, the power of the PETSc design and the algorithms it incorporates may make the efficient implementation of many application codes simpler than “rolling them” yourself. For many tasks a package such as MATLAB is often the best tool; PETSc is not intended for the classes of problems for which effective MATLAB code can be written. There are several packages, built on PETSc, that may satisfy your needs without requiring directly using PETSc. We recommend reviewing these packages functionality before starting to code directly with PETSc. PETSc can be used to provide a “MPI parallel linear solver” in an otherwise sequential, or OpenMP parallel code. This approach cannot provide extremely large improvements in the application time by utilizing large numbers of MPI processes but can still improve the performance. Certainly all parts of a previously sequential code need not be parallelized but the matrix generation portion must be parallelized to expect true scalability to large numbers of MPI processes. See PCMPI for details on how to utilize the PETSc MPI linear solver server. Since PETSc is under continued development, small changes in usage and calling sequences of routines will occur. PETSc has been supported for twenty-five years; see mailing list information on our website for information on contacting support.

97 MATHEMATICS AND COMPUTING↗

On the Convergence of Overlapping Schwarz Decomposition for Nonlinear Optimal Control

Here, we study the convergence properties of an overlapping Schwarz decomposition algorithm for solving nonlinear optimal control problems (OCPs). The algorithm decomposes the time domain into a set of overlapping subdomains, and solves all subproblems defined over subdomains in parallel. The convergence is attained by updating primal-dual information at the boundaries of overlapping subdomains. We show that the algorithm exhibits local linear convergence, and that the convergence rate improves exponentially with the overlap size. We also establish global convergence results for a general quadratic programming, which enables the application of the Schwarz scheme inside second-order optimization algorithms (e.g., sequential quadratic programming). The theoretical foundation of our convergence analysis is a sensitivity result of nonlinear OCPs, which we call "exponential decay of sensitivity" (EDS). Intuitively, EDS states that the impact of perturbations at domain boundaries (i.e., initial and terminal time) on the solution decays exponentially as one moves into the domain. Here, we expand a previous analysis available in the literature by showing that EDS holds for both primal and dual solutions of nonlinear OCPs, under uniform second-order sufficient condition, controllability condition, and boundedness condition. We conduct experiments with a quadrotor motion planning problem and a partial differential equations (PDE) control problem to validate our theory, and show that the approach is significantly more efficient than alternating direction method of multipliers and as efficient as the centralized interior-point solver.

42 ENGINEERING↗

KiT-RT: An Extendable Framework for Radiative Transfer and Therapy

Here, in this article, we present Kinetic Transport Solver for Radiation Therapy (KiT-RT), an open-source C++-based framework for solving kinetic equations in therapy applications available at https://github.com/CSMMLab/KiT-RT . This software framework aims to provide a collection of classical deterministic solvers for unstructured meshes that allow for easy extendability. Therefore, KiT-RT is a convenient base to test new numerical methods in various applications and compare them against conventional solvers. The implementation includes spherical harmonics, minimal entropy, neural minimal entropy, and discrete ordinates methods. Solution characteristics and efficiency are presented through several test cases ranging from radiation transport to electron radiation therapy. Due to the variety of included numerical methods and easy extendability, the presented open-source code is attractive for both developers, who want a basis to build their numerical solvers, and users or application engineers, who want to gain experimental insights without directly interfering with the codebase.

97 MATHEMATICS AND COMPUTING↗