Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “differential equation solver”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

A semi-direct procedure using a local relaxation factor and its application to an internal flow problem

Generally, fast direct solvers are not directly applicable to a nonseparable elliptic partial differential equation. This limitation, however, is circumvented by a semi-direct procedure, i.e., an iterative procedure using fast direct solvers. An efficient semi-direct procedure which is easy to implement and applicable to a variety of boundary conditions is presented. The current procedure also possesses other highly desirable properties, i.e.: (1) the convergence rate does not decrease with an increase of grid cell aspect ratio, and (2) the convergence rate is estimated using the coefficients of the partial differential equation being solved.

Chang, S. C.↗

Block smoothers and generalized ideal interpolation in AMG (Final Report)

The Pennsylvania State University (“Subcontractor”) worked on developing new parallel algebraic multilevel methods suitable for solving PDEs. Specifically, work on the design of multigrid solvers for coupled systems of partial differential equations arising in numerical modeling of various applications was completed. A main emphasis was on the design of new ideal algebraic multigrid interpolation for problems such as Maxwell’s equations where block smoothers are needed and the standard form of ideal interpolation is not an effective choice.

97 MATHEMATICS AND COMPUTING↗

Enhancing high-fidelity nonlinear solver with reduced order model

Abstract We propose the use of reduced order modeling (ROM) to reduce the computational cost and improve the convergence rate of nonlinear solvers of full order models (FOM) for solving partial differential equations. In this study, a novel ROM-assisted approach is developed to improve the computational efficiency of FOM nonlinear solvers by using ROM’s prediction as an initial guess. We hypothesize that the nonlinear solver will take fewer steps to the converged solutions with an initial guess that is closer to the real solutions. To evaluate our approach, four physical problems with varying degrees of nonlinearity in flow and mechanics have been tested: Richards’ equation of water flow in heterogeneous porous media, a contact problem in a hyperelastic material, two-phase flow in layered porous media, and fracture propagation in a homogeneous material. Overall, our approach maintains the FOM’s accuracy while speeding up nonlinear solver by 18–73% (through suitable ROM-assisted FOMs). More importantly, the proximity of ROM’s prediction to the solution space leads to the improved convergence of FOMs that would have otherwise diverged with default initial guesses. We demonstrate that the ROM’s accuracy can impact the computational efficiency with more accurate ROM solutions, resulting in a better cost reduction. We also illustrate that this approach could be used in many FOM discretizations (e.g., finite volume, finite element, or a combination of those). Since our ROMs are data-driven and non-intrusive, the proposed procedure can easily lend itself to any nonlinear physics-based problem.

97 MATHEMATICS AND COMPUTING↗

Direct numerical solution of three-dimensional equations containing elliptic operators.

A direct three-dimensional elliptic solver is presented for application in a wide class of numerical methods for solving partial differential equations in physics and engineering. The derived algorithm and FORTRAN code implement Buzbee, Golub and Nielson's proposed extension of Buneman's Cyclic-Reduction Poisson solver to three dimensions. Both a 'most direct' cyclic reduction and a revised method (to eliminate roundoff error difficulties) are derived. Tests on an IBM 360/67 computer, using various optional combinations of subroutines, showed significant differences in accuracy and computing time, with the optimum subroutine combination depending on mesh size.

Martin, E. D.↗

Unsupervised discovery of nonlinear plasma physics using differentiable kinetic simulations

Plasma supports collective modes and particle–wave interactions that lead to complex behaviour in, for example, inertial fusion energy applications. While plasma can sometimes be modelled as a charged fluid, a kinetic description is often crucial for studying nonlinear effects in the higher-dimensional momentum–position phase space that describes the full complexity of the plasma dynamics. We create a differentiable solver for the three-dimensional partial-differential equation describing the plasma kinetics and introduce a domain-specific objective function. Using this framework, we perform gradient-based optimization of neural networks that provide forcing function parameters to the differentiable solver given a set of initial conditions. We apply this to an inertial-fusion-relevant configuration and find that the optimization process exploits a novel physical effect.

plasma nonlinear phenomena↗

Carbon Organisms Rhizosphere and Protection in Soil Environment model script and input data for soil moisture-respiration responses in tropical forests

Objectives: Climatic drying is predicted for many tropical forests, yet models remain poorly parameterized for tropical forests, hampering predictions of forest-climate feedbacks. We applied an integrated model–experiment approach, parameterizing an ecosystem model Carbon Organisms Rhizosphere and Protection in the Soil Environment (CORPSE) with tropical forest observational data, and comparing model predictions with a field drying manipulation. We hypothesized that drying would suppress soil CO2 fluxes (i.e., respiration) in already-drier tropical forests, but increases CO2 fluxes in wetter tropical forests by alleviating anaerobiosis. We measured soil CO2 fluxes, soil moisture, soil temperature, and forest floor biomass during wet-dry cycles (2015 – 2022) in four Panamanian forests that vary in rainfall and soil fertility. We used the field data to parameterize and run tests in the model.Results: Measured CO2 fluxes declined in the dry season and peaked in the early wet season ahead of peak soil moisture, resulting in a lower soil moisture optimum for respiration than previously modeled. We used this data to parameterize the model, which then predicted increased soil CO2 fluxes in wetter and fertile forests with drying, and decreased fluxes in drier, infertile forests. In contrast to model predictions, a chronic throughfall exclusion experiment in the forests initially suppressed soil CO2 fluxes across forests, with sustained suppression after four years in the wettest forest only (-28 ± 4% during the dry season), but elevated soil CO2 fluxes in a fertile forest after four years (+75 ± 28% during the late wet season), as predicted by the model. The unexpected negative drying effect in the wettest, most infertile forest could have resulted from reduced vertical flushing of nutrients into soils. Including hydro-nutrient interactions in ecosystem models could improve predictions of tropical forest-climate feedbacks (results presented in Cusack et al. 2023). Datasets included: Code files:CORPSE_array.py: Defines the equations of the CORPSE modelCORPSE_solvers: Functions for running the CORPSE model using either iterative or ordinary differential equation (ODE) solversrun_Panama_sims.py: Read in datasets and run the model simulations for this studyInput data:PanamaGradientEcosystemChem_BT_CPools_20152016CO2_DC_20190615.xlsx: Plot characteristics used in running model simulationsLiCor compiled surface flux only to 2020_03 DC_20200825.xlsx: Surface gas exchange fluxes used in model-data comparisonsPARCHED litterfall data for Ben Sulman LD 20200902.xlsx: Litterfall data used to drive model simulationsInitialization data:state_500y_20190823.csv: Initial state of model pools based on previous spinup runsOutput data:Outputs/prev_moisture_response.csv: Simulations of multiple sites using original model moisture response function.Outputs/updated_moisture_response.csv: Simulations of multiple sites using updated model moisture response function.Outputs/dry15_prev_moisture_response.csv: Simulations with soil moisture reduced by 15%, using original moisture response function.Outputs/dry15_updated_moisture_response.csv: Simulations with soil moisture reduced by 15%, using updated moisture response function.Outputs/dry30_prev_moisture_response.csv: Simulations with soil moisture reduced by 30%, using original moisture response function.Outputs/dry30_updated_moisture_response.csv: Simulations with soil moisture reduced by 30%, using updated moisture response function.Outputs/latestart_prev_moisture_response.csv: Simulations with extended dry season, using original moisture response function.Outputs/latestart_updated_moisture_response.csv: Simulations with extended dry season, using updated moisture response function.Outputs/[site name]_oneyear.csv: One-year simulation for each site in expanded site list using original moisture response function.Outputs/[site name]_oneyear_dried.csv: One-year simulation for each site in expanded site list using original moisture response function, with soil moisture reduced by 25%.Outputs/[site name]_oneyear_updated_moisture_response.csv: One-year simulation for each site in expanded site list using updated moisture response function.Outputs/[site name]_oneyear_updated_moisture_response_dried.csv: One-year simulation for each site in expanded site list using updated moisture response function, with soil moisture reduced by 25%.Field plot location data:There is also a .kml file that includes coordinates for all 32 plots included in the study of four forests (n = 4 throughfall reduction and n = 4 control plots per site).

54 ENVIRONMENTAL SCIENCES↗

Discrete sensitivity derivatives of the Navier-Stokes equations with a parallel Krylov solver

This paper solves an 'incremental' form of the sensitivity equations derived by differentiating the discretized thin-layer Navier Stokes equations with respect to certain design variables of interest. The equations are solved with a parallel, preconditioned Generalized Minimal RESidual (GMRES) solver on a distributed-memory architecture. The 'serial' sensitivity analysis code is parallelized by using the Single Program Multiple Data (SPMD) programming model, domain decomposition techniques, and message-passing tools. Sensitivity derivatives are computed for low and high Reynolds number flows over a NACA 1406 airfoil on a 32-processor Intel Hypercube, and found to be identical to those computed on a single-processor Cray Y-MP. It is estimated that the parallel sensitivity analysis code has to be run on 40-50 processors of the Intel Hypercube in order to match the single-processor processing time of a Cray Y-MP.

Ajmani, Kumud↗

On the structure of parallelism in a highly concurrent PDE solver

A parallel multigrid algorithm for solving elliptic partial differential equations is developed and evaluated. A V-cycle multigrid method is altered to increase the degree of parallelism. A numerical analysis of the resulting concurrent-iteration multigrid algorithm is performed; its architectural implications are considered; highly parallel systems without shared memory are examined (including mesh-connected arrays, mesh-shuffle-connected systems, permutation networks, and direct VLSI embeddings); and the results of numerical experiments are presented in tables and graphs.

Gannon, D.↗

Learning Constitutive Relations From Soil Moisture Data via Physically Constrained Neural Networks

Abstract The constitutive relations of the Richardson‐Richards equation encode the macroscopic properties of soil water retention and conductivity. These soil hydraulic functions are commonly represented by models with a handful of parameters. The limited degrees of freedom of such soil hydraulic models constrain our ability to extract soil hydraulic properties from soil moisture data via inverse modeling. We present a new free‐form approach to learning the constitutive relations using physically constrained neural networks. We implemented the inverse modeling framework in a differentiable modeling framework, JAX, to ensure scalability and extensibility. For efficient gradient computations, we implemented implicit differentiation through a nonlinear solver for the Richardson‐Richards equation. We tested the framework against synthetic noisy data and demonstrated its robustness against varying magnitudes of noise and degrees of freedom of the neural networks. We applied the framework to soil moisture data from an upward infiltration experiment and demonstrated that the neural network‐based approach was better fitted to the experimental data than a parametric model and that the framework can learn the constitutive relations.

54 ENVIRONMENTAL SCIENCES↗

Leveraging Multitime Hamilton–Jacobi PDEs for Certain Scientific Machine Learning Problems

Hamilton-Jacobi partial differential equations (HJ PDEs) have deep connections with a wide range of fields, including optimal control, differential games, and imaging sciences. By considering the time variable to be a higher dimensional quantity, HJ PDEs can be extended to the multi-time case. In this paper, we establish a novel theoretical connection between specific optimization problems arising in machine learning and the multi-time Hopf formula, which corresponds to a representation of the solution to certain multi-time HJ PDEs. Through this connection, we increase the interpretability of the training process of certain machine learning applications by showing that when we solve these learning problems, we also solve a multi-time HJ PDE and, by extension, its corresponding optimal control problem. As a first exploration of this connection, we develop the relation between the regularized linear regression problem and the Linear Quadratic Regulator (LQR). We then leverage our theoretical connection to adapt standard LQR solvers (namely, those based on the Riccati ordinary differential equations) to design new training approaches for machine learning. Lastly, we provide some numerical examples that demonstrate the versatility and possible computational advantages of our Riccati-based approach in the context of continual learning, post-training calibration, transfer learning, and sparse dynamics identification.

97 MATHEMATICS AND COMPUTING↗

PETSc/TAO developments for GPU-based early exascale systems

The Portable Extensible Toolkit for Scientific Computation (PETSc) library provides scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization via the Toolkit for Advanced Optimization (TAO). PETSc is used in dozens of scientific fields and is an important building block for many simulation codes. During the U.S. Department of Energy’s Exascale Computing Project, the PETSc team has made substantial efforts to enable efficient utilization of the massive fine-grain parallelism present within exascale compute nodes and to enable performance portability across exascale architectures. We recap some of the challenges that designers of numerical libraries face in such an endeavor, and then discuss the many developments we have made, which include the addition of new GPU backends, features supporting efficient on-device matrix assembly, better support for asynchronicity and GPU kernel concurrency, and new communication infrastructure. In conclusion, we evaluate the performance of these developments on some pre-exascale systems as well as the early exascale systems Frontier and Aurora, using compute kernel, communication layer, solver, and mini-application benchmark studies, and then close with a few observations drawn from our experiences on the tension between portable performance and other goals of numerical libraries.

Exascale Computing Project (ECP)↗

Parallel-in-Time Solution of Allen-Cahn Equations by Integrating Operator Learning into the Parareal Method

While recent advances in deep learning have shown promising efficiency gains in solving time-dependent partial differential equations (PDEs), matching the accuracy of conventional numerical solvers still remains a challenge. One strategy to improve the accuracy of deep learning-based solutions for time-dependent PDEs is to use the learned model as the coarse propagator in the Parareal method and a traditional numerical method as the fine solver. However, successful integration of deep learning into the Parareal method requires consistency between the coarse and fine solvers, particularly for PDEs exhibiting rapid changes such as sharp transitions. Here, to ensure this consistency, we propose using convolutional neural networks (CNNs) to learn the fully discrete time-stepping operator defined by the same numerical scheme employed as the fine solver. We demonstrate the effectiveness of the proposed method in solving the classical and mass-conservative Allen–Cahn (AC) equations. Through iterative updates in the Parareal algorithm, our approach achieves a significant computational speedup compared to traditional fine solvers while converging to high-accuracy solutions. Our results highlight that the proposed hybrid Parareal algorithm effectively accelerates simulations, particularly when implemented on multiple GPUs, and converges to the desired accuracy in only a few iterations. Another advantage of our method is that the CNN model is trained on trajectory-based data generated from random initial conditions, such that the trained model can be used to solve the AC equations with various initial conditions without retraining. This work demonstrates the potential of integrating neural network methods into parallel-in-time frameworks for efficient and accurate simulations of time-dependent PDEs.

97 MATHEMATICS AND COMPUTING↗

Toward performance-portable PETSc for GPU-based exascale systems

The Portable Extensible Toolkit for Scientific computation (PETSc) library delivers scalable solvers for nonlinear time-dependent differential and algebraic equations and for numerical optimization. The PETSc design for performance portability addresses fundamental GPU accelerator challenges and stresses flexibility and extensibility by separating the programming model used by the application from that used by the library, and it enables application developers to use their preferred programming model, such as Kokkos, RAJA, SYCL, HIP, CUDA, or OpenCL, on upcoming exascale systems. Furthermore, a blueprint for using GPUs from PETSc-based codes is provided, and case studies emphasize the flexibility and high performance achieved on current GPU-based systems.

97 MATHEMATICS AND COMPUTING↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Crowley, Kay↗

Run-time scheduling and execution of loops on message passing machines

Sparse system solvers and general purpose codes for solving partial differential equations are examples of the many types of problems whose irregularity can result in poor performance on distributed memory machines. Often, the data structures used in these problems are very flexible. Crucial details concerning loop dependences are encoded in these structures rather than being explicitly represented in the program. Good methods for parallelizing and partitioning these types of problems require assignment of computations in rather arbitrary ways. Naive implementations of programs on distributed memory machines requiring general loop partitions can be extremely inefficient. Instead, the scheduling mechanism needs to capture the data reference patterns of the loops in order to partition the problem. First, the indices assigned to each processor must be locally numbered. Next, it is necessary to precompute what information is needed by each processor at various points in the computation. The precomputed information is then used to generate an execution template designed to carry out the computation, communication, and partitioning of data, in an optimized manner. The design is presented for a general preprocessor and schedule executer, the structures of which do not vary, even though the details of the computation and of the type of information are problem dependent.

Saltz, Joel↗

Aerodynamic Design Optimization on Unstructured Meshes Using the Navier-Stokes Equations

A discrete adjoint method is developed and demonstrated for aerodynamic design optimization on unstructured grids. The governing equations are the three-dimensional Reynolds-averaged Navier-Stokes equations coupled with a one-equation turbulence model. A discussion of the numerical implementation of the flow and adjoint equations is presented. Both compressible and incompressible solvers are differentiated and the accuracy of the sensitivity derivatives is verified by comparing with gradients obtained using finite differences. Several simplifying approximations to the complete linearization of the residual are also presented, and the resulting accuracy of the derivatives is examined. Demonstration optimizations for both compressible and incompressible flows are given.

Nielsen, Eric J.↗

Variants and extensions of a fast direct numerical cauchy-riemann solver, with illustrative applications

Revised and extended versions of a fast, direct (noniterative) numerical Cauchy-Riemann solver are presented for solving finite difference approximations of first order systems of partial differential equations. Although the difference operators treated are linear and elliptic, one significant application of these extended direct Cauchy-Riemann solvers is in the fast, semidirect (iterative) solution of fluid dynamic problems governed by the nonlinear mixed elliptic-hyperbolic equations of transonic flow. Different versions of the algorithms are derived and the corresponding FORTRAN computer programs for a simple example problem are described and listed. The algorithms are demonstrated to be efficient and accurate.

Martin, E. D.↗

Large liquid rocket engine transient performance simulation system

A simulation system, ROCETS, was designed and developed to allow cost-effective computer predictions of liquid rocket engine transient performance. The system allows a user to generate a simulation of any rocket engine configuration using component modules stored in a library through high-level input commands. The system library currently contains 24 component modules, 57 sub-modules and maps, and 33 system routines and utilities. FORTRAN models from other sources can be operated in the system upon inclusion of interface information on comment cards. Operation of the simulation is simplified for the user by run, execution, and output processors. The simulation system makes available steady-state trim balance, transient operation, and linear partial generation. The system utilizes a modern equation solver for efficient operation of the simulations. Transient integration methods include integral and differential forms for the trapezoidal, first order Gear, and second order Gear corrector equations. A detailed technology test bed engine (TTBE) model was generated to be used as the acceptance test of the simulation system. The general level of model detail was that reflected in the Space Shuttle Main Engine DTM. The model successfully obtained steady-state balance in main stage operation and simulated throttle transients, including engine starts and shutdown. A NASA FORTRAN control model was obtained, ROCETS interface installed in comment cards, and operated with the TTBE model in closed-loop transient mode.

Mason, J. R.↗