Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Linear Solvers”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Quantum microgrid state estimation

This paper investigates the feasibility and efficiency of quantum-circuit-based algorithms for microgrid state estimation. Here, our new contributions include: (1) a general quantum state estimation (GQSE) formulation is devised for swing-bus-contained microgrids through the quantized Gaussian–Newton iteration, (2) a preconditioned quantum linear solver (PQLS) is developed for tackling the ill-conditioned GQSE with limited quantum resources, and (3) an enhanced quantum state estimation (EQSE) algorithm is further established for hierarchical-control-based microgrids with exogenous disturbances. Extensive case studies demonstrate the correctness of GQSE, PQLS and EQSE in two typical microgrids. The robustness and convergence performance of EQSE are also verified.

29 ENERGY PLANNING, POLICY, AND ECONOMY↗

Bridging paradigms: Designing for HPC-Quantum convergence

Here, this paper presents a comprehensive software stack architecture for integrating quantum computing (QC) capabilities with High-Performance Computing (HPC) environments. While quantum computers show promise as specialized accelerators for scientific computing, their effective integration with classical HPC systems presents significant technical challenges. We propose a hardware-agnostic software framework that supports both current noisy intermediate-scale quantum devices and future fault-tolerant quantum computers, while maintaining compatibility with existing HPC workflows. The architecture includes a quantum gateway interface, standardized APIs for resource management, and robust scheduling mechanisms to handle both simultaneous and interleaved quantum–classical workloads. Key innovations include: (1) a unified resource management system that efficiently coordinates quantum and classical resources, (2) a flexible quantum programming interface that abstracts hardware-specific details, (3) A Quantum Platform Manager API that simplifies the integration of various quantum hardware systems, and (4) a comprehensive tool chain for quantum circuit optimization and execution. We demonstrate our architecture through implementation of quantum–classical algorithms, including the variational quantum linear solver, showcasing the framework’s ability to handle complex hybrid workflows while maximizing resource utilization. This work provides a foundational blueprint for integrating QC capabilities into existing HPC infrastructures, addressing critical challenges in resource management, job scheduling, and efficient data movement between classical and quantum resources.

97 MATHEMATICS AND COMPUTING↗

A sharp interface Lagrangian-Eulerian method for flexible-body fluid-structure interaction

This paper introduces a sharp-interface approach to simulating fluid-structure interaction (FSI) involving flexible bodies described by general nonlinear material models and across a broad range of mass density ratios. This new flexible-body immersed Lagrangian-Eulerian (ILE) scheme extends our prior work on integrating partitioned and immersed approaches to rigid-body FSI. Our numerical approach incorporates the geometrical and domain solution flexibility of the immersed boundary (IB) method with an accuracy comparable to body-fitted approaches that sharply resolve flows and stresses up to the fluid-structure interface. Unlike many IB methods, our ILE formulation uses distinct momentum equations for the fluid and solid subregions with a Dirichlet-Neumann coupling strategy that connects fluid and solid subproblems through simple interface conditions. As in earlier work, we use approximate Lagrange multiplier forces to treat the kinematic interface conditions along the fluid-structure interface. This penalty approach simplifies the linear solvers needed by our formulation by introducing two representations of the fluid-structure interface, one that moves with the fluid and another that moves with the structure, that are connected by stiff springs. This approach also enables the use of multi-rate time stepping, which allows us to use different time step sizes for the fluid and structure subproblems. Our fluid solver relies on an immersed interface method (IIM) for discrete surfaces to impose stress jump conditions along complex interfaces while enabling the use of fast structured-grid solvers for the incompressible Navier-Stokes equations. The dynamics of the volumetric structural mesh are determined using a standard finite element approach to large-deformation nonlinear elasticity via a nearly incompressible solid mechanics formulation. This formulation also readily accommodates compressible structures with a constant total volume, and it can handle fully compressible solid structures for cases in which at least part of the solid boundary does not contact the incompressible fluid. Selected grid convergence studies demonstrate second-order convergence in volume conservation and in the pointwise discrepancies between corresponding positions of the two interface representations as well as between first and second-order convergence in the structural displacements. The time stepping scheme is also demonstrated to yield second-order convergence. To assess and validate the robustness and accuracy of the new algorithm, comparisons are made with computational and experimental FSI benchmarks. Test cases include both smooth and sharp geometries in various flow conditions. Furthermore, we also demonstrate the capabilities of this methodology by applying it to model the transport and capture of a geometrically realistic, deformable blood clot in an inferior vena cava filter.

97 MATHEMATICS AND COMPUTING↗

Pseudospectral convex optimization for on-ramp merging control of connected vehicles

It can be a daunting task for human drivers to merge into highways because of the intricate vehicle negotiations and potential risk within limited time and space. Connected vehicle (CV) technologies could be a solution to this problem and offer many benefits to the road safety, traffic mobility, and energy efficiency. However, real-time optimal control of CVs is still an open challenge, due to the nonlinear vehicle dynamics, non-convex fuel consumption model, and highly dynamic uncertain inter-vehicle interactions. To tackle these issues, a novel real-time optimal control approach that balances the computational efficiency and solution optimality is proposed for the purpose of onboard application. To this end, the pseudospectral collocation method is integrated with a sequential convex programming approach to develop two new optimization algorithms, which are implemented within a model predictive control (MPC) framework to allow for real-time generation of optimal merging speed profiles. One algorithm leverages the line search technique to improve convergence, and the other benefits from the trust region method for better computational efficiency. The optimality and convergence process of both proposed algorithms are investigated by comparing their solutions with a popular non-linear solver. Furthermore, simulation results show that the proposed methods outperform the benchmark in terms of computational cost, fuel consumption, and traffic efficiency. In particular, the proposed fuel-economy merging rule can save 57.1% fuel consumption on average on four different traffic volumes. Meanwhile, the proposed optimal control algorithms can reduce 2.2% travel time on average comparing to the “first-in-first-out” merging rule.

33 ADVANCED PROPULSION SYSTEMS↗

Porting hypre to heterogeneous computer architectures: Strategies and experiences

We report that linear systems are occurring in many applications, and solving them can take a large amount of the total simulation time. The high performance library hypre provides a variety of interfaces and linear solvers, including various multigrid methods, that have achieved good scalability on a variety of homogeneous parallel computer architectures. Heterogeneous architectures with nodes that have both CPUs and accelerators provide new challenges, since they require more fine-grained parallelism and reduced data movement between different memories on a single node as well as across nodes. We will discuss our experiences and strategies to port hypre to heterogeneous computers with accelerators, including the design of a new memory model, the use of abstractions, the BoxLoop macros in the structured and semi-structured interfaces, and the restructuring of algebraic multigrid (AMG) into modular components. We present numerical experiments comparing CPU and GPU performance for several test problems.

97 MATHEMATICS AND COMPUTING↗

Development of a continuous synthesis process for carbamazepine using validated in-line Raman spectroscopy and kinetic modelling for disturbance simulation

Mitigation of failure modes in the continuous synthesis (CS) of a drug substance (DS) has the potential to widen the adoption of continuous manufacturing (CM) technologies by the pharmaceutical industry. Here, this work demonstrates the development of a robust continuous process for the synthesis of carbamazepine (CBZ), an essential medicine as per the World Health Organization (WHO), facilitated by kinetic modelling and monitored by in-line Raman spectroscopy. Accurate kinetic modelling and the use of validated process analytical technology (PAT) models for quantitative measurement were found to play an important role in developing CS of drug substances. Kinetic data for the formation of CBZ from iminostilbene (ISB) were collected by batch reaction sampling and high-performance liquid chromatography (HPLC) analysis. A non-linear solver and iterative method was applied to determine two sets of Arrhenius parameters simultaneously for the reaction system by minimizing the standard error of the model fit. The start-up and dynamic equilibrium stages for the CS of CBZ using a continuous stirred tank reactor (CSTR) were modelled based on the batch kinetic data and employed to optimize conversion and simulate process disturbances. An in-line Raman spectroscopy method was successfully developed, validated, and integrated to determine the concentrations of CBZ and ISB within the operating range for the CS. The CS kinetic model was evaluated experimentally from startup to dynamic equilibrium over 10 residence times with monitoring by HPLC and in-line Raman spectroscopy. The developed kinetic model in tandem with in-line Raman spectroscopy successfully predicted disturbances due to changes in process variables and can serve as a useful tool in the future design of advanced process control strategies for the continuous synthesis of CBZ.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Isotope effects on energy transport in the core of ASDEX-Upgrade tokamak plasmas: Turbulence measurements and model validation

Design and operation of future tokamak fusion reactors using a deuterium–tritium 50:50 mix requires a solid understanding of how energy confinement properties change with ion mass. This study looks at how turbulence and energy transport change in L-mode plasmas in the ASDEX Upgrade tokamak when changing ion species between hydrogen and deuterium. For this purpose, both experimental turbulence measurements and modeling are employed. Local measurements of ion-scale (with wavevector of fluctuations perpendicular to the B-field k⊥< 2 cm−1, k⊥ρs< 0.2, where ρs is the ion sound Larmor radius using the deuterium ion mass) electron temperature fluctuations have been performed in the outer core (normalized toroidal flux ρTor=0.65−0.8) using a multi-channel correlation electron cyclotron emission diagnostic. Lower root mean square perpendicular fluctuation amplitudes and radial correlation lengths have been measured in hydrogen vs deuterium. Measurements of the cross-phase angle between a normal-incidence reflectometer and an ECE signal were made to infer the cross-phase angle between density and temperature fluctuations. The magnitude of the cross-phase angle was found larger (more out-of-phase) in hydrogen than in deuterium. TRANSP power balance simulations show a larger ion heat flux in hydrogen where the electron-ion heat exchange term is found to play an important role. These experimental observations were used as the basis of a validation study of both quasilinear gyrofluid trapped gyro-Landau fluid-SAT2 and nonlinear gyrokinetic GENE codes. Linear solvers indicate that, at long wavelengths (k⊥ρs<1), energy transport in the deuterium discharge is dominated by a mixed ion-temperature-gradient (ITG) and trapped-electron mode turbulence while in hydrogen transport is exclusively and more strongly driven by ITG turbulence. The Ricci validation metric has been used to quantify the agreement between experiments and simulations taking into account both experimental and simulation uncertainties as well as four different observables across different levels of the primacy hierarchy.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Stability of superthermal strahl electrons in the solar wind

We present a kinetic stability analysis of the solar wind electron distribution function consisting of the Maxwellian core and the magnetic-field aligned strahl, a superthermal electron beam propagating away from the sun. We use an electron strahl distribution function obtained as a solution of a weakly collisional drift-kinetic equation, representative of a strahl affected by Coulomb collisions but unadulterated by possible broadening from turbulence. This distribution function is essentially non-Maxwellian and varies with the heliospheric distance. The stability analysis is performed with the Vlasov–Maxwell linear solver leopard. We find that depending on the heliospheric distance, the core-strahl electron distribution becomes unstable with respect to sunward-propagating kinetic-Alfvén, magnetosonic, and whistler modes, in a broad range of propagation angles. The wavenumbers of the unstable modes are close to the ion inertial scales, and the radial distances at which the instabilities first appear are on the order of 1 au. However, we have not detected any instabilities driven by resonant wave interactions with the superthermal strahl electrons. Instead, the observed instabilities are triggered by a relative drift between the electron and ion cores necessary to maintain zero electric current in the solar wind frame (ion frame). Contrary to strahl distributions modelled by shifted Maxwellians, the electron strahl obtained as a solution of the kinetic equation is stable. Our results are consistent with the previous studies based on a more restricted solution for the electron strahl.

79 ASTRONOMY AND ASTROPHYSICS↗

Fast and scalable quantum Monte Carlo simulations of electron-phonon models

We introduce methodologies for highly scalable quantum Monte Carlo simulations of electron-phonon models, and report benchmark results for the Holstein model on the square lattice. The determinant quantum Monte Carlo (DQMC) method is a widely used tool for simulating simple electron-phonon models at finite temperatures, but incurs a computational cost that scales cubically with system size. Alternatively, near-linear scaling with system size can be achieved with the hybrid Monte Carlo (HMC) method and an integral representation of the Fermion determinant. Here, we introduce a collection of methodologies that make such simulations even faster. To combat "stiffness" arising from the bosonic action, we review how Fourier acceleration can be combined with time-step splitting. To overcome phonon sampling barriers associated with strongly-bound bipolaron formation, we design global Monte Carlo updates that approximately respect particle-hole symmetry. To accelerate the iterative linear solver, we introduce a preconditioner that becomes exact in the adiabatic limit of infinite atomic mass. Finally, we demonstrate how stochastic measurements can be accelerated using fast Fourier transforms. Here, these methods are all complementary and, combined, may produce multiple orders of magnitude speedup, depending on model details.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

PLEXUS: A Pattern-Oriented Runtime System Architecture for Resilient Extreme-Scale High-Performance Computing Systems

For high-performance computing (HPC) system designers and users, meeting the myriad challenges of next-generation exascale supercomputing systems requires rethinking their approach to application and system software design. Among these challenges, providing resiliency and stability to the scientific applications in the presence of high fault rates requires new approaches to software architecture and design. As HPC systems become increasingly complex, they require intricate solutions for detection and mitigation for various modes of faults and errors that occur in these large-scale systems, as well as solutions for failure recovery. These resiliency solutions often interact with and affect other system properties, including application scalability, power and energy efficiency. Therefore, resilience solutions for HPC systems must be thoughtfully engineered and deployed.In previous work, we developed the concept of resilience design patterns, which consist of templated solutions based on well-established techniques for detection, mitigation and recovery. In this paper, we use these patterns as the foundation to propose new approaches to designing runtime systems for HPC systems. The instantiation of these patterns within a runtime system enables flexible and adaptable end-to-end resiliency solutions for HPC environments. The paper describes the architecture of the runtime system, named Plexus, and the strategies for dynamically composing and adapting pattern instances under runtime control. This runtime-based approach enables actively balancing the cost-benefit trade-off between performance overhead and protection coverage of the resilience solutions. Based on a prototype implementation of PLEXUS, we demonstrate the resiliency and performance gains achieved by the pattern-based runtime system for a parallel linear solver application.

Hukerikar, Saurabh↗

High-Performance Computing Based EMT Simulation: Power Grid with IBRs

Electromagnetic transient (EMT) simulation of power grids with high-fidelity models of inverter-based resources (IBRs) is time-consuming and difficult to scale. The necessity for high-fidelity models of IBRs that incorporate the dynamics of individual inverters within IBRs has been showcased in recent studies. These studies focused on events with partial power reduction in each IBR during a transmission line fault in the power grid. These types of events have been documented in multiple North American Electric Reliability Council (NERC) reports in the past decade. It is imperative then to find solutions to speed-up EMT simulations and scale the size of the region with IBRs studied in EMT simulations. In this paper, a combination of numerical simulation algorithms with high-performance computing techniques are employed in discretization and linear solvers employed in the proposed RE-INTEGRATE EMT simulation platform for power grid with IBRs. For ease of scalability, modular and object-oriented programming is used as these techniques are implemented. Additionally, automation software is developed to convert legacy software codes to the proposed RE-INTEGRATE EMT simulation platform. Thereafter, this platform is evaluated on multi-core central processing units (CPUs). Finally, scale-up tests are performed to showcase the scalability that is possible.

Marthi, Phani Ratna Vanamali [ORNL] (ORCID:0000000↗

Numerical Solution of the Steady-State Network Flow Equations for a Non-Ideal Gas

Herein we formulate a steady-state network flow problem for non-ideal gas that relates injection rates and nodal pressures in the network to flows in pipes. For this problem, we present and prove a theorem on uniqueness of generalized solution for a broad class of non-ideal pressure-density relations that satisfy a monotonicity property. Further, we develop a Newton-Raphson algorithm for numerical solution of the steady-state problem, which is made possible by a systematic non-dimensionalization of the equations. The developed algorithm has been extensively tested on benchmark instances and shown to converge robustly to a generalized solution. Previous results [1]-[4], indicate that the steady-state network flow equations for an ideal gas are difficult to solve by the Newton-Raphson method because of its extreme sensitivity to the initial guess. In contrast, we find that non-dimensionalization of the steady-state problem is key to robust convergence of the Newton-Raphson method. We identify criteria based on the uniqueness of solutions under which the existence of a non-physical generalized solution found by a non-linear solver implies non-existence of a physical solution, i.e., infeasibility of the problem. Finally, we compare pressure and flow solutions based on ideal and non-ideal equations of state to demonstrate the need to apply the latter in practice. The solver developed in this article is open-source and is made available for both the academic and research communities as well as the industry.

97 MATHEMATICS AND COMPUTING↗

ALESQP: An Augmented Lagrangian Equality-Constrained SQP Method for Optimization with General Constraints

Here we present a new algorithm for infinite-dimensional optimization with general constraints, called ALESQP. In short, ALESQP is an augmented Lagrangian method that penalizes inequality constraints and solves equality-constrained nonlinear optimization subproblems at every iteration. The subproblems are solved using a matrix-free trust-region sequential quadratic programming (SQP) method that takes advantage of iterative, i.e., inexact linear solvers, and is suitable for large-scale applications. A key feature of ALESQP is a constraint decomposition strategy that allows it to exploit problem-specific variable scalings and inner products. We analyze convergence of ALESQP under different assumptions. We show that strong accumulation points are stationary. Consequently, in finite dimensions ALESQP converges to a stationary point. In infinite dimensions we establish that weak accumulation points are feasible in many practical situations. Under additional assumptions we show that weak accumulation points are stationary. We present several infinite-dimensional examples where ALESQP shows remarkable discretization-independent performance in all of its iterative components, requiring a modest number of iterations to meet constraint tolerances at the level of machine precision. Also, we demonstrate a fully matrix-free solution of an infinite-dimensional problem with nonlinear inequality constraints.

97 MATHEMATICS AND COMPUTING↗

Reproduced Computational Results Report for “Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing”

The article titled “Ginkgo: A Modern Linear Operator Algebra Framework for High Performance Computing” by Anzt et al. presents a modern, linear operator centric, C++ library for sparse linear algebra. Experimental results in the article demonstrate that Ginkgo is a flexible and user-friendly framework capable of achieving high-performance on state-of-the-art GPU architectures. In this report, the Ginkgo library is installed and a subset of the experimental results are reproduced. Specifically, the experiment that shows the achieved memory bandwidth of the Ginkgo Krylov linear solvers on NVIDIA A100 and AMD MI100 GPUs is redone and the results are compared to what presented in the published article. Upon completion of the comparison, the published results are deemed reproducible.

97 MATHEMATICS AND COMPUTING↗

EKAT v.1.0

E3SM Kokkos Application Toolkit (EKAT) is a collection of C++, Fortran, and CMake utilities for providing a single implementation of common kernels based on the Kokkos programming model. The library contains utilities for vectorization, tridiagonal linear system solvers, and linear interpolation as well as some general-purpose utilities such as testing utilities, parameter lists, representation of physical units, and additional interfaces. The goal is to provide a centralized implementation for high-performance computing structures and common utilities that reduce code duplication and streamline maintenance efforts. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy's National Nuclear Security Administration under contract DE-NA0003525. SAND2022-1327 O

Bertagna, Luca↗

ReSolve

Library of GPU-resident linear solvers

Swirydowicz, Kasia↗

AMReX v2024

The software framework, AMReX, supports the development of block-structured adaptive mesh refinement (AMR) algorithms for solving systems of partial differential equations. AMR reduces the computational cost and memory footprint compared to a uniform mesh while preserving the essential local descriptions of different physical processes in complex multiphysics algorithms. AMR uses a hierarchical representation of the solution at multiple levels of resolution where the solution on each level is defined on the union of data containers at that resolution. These data containers, which represent the solution over a logically rectangular subregion of the domain, can contain field data defined on a mesh, Lagrangian particles or combinations of both. In addition to these basic data types, AMReX supports a multilevel embedded boundary representation of complex geometry; linear solvers for cell-centered and nodal data; asynchronous I/O in a native format readable by ParaView, VisIt and yt; and interfaces to hypre and PETSc solvers. AMReX enables applications to run on distributed memory architectures with multicore CPUs and with GPU accelerators. AMReX uses a lightweight abstraction layer that effectively hides the details of the architecture from the application. The framework currently supports CUDA, HIP and SYCL for GPU acceleration and OpenMP for multi-core CPU architectures.

Almgren, Ann↗

pnnl/LAP

A software framework to study, from the performance and energy perspective, the efficacy of GPU-resident parallel Conjugate Gradient (CG) linear solver with different preconditioner options, including Gauss-Seidel, Jacobi, and incomplete Cholesky. We also propose a novel GPU-based preconditioner, in which the triangular solves are approximated by an iterative process

Swirydowicz, Kasia↗