Optimal Scaling Quantum Linear-Systems Solver via Discrete Adiabatic Theorem
Not Available
SEARCH · Engineering Papers
Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.
Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.
Not Available
Concurrent computing environments provide the means to achieve very high performance for finite element analysis of systems, provided the algorithms take advantage of multiple processors. The authors have examined several algorithms for both linear and nonlinear finite element analysis. The performance of these algorithms on an Alliant FX/80 parallel supercomputer has been studied. For single load case linear analysis, the optimal solution algorithm is strongly problem dependent. For multiple load cases or nonlinear analysis through a modified Newton-Raphson method, decomposition algorithms are shown to have a decided advantage over element-by-element preconditioned conjugate gradient algorithms.
Abstract not provided.
Quantum Phase Estimation (QPE) is a cardinal algorithm in quantum computing that plays a crucial role in various applications, including cryptography, molecular simulation, and solving systems of linear equations. However, the standard implementation of QPE faces challenges related to time complexity and circuit depth, which limit its practicality for large-scale computations. We introduce LuGo, a novel framework designed to enhance the performance of QPE by reducing circuit duplication, as well as using parallelization techniques to achieve faster generation of the QPE circuit and gate reduction. We validate the effectiveness of our framework by generating quantum linear solver circuits, which require both QPE and inverse QPE, to solve linear systems of equations. LuGo achieves significant improvements in both computational efficiency and hardware requirements without compromising on accuracy. Compared to a standard QPE implementation, LuGo reduces time consumption to generate a circuit that solves a 2 6 × 2 6 system matrix by a factor of 50.68 and over 31× reduction of quantum gates and circuit depth, with no fidelity loss on an ideal quantum simulator. Furthermore, we demonstrated the versatility and scalability of LuGo enabled HHL algorithm by simulating a canonical Hele-Shaw fluid problem using a quantum simulator. With these advantages, LuGo paves the way for more efficient implementations of QPE, enabling broader applications across several quantum computing domains.
We present a Graphics Processing Unit (GPU)-accelerated version of the real-space SPARC electronic structure code for performing hybrid functional calculations in generalized Kohn–Sham density functional theory. In particular, we develop a batch variant of the recently formulated Kronecker product-based linear solver for the simultaneous solution of multiple linear systems. We then develop a modular, math kernel based implementation for hybrid functionals on NVIDIA architectures, where computationally intensive operations are offloaded to the GPUs, while the remaining workload is handled by the central processing units (CPUs). Considering bulk and slab examples, we demonstrate that GPUs enable up to 8× speedup in node-hours and 80× in core-hours compared to CPU-only execution, reducing the time to solution on V100 GPUs to around 300 s for a metallic system with over 6000 electrons, and significantly reducing the computational resources required for a given wall time.
Here we introduce the concept of half-closed nodes for nodal discontinuous Galerkin (DG) discretisations. Unlike more commonly used closed nodes in DG, where on every element nodes are placed on all of its boundaries, half-closed nodes only require nodes to be placed on a subset of the element's boundaries. The effect of using different nodes on DG operator sparsity is studied and we find in particular for there to be no difference in the sparsity pattern of the Laplace operator whether closed or half-closed nodes are used. On quadrilateral/hexahedral elements we use the Gauss-Radau points as the half-closed nodes of choice, which we demonstrate is able to speed up DG operator assembly in addition to leverage previously known superconvergence results. We also discuss in this work some linear solver techniques commonly used for Finite Element or discontinuous Galerkin methods such as static condensation and block-based methods, and how they can be applied to half-closed DG discretisations.
The goal of this work is to present a fast and viable approach for the numerical solution of the high-contrast state problems arising in topology optimization. The optimization process is iterative, and the gradients are obtained by an adjoint analysis, which requires the numerical solution of large high-contrast linear elastic problems with features spanning several length scales. The size of the discretized problems forces the utilization of iterative linear solvers with solution time dependent on the quality of the preconditioner. The lack of clear separation between the scales, as well as the high-contrast, imposes severe challenges on the standard preconditioning techniques. Thus, here we propose new methods for the high-contrast elasticity equation with performance independent of the high-contrast and the multi-scale structure of the elasticity problem. The solvers are based on two-levels domain decomposition techniques with a carefully constructed coarse level to deal with the high-contrast and multi-scale nature of the problem. The construction utilizes spectral equivalence between scalar diffusion and each displacement block of the elasticity problems and, in contrast to previous solutions proposed in the literature, is able to select the appropriate dimension of the coarse space automatically. The new methods inherit the advantages of domain decomposition techniques, such as easy parallelization and scalability. Finally, the presented numerical experiments demonstrate the excellent performance of the proposed methods.
Quantum-empowered electromagnetic transients program (QEMTP) is a promising paradigm for tackling EMTP's computational burdens. Nevertheless, no existing studies truly achieve a practical and scalable QEMTP operable on today's noisy-intermediate-scale quantum (NISQ) computers. The strong reliance on noise-free and fault-tolerant quantum devices--which appears to be decades away--hinder practical applications of current QEMTP methods. Here, we devise a NISQ-QEMTP methodology which for the first time transitions the QEMTP operations from ideal, noise-free quantum simulators to real, noisy quantum computers. The main contributions lie in: (1) a shallow-depth QEMTP quantum circuit for mitigating noises on NISQ quantum devices; (2) practical QEMTP linear solvers incorporating executable quantum state preparation and measurements for nodal voltage computations; (3) a noise-resilient QEMTP algorithm leveraging quantum resources logarithmically scaled with power system dimension; (4) a quantum shifted frequency analysis (QSFA) for accelerating QEMTP by exploiting dynamic phasor simulations with larger time steps; (5) a systematical analysis on QEMTPs performance under various noisy quantum environments. Extensive experiments systematically verify the accuracy, efficacy, universality and noise-resilience of QEMTP on both noise-free simulators and IBM real quantum computers.
This software implements parallel versions of an interior-point solver, based on the publicly available ipopt solver. Here we have full control over the linear solver and our algorithm is fully parallel thus enabling scalability to large-scale optimization problems. This package also has a parallel implementation of a homotopy solver developed under the scalable methods for contact LDRD project 23-ERD-017. This solver is an mfem-based implementation of algorithm described in ``A filter trust-region Newton continuation method for nonlinear complementarity problems''. Cosmin G. Petra, Nai-Yuan Chiang, Jingyi Wang, Tucker Hartland, and Michael Puso (submitted), LLNL-JRNL-869761.
The Challenges of Unsteady Adjoint-Based Design: - The adjoint approach still provides all of the sensitivities at the same cost as analysis, and the 20x estimate still applies for the expense of an optimization - But every simulation is now an unsteady problem - Where the steady adjoint solver linearized about a single solution (the steady-state), the unsteady adjoint solver must essentially do this at every physical time step
The efficiency gains obtained using higher-order implicit Runge-Kutta schemes as compared with the second-order accurate backward difference schemes for the unsteady Navier-Stokes equations are investigated. Three different algorithms for solving the nonlinear system of equations arising at each timestep are presented. The first algorithm (NMG) is a pseudo-time-stepping scheme which employs a non-linear full approximation storage (FAS) agglomeration multigrid method to accelerate convergence. The other two algorithms are based on Inexact Newton's methods. The linear system arising at each Newton step is solved using iterative/Krylov techniques and left preconditioning is used to accelerate convergence of the linear solvers. One of the methods (LMG) uses Richardson's iterative scheme for solving the linear system at each Newton step while the other (PGMRES) uses the Generalized Minimal Residual method. Results demonstrating the relative superiority of these Newton's methods based schemes are presented. Efficiency gains as high as 10 are obtained by combining the higher-order time integration schemes with the more efficient nonlinear solvers.
We present an efficient implementation of the Generalized Green’s function Cluster Expansion (GGCE), which is a new method for computing the ground-state properties and dynamics of polarons (single electrons coupled to lattice vibrations) in model electron-phonon systems. The GGCE works at arbitrary temperature and is well suited for a variety of electron-phonon couplings, including, but not limited to, site and bond Holstein and Peierls (Su-SchriefferHeeger) couplings, and couplings to multiple phonon modes with different energy scales and coupling strengths. Quick calculations can be performed efficiently on a laptop using solvers from NumPy and SciPy, or in parallel at scale using the PETSc sparse linear solver engine.
This work considers static and dynamic aeroelastic optimization of a cantilevered platewing in transonic flow. Low- and high-fidelity aeroelastic predictions are fed into a standard trust region model management scheme in order to efficiently solve this expensive optimization problem with multifidelity methods. Differences in fidelity are entirely driven by different aerodynamic solvers: linear panel methods for the low-fidelity method, and inviscid CFD-based solvers for the high-fidelity response. The optimization problem utilizes shape, sizing, and trim variables to satisfy stress, trim, and flutter constraints, and is solved across a range of subsonic and transonic Mach numbers. The multifidelity method is able to provide a speed-up for all cases considered here, despite sizable inaccuracies in the low-fidelity response for transonic flows.
The development of quantum processors for practical fluid flow problems is a promising yet distant goal. Recent advances in quantum linear solvers have highlighted their potential for classical fluid dynamics. In this study, we evaluate the Harrow–Hassidim–Lloyd (HHL) quantum linear systems algorithm (QLSA) for solving the idealized Hele–Shaw flow. Our focus is on the accuracy and computational cost of the HHL solver, which we find to be sensitive to the condition number, scaling exponentially with problem size. This emphasizes the need for preconditioning to enhance the practical use of QLSAs in fluid flow applications. Moreover, we perform shots-based simulations on quantum simulators and test the HHL solver on superconducting quantum devices, where noise, large circuit depths, and gate errors limit performance. Error suppression and mitigation techniques improve accuracy, suggesting that such fluid flow problems can benchmark noise mitigation efforts. Finally, our findings provide a foundation for future, more complex application of QLSAs in fluid flow simulations.
We present a family of discretizations for the Variable Eddington Factor (VEF) equations that have high-order accuracy on curved meshes and efficient preconditioned iterative solvers. The VEF discretizations are combined with the Discontinuous Galerkin transport discretization from to form effective high-order, linear transport methods. The VEF discretizations are derived by extending the unified analysis of Discontinuous Galerkin methods for elliptic problems presented by Arnold et al. to the VEF equations. This framework is used to define analogs of the interior penalty, second method of Bassi and Rebay, minimal dissipation local Discontinuous Galerkin, and continuous finite element methods. The analysis of subspace correction preconditioners, which use a continuous operator to iteratively precondition the discontinuous discretization, is extended to the case of the non-symmetric VEF system. Numerical results demonstrate that the VEF discretizations have arbitrary-order accuracy on curved meshes, preserve the thick diffusion limit, and are effective on a proxy problem from thermal radiative transfer in both outer transport iterations and inner preconditioned linear solver iterations. We demonstrate that the VEF solution converges to the S N transport solution as the mesh is refined on both problems with smooth and non-smooth behavior in angle. Parallel performance studies show that the interior penalty VEF discretization's linear solve weak scales out to 1024 processors and strong scales well on a single node. Particular attention is paid to the parallel performance of the VEF algorithm when used in combination with a parallel block Jacobi transport sweep.
Here, we demonstrate use of a modern Fortran solver interface to manage solver algorithms for an implicit barotropic mode solver in the Model for Predictions Across Scales-Ocean (MPAS-O). ForTrilinos, a Fortran interface to Trilinos that contains a large collection of solver capabilities written in C++, has been implemented in MPAS-O to provide access to a suite of linear solver options. By virtue of the simplified wrapper and interface generator (SWIG) automation tool that generates modern Fortran interfaces to C++ code, we were able to implement the Fortran solver interface in MPAS-O using a familiar Fortran coding style while minimizing performance degradation. The ForTrilinos solver interface is written within MPAS-O’s time stepping modules as a subroutine in conjunction with MPAS-O code. Applied to an idealized ocean and a high-resolution realistic ocean test case, parallel performance of ForTrilinos solvers is examined. It is found that parallel scalability of the ForTrilinos solvers is highly dependent on the number of global synchronization points per solver iteration in each iterative solver algorithm. ForTrilinos solvers perform best compared to the Fortran hand-crafted (FHC) solver when the amount of work per processor is large enough. However, parallel scalability is better with the FHC solver and so when the work per core is modest FHC outperforms ForTrilinos. The intercomparison between the ForTrilinos and FHC solvers reveals that this performance hit in the ForTrilinos solver mostly comes from the global synchronization process, while suggesting that the matrix-vector multiplication process in the FHC solver needs to be optimized for better performance.
Two linearized solvers (time and frequency domain) based on a high resolution numerical scheme are presented. The basic approach is to linearize the flux vector by expressing it as a sum of a mean and a perturbation. This allows the governing equations to be maintained in conservation law form. A key difference between the time and frequency domain computations is that the frequency domain computations require only one grid block irrespective of the interblade phase angle for which the flow is being computed. As a result of this and due to the fact that the governing equations for this case are steady, frequency domain computations are substantially faster than the corresponding time domain computations. The linearized equations are used to compute flows in turbomachinery blade rows (cascades) arising due to blade vibrations. Numerical solutions are compared to linear theory (where available) and to numerical solutions of the nonlinear Euler equations.
Kynema-SGF (formerly AMR-Wind), wherein SGF stands for structured-grid fluid dynamics, is a massively parallel, block-structured adaptive-mesh, incompressible flow solver. The codebase was initiated in 2019 from incflo. The solver is built on top of the AMReX library. AMReX library provides the mesh data structures, mesh adaptivity, as well as the linear solvers used for solving the governing equations. Kynema-SGF is actively developed and maintained by a dedicated multi-institutional team from Lawrence Berkeley National Laboratory, National Laboratory of the Rockies, and Sandia National Laboratories. The primary applications for Kynema-SGF are: performing large-eddy simulations (LES) of atmospheric boundary layer (ABL) flows, simulating wind farm turbine-wake interactions using actuator disk or actuator line models for turbines, and as a background solver when coupled with a near-body solver (e.g., Kynema-UGF) with overset methodology to perform blade-resolved simulations of multiple wind turbines within a wind farm. For offshore applications, the ability to model the air-sea interaction effects and its impact on the ABL characteristics is another focus for the code development effort. As with other codes in the Kynema ecosystem, Kynema-SGF shares the following objectives: *an open, well-documented implementation of the state-of-the-art computational models for modeling wind farm flow physics at various fidelities that are backed by a comprehensive verification and validation (V&V) process; *be capable of performing the highest-fidelity simulations of flow fields within wind farms; and *be able to leverage the high-performance leadership class computing facilities available at DOE national laboratories.