Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “conjugate gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Forward and inverse modeling of fault transmissibility in subsurface flows

Characterizing physical properties of faults, such as their transmissibility, is crucial for performing predictive numerical simulation of subsurface flows, such as those encountered in petroleum engineering and remediation of subsurface contamination. Here, this paper provides a complete investigation of the inverse problem for fault transmissibility in subsurface flow models, under appropriate assumptions on fault structure. In particular, the following aspects are considered: 1) fault modeling and well-posedness of the forward problem; 2) finite element (FEM) discretizations of the forward problem and their rigorous a priori convergence analysis; 3) Well-posedness of the Bayesian inverse problem, FEM discretization of the infinite dimensional Bayesian inverse formulation, and its rigorous a priori analysis. Moreover, computation of the maximum a posteriori (MAP) point via fast inexact Newton-conjugate gradient optimization and a Laplace approximation of the Bayesian posterior are also presented. Numerical results illustrate the use of the proposed fault model in forward and inverse problems for subsurface flows in two dimensional domains with multiple faults.

97 MATHEMATICS AND COMPUTING↗

Integral boundary conditions in phase field models

Modeling the chemical, electric and thermal transport as well as phase transitions and the accompanying mesoscale microstructure evolution within a material in an electronic device setting involves the solution of partial differential equations often with integral boundary conditions. Employing the familiar Poisson equation describing the electric potential evolution in a material exhibiting insulator to metal transitions, we exploit a special property of such an integral boundary condition, and we properly formulate the variational problem and establish its well-posedness. Next, we compare our method with the commonly-used Lagrange multiplier method that can also handle such boundary conditions. Numerical experiments demonstrate that our new method achieves optimal convergence rate in contrast to the conventional Lagrange multiplier method. Furthermore, the linear system derived from our method is symmetric positive definite, and can be efficiently solved by Conjugate Gradient method with algebraic multigrid preconditioning.

97 MATHEMATICS AND COMPUTING↗

Simulation of multi-shell fullerenes using Machine-Learning Gaussian Approximation Potential

Multi-shell fullerenes ”buckyonions ” were simulated, starting from initially random configurations, using a density-functional-theory (DFT)-trained machine-learning carbon potential within the Gaussian Approximation Potential (GAP) Framework [Volker L. Deringer and Gábor Csányi, Phys. Rev. B 95, 094203 (2017)]. Fullerenes formed from seven different system sizes, ranging from 60 ~ 3774 atoms, were considered. The buckyonions are formed by clustering and layering starting from the outermost shell and proceeding inward. Inter-shell cohesion is partly due to interaction between delocalized π electrons protruding into the gallery. The energies of the models were validated ex post facto using density functional codes, VASP and SIESTA , revealing an energy difference within the range of 0.02 - 0.08 eV/atom after conjugate gradient energy convergence of the models was achieved with both methods.

74 ATOMIC AND MOLECULAR PHYSICS↗

Simulation toolkit for digital material characterization of large image-based microstructures

In this paper, an efficient image-based simulation toolkit for material characterization is presented, which is scalable to work from personal computers to workstations. The effective thermal conductivity, elasticity, and permeability are evaluated employing a computational homogenization framework based on the Finite Element Method (FEM). Two complementary open-source packages are presented: one developed in Python, which can convert digital images into voxel meshes (pyTomoviewer); the other developed in Julia, that can run numerical simulations to compute effective material properties (chpack). Also, a CUDA C version of chpack is provided (chfem_gpu). They were designed to deal with large multi-phase models, so strategies were devised to minimize their memory footprint, while avoiding a high toll on execution time. The voxel-based approach significantly simplifies the FEM meshes and allows efficient matrix-free implementations. In that sense, to handle large linear systems of equations, the element-by-element (EBE) technique is adopted, in conjunction with a low-memory implementation of the Preconditioned Conjugate Gradient (PCG) method. Finally, the code was thoroughly tested on an artificial geometry made of a square array of cylinders, for which analytical solutions exist, as well as on a real micro-tomographic reconstruction of FiberForm TM , a carbon preform commonly used in thermal protection systems.

36 MATERIALS SCIENCE↗

Hybrid eigensolvers for nuclear configuration interaction calculations

We examine and compare several iterative methods for solving large-scale eigenvalue problems arising from nuclear structure calculations. In particular, we discuss the possibility of using block Lanczos method, a Chebyshev filtering based subspace iterations and the residual minimization method accelerated by direct inversion of iterative subspace (RMM-DIIS) and describe how these algorithms compare with the standard Lanczos algorithm and the locally optimal block preconditioned conjugate gradient (LOBPCG) algorithm. Although the RMM-DIIS method does not exhibit rapid convergence when the initial approximations to the desired eigenvectors are not sufficiently accurate, it can be effectively combined with either the block Lanczos or the LOBPCG method to yield a hybrid eigensolver that has several desirable properties. We will describe a few practical issues that need to be addressed to make the hybrid solver efficient and robust.

97 MATHEMATICS AND COMPUTING↗

Optimal design of chemoepitaxial guideposts for the directed self-assembly of block copolymer systems using an inexact Newton algorithm

Directed self-assembly (DSA) of block copolymers (BCPs) is one of the most promising developments in the cost-effective production of nanoscale devices. The process makes use of the natural tendency for BCP melts to form nanoscale structures upon phase separation. The phase separation can be directed through the use of chemically patterned substrates to promote the formation of morphologies that are essential to the production of semiconductor devices. Moreover, the design of substrate pattern can be formulated as an optimization problem for which we seek optimal substrate designs that effectively produce given target morphologies. In this paper, we adopt a phase field model given by a nonlocal Cahn–Hilliard partial differential equation (PDE) based on the minimization of the Ohta–Kawasaki free energy, and present an efficient PDE-constrained optimization framework for the optimal design problem. The design variables are the locations of circular- or strip-shaped guiding posts that are used to model the substrate chemical pattern. To solve the ensuing optimization problem, we propose a variant of an inexact Newton conjugate gradient algorithm tailored to this problem. Additionally, we demonstrate the effectiveness of our computational strategy on numerical examples that span a range of target morphologies. Owing to our second-order optimizer and fast state solver, the numerical results demonstrate five orders of magnitude reduction in computational cost over previous work. The efficiency of our framework and the fast convergence of our optimization algorithm enable us to rapidly solve the optimal design problem in not only two, but also three spatial dimensions.

97 MATHEMATICS AND COMPUTING↗

Accelerating eigenvalue computation for nuclear structure calculations via perturbative corrections

Subspace projection methods utilizing perturbative corrections have been proposed for computing the lowest few eigenvalues and corresponding eigenvectors of large Hamiltonian matrices. In this paper, we build upon these methods and introduce the term Subspace Projection with Perturbative Corrections (SPPC) method to refer to this approach. We tailor the SPPC for nuclear many-body Hamiltonians represented in a truncated configuration interaction subspace, i.e., the no-core shell model (NCSM). We use the hierarchical structure of the NCSM Hamiltonian to partition the Hamiltonian as the sum of two matrices. The first matrix corresponds to the Hamiltonian represented in a small configuration space, whereas the second is viewed as the perturbation to the first matrix. Eigenvalues and eigenvectors of the first matrix can be computed efficiently. Because of the split, perturbative corrections to the eigenvectors of the first matrix can be obtained efficiently from the solutions of a sequence of linear systems of equations defined in the small configuration space. These correction vectors can be combined with the approximate eigenvectors of the first matrix to construct a subspace from which more accurate approximations of the desired eigenpairs can be obtained. We show by numerical examples that the SPPC method can be more efficient than conventional iterative methods for solving large-scale eigenvalue problems such as the Lanczos, block Lanczos and the locally optimal block preconditioned conjugate gradient (LOBPCG) method. The method can also be combined with other methods to avoid convergence stagnation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Random Phase Approximation Correlation Energy Using Real-Space Density Functional Perturbation Theory

We present a real-space method for computing the random phase approximation (RPA) correlation energy within Kohn–Sham density functional theory, leveraging the low-rank nature of the frequency-dependent density response operator. In particular, we employ a cubic-scaling formalism based on density functional perturbation theory that circumvents the calculation of the response function matrix, instead relying on the ability to compute its product with a vector through the solution of the associated Sternheimer linear systems. We develop a large-scale parallel implementation of this formalism using the subspace iteration method in conjunction with the spectral quadrature method while employing the Kronecker product-based method for the application of the Coulomb operator and the conjugate orthogonal conjugate gradient method for the solution of the linear systems. We demonstrate convergence with respect to key parameters and verify the method’s accuracy by comparing with plane-wave results. We show that the framework achieves good strong scaling to many thousands of processors, reducing the time to solution for a lithium hydride system with 128 electrons to around 150 s on 4608 processors.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A Scalable Semi-Implicit Barotropic Mode Solver for the MPAS-Ocean

A scalable semi-implicit barotropic mode solver for the ocean component of the model for prediction across scales has been implemented as a competitor to an existing explicit-subcycling scheme to allow faster and more stable simulations while not sacrificing accuracy. The semi-implicit solver adopts the pipelined preconditioned bi-conjugate gradient stabilization algorithm as an iterative solver in conjunction with the restricted additive Schwarz preconditioner that accelerates the convergence rate of the iterative solver. The preconditioner is constructed from a linearized barotropic system that also reorders the system for optimal performance, while the semi-implicit solver deals with the fully nonlinear barotropic system that requires reassembly of the coefficient matrix for every time step. Several numerical experiments, from simple one-dimensional tests to three-dimensional real-world tests, demonstrate that the semi-implicit solver has almost the same accuracy and better parallel scalability compared with the existing scheme while allowing faster and more stable simulations. Furthermore, the semi-implicit solver accelerates the barotropic mode up to 2.9 times faster than the existing scheme on 16,320 processors, leading to an overall runtime speedup of 1.9.

97 MATHEMATICS AND COMPUTING↗

A Comparison of Linear Solvers for Resolving Flow in Three-Dimensional Discrete Fracture Networks

We compare various methods for resolving steady flow within three-dimensional discrete fracture networks, including direct methods, Krylov subspace methods with and without preconditioning, and multi-grid methods. We compared the performance of the methods based on compute times and scaling of the solution as a function of the number of grid nodes and log-variance of the hydraulic aperture. The methods are applied to three test cases: (a) variable density of networks with a truncated power-law distribution of fracture lengths, (b) a fixed network composed of monodisperse fracture sizes but varied permeability/aperture heterogeneity, (c) and a network based on field site in Nevada, US. We chose these cases to allow us to study the impact of the mesh size and flow properties, as well as to demonstrate our conclusions on a large-scale, realistic problem (more than 40 million mesh nodes). A direct solution using Cholesky factorization outperformed other methods for every example but was closely followed in performance by some algebraic multigrid (AMG) preconditioned Krylov subspace methods. Among the Krylov methods, conjugate gradients (CG) with an AMG preconditioner performs the best. Generally, Cholesky factorization is recommended, but CG with an AMG preconditioner may be suitable for very large problems beyond 40 million nodes where the entire linear system cannot reside in memory.

58 GEOSCIENCES↗

Predicting Flow in Fracture Networks With Quantum Algorithms

Uncertainty quantification plays a crucial role in the modeling of subsurface flow. For instance, uncertainties in the properties of geologic fracture networks significantly impact flow, requiring numerous simulations to accurately estimate quantities of interest. However, each simulation is computationally expensive because it requires solving a large linear system to capture features that involve both small and large fractures. An example is in percolation, where the interaction of many small fractures (which cumulatively can have a large surface area) with the rock matrix must be modeled precisely. Quantum computing is an emerging tool with the potential to address this issue. Quantum algorithms offer a significant speedup in solving linear systems, achieving efficiencies that are challenging to match with classical approaches. These classical approaches include direct solvers, such as LU decomposition, and iterative methods, notably preconditioned conjugate gradient, commonly used in subsurface modeling to solve large sparse systems. However, applying quantum algorithms to geologic fracture flow requires careful attention to algorithmic and problem-specific constraints to fully realize this quantum advantage. In this work we describe a quantum algorithm for generalized Monte Carlo applications with a quadratic speedup over the classical approaches which can be combined with the quantum speedup, currently under investigation, for solving quantum linear systems for subsurface flow. We show that for quantum algorithms the computational cost of estimating a quantity of interest for a statistical ensemble of networks is roughly the same as that of a single realization, essentially implying that one can get uncertainty quantification for free.

58 GEOSCIENCES↗

Scalable and accurate multi-GPU-based image reconstruction of large-scale ptychography data

Abstract While the advances in synchrotron light sources, together with the development of focusing optics and detectors, allow nanoscale ptychographic imaging of materials and biological specimens, the corresponding experiments can yield terabyte-scale volumes of data that can impose a heavy burden on the computing platform. Although graphics processing units (GPUs) provide high performance for such large-scale ptychography datasets, a single GPU is typically insufficient for analysis and reconstruction. Several works have considered leveraging multiple GPUs to accelerate the ptychographic reconstruction. However, most of these works utilize only the Message Passing Interface to handle the communications between GPUs. This approach poses inefficiency for a hardware configuration that has multiple GPUs in a single node, especially while reconstructing a single large projection, since it provides no optimizations to handle the heterogeneous GPU interconnections containing both low-speed (e.g., PCIe) and high-speed links (e.g., NVLink). In this paper, we provide an optimized intranode multi-GPU implementation that can efficiently solve large-scale ptychographic reconstruction problems. We focus on the maximum likelihood reconstruction problem using a conjugate gradient (CG) method for the solution and propose a novel hybrid parallelization model to address the performance bottlenecks in the CG solver. Accordingly, we have developed a tool, called PtyGer ( Pty chographic G PU(multipl e )-based r econstruction), implementing our hybrid parallelization model design. A comprehensive evaluation verifies that PtyGer can fully preserve the original algorithm’s accuracy while achieving outstanding intranode GPU scalability.

97 MATHEMATICS AND COMPUTING↗

BEYONDPLANCK II. CMB mapmaking through Gibbs sampling

We present a Gibbs sampling solution to the mapmaking problem for cosmic microwave background (CMB) measurements that builds on existing destriping methodology. Gibbs sampling breaks the computationally heavy destriping problem into two separate steps: noise filtering and map binning. Considered as two separate steps, both are computationally much cheaper than solving the combined problem. This provides a huge performance benefit as compared to traditional methods and it allows us, for the first time, to bring the destriping baseline length to a single sample. Here, we applied the Gibbs procedure to simulated Planck 30 GHz data. We find that gaps in the time-ordered data are handled efficiently by filling them in with simulated noise as part of the Gibbs process. The Gibbs procedure yields a chain of map samples, from which we are able to compute the posterior mean as a best-estimate map. The variation in the chain provides information on the correlated residual noise, without the need to construct a full noise covariance matrix. However, if only a single maximum-likelihood frequency map estimate is required, we find that traditional conjugate gradient solvers converge much faster than a Gibbs sampler in terms of the total number of iterations. The conceptual advantages of the Gibbs sampling approach lies in statistically well-defined error propagation and systematic error correction. This methodology thus forms the conceptual basis for the mapmaking algorithm employed in the BEYONDPLANCK framework, which implements the first end-to-end Bayesian analysis pipeline for CMB observations.

79 ASTRONOMY AND ASTROPHYSICS↗

Efficient optimization method for finding minimum energy paths of magnetic transitions

Here, efficient algorithms for the calculation of minimum energy paths of magnetic transitions are implemented within the geodesic nudged elastic band (GNEB) approach. While an objective function is not available for GNEB and a traditional line search can, therefore, not be performed, the use of limited memory Broyden–Fletcher–Goldfarb–Shanno (LBFGS) and conjugate gradient algorithms in conjunction with orthogonal spin optimization (OSO) approach is shown to greatly outperform the previously used velocity projection and dissipative Landau–Lifschitz dynamics optimization methods. The implementation makes use of energy weighted springs for the distribution of the discretization points along the path and this is found to improve performance significantly. The various methods are applied to several test problems using a Heisenberg-type Hamiltonian, extended in some cases to include Dzyaloshinskii–Moriya and exchange interactions beyond nearest neighbours. Minimum energy paths are found for magnetization reversals in a nano-island, collapse of skyrmions in two-dimensional layers and annihilation of a chiral bobber near the surface of a three-dimensional magnet. The LBFGS-OSO method is found to outperform the dynamics based approaches by up to a factor of 8 in some cases.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

A rotationally invariant approach based on Gutzwiller wave function for correlated electron systems

Here, we introduce a rotationally invariant approach combined with the Gutzwiller conjugate gradient minimization method to study correlated electron systems. In the approach, the Gutzwiller projector is parametrized based on the number of electrons occupying the onsite orbitals instead of the onsite configurations. The approach efficiently groups the onsite orbitals according to their symmetry and greatly reduces the computational complexity, which yields a speedup of $20 \sim 50 \times $ in the minimal basis energy calculation of dimers. The computationally efficient approach promotes more accurate calculations beyond the minimal basis that is inapplicable in the original approach. A large-basis energy calculation of F 2 demonstrates favorable agreements with standard quantum-chemical calculations Bytautas et al (2007 J. Chem. Phys. 127 164317).

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Ab initio calculation of atomic solid hydrogen phases based on Gutzwiller many-body wave functions

We apply two ab initio many-body methods based on Gutzwiller wave functions, i.e., correlation matrix renormalization theory (CMRT) and Gutzwiller conjugate gradient minimization (GCGM), to the study of crystalline phases of atomic hydrogen. Both methods avoid empirical Hubbard U parameters and are free from double-counting issues. CMRT employs a Gutzwiller-type approximation that enables efficient calculations, while GCGM goes beyond this approximation to achieve higher accuracy at higher computational cost. By benchmarking against available quantum Monte Carlo (QMC) results, we demonstrate that while both methods are more accurate than the widely used density-functional theory, GCGM systematically captures additional correlation energy missing in CMRT, leading to significantly improved total energy predictions. We also show that by including the correlation energy Ec from local density approximation in the CMRT calculation, CMRT + E c produces energy in better agreement with the QMC results in these hydrogen lattice systems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

Light in the dark forest. Part I. An efficient optimal estimator for 3D Lyman-alpha forest power spectrum

The highly anisotropic nature of the Lyman-alpha (Lyα) forest data introduces a complex survey window function that complicates the measurement of the three-dimensional power spectrum ( P 3D ). In this paper, we present the first fully optimal estimator for P 3D , which exactly deconvolves the survey window function and marginalizes contaminated modes that distort the power spectrum. Our approach adapts optimal estimator techniques developed for the 2D cosmic microwave background data to the 3D case. To achieve computational feasibility, we employ the conjugate gradient method and implement the P 3 M formalism to handle large-scale and small-scale operations separately and efficiently. We validate our estimator using Monte Carlo mocks and Gaussian simulations, demonstrating its accuracy and computational efficiency. We confirm that mode marginalization eliminates distortions arising from quasar continuum errors and delivers robust power spectrum estimation, though it also inflates errors at large scales. This first implementation works in the flat-sky case; we discuss the remaining steps needed to generalize it to the curved-sky case. This formalism offers a foundation for the Lyα forest P 3D measurements and a new path toward cosmological constraints from the Lyα forest data.

Lyman alpha forest↗

Nucleon-pair coupling scheme in Elliott's SU(3) model

Elliott's SU(3) model is at the basis of the shell-model description of rotational motion in atomic nuclei. Here we demonstrate that SU(3) symmetry can be realized in a truncated shell-model space if constructed in terms of a sufficient number of collective S, D, G,...pairs (i.e., with angular momentum zero, two, four,...) and if the structure of the pairs is optimally determined either by a conjugate-gradient minimization method or from a Hartree-Fock intrinsic state. We illustrate the procedure for six protons and six neutrons in the pf (sdg) shell and exactly reproduce the level energies and electric quadrupole properties of the ground-state rotational band with SDG (SDGI) pairs. The SD-pair approximation without significant renormalization, on the other hand, cannot describe the full SU(3) collectivity. A mapping from Elliott's fermionic SU(3) model to systems with s, d, g,... bosons provides insight into the existence of a decoupled collective subspace in terms of S, D, G,... pairs.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗