Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Iteration method”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Iterative methods in GPU-resident linear solvers for nonlinear constrained optimization

Linear solvers are major computational bottlenecks in a wide range of decision support and optimization computations. The challenges become even more pronounced on heterogeneous hardware, where traditional sparse numerical linear algebra methods are often inefficient. For example, methods for solving ill-conditioned linear systems have relied on conditional branching, which degrades performance on hardware accelerators such as graphical processing units (GPUs). To improve the efficiency of solving ill-conditioned systems, our computational strategy separates computations that are efficient on GPUs from those that need to run on traditional central processing units (CPUs). Our strategy maximizes the reuse of expensive CPU computations. Iterative methods, which thus far have not been broadly used for ill-conditioned linear systems, play an important role in our approach. In particular, we extend ideas from Arioli et al., (2007) to implement iterative refinement using inexact LU factors and flexible generalized minimal residual (FGMRES), with the aim of efficient performance on GPUs. In conclusion, we focus on solutions that are effective within broader application contexts, and discuss how early performance tests could be improved to be more predictive of the performance in a realistic environment.

97 MATHEMATICS AND COMPUTING↗

An iterative method to deblend AGN-Host contributions for Integral Field spectroscopic observations

ABSTRACT We present a new iterative deblending method to separate the host galaxy (HG) and their Active Galactic Nuclei (AGNs) emission with the use of Integral Field spectroscopic (IFS) data. The method decomposes the resolved HG emission from the unresolved AGN emission by modelling the two-dimensional surface brightness (SB) profile of the point-spread function (PSF) and the two-dimensional SB HG continuum simultaneously per each monochromatic slide. Our method does not require any prior information about the observed SB profile or a detailed fitting of the PSF, making it ideal for the automatic analysis of large galaxy samples. In this work, we test the quality of our method, its advantages, and its disadvantages. We test our method by using a set of IFS mock data cubes to quantify the reliability of our deblending process and further compare our method with the qdblend3d analysis tool. Furthermore, we applied our method to three data cubes selected from the MaNGA survey according to the dominance of either its HG or its AGN. We show that our deblending method is capable of disengaging the bright, non-resolved AGN emission from the HG continuum and its narrow emission lines. However, the decoupling depends on how well the IFS spatially resolves the PSF, and on the relative flux intensity of the HG-AGN. Therefore, the method is ideal for disentangling the bright-flux contribution from AGN-dominated spectra.

Ibarra-Medel, H. (ORCID:0000000297906313)↗

Robust Iterative Method for Symmetric Quantum Signal Processing in All Parameter Regimes

Here, this paper addresses the problem of solving nonlinear systems in the context of symmetric quantum signal processing (QSP), a powerful technique for implementing matrix functions on quantum computers. Symmetric QSP focuses on representing target polynomials as products of matrices in SU(2) that possess symmetry properties. We present a novel Newton’s method tailored for efficiently solving the nonlinear system involved in determining the phase factors within the symmetric QSP framework. Our method demonstrates rapid and robust convergence in all parameter regimes, including the challenging scenario with ill-conditioned Jacobian matrices, using standard double precision arithmetic operations. For instance, solving symmetric QSP for a highly oscillatory target function α cos(1000x) (polynomial degree ≈ 1433) takes 6 iterations to converge to machine precision when α = 0.9, and the number of iterations only increases to 18 iterations when α = 1 – 10 -9 with a highly ill-conditioned Jacobian matrix. Leveraging the matrix product state structure of symmetric QSP, the computation of the Jacobian matrix incurs a computational cost comparable to a single function evaluation. Moreover, we introduce a reformulation of symmetric QSP using real-number arithmetics, further enhancing the method’s efficiency. Extensive numerical tests validate the effectiveness and robustness of our approach, which has been implemented in the QSPPACK software package.

97 MATHEMATICS AND COMPUTING↗

An efficient iterative method for dynamical Ginzburg-Landau equations

Here we propose a new finite element approach to solving the time-dependent Ginzburg-Landau equations for superconductivity under the temporal gauge and design an efficient preconditioner for the Newton iteration of the resulting discrete system. It solves the vector magnetic potential in $H$(curl) space by the lowest order of the second kind N´ed´elec element. This approach offers a simple way to deal with the boundary conditions, leading to a stable and reliable performance for superconductor geomtry with reentrant corners. The bench-marking in numerical simulations verifies the efficiency of the proposed preconditioner, which can be employed to significantly speed up in large-scale computations.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗

HyKKT: a hybrid direct-iterative method for solving KKT linear systems

Here, we propose a solution strategy for the large indefinite linear systems arising in interior methods for nonlinear optimization. The method is suitable for implementation on hardware accelerators such as graphical processing units (GPUs). The current gold standard for sparse indefinite systems is the LBLT factorization where L is a lower triangular matrix and B is 1×1 or 2×2 block diagonal. However, this requires pivoting, which substantially increases communication cost and degrades performance on GPUs. Our approach solves a large indefinite system by solving multiple smaller positive definite systems, using an iterative solver on the Schur complement and an inner direct solve (via Cholesky factorization) within each iteration. Cholesky is stable without pivoting, thereby reducing communication and allowing reuse of the symbolic factorization. We demonstrate the practicality of our approach on large optimal power flow problems and show that it can efficiently utilize GPUs and outperform LBL T factorization of the full system.

97 MATHEMATICS AND COMPUTING↗

Low-synch Gram–Schmidt with delayed reorthogonalization for Krylov solvers

The parallel strong-scaling of iterative methods is often determined by the number of global reductions at each iteration. Low-synch Gram-Schmidt algorithms are applied here to the Arnoldi algorithm to reduce the number of global reductions and therefore to improve the parallel strong-scaling of iterative solvers for nonsymmetric matrices such as the GMRES and the Krylov-Schur iterative methods. In the Arnoldi context, the factorization is "left-looking" and processes one column at a time. Among the methods for generating an orthogonal basis for the Arnoldi algorithm, the classical Gram-Schmidt algorithm, with reorthogonalization (CGS2) requires three global reductions per iteration. A new variant of CGS2 that requires only one reduction per iteration is presented and applied to the Arnoldi algorithm. Delayed CGS2 (DCGS2) employs the minimum number of global reductions per iteration (one) for a one-column at-a-time algorithm. The main idea behind the new algorithm is to group global reductions by rearranging the order of operations. DCGS2 must be carefully integrated into an Arnoldi expansion or a GMRES solver. Numerical stability experiments assess robustness for Krylov-Schur eigenvalue computations. Performance experiments on the ORNL Summit supercomputer then establish the superiority of DCGS2 over CGS2.

97 MATHEMATICS AND COMPUTING↗

A simulation study of the ability to detect power distribution perturbations in the texas A&M TRIGA reactor with self-powered neutron detectors

Given the variety of ways that nuclear reactor core power may be perturbed, reactor operators and developers are keen on understanding the accuracy and convergence time during which perturbations in reactor power distribution may be synthesized (i.e., inferred) from an array of in-core radiation detectors. A simulation study was conducted as described herein using a highly detailed model of the Texas A&M Training, Research, Isotopes, General Atomics Reactor, in which an array of self-powered neutron detectors (SPNDs) was considered for input to the power synthesis methodology. The core power synthesis is conducted using a point-based iterative method with an iterative loop built in to ensure working equation consistency. The forward problem of SPND response to simulated perturbations in reactor power was solved for Gaussian peak-type perturbations in the reactor power distribution. These perturbations varied in variance, amplitude, and core location to assess their impact on synthesis error and to determine the number of iterations required for convergence. A relation between the unique resolvability limit and perturbation width was identified such that the maximum synthesis error increased rapidly when the peak width went beneath this limit (a width approximating half the reactor’s fuel pin-to-pin pitch); this resolvability limit is specific to the SPND configuration and fuel segmentation considered herein. The synthesis error increased linearly with perturbation peak amplitude, whereas the convergence time increased nonlinearly. Perturbations located closer to the center of the core were synthesized more accurately, albeit with a higher number of required iterations. These findings provide a qualitative and quantitative understanding of the accuracy and speed at which different types of spatial power perturbations can be resolved in light-water reactors.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Hybrid eigensolvers for nuclear configuration interaction calculations

We examine and compare several iterative methods for solving large-scale eigenvalue problems arising from nuclear structure calculations. In particular, we discuss the possibility of using block Lanczos method, a Chebyshev filtering based subspace iterations and the residual minimization method accelerated by direct inversion of iterative subspace (RMM-DIIS) and describe how these algorithms compare with the standard Lanczos algorithm and the locally optimal block preconditioned conjugate gradient (LOBPCG) algorithm. Although the RMM-DIIS method does not exhibit rapid convergence when the initial approximations to the desired eigenvectors are not sufficiently accurate, it can be effectively combined with either the block Lanczos or the LOBPCG method to yield a hybrid eigensolver that has several desirable properties. We will describe a few practical issues that need to be addressed to make the hybrid solver efficient and robust.

97 MATHEMATICS AND COMPUTING↗

An Efficient High-Order Solver for Diffusion Equations with Strong Anisotropy on Non-Anisotropy-Aligned Meshes

This paper concerns numerical solution of the diffusion equation with strong anisotropy on meshes not aligned with the anisotropic vector field. In order to resolve the numerical pollution for simulations on a non-anisotropy-aligned mesh and reduce the associated high computational cost we propose an effective preconditioner, extending our previous work. Similar to the anisotropy-aligned mesh case, we apply the auxiliary space preconditioning framework to design a preconditioner where a continuous finite element space is used as the auxiliary space for the discontinuous finite element space. The key component is an effective line smoother that can mitigate the high-frequency errors perpendicular to the magnetic field. We design a graph-based approach to find such a line smoother that is approximately perpendicular to the vector fields when the mesh does not align with the anisotropy. Finally, numerical experiments for several benchmark problems are presented, demonstrating the effectiveness and robustness of the proposed preconditioner when applied to Krylov iterative methods.

97 MATHEMATICS AND COMPUTING↗

Detailed Characterization of CZT Detector Response for Improved Coded-Aperture Imaging Performance

Gamma-ray imaging is a powerful method for locating and quantifying sources of radiation. The coded-aperture technique demonstrates superior angular resolution in comparison to other methods (e.g., Compton reconstruction). In this method, a mask constructed of highly attenuating material encodes the scene as a shadow pattern on a position-sensitive detector; this pattern can then be used to recreate the origin(s) of incident radiation. This is typically done through convolution of the mask and shadow patterns. Iterative methods which attempt to reconstruct the observed shadow pattern using a weighted combination of simulated patterns may also be employed. In either case, errors in event position reconstruction due to detector imperfections alter the shadow pattern and will therefore degrade system performance and may introduce imaging artifacts. These effects can be mitigated with a detailed understanding of such errors – allowing for the generation of representative simulations that include the errors and/or correction of raw imager data to remove the errors. We present a calibration process for a commercially available cadmium zinc telluride (CZT) gamma imager which provides a comprehensive characterization of the spatial and energy dependence of event reconstruction. By illuminating a mask featuring a regular grid of pinholes with a calibration source, the localized response of the detector can be measured with fine granularity. These local responses are combined to generate a full detector response map which can be used to distort simulations in a manner that is representative of the observed detector data. Details of the calibration procedure and an assessment of the impact of its end products on the performance of iterative imaging methods will be presented.

Ziock, Klaus-Peter↗

Accelerating eigenvalue computation for nuclear structure calculations via perturbative corrections

Subspace projection methods utilizing perturbative corrections have been proposed for computing the lowest few eigenvalues and corresponding eigenvectors of large Hamiltonian matrices. In this paper, we build upon these methods and introduce the term Subspace Projection with Perturbative Corrections (SPPC) method to refer to this approach. We tailor the SPPC for nuclear many-body Hamiltonians represented in a truncated configuration interaction subspace, i.e., the no-core shell model (NCSM). We use the hierarchical structure of the NCSM Hamiltonian to partition the Hamiltonian as the sum of two matrices. The first matrix corresponds to the Hamiltonian represented in a small configuration space, whereas the second is viewed as the perturbation to the first matrix. Eigenvalues and eigenvectors of the first matrix can be computed efficiently. Because of the split, perturbative corrections to the eigenvectors of the first matrix can be obtained efficiently from the solutions of a sequence of linear systems of equations defined in the small configuration space. These correction vectors can be combined with the approximate eigenvectors of the first matrix to construct a subspace from which more accurate approximations of the desired eigenpairs can be obtained. We show by numerical examples that the SPPC method can be more efficient than conventional iterative methods for solving large-scale eigenvalue problems such as the Lanczos, block Lanczos and the locally optimal block preconditioned conjugate gradient (LOBPCG) method. The method can also be combined with other methods to avoid convergence stagnation.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Anderson acceleration stability in NDA-accelerated k-eigenvalue problems

Anderson acceleration (AA) has been used to improve the stability and convergence rate of multiphysics iterative methods for reactor analysis. Most applications studied assume a tightly converged solution for the different physics problems, and AA is usually applied to state variables like temperature, density, and heat generation rate. In this paper, we study the theoretical performance of AA in NDA-accelerated k-eigenvalue problems. The problems and algorithms studied are simplified from the coupled iteration scheme adopted by MPACT and many other high-fidelity whole-core reactor codes. Compared to previous analyses of AA for these iteration schemes, we study the case with a partially converged neutronics solution and possibly partially converged nonlinear diffusion acceleration (NDA)/coarse mesh finite difference (CMFD) solutions. We observe that the performance of the iteration scheme with AA is very sensitive to the initial guess and is affected by the partially converged CMFD solutions. When the NDA solution is fully converged, using AA cannot achieve the optimal convergence rate in large-sized problems. Conversely, if the NDA solution is partially converged, the iteration scheme with AA can diverge or converge extremely slowly. It is found that the loss of robustness for AA is due to the fact that it is applied to the iterative subspace of state variables rather than the fundamental unknowns of the governing equations. To improve the robustness, the scalar flux should also be considered in the implementation of AA. After considering the residuals of flux, we observe that the stability is regardless of the partial convergence of NDA solutions. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Seven open problems in applied combinatorics

We present and discuss seven different open problems in applied combinatorics. Additionally, the application areas relevant to this compilation include quantum computing, algorithmic differentiation, topological data analysis, iterative methods, hypergraph cut algorithms, and power systems.

97 MATHEMATICS AND COMPUTING↗

2022 AI Testbed Expeditions Report

By exploiting the coherent properties of a light source, coherent diffraction imaging (CDI) is able to obtain the sample image at a nanoscale resolution using the measured diffraction pattern. Bragg Coherent Diffraction Imaging (BCDI) has become valuable for recovering the displacement and strain field of crystals, providing a valuable tool in material science and solid-state physics. X-ray ptychography is another emerging CDI technique that can produce a high-resolution image of the extended sample and has become popular in many research areas (e.g., materials science, biology, electronics, and optics characterization). CDI including BCDI and ptychography has become an established technique in Synchrotron Facilities including the Advanced Photon Source (APS) and will greatly benefit from the 100x coherent flux increase of the upcoming APS Upgrade (APSU). The current image formation process in CDI employs iterative phase retrieval algorithms, which is a time-consuming and computationally expensive process. Especially after APSU, the traditional iterative methods will not be able to match the experimental data acquisition speed. We employ deep learning (DL) approach to replace the iterative approaches, therefore allowing hundreds of times faster recovery of the object. We developed AutoPhaseNN, a DL-based approach which learns to solve the inverse problem without labeled data. Taking 3D BCDI as a representative technique, AutoPhaseNN has been demonstrated to be one hundred times faster than traditional iterative phase retrieval methods while providing comparable image quality. The current network is trained with 64 x 64 x 64 data size, to achieve higher resolution imaging, we will need to scale the network to input and train/infer 3D arrays of size 256 x 256 x 256 (today) and of size 2560x2560x2560 (APSU). However, the scalability of the network is restricted due to the memory-intensive training process. To perform the training for a 256 x 256 x 256 data size, the required memory exceeds the capacity of the current machine. In this project, we explore using Sambanova system to train the network for the direct data inversion for CDI.

36 MATERIALS SCIENCE↗