Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “convergence rate”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Stochastic Approximation for Multi-period Simulation Optimization with Streaming Input Data

We consider a continuous-valued simulation optimization (SO) problem, where a simulator is built to optimize an expected performance measure of a real-world system while parameters of the simulator are estimated from streaming data collected periodically from the system. At each period, a new batch of data is combined with the cumulative data and the parameters are re-estimated with higher precision. The system requires the decision variable to be selected in all periods. Therefore, it is sensible for the decision-maker to update the decision variable at each period by solving a more precise SO problem with the updated parameter estimate to reduce the performance loss with respect to the target system. We define this decision-making process as the multi-period SO problem and introduce a multi-period stochastic approximation (SA) framework that generates a sequence of solutions. Two algorithms are proposed: Re-start SA (ReSA) reinitializes the stepsize sequence in each period, whereas Warm-start SA (WaSA) carefully tunes the stepsizes, taking both fewer and shorter gradient-descent steps in later periods as parameter estimates become increasingly more precise. We show that under suitable strong convexity and regularity conditions, ReSA and WaSA achieve the best possible convergence rate in expected sub-optimality either when an unbiased or a simultaneous perturbation gradient estimator is employed, while WaSA accrues significantly lower computational cost as the number of periods increases. In addition, we present the regularized ReSA, which obviates the need to know the strong convexity constant and achieves the same convergence rate at the expense of additional computation.

Computer Science↗

Addendum to SAND2023-09604 Xyce lumped-element transmission line model verification to support Empire-Cable cable SGEMP analyses

This report supplements the Verification of Empire-Cable SAND report by expanding on the use of Xyce to simulate the coupling to a transmission line cable model. While Empire-Cable solves its governing equations on a high-order, finite-element mesh with an an implicit-in-time formulation, Xyce must use a first order graph for the circuit and explicit-in-time approach to be compatible with non-linear electrical device models. Thus, given the different solution methodologies in Xyce as compared to Empire-Cable, the convergence rates are expected to be different but the overall quality of the solution should be the same. The original four canonical problems studied in the Empire-Cable verification report are replicated here running in Xyce using transmission line modeling parameters from the verification report. Overall, agreement between the codes is excellent with Xyce’s convergence rates being limited mostly to first order due to the circuit network approximation of a transmission line being a first order approximation.

42 ENGINEERING↗

Addressing Load Imbalance in Bioinformatics and Biomedical Applications: Efficient Scheduling across Multiple GPUs

Computational bioinformatics and biomedical applications frequently contain heterogeneously sized units of work or tasks, for instance due to variability in the sizes of biological sequences and molecules. Variable-sized workloads lead to load imbalances in parallel implementations which detract from efficiency and performance. Many modern computing resources now have multiple graphics processing units(GPUs) per computer for acceleration. These multiple GPU resources need to be used efficiently through balancing of workloads across the GPUs. OpenMP is a portable directive-based parallel programming API used ubiquitously in bioscience applications to program CPUs; recently, the use of OpenMP directives for GPU acceleration has become possible. Here, motivated by experiences with imbalanced loads in GPU-accelerated bioinformatics applications, we address the load balancing problem using OpenMP task-to-GPU scheduling combined with OpenMP GPU offloading for multiply heterogeneous workloads – loads with both variable input sizes, and simultaneously, variable convergence rates for algorithms with a stochastic component – scheduled across multiple GPUs. We aim to develop strategies which are both easy to use and have lower overheads, and may be incorporated incrementally in existing programs which already make use of OpenMP for CPU-based threading in order to make use of multi-GPU computers. We test different combinations of input size variability and convergence rate variability, and characterize the effects of these different scenarios on the performance of scheduling strategies across multiple GPUs with OpenMP. We present several dynamic scheduling solutions for different parallel patterns, explore optimizations, and provide publicly available example computational kernels to make these strategies easy to use in programs. This work will enable application developers to efficiently and easily use multiple GPUs for imbalanced workloads found in bioinformatics and biomedical applications.

Thavappiragasam, Mathialakan↗

On the numerical accuracy in finite-volume methods to accurately capture turbulence in compressible flows

The goal of the present article is to understand the impact of numerical schemes for the reconstruction of data at cell faces in finite-volume methods, and to assess their interaction with the quadrature rule used to compute the average over the cell volume. Here, third-, fifth- and seventh-order WENO-Z schemes are investigated. On a problem with a smooth solution, the theoretical order of convergence rate for each method is retrieved, and changing the order of the reconstruction at cell faces does not impact the results, whereas for a shock-driven problem all the methods collapse to first-order. Here, study of the decay of compressible homogeneous isotropic turbulence reveals that using a high-order quadrature rule to compute the average over a finite-volume cell does not improve the spectral accuracy and that all methods present a second-order convergence rate. However the choice of the numerical method to reconstruct data at cell faces is found to be critical to correctly capture turbulent spectra. In the context of simulations with finite-volume methods of practical flows encountered in engineering applications, it becomes apparent that an efficient strategy is to perform the average integration with a low-order quadrature rule on a fine mesh resolution, whereas high-order schemes should be used to reconstruct data at cell faces.

97 MATHEMATICS AND COMPUTING↗

Distributed and communication-efficient solutions to linear equations with special sparse structure

In this paper we report two distributed and communication-efficient algorithms based on the multi-agent system are proposed to solve a system of linear equations with the Laplacian sparse system matrix. One algorithm is based on the gradient descent method in optimization. In this algorithm, the agents only share partial information instead of all of their collective state vectors to save significant communication. The other algorithm is obtained by approximating Newton’s method for a faster convergence rate. Although it requires twice as much communication as the first one, it is still communication-efficient given the low dimension of the information shared among agents. The convergence at a linear rate is proved for both algorithms, and a comprehensive comparison of their convergence rate, communication burden, and computation costs is also performed. The proposed algorithms can be applied to various systems to solve those problems that can be modeled as a system of linear equations with a Laplacian sparse system matrix. Simulation results with the electric power system illustrate their effectiveness.

42 ENGINEERING↗

Anderson acceleration stability in NDA-accelerated k-eigenvalue problems

Anderson acceleration (AA) has been used to improve the stability and convergence rate of multiphysics iterative methods for reactor analysis. Most applications studied assume a tightly converged solution for the different physics problems, and AA is usually applied to state variables like temperature, density, and heat generation rate. In this paper, we study the theoretical performance of AA in NDA-accelerated k-eigenvalue problems. The problems and algorithms studied are simplified from the coupled iteration scheme adopted by MPACT and many other high-fidelity whole-core reactor codes. Compared to previous analyses of AA for these iteration schemes, we study the case with a partially converged neutronics solution and possibly partially converged nonlinear diffusion acceleration (NDA)/coarse mesh finite difference (CMFD) solutions. We observe that the performance of the iteration scheme with AA is very sensitive to the initial guess and is affected by the partially converged CMFD solutions. When the NDA solution is fully converged, using AA cannot achieve the optimal convergence rate in large-sized problems. Conversely, if the NDA solution is partially converged, the iteration scheme with AA can diverge or converge extremely slowly. It is found that the loss of robustness for AA is due to the fact that it is applied to the iterative subspace of state variables rather than the fundamental unknowns of the governing equations. To improve the robustness, the scalar flux should also be considered in the implementation of AA. After considering the residuals of flux, we observe that the stability is regardless of the partial convergence of NDA solutions. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Detecting macroevolutionary genotype–phenotype associations using error-corrected rates of protein convergence

On macroevolutionary timescales, extensive mutations and phylogenetic uncertainty mask the signals of genotype–phenotype associations underlying convergent evolution. To overcome this problem, we extended the widely used framework of non-synonymous to synonymous substitution rate ratios and developed the novel metric ω C , which measures the error-corrected convergence rate of protein evolution. While ω C distinguishes natural selection from genetic noise and phylogenetic errors in simulation and real examples, its accuracy allows an exploratory genome-wide search of adaptive molecular convergence without phenotypic hypothesis or candidate genes. Using gene expression data, we explored over 20 million branch combinations in vertebrate genes and identified the joint convergence of expression patterns and protein sequences with amino acid substitutions in functionally important sites, providing hypotheses on undiscovered phenotypes. We further extended our method with a heuristic algorithm to detect highly repetitive convergence among computationally non-trivial higher-order phylogenetic combinations. Our approach allows bidirectional searches for genotype–phenotype associations, even in lineages that diverged for hundreds of millions of years.

59 BASIC BIOLOGICAL SCIENCES↗

Stability Quantification for Consensus-Based Power Flow Control between Transmission and Distribution Power Systems

With increasing integration of distributed energy resources (DERs), distribution systems (DS) with DERs are expected to provide proactive grid services. The result is that power references, or known as dispatch signals, required for DS can become faster changing than in legacy power system. Consensus-based integral controls have been proposed for the purpose of coordinating power generations of DERs in DS to match the power references. These existing controls assume the integral control signal to be sufficiently slow or constant, and ignore the potential dynamics of the integral controller when tracking more varying power references. Therefore, in this paper we present an improvement for such controls by deriving the stability condition utilizing generalized Nyquist criterion. The stability condition is quantified by a set of integral gains that guarantee stability in closed-loop system without assuming a constant integral signal. A rule-of-thumb criterion is also derived to instruct the design of the consensus topology that can provide faster convergence rate for the closed-loop system. Here, the stability and convergence improvements developed in this paper are demonstrated and verified through numerical examples.

42 ENGINEERING↗

On the Convergence of Overlapping Schwarz Decomposition for Nonlinear Optimal Control

Here, we study the convergence properties of an overlapping Schwarz decomposition algorithm for solving nonlinear optimal control problems (OCPs). The algorithm decomposes the time domain into a set of overlapping subdomains, and solves all subproblems defined over subdomains in parallel. The convergence is attained by updating primal-dual information at the boundaries of overlapping subdomains. We show that the algorithm exhibits local linear convergence, and that the convergence rate improves exponentially with the overlap size. We also establish global convergence results for a general quadratic programming, which enables the application of the Schwarz scheme inside second-order optimization algorithms (e.g., sequential quadratic programming). The theoretical foundation of our convergence analysis is a sensitivity result of nonlinear OCPs, which we call "exponential decay of sensitivity" (EDS). Intuitively, EDS states that the impact of perturbations at domain boundaries (i.e., initial and terminal time) on the solution decays exponentially as one moves into the domain. Here, we expand a previous analysis available in the literature by showing that EDS holds for both primal and dual solutions of nonlinear OCPs, under uniform second-order sufficient condition, controllability condition, and boundedness condition. We conduct experiments with a quadrotor motion planning problem and a partial differential equations (PDE) control problem to validate our theory, and show that the approach is significantly more efficient than alternating direction method of multipliers and as efficient as the centralized interior-point solver.

42 ENGINEERING↗

Lossy checkpoint compression in full waveform inversion: a case study with ZFPv0.5.5 and the overthrust model

This paper proposes a new method that combines checkpointing methods with error-controlled lossy compression for large-scale high-performance full-waveform inversion (FWI), an inverse problem commonly used in geophysical exploration. This combination can significantly reduce data movement, allowing a reduction in run time as well as peak memory. In the exascale computing era, frequent data transfer (e.g., memory bandwidth, PCIe bandwidth for GPUs, or network) is the performance bottleneck rather than the peak FLOPS of the processing unit. Like many other adjoint-based optimization problems, FWI is costly in terms of the number of floating-point operations, large memory footprint during backpropagation, and data transfer overheads. Past work for adjoint methods has developed checkpointing methods that reduce the peak memory requirements during backpropagation at the cost of additional floating-point computations. Combining this traditional checkpointing with error-controlled lossy compression, we explore the three-way tradeoff between memory, precision, and time to solution. We investigate how approximation errors introduced by lossy compression of the forward solution impact the objective function gradient and final inverted solution. Empirical results from these numerical experiments indicate that high lossy-compression rates (compression factors ranging up to 100) have a relatively minor impact on convergence rates and the quality of the final solution.

58 GEOSCIENCES↗

Reducing measurement costs by recycling the Hessian in adaptive variational quantum algorithms

Abstract Adaptive protocols enable the construction of more efficient state preparation circuits in variational quantum algorithms (VQAs) by utilizing data obtained from the quantum processor during the execution of the algorithm. This idea originated with Adaptive Derivative-Assembled Problem-Tailored variational quantum eigensolver (ADAPT-VQE), an algorithm that iteratively grows the state preparation circuit operator by operator, with each new operator accompanied by a new variational parameter, and where all parameters acquired thus far are optimized in each iteration. In ADAPT-VQE and other adaptive VQAs that followed it, it has been shown that initializing parameters to their optimal values from the previous iteration speeds up convergence and avoids shallow local traps in the parameter landscape. However, no other data from the optimization performed at one iteration is carried over to the next. In this work, we propose an improved quasi-Newton optimization protocol specifically tailored to adaptive VQAs. The distinctive feature in our proposal is that approximate second derivatives of the cost function are recycled across iterations in addition to optimal parameter values. We implement a quasi-Newton optimizer where an approximation to the inverse Hessian matrix is continuously built and grown across the iterations of an adaptive VQA. The resulting algorithm has the flavor of a continuous optimization where the dimension of the search space is augmented when the gradient norm falls below a given threshold. We show that this inter-optimization exchange of second-order information leads the approximate Hessian in the state of the optimizer to be consistently closer to the exact Hessian. As a result, our method achieves a superlinear convergence rate even in situations where the typical implementation of a quasi-Newton optimizer converges only linearly. Our protocol decreases the measurement costs in implementing adaptive VQAs on quantum hardware as well as the runtime of their classical simulation.

Ramôa, Mafalda (ORCID:0000000302187801)↗

Towards sharp error analysis of extended Lagrangian molecular dynamics

The extended Lagrangian molecular dynamics (XLMD) method provides a useful framework for reducing the computational cost of a class of molecular dynamics simulations with constrained latent variables. The XLMD method relaxes the constraints by introducing a fictitious mass ε for the latent variables and solving a set of singularly perturbed ordinary differential equations. While favorable numerical performance of XLMD has been demonstrated in several different contexts in the past decade, mathematical analysis of the method remains scarce. Here, we propose the first error analysis of the XLMD method in the context of a classical polarizable force field model. While the dynamics with respect to the atomic degrees of freedom are general and nonlinear, the key mathematical simplification of the polarizable force field model is that the constraints on the latent variables are given by a linear system of equations. We prove that when the initial value of the latent variables is compatible in a sense that we define, XLMD converges as the fictitious mass ε is made small with $\mathscr{O}$(ε) error for the atomic degrees of freedom and with $\mathscr{O}$($\sqrt{ε}$) error for the latent variables, when the dimension of the latent variable d' is 1. Furthermore, when the initial value of the latent variables is improved to be optimally compatible in a certain sense, we prove that the convergence rate can be improved to $\mathscr{O}$(ε) for the latent variables as well. Numerical results verify that both estimates are sharp not only for d'=1, but also for arbitrary d'. In the setting of general d', we do obtain convergence, but with the non-sharp rate of $\mathscr{O}$($\sqrt{ε}$) for both the atomic and latent variables.

74 ATOMIC AND MOLECULAR PHYSICS↗

A deterministic gradient-based approach to avoid saddle points

Abstract Loss functions with a large number of saddle points are one of the major obstacles for training modern machine learning (ML) models efficiently. First-order methods such as gradient descent (GD) are usually the methods of choice for training ML models. However, these methods converge to saddle points for certain choices of initial guesses. In this paper, we propose a modification of the recently proposed Laplacian smoothing gradient descent (LSGD) [Osher et al., arXiv:1806.06317 ], called modified LSGD (mLSGD), and demonstrate its potential to avoid saddle points without sacrificing the convergence rate. Our analysis is based on the attraction region, formed by all starting points for which the considered numerical scheme converges to a saddle point. We investigate the attraction region’s dimension both analytically and numerically. For a canonical class of quadratic functions, we show that the dimension of the attraction region for mLSGD is $\lfloor (n-1)/2\rfloor$ , and hence it is significantly smaller than that of GD whose dimension is $n-1$ .

Mathematics↗

Direct Discontinuous Galerkin methods for the reacting multi-component flow equations

The Direct Discontinuous Galerkin (DDG (Liu and Yan, 2008)) method and a counterpart with Interface Correction (DDGIC (Danis and Yan, 2022)) are extended to compute diffusion terms that arise when solving the compressible multi-component flow equations in thermochemical nonequilibrium. Thermodynamic properties, transport properties, chemical reaction rates, and energy exchange terms are computed using Mutation++ (Scoggins et al., 2020). The DG method is applied on unstructured grids, where the accuracy and convergence rates can be sensitive to the numerical method chosen for parabolic terms. A method for determining the homogeneity tensor of the flow equations required for DDGIC is shown. The convergence properties of the DDG methods are studied and compared to the Interior Penalty (IP) method. A number of numerical experiments are conducted to assess the accuracy and performance of the method. The numerical results and convergence studies indicate that DDG and DDGIC provide accurate solutions and perform well for general flows in thermochemical nonequilibrium.

Diffusion↗

Optimization-based algorithms for nonlinear mechanics and frictional contact

An optimization-based strategy for solving nonlinear mechanics problems is proposed. In contrast to typical nonlinear equation solver algorithms that aim to find zeros in the residual force function, we minimize an energy (or energy-like) function to encourage solutions which are locally stable equilibria. These smooth and potentially non-convex objective functions are minimized using a preconditioned conjugate-gradient trust-region algorithm. Contact is formulated as an inequality constrained minimization problem, and is solved with an augmented Lagrangian algorithm. Friction is included in the approach via a regularized quasi-potential energy, and other dissipative behavior is included through the use of variational constitutive updates. Finally, to accelerate convergence rates for the Lagrange multipliers, we propose a novel multiplier update algorithm utilizing the Fischer-Burmeister function, and demonstrate super-linear solver convergence for some applications.

42 ENGINEERING↗

Fair Concurrent Training of Multiple Models in Federated Learning

Federated learning (FL) enables collaborative learning across multiple clients. In most FL work, all clients train a single learning task. However, the recent proliferation of FL applications may increasingly require multiple FL tasks to be trained simultaneously, sharing clients’ computing resources, which we call Multiple-Model Federated Learning (MMFL). Current MMFL algorithms use naïve average-based client-task allocation schemes that often lead to unfair performance when FL tasks have heterogeneous difficulty levels, as the more difficult tasks may need more client participation to train effectively. Furthermore, in the MMFL setting, we face a further challenge that some clients may prefer training specific tasks to others, and may not even be willing to train other tasks, e.g., due to high computational costs, which may exacerbate unfairness in training outcomes across tasks. We address both challenges by firstly designing FedFairMMFL, a difficulty-aware algorithm that dynamically allocates clients to tasks in each training round, based on the tasks’ current performance levels. We provide guarantees on the resulting task fairness and FedFairMMFL’s convergence rate. We then propose novel auction designs that incentivizes clients to train multiple tasks, so as to fairly distribute clients’ training efforts across the tasks, and extend our convergence guarantees to this setting. Here, we finally evaluate our algorithm with multiple sets of learning tasks on real world datasets, showing that our algorithm improves fairness by improving the final model accuracy and convergence speed of the worst performing tasks, while maintaining the average accuracy across tasks.

Federated learning↗

Adaptive Mesh Refinement for Parallel in Time Methods

The project applied the multigrid-reduction-in-time (MGRIT) algorithm to an existing sub-cycled adaptive mesh refinement (AMR) code to investigate the performance of flows dominated by inertial physics. Previous work demonstrated good performance from MGRIT+AMR applied to flows dominated by diffusive physics. Consistent with previous experience, inertial physics negatively affected convergence rates and performance. Efforts to circumvent this issue by appealing to the physics of turbulence were investigated. It has been demonstrated that scales can be effectively transferred between multigrid levels for a turbulent flow resulting in a) partial convergence observed and b) nearly identical results to sequential time-stepping. Performance improvements have not yet been demonstrated - attempts at coarsening the grid on coarser MG levels compromises the solution quality and leads to divergence. This report summarize the accomplishments for the time-frame from 10/5/2020 to 12/31/2020 with an informal no-cost-extension to 05/20/2021

97 MATHEMATICS AND COMPUTING↗