Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Laplacian Smoothing Stochastic Gradient Markov Chain Monte Carlo

As an important Markov chain Monte Carlo (MCMC) method, the stochastic gradient Langevin dynamics (SGLD) algorithm has achieved great success in Bayesian learning and posterior sampling. Furthermore, SGLD typically suffers from a slow convergence rate due to its large variance caused by the stochastic gradient. In order to alleviate these drawbacks, we leverage the recently developed Laplacian smoothing technique and propose a Laplacian smoothing stochastic gradient Langevin dynamics (LS-SGLD) algorithm. We prove that for sampling from both log-concave and non-log-concave densities, LS-SGLD achieves strictly smaller discretization error in 2-Wasserstein distance, although its mixing rate can be slightly slower. Experiments on both synthetic and real datasets verify our theoretical results and demonstrate the superior performance of LS-SGLD on different machine learning tasks including posterior sampling, Bayesian logistic regression, and training Bayesian convolutional neural networks.

97 MATHEMATICS AND COMPUTING↗

A Stochastic Gradient Descent Approach for Stochastic Optimal Control

In this work, we introduce a stochastic gradient descent approach to solve the stochastic optimal control problem through stochastic maximum principle. The motivation that drives our method is the gradient of the cost functional in the stochastic optimal control problem is under expectation, and numerical calculation of such an expectation requires fully computation of a system of forward backward stochastic differential equations, which is computationally expensive. By evaluating the expectation with single-sample representation as suggested by the stochastic gradient descent type optimisation, we could save computational efforts in solving FBSDEs and only focus on the optimisation task which aims to determine the optimal control process.

97 MATHEMATICS AND COMPUTING↗

Efficient Ensemble-Based Stochastic Gradient Methods for Optimization Under Geological Uncertainty

Ensemble-based stochastic gradient methods, such as the ensemble optimization (EnOpt) method, the simplex gradient (SG) method, and the stochastic simplex approximate gradient (StoSAG) method, approximate the gradient of an objective function using an ensemble of perturbed control vectors. These methods are increasingly used in solving reservoir optimization problems because they are not only easy to parallelize and couple with any simulator but also computationally more efficient than the conventional finite-difference method for gradient calculations. In this work, we show that EnOpt may fail to achieve sufficient improvement of the objective function when the differences between the objective function values of perturbed control variables and their ensemble mean are large. On the basis of the comparison of EnOpt and SG, we propose a hybrid gradient of EnOpt and SG to save on the computational cost of SG. We also suggest practical ways to reduce the computational cost of EnOpt and StoSAG by approximating the objective function values of unperturbed control variables using the values of perturbed ones. We first demonstrate the performance of our improved ensemble schemes using a benchmark problem. Results show that the proposed gradients saved about 30–50% of the computational cost of the same optimization by using EnOpt, SG, and StoSAG. As a real application, we consider pressure management in carbon storage reservoirs, for which brine extraction wells need to be optimally placed to reduce reservoir pressure buildup while maximizing the net present value. Results show that our improved schemes reduce the computational cost significantly.

58 GEOSCIENCES↗

Convergence of Hyperbolic Neural Networks Under Riemannian Stochastic Gradient Descent

Abstract We prove, under mild conditions, the convergence of a Riemannian gradient descent method for a hyperbolic neural network regression model, both in batch gradient descent and stochastic gradient descent. We also discuss a Riemannian version of the Adam algorithm. We show numerical simulations of these algorithms on various benchmarks.

Whiting, Wes (ORCID:0000000247505060)↗

An adaptive Hessian approximated stochastic gradient MCMC method

Bayesian approaches have been successfully integrated into training deep neural networks. One popular family is stochastic gradient Markov chain Monte Carlo methods (SG-MCMC), which have gained increasing interest due to their ability to handle large datasets and the potential to avoid overfitting. Although standard SG-MCMC methods have shown great performance in a variety of problems, they may be inefficient when the random variables in the target posterior densities have scale differences or are highly correlated. Here, we present an adaptive Hessian approximated stochastic gradient MCMC method to incorporate local geometric information while sampling from the posterior. The idea is to apply stochastic approximation (SA) to sequentially update a preconditioning matrix at each iteration. The preconditioner possesses second-order information and can guide the random walk of a sampler efficiently. Instead of computing and saving the full Hessian of the log posterior, we use limited memory of the samples and their stochastic gradients to approximate the inverse Hessian-vector multiplication in the updating formula. Moreover, by smoothly optimizing the preconditioning matrix via SA, our proposed algorithm can asymptotically converge to the target distribution with a controllable bias under mild conditions. To reduce the training and testing computational burden, we adopt a magnitude-based weight pruning method to enforce the sparsity of the network. Our method is user-friendly and demonstrates better learning results compared to standard SG-MCMC updating rules. The approximation of inverse Hessian alleviates storage and computational complexities for large dimensional models. Numerical experiments are performed on several problems, including sampling from 2D correlated distribution, synthetic regression problems, and learning the numerical solutions of heterogeneous elliptic PDE. The numerical results demonstrate great improvement in both the convergence rate and accuracy.

97 MATHEMATICS AND COMPUTING↗

Stochastic Gradients for Large-Scale Tensor Decomposition

Tensor decomposition is a well-known tool for multiway data analysis. This work proposes using stochastic gradients for efficient generalized canonical polyadic (GCP) tensor decomposition of large-scale tensors. GCP tensor decomposition is a recently proposed version of tensor decomposition that allows for a variety of loss functions such as Bernoulli loss for binary data or Huber loss for robust estimation. Here, the stochastic gradient is formed from randomly sampled elements of the tensor and is efficient because it can be computed using the sparse matricized-tensor times Khatri--Rao product tensor kernel. For dense tensors, we simply use uniform sampling. For sparse tensors, we propose two types of stratified sampling that give precedence to sampling nonzeros. Numerical results demonstrate the advantages of the proposed approach and its scalability to large-scale problems.

97 MATHEMATICS AND COMPUTING↗

Bayesian sparse learning with preconditioned stochastic gradient MCMC and its applications

Deep neural networks have been successfully employed in an extensive variety of research areas, including solving partial differential equations. Despite its significant success, there are some challenges in effectively training DNN, such as avoiding overfitting in over-parameterized DNNs and accelerating the optimization in DNNs with pathological curvature. Here, we propose a Bayesian type sparse deep learning algorithm. The algorithm utilizes a set of spike-and-slab priors for the parameters in the deep neural network. The hierarchical Bayesian mixture will be trained using an adaptive empirical method. That is, one will alternatively sample from the posterior using preconditioned stochastic gradient Langevin Dynamics (PSGLD), and optimize the latent variables via stochastic approximation. The sparsity of the network is achieved while optimizing the hyperparameters with adaptive searching and penalizing. A popular SG-MCMC approach is Stochastic gradient Langevin dynamics (SGLD). However, considering the complex geometry in the model parameter space in nonconvex learning, updating parameters using a universal step size in each component as in SGLD may cause slow mixing. To address this issue, we apply a computationally manageable preconditioner in the updating rule, which provides a step-size parameter to adapt to local geometric properties. Moreover, by smoothly optimizing the hyperparameter in the preconditioning matrix, our proposed algorithm ensures a decreasing bias, which is introduced by ignoring the correction term in the preconditioned SGLD. According to the existing theoretical framework, we show that the proposed algorithm can asymptotically converge to the correct distribution with a controllable bias under mild conditions. Numerical tests are performed on both synthetic regression problems and learning solutions of elliptic PDE, which demonstrate the accuracy and efficiency of the present work.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Multi-Pass Sequential Mini-Batch Stochastic Gradient Descent Algorithms for Noise Covariance Estimation in Adaptive Kalman Filtering

Estimation of unknown noise covariances in a Kalman filter is a problem of significant practical interest in a wide array of applications. Although this problem has a long history, reliable algorithms for their estimation were scant, and necessary and sufficient conditions for identifiability of the covariances were in dispute until recently. Necessary and sufficient conditions for covariance estimation and a batch estimation algorithm were presented in our previous study. This paper presents stochastic gradient descent algorithms for noise covariance estimation in adaptive Kalman filters that are an order of magnitude faster than the batch method for similar or better root mean square error. More significantly, these algorithms are applicable to non-stationary systems where the noise covariances can occasionally jump up or down by an unknown magnitude. The computational efficiency of the new algorithms stems from adaptive thresholds for convergence, recursive fading memory estimation of the sample cross-correlations of the innovations, and accelerated stochastic gradient descent algorithms. The comparative evaluation of the proposed methods on a number of test cases demonstrates their computational efficiency and accuracy.

Adaptive Kalman filtering↗

Online Distribution System State Estimation via Stochastic Gradient Algorithm

Distribution network operation is becoming more challenging because of the growing integration of intermittent and volatile distributed energy resources (DERs). This motivates the development of new distribution system state estimation (DSSE) paradigms that can operate at fast timescale based on real-time data stream of asynchronous measurements enabled by modern information and communications technology. To solve the real-time DSSE with asynchronous measurements effectively and accurately, this paper formulates a weighted least squares DSSE problem and proposes an online stochastic gradient algorithm to solve it. The performance of the proposed scheme is analytically guaranteed and is numerically corroborated with realistic data on IEEE 123-bus feeder.

distribution system state estimation↗

Stochastic gradient descent for optimization for nuclear systems

The use of gradient descent methods for optimizing k-eigenvalue nuclear systems has been shown to be useful in the past, but the use of k-eigenvalue gradients have proved computationally challenging due to their stochastic nature. ADAM is a gradient descent method that accounts for gradients with a stochastic nature. This analysis uses challenge problems constructed to verify if ADAM is a suitable tool to optimize k-eigenvalue nuclear systems. ADAM is able to successfully optimize nuclear systems using the gradients of k-eigenvalue problems despite their stochastic nature and uncertainty. Furthermore, it is clearly demonstrated that low-compute time, high-variance estimates of the gradient lead to better performance in the optimization challenge problems tested here.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Implementation of Stochastic Gradient Descent in an Automated Glow Peak Identification Software for Multiple Thermoluminescent Dosimeter Types

A glow-curve analysis code was previously developed in C++ to analyze thermoluminescent dosimeter glow curves using automated peak detection while applying a first-order kinetics model. A newer version of this code was implemented to improve the automated peak detection and curve fitting models. The Stochastic Gradient Descent Algorithm was introduced to replace the prior approach of taking first and second-order derivatives for peak detection. Additionally, early stopping mechanisms were invoked to improve the previously used Levenberg-Marquardt Algorithm employed for curve fitting. The two software versions were compared through glow curve analysis of different thermoluminescent dosimeter materials and calculation of the corresponding figures of merit. Altogether improvements were shown, namely an increase in the number of peaks detected and a reduction of the mean figure of merit by approximately 46%.

137Cs↗

A backward SDE method for uncertainty quantification in deep learning

Here, we develop a backward stochastic differential equation based probabilistic machine learning method, which formulates a class of stochastic neural networks as a stochastic optimal control problem. An efficient stochastic gradient descent algorithm is introduced with the gradient computed through a backward stochastic differential equation. Convergence analysis for stochastic gradient descent optimization and numerical experiments for applications of stochastic neural networks are carried out to validate our methodology in both theory and performance.

97 MATHEMATICS AND COMPUTING↗

Stochastic Gradient-Based Distributed Bayesian Estimation in Cooperative Sensor Networks

Distributed Bayesian inference provides a full quantification of uncertainty offering numerous advantages over point estimates that autonomous sensor networks are able to exploit. However, fully-decentralized Bayesian inference often requires large communication overheads and low network latency, resources that are not typically available in practical applications. In this paper, we propose a decentralized Bayesian inference approach based on stochastic gradient Langevin dynamics, which produces full posterior distributions at each of the nodes with significantly lower communication overhead. We provide analytical results on convergence of the proposed distributed algorithm to the centralized posterior, under typical network constraints. Finally, we also provide extensive simulation results to demonstrate the validity of the proposed approach.

42 ENGINEERING↗

Mode connectivity in the loss landscape of parameterized quantum circuits

Variational training of parameterized quantum circuits (PQCs) underpins many workflows employed on near-term noisy intermediate scale quantum (NISQ) devices. It is a hybrid quantum-classical approach that minimizes an associated cost function in order to train a parameterized ansatz. In this work we adapt the qualitative loss landscape characterization for neural networks introduced in Goodfellow et al. (2014); Li et al. (2017) and tests for connectivity used in Draxler et al. (2018) to study the loss landscape features in PQC training. We present results for PQCs trained on a simple regression task, using the bilayer circuit ansatz, which consists of alternating layers of parameterized rotation gates and entangling gates. Multiple circuits are trained with 3 different batch gradient optimizers: stochastic gradient descent, the quantum natural gradient, and Adam. We identify large features in the landscape that can lead to faster convergence in training workflows.

97 MATHEMATICS AND COMPUTING↗

Structure optimization with stochastic density functional theory

Linear-scaling techniques for Kohn–Sham density functional theory are essential to describe the ground state properties of extended systems. Still, these techniques often rely on the localization of the density matrix or accurate embedding approaches, limiting their applicability. In contrast, stochastic density functional theory (sDFT) achieves linear- and sub-linear scaling by statistically sampling the ground state density without relying on embedding or imposing localization. In return, ground state observables, such as the forces on the nuclei, fluctuate in sDFT, making optimizing the nuclear structure a highly non-trivial problem. In this work, we combine the most recent noise-reduction schemes for sDFT with stochastic optimization algorithms to perform structure optimization within sDFT. We compare the performance of the stochastic gradient descent approach and its variations (stochastic gradient descent with momentum) with stochastic optimization techniques that rely on the Hessian, such as the stochastic Broyden–Fletcher–Goldfarb–Shanno algorithm. In conclusion, we further provide a detailed assessment of the computational efficiency and its dependence on the optimization parameters of each method for determining the ground state structure of bulk silicon with varying supercell dimensions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗