Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Stochastic Gradient Descent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Convergence of Hyperbolic Neural Networks Under Riemannian Stochastic Gradient Descent

Abstract We prove, under mild conditions, the convergence of a Riemannian gradient descent method for a hyperbolic neural network regression model, both in batch gradient descent and stochastic gradient descent. We also discuss a Riemannian version of the Adam algorithm. We show numerical simulations of these algorithms on various benchmarks.

Whiting, Wes (ORCID:0000000247505060)↗

Multi-Pass Sequential Mini-Batch Stochastic Gradient Descent Algorithms for Noise Covariance Estimation in Adaptive Kalman Filtering

Estimation of unknown noise covariances in a Kalman filter is a problem of significant practical interest in a wide array of applications. Although this problem has a long history, reliable algorithms for their estimation were scant, and necessary and sufficient conditions for identifiability of the covariances were in dispute until recently. Necessary and sufficient conditions for covariance estimation and a batch estimation algorithm were presented in our previous study. This paper presents stochastic gradient descent algorithms for noise covariance estimation in adaptive Kalman filters that are an order of magnitude faster than the batch method for similar or better root mean square error. More significantly, these algorithms are applicable to non-stationary systems where the noise covariances can occasionally jump up or down by an unknown magnitude. The computational efficiency of the new algorithms stems from adaptive thresholds for convergence, recursive fading memory estimation of the sample cross-correlations of the innovations, and accelerated stochastic gradient descent algorithms. The comparative evaluation of the proposed methods on a number of test cases demonstrates their computational efficiency and accuracy.

Adaptive Kalman filtering↗

Implementation of Stochastic Gradient Descent in an Automated Glow Peak Identification Software for Multiple Thermoluminescent Dosimeter Types

A glow-curve analysis code was previously developed in C++ to analyze thermoluminescent dosimeter glow curves using automated peak detection while applying a first-order kinetics model. A newer version of this code was implemented to improve the automated peak detection and curve fitting models. The Stochastic Gradient Descent Algorithm was introduced to replace the prior approach of taking first and second-order derivatives for peak detection. Additionally, early stopping mechanisms were invoked to improve the previously used Levenberg-Marquardt Algorithm employed for curve fitting. The two software versions were compared through glow curve analysis of different thermoluminescent dosimeter materials and calculation of the corresponding figures of merit. Altogether improvements were shown, namely an increase in the number of peaks detected and a reduction of the mean figure of merit by approximately 46%.

137Cs↗

Stochastic gradient descent for optimization for nuclear systems

The use of gradient descent methods for optimizing k-eigenvalue nuclear systems has been shown to be useful in the past, but the use of k-eigenvalue gradients have proved computationally challenging due to their stochastic nature. ADAM is a gradient descent method that accounts for gradients with a stochastic nature. This analysis uses challenge problems constructed to verify if ADAM is a suitable tool to optimize k-eigenvalue nuclear systems. ADAM is able to successfully optimize nuclear systems using the gradients of k-eigenvalue problems despite their stochastic nature and uncertainty. Furthermore, it is clearly demonstrated that low-compute time, high-variance estimates of the gradient lead to better performance in the optimization challenge problems tested here.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Structure optimization with stochastic density functional theory

Linear-scaling techniques for Kohn–Sham density functional theory are essential to describe the ground state properties of extended systems. Still, these techniques often rely on the localization of the density matrix or accurate embedding approaches, limiting their applicability. In contrast, stochastic density functional theory (sDFT) achieves linear- and sub-linear scaling by statistically sampling the ground state density without relying on embedding or imposing localization. In return, ground state observables, such as the forces on the nuclei, fluctuate in sDFT, making optimizing the nuclear structure a highly non-trivial problem. In this work, we combine the most recent noise-reduction schemes for sDFT with stochastic optimization algorithms to perform structure optimization within sDFT. We compare the performance of the stochastic gradient descent approach and its variations (stochastic gradient descent with momentum) with stochastic optimization techniques that rely on the Hessian, such as the stochastic Broyden–Fletcher–Goldfarb–Shanno algorithm. In conclusion, we further provide a detailed assessment of the computational efficiency and its dependence on the optimization parameters of each method for determining the ground state structure of bulk silicon with varying supercell dimensions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A backward SDE method for uncertainty quantification in deep learning

Here, we develop a backward stochastic differential equation based probabilistic machine learning method, which formulates a class of stochastic neural networks as a stochastic optimal control problem. An efficient stochastic gradient descent algorithm is introduced with the gradient computed through a backward stochastic differential equation. Convergence analysis for stochastic gradient descent optimization and numerical experiments for applications of stochastic neural networks are carried out to validate our methodology in both theory and performance.

97 MATHEMATICS AND COMPUTING↗

Elastic distributed training with fast convergence and efficient resource utilization

Distributed learning is now routinely conducted on cloud as well as dedicated clusters. Training with elastic resources brings new challenges and design choices. Prior studies focus on runtime performance and assume a static algorithmic behavior. In this work, by analyzing the impact of of resource scaling on convergence, we introduce schedules for synchronous stochastic gradient descent that proactively adapt the number of learners to reduce training time and improve convergence. Our approach no longer assumes a constant number of processors throughout training. In our experiment, distributed stochastic gradient descent with dynamic schedules and reduction momentum achieves better convergence and significant speedups over prior static ones. Numerous distributed training jobs running on cloud may benefit from our approach.

Cong, Guojing↗

Towards provably efficient quantum algorithms for large-scale machine-learning models

Large machine learning models are revolutionary technologies of artificial intelligence whose bottlenecks include huge computational expenses, power, and time used both in the pre-training and fine-tuning process. In this work, we show that fault-tolerant quantum computing could possibly provide provably efficient resolutions for generic (stochastic) gradient descent algorithms, scaling as $\mathcal{O}$(T 2 x polylog($n$)), where n is the size of the models and T is the number of iterations in the training, as long as the models are both sufficiently dissipative and sparse, with small learning rates. Based on earlier efficient quantum algorithms for dissipative differential equations, we find and prove that similar algorithms work for (stochastic) gradient descent, the primary algorithm for machine learning. In practice, we benchmark instances of large machine learning models from 7 million to 103 million parameters. We find that, in the context of sparse training, a quantum enhancement is possible at the early stage of learning after model pruning, motivating a sparse parameter download and re-upload scheme. Our work shows solidly that fault-tolerant quantum algorithms could potentially contribute to most state-of-the-art, large-scale machine-learning problems.

97 MATHEMATICS AND COMPUTING↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Stochastic Optimization and Uncertainty Quantification of Natrium-based Nuclear-Renewable Energy Systems for Flexible Power Applications in Deregulated Markets

Rapid integration of variable renewable energy sources (VRES) has made modeling and stochastic optimization of hybrid energy systems crucial for studying their long-term performance and viability. However, most studies have focused on just historical data, which may be unreliable for capturing short-term fluctuations, rare events, and long-term patterns of energy demand, price, and the variability of renewable energy sources. For this study, optimal synthetic time series models were developed using Wasserstein distance. The models were validated by comparing the key statistical measures against those of the historical data. They were then used to optimize the integrated Natrium-style advanced energy systems and their long-term (30 years) economics. The stochastic model performs bi-level optimization to find the optimal sizes for the balance of plant and thermal energy storage, while also optimizing energy dispatch to achieve the maximum net present value. In studies of two deregulated markets (California ISO and the Electric Reliability Council of Texas), the integrated Natrium-style system performed better in CAISO than in ERCOT, given higher and more consistent electricity prices during peak-demand periods. The potentially enlarged cost associated with the variable operation and maintenance of the TES system also plays a significant role in driving the system sizing, thus its impacts on the system are investigated in detail through comparison against a baseline case. The study also finds that the bi-level optimization results based on stochastic gradient descent closely match the grid search results. The uncertainty quantification of the stochastic signals provides further NPV-related insights and probability distributions for the case studies. The normal standard error of the mean of NPV for the case with and without TES VOM for CAISO were found to be 7.73M (plus-minus sign) 1.09M USD and 104.99M (plus-minus sign) 1.25M USD, respectively based on a 95% confidence. Given the relatively small NPV variance based on 150 samples, the analysis affords the most robust possible prediction of the techno-economic performance of the integrated Natrium-style energy systems.

25 ENERGY STORAGE↗

Learning generative neural networks with physics knowledge

Deep generative neural networks have enabled modeling complex distributions, but incorporating physics knowledge into the neural networks is still challenging and is at the core of current physics-based machine learning research. To this end, we propose a physics generative neural network (PhysGNN), a new class of generative neural networks for learning unknown distributions in a physical system described by partial differential equations (PDE). PhysGNN couples PDE systems with generative neural networks. It is a fully differentiable model that allows back-propagation of gradients through both numerical PDE solvers and generative neural networks, and is trained by minimizing the discrete Wasserstein distance between generated and observed probability distributions of the PDE outputs using the stochastic gradient descent method. Moreover, PhysGNN does not require adversarial training like standard generative neural networks, which offers better stability than adversarial training. We show that PhysGNN can learn complex distributions in stochastic inverse problems, where conventional methods such as maximum likelihood estimation and momentum matching methods may be inapplicable when little knowledge is known about the form of unknown distributions or the physical model is too complex. Furthermore, our method allows physics-based generative neural network training for learning complex distributions in the context of differential equations.

97 MATHEMATICS AND COMPUTING↗

Prediction of gas hydrate saturation using machine learning and optimal set of well-logs

We report resistivity and acoustic logs are widely used to estimate gas hydrate saturation in various sedimentary systems using one of the two popular methods ((1) acoustic velocity and (2) electrical resistivity), but the limitations of these two methods are often overlooked, which include (i) well-specific calibration of empirical exponents in the electrical resistivity method, (ii) assumption of known pore morphology for gas hydrates in the acoustic velocity method, and (iii) presence of unknown mineralogy and bulk modulus terms in the acoustic velocity method. NMR-density porosity-derived gas hydrate saturation based on the analysis of the transverse magnetization relaxation time (T2) is considered the most precise method, but acquisition of NMR-based logs is limited at relatively recent drilled sites; additionally, its use in conventional oil and gas reservoirs is not that common due to higher cost and operational deployment limitations associated with acquiring NMR well-logs. This study proposes a new method that predicts gas hydrate saturation (S h ) for any well using porosity, bulk density, and compressional wave (P wave) velocity well-logs with neural network (or stochastic gradient descent regression) without any well-specific calibration and/or other aforementioned shortcomings of the existing methods. The method is developed by examining the underlying dependency between S h and different combinations of well-logs, chosen from 6 routine logs, with 12 different machine learning (ML) algorithms. The accuracy of the proposed method in predicting S h is ~ 84%, which is better than the accuracy of seismic and electrical resistivity methods (≤ 75%) per the results reported by three different studies. The robustness of the method in the specific case of permafrost-associated gas hydrates is demonstrated with well-log data from two wells drilled on the Alaska North Slope.

58 GEOSCIENCES↗

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗

Mode connectivity in the loss landscape of parameterized quantum circuits

Variational training of parameterized quantum circuits (PQCs) underpins many workflows employed on near-term noisy intermediate scale quantum (NISQ) devices. It is a hybrid quantum-classical approach that minimizes an associated cost function in order to train a parameterized ansatz. In this work we adapt the qualitative loss landscape characterization for neural networks introduced in Goodfellow et al. (2014); Li et al. (2017) and tests for connectivity used in Draxler et al. (2018) to study the loss landscape features in PQC training. We present results for PQCs trained on a simple regression task, using the bilayer circuit ansatz, which consists of alternating layers of parameterized rotation gates and entangling gates. Multiple circuits are trained with 3 different batch gradient optimizers: stochastic gradient descent, the quantum natural gradient, and Adam. We identify large features in the landscape that can lead to faster convergence in training workflows.

97 MATHEMATICS AND COMPUTING↗

Concurrent multi-peak Bragg coherent x-ray diffraction imaging of 3D nanocrystal lattice displacement via global optimization

Abstract In this paper we demonstrated a method to reconstruct vector-valued lattice distortion fields within nanoscale crystals by optimization of a forward model of multi-reflection Bragg coherent diffraction imaging (MR-BCDI) data. The method flexibly accounts for geometric factors that arise when making BCDI measurements, is amenable to efficient inversion with modern optimization toolkits, and allows for globally constraining a single image reconstruction to multiple Bragg peak measurements. This is enabled by a forward model that emulates the multiple Bragg peaks of a MR-BCDI experiment from a single estimate of the 3D crystal sample. We present this forward model, we implement it within the stochastic gradient descent optimization framework, and we demonstrate it with simulated and experimental data of nanocrystals with inhomogeneous internal lattice displacement. We find that utilizing a global optimization approach to MR-BCDI affords a reliable path to convergence of data which is otherwise challenging to reconstruct.

36 MATERIALS SCIENCE↗