Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradient descent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Convergence of Hyperbolic Neural Networks Under Riemannian Stochastic Gradient Descent

Abstract We prove, under mild conditions, the convergence of a Riemannian gradient descent method for a hyperbolic neural network regression model, both in batch gradient descent and stochastic gradient descent. We also discuss a Riemannian version of the Adam algorithm. We show numerical simulations of these algorithms on various benchmarks.

Whiting, Wes (ORCID:0000000247505060)↗

Implementation of Stochastic Gradient Descent in an Automated Glow Peak Identification Software for Multiple Thermoluminescent Dosimeter Types

A glow-curve analysis code was previously developed in C++ to analyze thermoluminescent dosimeter glow curves using automated peak detection while applying a first-order kinetics model. A newer version of this code was implemented to improve the automated peak detection and curve fitting models. The Stochastic Gradient Descent Algorithm was introduced to replace the prior approach of taking first and second-order derivatives for peak detection. Additionally, early stopping mechanisms were invoked to improve the previously used Levenberg-Marquardt Algorithm employed for curve fitting. The two software versions were compared through glow curve analysis of different thermoluminescent dosimeter materials and calculation of the corresponding figures of merit. Altogether improvements were shown, namely an increase in the number of peaks detected and a reduction of the mean figure of merit by approximately 46%.

137Cs↗

Stochastic gradient descent for optimization for nuclear systems

The use of gradient descent methods for optimizing k-eigenvalue nuclear systems has been shown to be useful in the past, but the use of k-eigenvalue gradients have proved computationally challenging due to their stochastic nature. ADAM is a gradient descent method that accounts for gradients with a stochastic nature. This analysis uses challenge problems constructed to verify if ADAM is a suitable tool to optimize k-eigenvalue nuclear systems. ADAM is able to successfully optimize nuclear systems using the gradients of k-eigenvalue problems despite their stochastic nature and uncertainty. Furthermore, it is clearly demonstrated that low-compute time, high-variance estimates of the gradient lead to better performance in the optimization challenge problems tested here.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

Structure optimization with stochastic density functional theory

Linear-scaling techniques for Kohn–Sham density functional theory are essential to describe the ground state properties of extended systems. Still, these techniques often rely on the localization of the density matrix or accurate embedding approaches, limiting their applicability. In contrast, stochastic density functional theory (sDFT) achieves linear- and sub-linear scaling by statistically sampling the ground state density without relying on embedding or imposing localization. In return, ground state observables, such as the forces on the nuclei, fluctuate in sDFT, making optimizing the nuclear structure a highly non-trivial problem. In this work, we combine the most recent noise-reduction schemes for sDFT with stochastic optimization algorithms to perform structure optimization within sDFT. We compare the performance of the stochastic gradient descent approach and its variations (stochastic gradient descent with momentum) with stochastic optimization techniques that rely on the Hessian, such as the stochastic Broyden–Fletcher–Goldfarb–Shanno algorithm. In conclusion, we further provide a detailed assessment of the computational efficiency and its dependence on the optimization parameters of each method for determining the ground state structure of bulk silicon with varying supercell dimensions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A backward SDE method for uncertainty quantification in deep learning

Here, we develop a backward stochastic differential equation based probabilistic machine learning method, which formulates a class of stochastic neural networks as a stochastic optimal control problem. An efficient stochastic gradient descent algorithm is introduced with the gradient computed through a backward stochastic differential equation. Convergence analysis for stochastic gradient descent optimization and numerical experiments for applications of stochastic neural networks are carried out to validate our methodology in both theory and performance.

97 MATHEMATICS AND COMPUTING↗

Elastic distributed training with fast convergence and efficient resource utilization

Distributed learning is now routinely conducted on cloud as well as dedicated clusters. Training with elastic resources brings new challenges and design choices. Prior studies focus on runtime performance and assume a static algorithmic behavior. In this work, by analyzing the impact of of resource scaling on convergence, we introduce schedules for synchronous stochastic gradient descent that proactively adapt the number of learners to reduce training time and improve convergence. Our approach no longer assumes a constant number of processors throughout training. In our experiment, distributed stochastic gradient descent with dynamic schedules and reduction momentum achieves better convergence and significant speedups over prior static ones. Numerous distributed training jobs running on cloud may benefit from our approach.

Cong, Guojing↗

Towards provably efficient quantum algorithms for large-scale machine-learning models

Large machine learning models are revolutionary technologies of artificial intelligence whose bottlenecks include huge computational expenses, power, and time used both in the pre-training and fine-tuning process. In this work, we show that fault-tolerant quantum computing could possibly provide provably efficient resolutions for generic (stochastic) gradient descent algorithms, scaling as $\mathcal{O}$(T 2 x polylog($n$)), where n is the size of the models and T is the number of iterations in the training, as long as the models are both sufficiently dissipative and sparse, with small learning rates. Based on earlier efficient quantum algorithms for dissipative differential equations, we find and prove that similar algorithms work for (stochastic) gradient descent, the primary algorithm for machine learning. In practice, we benchmark instances of large machine learning models from 7 million to 103 million parameters. We find that, in the context of sparse training, a quantum enhancement is possible at the early stage of learning after model pruning, motivating a sparse parameter download and re-upload scheme. Our work shows solidly that fault-tolerant quantum algorithms could potentially contribute to most state-of-the-art, large-scale machine-learning problems.

97 MATHEMATICS AND COMPUTING↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Stochastic Optimization and Uncertainty Quantification of Natrium-based Nuclear-Renewable Energy Systems for Flexible Power Applications in Deregulated Markets

Rapid integration of variable renewable energy sources (VRES) has made modeling and stochastic optimization of hybrid energy systems crucial for studying their long-term performance and viability. However, most studies have focused on just historical data, which may be unreliable for capturing short-term fluctuations, rare events, and long-term patterns of energy demand, price, and the variability of renewable energy sources. For this study, optimal synthetic time series models were developed using Wasserstein distance. The models were validated by comparing the key statistical measures against those of the historical data. They were then used to optimize the integrated Natrium-style advanced energy systems and their long-term (30 years) economics. The stochastic model performs bi-level optimization to find the optimal sizes for the balance of plant and thermal energy storage, while also optimizing energy dispatch to achieve the maximum net present value. In studies of two deregulated markets (California ISO and the Electric Reliability Council of Texas), the integrated Natrium-style system performed better in CAISO than in ERCOT, given higher and more consistent electricity prices during peak-demand periods. The potentially enlarged cost associated with the variable operation and maintenance of the TES system also plays a significant role in driving the system sizing, thus its impacts on the system are investigated in detail through comparison against a baseline case. The study also finds that the bi-level optimization results based on stochastic gradient descent closely match the grid search results. The uncertainty quantification of the stochastic signals provides further NPV-related insights and probability distributions for the case studies. The normal standard error of the mean of NPV for the case with and without TES VOM for CAISO were found to be 7.73M (plus-minus sign) 1.09M USD and 104.99M (plus-minus sign) 1.25M USD, respectively based on a 95% confidence. Given the relatively small NPV variance based on 150 samples, the analysis affords the most robust possible prediction of the techno-economic performance of the integrated Natrium-style energy systems.

25 ENERGY STORAGE↗

Learning generative neural networks with physics knowledge

Deep generative neural networks have enabled modeling complex distributions, but incorporating physics knowledge into the neural networks is still challenging and is at the core of current physics-based machine learning research. To this end, we propose a physics generative neural network (PhysGNN), a new class of generative neural networks for learning unknown distributions in a physical system described by partial differential equations (PDE). PhysGNN couples PDE systems with generative neural networks. It is a fully differentiable model that allows back-propagation of gradients through both numerical PDE solvers and generative neural networks, and is trained by minimizing the discrete Wasserstein distance between generated and observed probability distributions of the PDE outputs using the stochastic gradient descent method. Moreover, PhysGNN does not require adversarial training like standard generative neural networks, which offers better stability than adversarial training. We show that PhysGNN can learn complex distributions in stochastic inverse problems, where conventional methods such as maximum likelihood estimation and momentum matching methods may be inapplicable when little knowledge is known about the form of unknown distributions or the physical model is too complex. Furthermore, our method allows physics-based generative neural network training for learning complex distributions in the context of differential equations.

97 MATHEMATICS AND COMPUTING↗

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗

Mode connectivity in the loss landscape of parameterized quantum circuits

Variational training of parameterized quantum circuits (PQCs) underpins many workflows employed on near-term noisy intermediate scale quantum (NISQ) devices. It is a hybrid quantum-classical approach that minimizes an associated cost function in order to train a parameterized ansatz. In this work we adapt the qualitative loss landscape characterization for neural networks introduced in Goodfellow et al. (2014); Li et al. (2017) and tests for connectivity used in Draxler et al. (2018) to study the loss landscape features in PQC training. We present results for PQCs trained on a simple regression task, using the bilayer circuit ansatz, which consists of alternating layers of parameterized rotation gates and entangling gates. Multiple circuits are trained with 3 different batch gradient optimizers: stochastic gradient descent, the quantum natural gradient, and Adam. We identify large features in the landscape that can lead to faster convergence in training workflows.

97 MATHEMATICS AND COMPUTING↗

Concurrent multi-peak Bragg coherent x-ray diffraction imaging of 3D nanocrystal lattice displacement via global optimization

Abstract In this paper we demonstrated a method to reconstruct vector-valued lattice distortion fields within nanoscale crystals by optimization of a forward model of multi-reflection Bragg coherent diffraction imaging (MR-BCDI) data. The method flexibly accounts for geometric factors that arise when making BCDI measurements, is amenable to efficient inversion with modern optimization toolkits, and allows for globally constraining a single image reconstruction to multiple Bragg peak measurements. This is enabled by a forward model that emulates the multiple Bragg peaks of a MR-BCDI experiment from a single estimate of the 3D crystal sample. We present this forward model, we implement it within the stochastic gradient descent optimization framework, and we demonstrate it with simulated and experimental data of nanocrystals with inhomogeneous internal lattice displacement. We find that utilizing a global optimization approach to MR-BCDI affords a reliable path to convergence of data which is otherwise challenging to reconstruct.

36 MATERIALS SCIENCE↗

Criticality analysis of nuclear binding energy neural networks

Machine learning methods, in particular deep learning methods such as artificial neural networks (ANNs) with many layers, have become widespread and useful tools in nuclear physics. However, these ANNs are typically treated as ‘black boxes’, with their architecture (width, depth, and weight/bias initialization) and the training algorithm and parameters chosen empirically by optimizing learning based on limited exploration. We test a non-empirical approach to understanding and optimizing nuclear physics ANNs by adapting a criticality analysis based on renormalization group flows in terms of the hyperparameters for weight/bias initialization, training rates, and the ratio of depth to width. This treatment utilizes the statistical properties of neural network initialization to find a generating functional for network outputs at any layer, allowing for a path integral formulation of the ANN outputs as a Euclidean statistical field theory. We use a prototypical example to test the applicability of this approach: a simple ANN for nuclear binding energies. We find that with training using a stochastic gradient descent optimizer, the predicted criticality behavior is realized, and optimal performance is found with critical tuning. However, the use of an adaptive learning algorithm leads to somewhat superior results without concern for tuning and thus obscures the analysis. Nevertheless, the criticality analysis offers a way to look within the black box of ANNs, which is a first step towards potential improvements in network performance beyond using adaptive optimizers.

artificial neural network↗

Embedded training of neural-network subgrid-scale turbulence models

We report that the weights of a deep neural-network model are optimized in conjunction with the governing flow equations to provide a model for subgrid-scale stresses in a temporally developing plane turbulent jet at Reynolds number Re 0 = 6000 . The objective function for training is first based on the instantaneous filtered velocity fields from a corresponding direct numerical simulation, and the training is by a stochastic gradient descent method, which uses the adjoint Navier-Stokes equations to provide the end-to-end sensitivities of the model weights to the velocity fields. In-sample and out-of-sample testing on multiple dual-jet configurations show that its required mesh density in each coordinate direction for prediction of mean flow, Reynolds stresses, and spectra is half that needed by the dynamic Smagorinsky model for comparable accuracy. The same neural-network model trained directly to match filtered subgrid-scale stresses, without the constraint of being embedded within the flow equations during the training, fails to provide a qualitatively correct prediction. The coupled formulation is generalized to train based only on mean-flow and Reynolds stresses, which are more readily available in experiments. The mean-flow training provides a robust model, which is important, though a somewhat less accurate prediction for the same coarse meshes, as might be anticipated due to the reduced information available for training in this case. The anticipated advantage of the formulation is that the inclusion of resolved physics in the training increases its capacity to extrapolate. This is assessed for the case of passive scalar transport, for which it outperforms established models due to improved mixing predictions.

42 ENGINEERING↗