Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Phasing of seven-channel fibre laser radiation with dynamic turbulent phase distortions using a stochastic parallel gradient algorithm at a bandwidth of 450 kHz

We have demonstrated an experimental setup for the coherent phasing of a seven-channel fibre laser system ( λ = 1064 nm) in a scheme comprising a master oscillator and a set of parallel amplifiers with lithium niobate-based fibre-optic phase modulators. Using a stochastic parallel gradient algorithm, an instrumental phase modulator control unit ensures a bandwidth of the system up to 450 kHz. The effectiveness of phasing of light transmitted through a turbulent medium with a characteristic time scale τ{sub turb} has been studied experimentally as a function of phasing time τ{sub ph}. The results demonstrate that the average Strehl ratio begins to rise at τ{sub turb}/τ{sub ph} ⩾ 2 and that the effectiveness of compensation for dynamic phase distortions in the beam propagation path rises sharply at τ{sub turb}/τ{sub ph} ≈ 20. For τ{sub turb}/τ{sub ph} ⩾ 30 – 40, the average Strehl ratio remains constant at the level reached. (control of laser radiation parameters)

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scalable Techniques for Stochastic Power Flow Problems (Final Report)

The proposed research focuses on developing scalable algorithms for two-stage security-constrained OPF problems with AC power flow constraints, a class of problems complicated by (i) scale arising from a scenario representation; and (ii) the presence of nonlinearity, nonconvexity, and possibly second-stage discreteness or complementarity. Unfortunately, most existing solvers cannot contend with both challenges simultaneously; accordingly, the proposed research focuses on developing solution techniques that can both scale with the number of scenarios and contend with nonconvexity and second-stage complementarity. We consider three avenues for addressing such problems: (i) Variable sample-size SQP (VS-SQP) methods that combine sparse Quasi-Newton updates with a scalable variance-reduced stochastic gradient scheme for stochastic QP subproblems, allowing for contending with second-stage complementarity via regularization; (ii) Variable sample-size stochastic Interior-point (VS-sIP) schemes that propose a sampling-based regularized (to allow for contending with complementarity) interior-point schemes in which a Schur-complement technique is employed for decomposing the Newton direction computation step; (iii) Variable sample-size tractable ADMM (VS-tADMM) schemes combine variable sample-sizes with carefully designed techniques for resolving each of the nonconvex updates (by leveraging the QCQP structures). We intend to compare the three schemes using performance profiles in terms of solution quality, scalability, etc. and then select one scheme which will then be developed and further refined in Python for purposes of the GO competition.

42 ENGINEERING↗

Elastic distributed training with fast convergence and efficient resource utilization

Distributed learning is now routinely conducted on cloud as well as dedicated clusters. Training with elastic resources brings new challenges and design choices. Prior studies focus on runtime performance and assume a static algorithmic behavior. In this work, by analyzing the impact of of resource scaling on convergence, we introduce schedules for synchronous stochastic gradient descent that proactively adapt the number of learners to reduce training time and improve convergence. Our approach no longer assumes a constant number of processors throughout training. In our experiment, distributed stochastic gradient descent with dynamic schedules and reduction momentum achieves better convergence and significant speedups over prior static ones. Numerous distributed training jobs running on cloud may benefit from our approach.

Cong, Guojing↗

Towards provably efficient quantum algorithms for large-scale machine-learning models

Large machine learning models are revolutionary technologies of artificial intelligence whose bottlenecks include huge computational expenses, power, and time used both in the pre-training and fine-tuning process. In this work, we show that fault-tolerant quantum computing could possibly provide provably efficient resolutions for generic (stochastic) gradient descent algorithms, scaling as $\mathcal{O}$(T 2 x polylog($n$)), where n is the size of the models and T is the number of iterations in the training, as long as the models are both sufficiently dissipative and sparse, with small learning rates. Based on earlier efficient quantum algorithms for dissipative differential equations, we find and prove that similar algorithms work for (stochastic) gradient descent, the primary algorithm for machine learning. In practice, we benchmark instances of large machine learning models from 7 million to 103 million parameters. We find that, in the context of sparse training, a quantum enhancement is possible at the early stage of learning after model pruning, motivating a sparse parameter download and re-upload scheme. Our work shows solidly that fault-tolerant quantum algorithms could potentially contribute to most state-of-the-art, large-scale machine-learning problems.

97 MATHEMATICS AND COMPUTING↗

Multi-variance replica exchange SGMCMC for inverse and forward problems via Bayesian PINN

Physics-informed neural network (PINN) has been successfully applied in solving a variety of nonlinear non-convex forward and inverse problems. However, the training is challenging because of the non-convex loss functions and the multiple optima in the Bayesian inverse problem. In this work, we propose a multi-variance replica exchange stochastic gradient Langevin dynamics method to tackle the challenge of the multiple local optima in the optimization and the challenge of the multiple modal posterior distribution in the inverse problem. Replica exchange methods are capable of escaping from the local traps and accelerating the convergence; two chains with different temperatures are designed where the low temperature chain aims for the local convergence, and the target of the high temperature chain is to travel globally and explore the whole loss function entropy landscape. However, it may not be efficient to solve mathematical inversion problems by using the vanilla replica method directly since the method doubles the computational cost in evaluating the forward solvers (likelihood functions) in the two chains. To address this issue, we propose to make different assumptions on the energy function estimation and this facilities one to use solvers of different fidelities in the likelihood function evaluation. More precisely, one can use a solver with low fidelity in the high temperature chain while using a solver with high fidelity in the low temperature chain. Our proposed method significantly lowers the computational cost in the high temperature chain, meanwhile preserving the accuracy and converging very fast. Here we give an unbiased estimate of the swapping rate and give an estimation of the discretization error of the scheme. To verify our idea, we design and solve four inverse problems which have multiple modes. The proposed method is also employed to train the Bayesian PINN to solve the forward and inverse problems; faster and more accurate convergence has been observed when compared to the stochastic gradient Langevin dynamics (SGLD) method and vanilla replica exchange methods.

97 MATHEMATICS AND COMPUTING↗

Latency considerations for stochastic optimizers in variational quantum algorithms

Variational quantum algorithms, which have risen to prominence in the noisy intermediate-scale quantum setting, require the implementation of a stochastic optimizer on classical hardware. To date, most research has employed algorithms based on the stochastic gradient iteration as the stochastic classical optimizer. In this work we propose instead using stochastic optimization algorithms that yield stochastic processes emulating the dynamics of classical deterministic algorithms. This approach results in methods with theoretically superior worst-case iteration complexities, at the expense of greater per-iteration sample (shot) complexities. We investigate this trade-off both theoretically and empirically and conclude that preferences for a choice of stochastic optimizer should explicitly depend on a function of both latency and shot execution times.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Locally adaptive activation functions with slope recovery for deep and physics-informed neural networks

Here we propose two approaches of locally adaptive activation functions namely, layer-wise and neuron-wise locally adaptive activation functions, which improve the performance of deep and physics-informed neural networks. The local adaptation of activation function is achieved by introducing a scalable parameter in each layer (layer-wise) and for every neuron (neuron-wise) separately, and then optimizing it using a variant of stochastic gradient descent algorithm. In order to further increase the training speed, an activation slope-based slope recovery term is added in the loss function, which further accelerates convergence, thereby reducing the training cost. On the theoretical side, we prove that in the proposed method, the gradient descent algorithms are not attracted to sub-optimal critical points or local minima under practical conditions on the initialization and learning rate, and that the gradient dynamics of the proposed method is not achievable by base methods with any (adaptive) learning rates. We further show that the adaptive activation methods accelerate the convergence by implicitly multiplying conditioning matrices to the gradient of the base method without any explicit computation of the conditioning matrix and the matrix–vector product. The different adaptive activation functions are shown to induce different implicit conditioning matrices. Furthermore, the proposed methods with the slope recovery are shown to accelerate the training process.

97 MATHEMATICS AND COMPUTING↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Stochastic Optimization and Uncertainty Quantification of Natrium-based Nuclear-Renewable Energy Systems for Flexible Power Applications in Deregulated Markets

Rapid integration of variable renewable energy sources (VRES) has made modeling and stochastic optimization of hybrid energy systems crucial for studying their long-term performance and viability. However, most studies have focused on just historical data, which may be unreliable for capturing short-term fluctuations, rare events, and long-term patterns of energy demand, price, and the variability of renewable energy sources. For this study, optimal synthetic time series models were developed using Wasserstein distance. The models were validated by comparing the key statistical measures against those of the historical data. They were then used to optimize the integrated Natrium-style advanced energy systems and their long-term (30 years) economics. The stochastic model performs bi-level optimization to find the optimal sizes for the balance of plant and thermal energy storage, while also optimizing energy dispatch to achieve the maximum net present value. In studies of two deregulated markets (California ISO and the Electric Reliability Council of Texas), the integrated Natrium-style system performed better in CAISO than in ERCOT, given higher and more consistent electricity prices during peak-demand periods. The potentially enlarged cost associated with the variable operation and maintenance of the TES system also plays a significant role in driving the system sizing, thus its impacts on the system are investigated in detail through comparison against a baseline case. The study also finds that the bi-level optimization results based on stochastic gradient descent closely match the grid search results. The uncertainty quantification of the stochastic signals provides further NPV-related insights and probability distributions for the case studies. The normal standard error of the mean of NPV for the case with and without TES VOM for CAISO were found to be 7.73M (plus-minus sign) 1.09M USD and 104.99M (plus-minus sign) 1.25M USD, respectively based on a 95% confidence. Given the relatively small NPV variance based on 150 samples, the analysis affords the most robust possible prediction of the techno-economic performance of the integrated Natrium-style energy systems.

25 ENERGY STORAGE↗

A Hybrid Gradient Method to Designing Bayesian Experiments for Implicit Models

Bayesian experimental design (BED) aims at designing an experiment to maximize the information gathering from the collected data. The optimal design is usually achieved by maximizing the mutual information (MI) between the data and the model parameters. When the analytical expression of the MI is unavailable, e.g.,having implicit models with intractable data distributions, a neural network-based lower bound of the MI was recently proposed and a gradient ascent method was used to maximize the lower bound [1]. However, the approach in [1] requires a pathwise sampling path to compute the gradient of the MI lower bound with respect to the design variables, and such a pathwise sampling path is usually inaccessible for implicit models. In this work, we propose a hybrid gradient approach that leverages recent advances in variational MI estimator and evolution strategies (ES)combined with black-box stochastic gradient ascent (SGA) to maximize the MI lower bound. This allows the design process to be achieved through a unified scalable procedure for implicit models without sampling path gradients. Several experiments demonstrate that our approach significantly improves the scalability of BED for implicit models in high-dimensional design space.

Zhang, Jiaxin↗

Finding MIDDLE Ground: Scalable and Secure Distributed Learning

Edge computing methods allow devices to efficiently train a high-performing, robust, and personalized model for predictive tasks. However, these methods succumb to privacy and scalability concerns such as adversarial data recovery and expensive model communication. Furthermore, edge computing methods unrealistically assume that all devices train an identical model. In practice, edge devices have varying computational and memory constraints which may not allow certain devices to have the space or speed to train a specific model. To overcome these issues, we propose MIDDLE: a model independent distributed learning algorithm which allows heterogeneous edge devices to assist each other’s training while communicating only non-sensitive information. MIDDLE unlocks the ability for edge devices, regardless of computational or memory constraints, to assist each other even with completely different model architectures. Furthermore, MIDDLE does not require model or gradient communication which greatly reduces communication size and time. We prove that MIDDLE attains the optimal convergence rate O(1/sqrt(TM)) of stochastic gradient descent for convex and non-convex smooth optimization (for total iterations T and batch size M). Finally, our experimental results demonstrate that MIDDLE (even in non-IID data settings) attains robust and high-performing models without model or gradient communication.

Bornstein, Marc I.↗

Learning generative neural networks with physics knowledge

Deep generative neural networks have enabled modeling complex distributions, but incorporating physics knowledge into the neural networks is still challenging and is at the core of current physics-based machine learning research. To this end, we propose a physics generative neural network (PhysGNN), a new class of generative neural networks for learning unknown distributions in a physical system described by partial differential equations (PDE). PhysGNN couples PDE systems with generative neural networks. It is a fully differentiable model that allows back-propagation of gradients through both numerical PDE solvers and generative neural networks, and is trained by minimizing the discrete Wasserstein distance between generated and observed probability distributions of the PDE outputs using the stochastic gradient descent method. Moreover, PhysGNN does not require adversarial training like standard generative neural networks, which offers better stability than adversarial training. We show that PhysGNN can learn complex distributions in stochastic inverse problems, where conventional methods such as maximum likelihood estimation and momentum matching methods may be inapplicable when little knowledge is known about the form of unknown distributions or the physical model is too complex. Furthermore, our method allows physics-based generative neural network training for learning complex distributions in the context of differential equations.

97 MATHEMATICS AND COMPUTING↗

End-to-End Differentiable Modeling and Management of the Environment

Focal Area: (2) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization. We emphasize the importance of leveraging optimization techniques from AI/machine learning (ML) to solve challenging problems in Earth system modeling. Science Challenge: Automatic differentiation has had a transformative effect on ML by allowing the calculation of gradients of arbitrary functions in an incredibly large class of models. We can potentially realize similar improvements in parameter estimation and control for Earth system models (ESMs) by reimplementing them in computational frameworks from ML. Practitioners working with large (>10 7 parameters) models in ML and AI can obtain good predictive performance in a range of spatiotemporal tasks by making use of optimization via stochastic gradient descent and incorporating prior knowledge at multiple levels. We propose writing ESMs in open-source computational frameworks such as Torch, Tensorflow, and JAX to greatly expand the scope of environmental forecasting and management challenges, which can be addressed by leveraging automatic differentiation and gradient descent-like algorithms. We do not call for a wholesale replacement of physical models with data-driven surrogates, but rather advocate for interleaving physical and empirical equations in a manner that is most faithful to the extent of our scientific knowledge and observational data. Central to this topic is the merging of differentiable physical simulations with differentiable optimization layers, which are now both beginning to come to the forefront.

54 ENVIRONMENTAL SCIENCES↗

Improving Deep Neural Networks’ Training for Image Classification With Nonlinear Conjugate Gradient-Style Adaptive Momentum

Momentum is crucial in stochastic gradient-based optimization algorithms for accelerating or improving training deep neural networks (DNNs). In deep learning practice, the momentum is usually weighted by a well-calibrated constant. However, tuning the hyperparameter for momentum can be a significant computational burden. In this article, we propose a novel adaptive momentum for improving DNNs training; this adaptive momentum, with no momentum-related hyperparame- ter required, is motivated by the nonlinear conjugate gradient (NCG) method. Stochastic gradient descent (SGD) with this new adaptive momentum eliminates the need for the momentum hyperparameter calibration, allows using a significantly larger learning rate, accelerates DNN training, and improves the final accuracy and robustness of the trained DNNs. For example, SGD with this adaptive momentum reduces classification errors for training ResNet110 for CIFAR10 and CIFAR100 from 5.25% to 4.64% and 23.75% to 20.03%, respectively. Furthermore, SGD, with the new adaptive momentum, also benefits adversarial training and, hence, improves the adversarial robustness of the trained DNNs.

97 MATHEMATICS AND COMPUTING↗

Prediction of gas hydrate saturation using machine learning and optimal set of well-logs

We report resistivity and acoustic logs are widely used to estimate gas hydrate saturation in various sedimentary systems using one of the two popular methods ((1) acoustic velocity and (2) electrical resistivity), but the limitations of these two methods are often overlooked, which include (i) well-specific calibration of empirical exponents in the electrical resistivity method, (ii) assumption of known pore morphology for gas hydrates in the acoustic velocity method, and (iii) presence of unknown mineralogy and bulk modulus terms in the acoustic velocity method. NMR-density porosity-derived gas hydrate saturation based on the analysis of the transverse magnetization relaxation time (T2) is considered the most precise method, but acquisition of NMR-based logs is limited at relatively recent drilled sites; additionally, its use in conventional oil and gas reservoirs is not that common due to higher cost and operational deployment limitations associated with acquiring NMR well-logs. This study proposes a new method that predicts gas hydrate saturation (S h ) for any well using porosity, bulk density, and compressional wave (P wave) velocity well-logs with neural network (or stochastic gradient descent regression) without any well-specific calibration and/or other aforementioned shortcomings of the existing methods. The method is developed by examining the underlying dependency between S h and different combinations of well-logs, chosen from 6 routine logs, with 12 different machine learning (ML) algorithms. The accuracy of the proposed method in predicting S h is ~ 84%, which is better than the accuracy of seismic and electrical resistivity methods (≤ 75%) per the results reported by three different studies. The robustness of the method in the specific case of permafrost-associated gas hydrates is demonstrated with well-log data from two wells drilled on the Alaska North Slope.

58 GEOSCIENCES↗

A phase transition for finding needles in nonlinear haystacks with LASSO artificial neural networks

To fit sparse linear associations, a LASSO sparsity inducing penalty with a single hyperparameter provably allows to recover the important features (needles) with high probability in certain regimes even if the sample size is smaller than the dimension of the input vector (haystack). More recently learners known as artificial neural networks (ANN) have shown great successes in many machine learning tasks, in particular fitting nonlinear associations. Small learning rate, stochastic gradient descent algorithm and large training set help to cope with the explosion in the number of parameters present in deep neural networks. Yet few ANN learners have been developed and studied to find needles in nonlinear haystacks. Driven by a single hyperparameter, our ANN learner, like for sparse linear associations, exhibits a phase transition in the probability of retrieving the needles, which we do not observe with other ANN learners. To select our penalty parameter, we generalize the universal threshold of Donoho and Johnstone (Biometrika 81(3):425–455, 1994) which is a better rule than the conservative (too many false detections) and expensive cross-validation. In the spirit of simulated annealing, we propose a warm-start sparsity inducing algorithm to solve the high-dimensional, non-convex and non-differentiable optimization problem. We perform simulated and real data Monte Carlo experiments to quantify the effectiveness of our approach.

97 MATHEMATICS AND COMPUTING↗