Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradients”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Mode connectivity in the loss landscape of parameterized quantum circuits

Variational training of parameterized quantum circuits (PQCs) underpins many workflows employed on near-term noisy intermediate scale quantum (NISQ) devices. It is a hybrid quantum-classical approach that minimizes an associated cost function in order to train a parameterized ansatz. In this work we adapt the qualitative loss landscape characterization for neural networks introduced in Goodfellow et al. (2014); Li et al. (2017) and tests for connectivity used in Draxler et al. (2018) to study the loss landscape features in PQC training. We present results for PQCs trained on a simple regression task, using the bilayer circuit ansatz, which consists of alternating layers of parameterized rotation gates and entangling gates. Multiple circuits are trained with 3 different batch gradient optimizers: stochastic gradient descent, the quantum natural gradient, and Adam. We identify large features in the landscape that can lead to faster convergence in training workflows.

97 MATHEMATICS AND COMPUTING↗

Structure optimization with stochastic density functional theory

Linear-scaling techniques for Kohn–Sham density functional theory are essential to describe the ground state properties of extended systems. Still, these techniques often rely on the localization of the density matrix or accurate embedding approaches, limiting their applicability. In contrast, stochastic density functional theory (sDFT) achieves linear- and sub-linear scaling by statistically sampling the ground state density without relying on embedding or imposing localization. In return, ground state observables, such as the forces on the nuclei, fluctuate in sDFT, making optimizing the nuclear structure a highly non-trivial problem. In this work, we combine the most recent noise-reduction schemes for sDFT with stochastic optimization algorithms to perform structure optimization within sDFT. We compare the performance of the stochastic gradient descent approach and its variations (stochastic gradient descent with momentum) with stochastic optimization techniques that rely on the Hessian, such as the stochastic Broyden–Fletcher–Goldfarb–Shanno algorithm. In conclusion, we further provide a detailed assessment of the computational efficiency and its dependence on the optimization parameters of each method for determining the ground state structure of bulk silicon with varying supercell dimensions.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Fed-DeepONet: Stochastic Gradient-Based Federated Training of Deep Operator Networks

The Deep Operator Network (DeepONet) framework is a different class of neural network architecture that one trains to learn nonlinear operators, i.e., mappings between infinite-dimensional spaces. Traditionally, DeepONets are trained using a centralized strategy that requires transferring the training data to a centralized location. Such a strategy, however, limits our ability to secure data privacy or use high-performance distributed/parallel computing platforms. To alleviate such limitations, in this paper, we study the federated training of DeepONets for the first time. That is, we develop a framework, which we refer to as Fed-DeepONet, that allows multiple clients to train DeepONets collaboratively under the coordination of a centralized server. To achieve Fed-DeepONets, we propose an efficient stochastic gradient-based algorithm that enables the distributed optimization of the DeepONet parameters by averaging first-order estimates of the DeepONet loss gradient. Then, to accelerate the training convergence of Fed-DeepONets, we propose a moment-enhanced (i.e., adaptive) stochastic gradient-based strategy. Finally, we verify the performance of Fed-DeepONet by learning, for different configurations of the number of clients and fractions of available clients, (i) the solution operator of a gravity pendulum and (ii) the dynamic response of a parametric library of pendulums.

Moya, Christian↗

Modeling of terrain gradient for stochastically spaced rows of a measurement matrix

Terrain gradients are employed to evaluate passable regions for unmanned martian roving vehicle. Range data matrix is displaced randomly row wise at the shallow elevation angles near the skyline. The magnitude of the measurement noise in the elevation angles can approach that of the spacing of the same angle. By using a variable incremental data spacing scanning scheme, one can estimate this signal noise ratio. It is found that the error in slope estimate at far distance becomes large for a given elevation angle error. Evaluation of the in-path slopes can be expressed in terms of the inverse of the range slopes. This is because of the fact that the elevation angle is considered as a random variable while the range data are relatively less noisy. An error analysis is performed and it is found that the change of slope is a nonlinear function of the error in elevation angle.

Mediavilla, R.↗

Phasing of seven-channel fibre laser radiation with dynamic turbulent phase distortions using a stochastic parallel gradient algorithm at a bandwidth of 450 kHz

We have demonstrated an experimental setup for the coherent phasing of a seven-channel fibre laser system ( λ = 1064 nm) in a scheme comprising a master oscillator and a set of parallel amplifiers with lithium niobate-based fibre-optic phase modulators. Using a stochastic parallel gradient algorithm, an instrumental phase modulator control unit ensures a bandwidth of the system up to 450 kHz. The effectiveness of phasing of light transmitted through a turbulent medium with a characteristic time scale τ{sub turb} has been studied experimentally as a function of phasing time τ{sub ph}. The results demonstrate that the average Strehl ratio begins to rise at τ{sub turb}/τ{sub ph} ⩾ 2 and that the effectiveness of compensation for dynamic phase distortions in the beam propagation path rises sharply at τ{sub turb}/τ{sub ph} ≈ 20. For τ{sub turb}/τ{sub ph} ⩾ 30 – 40, the average Strehl ratio remains constant at the level reached. (control of laser radiation parameters)

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Scalable Techniques for Stochastic Power Flow Problems (Final Report)

The proposed research focuses on developing scalable algorithms for two-stage security-constrained OPF problems with AC power flow constraints, a class of problems complicated by (i) scale arising from a scenario representation; and (ii) the presence of nonlinearity, nonconvexity, and possibly second-stage discreteness or complementarity. Unfortunately, most existing solvers cannot contend with both challenges simultaneously; accordingly, the proposed research focuses on developing solution techniques that can both scale with the number of scenarios and contend with nonconvexity and second-stage complementarity. We consider three avenues for addressing such problems: (i) Variable sample-size SQP (VS-SQP) methods that combine sparse Quasi-Newton updates with a scalable variance-reduced stochastic gradient scheme for stochastic QP subproblems, allowing for contending with second-stage complementarity via regularization; (ii) Variable sample-size stochastic Interior-point (VS-sIP) schemes that propose a sampling-based regularized (to allow for contending with complementarity) interior-point schemes in which a Schur-complement technique is employed for decomposing the Newton direction computation step; (iii) Variable sample-size tractable ADMM (VS-tADMM) schemes combine variable sample-sizes with carefully designed techniques for resolving each of the nonconvex updates (by leveraging the QCQP structures). We intend to compare the three schemes using performance profiles in terms of solution quality, scalability, etc. and then select one scheme which will then be developed and further refined in Python for purposes of the GO competition.

42 ENGINEERING↗

Elastic distributed training with fast convergence and efficient resource utilization

Distributed learning is now routinely conducted on cloud as well as dedicated clusters. Training with elastic resources brings new challenges and design choices. Prior studies focus on runtime performance and assume a static algorithmic behavior. In this work, by analyzing the impact of of resource scaling on convergence, we introduce schedules for synchronous stochastic gradient descent that proactively adapt the number of learners to reduce training time and improve convergence. Our approach no longer assumes a constant number of processors throughout training. In our experiment, distributed stochastic gradient descent with dynamic schedules and reduction momentum achieves better convergence and significant speedups over prior static ones. Numerous distributed training jobs running on cloud may benefit from our approach.

Cong, Guojing↗

Towards provably efficient quantum algorithms for large-scale machine-learning models

Large machine learning models are revolutionary technologies of artificial intelligence whose bottlenecks include huge computational expenses, power, and time used both in the pre-training and fine-tuning process. In this work, we show that fault-tolerant quantum computing could possibly provide provably efficient resolutions for generic (stochastic) gradient descent algorithms, scaling as $\mathcal{O}$(T 2 x polylog($n$)), where n is the size of the models and T is the number of iterations in the training, as long as the models are both sufficiently dissipative and sparse, with small learning rates. Based on earlier efficient quantum algorithms for dissipative differential equations, we find and prove that similar algorithms work for (stochastic) gradient descent, the primary algorithm for machine learning. In practice, we benchmark instances of large machine learning models from 7 million to 103 million parameters. We find that, in the context of sparse training, a quantum enhancement is possible at the early stage of learning after model pruning, motivating a sparse parameter download and re-upload scheme. Our work shows solidly that fault-tolerant quantum algorithms could potentially contribute to most state-of-the-art, large-scale machine-learning problems.

97 MATHEMATICS AND COMPUTING↗

An investigation of Newton-Sketch and subsampled Newton methods

Sketching, a dimensionality reduction technique, has received much attention in the statistics community. In this paper, we study sketching in the context of Newton's method for solving finite-sum optimization problems in which the number of variables and data points are both large. In this work, we study two forms of sketching that perform dimensionality reduction in data space: Hessian subsampling and randomized Hadamard transformations. Each has its own advantages, and their relative tradeoffs have not been investigated in the optimization literature. Additionally, our study focuses on practical versions of the two methods in which the resulting linear systems of equations are solved approximately, at every iteration, using an iterative solver. The advantages of using the conjugate gradient method vs. a stochastic gradient iteration are revealed through a set of numerical experiments, and a complexity analysis of the Hessian subsampling method is presented.

97 MATHEMATICS AND COMPUTING↗

Multi-variance replica exchange SGMCMC for inverse and forward problems via Bayesian PINN

Physics-informed neural network (PINN) has been successfully applied in solving a variety of nonlinear non-convex forward and inverse problems. However, the training is challenging because of the non-convex loss functions and the multiple optima in the Bayesian inverse problem. In this work, we propose a multi-variance replica exchange stochastic gradient Langevin dynamics method to tackle the challenge of the multiple local optima in the optimization and the challenge of the multiple modal posterior distribution in the inverse problem. Replica exchange methods are capable of escaping from the local traps and accelerating the convergence; two chains with different temperatures are designed where the low temperature chain aims for the local convergence, and the target of the high temperature chain is to travel globally and explore the whole loss function entropy landscape. However, it may not be efficient to solve mathematical inversion problems by using the vanilla replica method directly since the method doubles the computational cost in evaluating the forward solvers (likelihood functions) in the two chains. To address this issue, we propose to make different assumptions on the energy function estimation and this facilities one to use solvers of different fidelities in the likelihood function evaluation. More precisely, one can use a solver with low fidelity in the high temperature chain while using a solver with high fidelity in the low temperature chain. Our proposed method significantly lowers the computational cost in the high temperature chain, meanwhile preserving the accuracy and converging very fast. Here we give an unbiased estimate of the swapping rate and give an estimation of the discretization error of the scheme. To verify our idea, we design and solve four inverse problems which have multiple modes. The proposed method is also employed to train the Bayesian PINN to solve the forward and inverse problems; faster and more accurate convergence has been observed when compared to the stochastic gradient Langevin dynamics (SGLD) method and vanilla replica exchange methods.

97 MATHEMATICS AND COMPUTING↗

Latency considerations for stochastic optimizers in variational quantum algorithms

Variational quantum algorithms, which have risen to prominence in the noisy intermediate-scale quantum setting, require the implementation of a stochastic optimizer on classical hardware. To date, most research has employed algorithms based on the stochastic gradient iteration as the stochastic classical optimizer. In this work we propose instead using stochastic optimization algorithms that yield stochastic processes emulating the dynamics of classical deterministic algorithms. This approach results in methods with theoretically superior worst-case iteration complexities, at the expense of greater per-iteration sample (shot) complexities. We investigate this trade-off both theoretically and empirically and conclude that preferences for a choice of stochastic optimizer should explicitly depend on a function of both latency and shot execution times.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Estimation of optical flow in airborne electro-optical sensors by stochastic approximation

The essence of motion or range estimation by passive electrooptical means is the ability to determine the correspondence of picture elements in pairs of image frames and to estimate their coordinates and their disparity (relative shifts) in the image plane of an electrooptical imaging sensor. The disparity can be in successive frames due to self-motion or in simultaneous frames of a stereo pair. A key issue is to provide these estimates on-line. This paper describes the theoretical background of such an interframe shift estimator. It is based on a stochastic gradient algorithm, specifically implementing a form of stochastic approximation, which can achieve rapid convergence of the shift estimate. Analytical and numerical simulation examples for random texture and isolated features validate the feasibility and the effectiveness of the estimator.

Merhav, S. J.↗

State Estimation for Distribution Networks with Asynchronous Sensors Using Stochastic Descent: Preprint

This paper investigates the problem of state estimation for distribution networks with asynchronous sensors comprising of a mix of smart meters and phasor measurement units (PMUs) with multiple sampling and reporting rates. We consider two independent scenarios of state estimation and tracking, with either voltages or currents as states. With these two sets, we investigate estimation under (a) full data, assuming all measurements are available and (b) limited data, where an online algorithmic approach is adopted to estimate the possibly time-varying states by processing measurements as and when available. The proposed algorithm, inspired by the classical Stochastic Gradient Descent (SGD) approach updates the states based on the previous estimate and the newly available measurements. Finally, we demonstrate the estimation and tracking efficacy through numerical simulations on the IEEE-37 test network, while also highlighting how estimation with currents as states leads to faster convergence.

asynchronous sensors↗

Locally adaptive activation functions with slope recovery for deep and physics-informed neural networks

Here we propose two approaches of locally adaptive activation functions namely, layer-wise and neuron-wise locally adaptive activation functions, which improve the performance of deep and physics-informed neural networks. The local adaptation of activation function is achieved by introducing a scalable parameter in each layer (layer-wise) and for every neuron (neuron-wise) separately, and then optimizing it using a variant of stochastic gradient descent algorithm. In order to further increase the training speed, an activation slope-based slope recovery term is added in the loss function, which further accelerates convergence, thereby reducing the training cost. On the theoretical side, we prove that in the proposed method, the gradient descent algorithms are not attracted to sub-optimal critical points or local minima under practical conditions on the initialization and learning rate, and that the gradient dynamics of the proposed method is not achievable by base methods with any (adaptive) learning rates. We further show that the adaptive activation methods accelerate the convergence by implicitly multiplying conditioning matrices to the gradient of the base method without any explicit computation of the conditioning matrix and the matrix–vector product. The different adaptive activation functions are shown to induce different implicit conditioning matrices. Furthermore, the proposed methods with the slope recovery are shown to accelerate the training process.

97 MATHEMATICS AND COMPUTING↗

Variance-Reduced Accelerated First-Order Methods: Central Limit Theorems and Confidence Statements

In this paper, we consider a strongly convex stochastic optimization problem and propose three classes of variable sample-size stochastic first-order methods: (i) the standard stochastic gradient descent method, (ii) its accelerated variant, and (iii) the stochastic heavy-ball method. In each scheme, the exact gradients are approximated by averaging across an increasing batch size of sampled gradients. We prove that when the sample size increases at a geometric rate, the generated estimates converge in mean to the optimal solution at an analogous geometric rate for schemes (i)–(iii). Based on this result, we provide central limit statements, whereby it is shown that the rescaled estimation errors converge in distribution to a normal distribution with the associated covariance matrix dependent on the Hessian matrix, the covariance of the gradient noise, and the step length. If the sample size increases at a polynomial rate, we show that the estimation errors decay at a corresponding polynomial rate and establish the associated central limit theorems (CLTs). Under certain conditions, we discuss how both the algorithms and the associated limit theorems may be extended to constrained and nonsmooth regimes. As a result, we provide an avenue to construct confidence regions for the optimal solution based on the established CLTs and test the theoretical findings on a stochastic parameter estimation problem.

Lei, Jinlong↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Stochastic Optimization and Uncertainty Quantification of Natrium-based Nuclear-Renewable Energy Systems for Flexible Power Applications in Deregulated Markets

Rapid integration of variable renewable energy sources (VRES) has made modeling and stochastic optimization of hybrid energy systems crucial for studying their long-term performance and viability. However, most studies have focused on just historical data, which may be unreliable for capturing short-term fluctuations, rare events, and long-term patterns of energy demand, price, and the variability of renewable energy sources. For this study, optimal synthetic time series models were developed using Wasserstein distance. The models were validated by comparing the key statistical measures against those of the historical data. They were then used to optimize the integrated Natrium-style advanced energy systems and their long-term (30 years) economics. The stochastic model performs bi-level optimization to find the optimal sizes for the balance of plant and thermal energy storage, while also optimizing energy dispatch to achieve the maximum net present value. In studies of two deregulated markets (California ISO and the Electric Reliability Council of Texas), the integrated Natrium-style system performed better in CAISO than in ERCOT, given higher and more consistent electricity prices during peak-demand periods. The potentially enlarged cost associated with the variable operation and maintenance of the TES system also plays a significant role in driving the system sizing, thus its impacts on the system are investigated in detail through comparison against a baseline case. The study also finds that the bi-level optimization results based on stochastic gradient descent closely match the grid search results. The uncertainty quantification of the stochastic signals provides further NPV-related insights and probability distributions for the case studies. The normal standard error of the mean of NPV for the case with and without TES VOM for CAISO were found to be 7.73M (plus-minus sign) 1.09M USD and 104.99M (plus-minus sign) 1.25M USD, respectively based on a 95% confidence. Given the relatively small NPV variance based on 150 samples, the analysis affords the most robust possible prediction of the techno-economic performance of the integrated Natrium-style energy systems.

25 ENERGY STORAGE↗