Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

Quantum optimization algorithms: Energetic implications

Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. Here in this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least 2x to 4x more energy efficient than their classical counterparts; why feedback-based quantum optimization is energy-inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of ≥2 0 x. Finally, we use the NchooseK high-level programming model to run optimization problems on both gate-based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate-based counterparts.

97 MATHEMATICS AND COMPUTING↗

Random coordinate descent: A simple alternative for optimizing parameterized quantum circuits

Variational quantum algorithms rely on the optimization of parameterized quantum circuits in noisy settings. The commonly used back-propagation procedure in classical machine learning is not directly applicable in this setting due to the collapse of quantum states after measurements. Thus, gradient estimations constitute a significant overhead in a gradient-based optimization of such quantum circuits. This paper introduces a random coordinate descent algorithm as a practical and easy-to-implement alternative to the full gradient descent algorithm. This algorithm only requires one partial derivative at each iteration. Motivated by the behavior of measurement noise in the practical optimization of parameterized quantum circuits, this paper presents an optimization problem setting that is amenable to analysis. Under this setting, the random coordinate descent algorithm exhibits the same level of stochastic stability as the full gradient approach, making it as resilient to noise. The complexity of the random coordinate descent method is generally no worse than that of the gradient descent and can be much better for various quantum optimization problems with anisotropic Lipschitz constants. Theoretical analysis and extensive numerical experiments validate our findings. Published by the American Physical Society 2024

Ding, Zhiyan (ORCID:000000018863403X)↗

Differential Equation Approximation Using Gradient-Boosted Quantile Regression

The operation of cyber-physical-human (CPH) systems is subject to various epistemic and aleatory uncertainties. Overall trustworthiness of CPH systems relies on the trustworthiness of its components and their interactions. It is important that computational models comprising the cyber component of CPH provide predictions accompanied by a measure of confidence in model outcomes. Uncertainty quantification (UQ) and propagation are especially important in safety critical CPH systems. Gradient-boosted trees is a modeling approach capable both of learning the dynamics of a system and performing UQ. In this paper, we devise a method for using gradient boosting to learn the dynamics of a second order differential equation and estimate uncertainty at the same time. We do this by creating a custom loss function that trains the model to approximate the second derivative of a noisy time series, and to penalize based on a parameter that corresponds to the desired quantile. The resulting gradient boosting model can simulate stochastic trajectories of the system given a single starting point, that is, it can estimate both the expected trajectory and its uncertainty. We show that the uncertainty estimation is well calibrated and that the model can learn the dynamics even in the presence of noise. We demonstrate the approach on a simple cartpole system.

Autonomous systems↗

Nonlinear Matrix Approximation with Radial Basis Function Components

We introduce and investigate matrix approximation by decomposition into a sum of radial basis function (RBF) components. An RBF component is a generalization of the outer product between a pair of vectors, where an RBF function replaces the scalar multiplication between individual vector elements. Even though the RBF functions are positive definite, the summation across components is not restricted to convex combinations and allows us to compute the decomposition for any real matrix that is not necessarily symmetric or positive definite. We formulate the problem of seeking such a decomposition as an optimization problem with a nonlinear and non-convex loss function. Several modern versions of the gradient descent method, including their scalable stochastic counterparts, are used to solve this problem. We provide extensive empirical evidence of the effectiveness of the RBF decomposition and that of the gradient-based fitting algorithm. While being conceptually motivated by singular value decomposition (SVD), our proposed nonlinear counterpart outperforms SVD by drastically reducing the memory required to approximate a data matrix with the same L2 error for a wide range of matrix types. For example, it leads to 2 to 6 times memory save for Gaussian noise, graph adjacency matrices, and kernel matrices. Moreover, this proximity-based decomposition can offer additional interpretability in applications that involve, e.g., capturing the inner low-dimensional structure of the data, retaining graph connectivity structure, and preserving the acutance of images.

Rebrova, Elizaveta↗

Machine learning unifies flexibility and efficiency of spinodal structure generation for stochastic biomaterial design

Abstract Porous biomaterials design for bone repair is still largely limited to regular structures (e.g. rod-based lattices), due to their easy parameterization and high controllability. The capability of designing stochastic structure can redefine the boundary of our explorable structure–property space for synthesizing next-generation biomaterials. We hereby propose a convolutional neural network (CNN) approach for efficient generation and design of spinodal structure—an intriguing structure with stochastic yet interconnected, smooth, and constant pore channel conducive to bio-transport. Our CNN-based approach simultaneously possesses the tremendous flexibility of physics-based model in generating various spinodal structures (e.g. periodic, anisotropic, gradient, and arbitrarily large ones) and comparable computational efficiency to mathematical approximation model. We thus successfully design spinodal bone structures with target anisotropic elasticity via high-throughput screening, and directly generate large spinodal orthopedic implants with desired gradient porosity. This work significantly advances stochastic biomaterials development by offering an optimal solution to spinodal structure generation and design.

59 BASIC BIOLOGICAL SCIENCES↗

Statistics of Experiments on Cluster Formation and Transport in a Gravitational Field

Metastable state relaxation in a gravitational field is investigated in the case of non-critical binary solutions. A relaxation description is presented in terms of the time-dependent Ginzburg-Landau formalism for a non-conserved order parameter. A new ansatz for solution of the corresponding partial nonlinear stochastic differential equation is discussed. It is proved that, for the supersaturated solution under consideration, the metastable state relaxation in a gravitational field leads to formation of solute concentration gradients due to the sedimentation of subcritical solute clusters. The pure discussion of the possible methods to compare theoretical results and experimental data related to solute sedimentation in a gravitational field is presented. It is shown that in order to describe these experiments it is necessary to deal both with the value of the solute concentration gradient and with its formation rate. The stochastic nature of the sedimentation process is shown.

Izmailov, Alexander F.↗

Plateau Phenomenon in Gradient Descent Training of RELU Networks: Explanation, Quantification, and Avoidance

The ability of neural networks to provide ‘best in class’ approximation across a wide range of applications is well-documented. Nevertheless, the powerful expressivity of neural networks comes to naught if one is unable to effectively train (choose) the parameters defining the network. In general, neural networks are trained by gradient descent type optimization methods,a stochastic variant thereof. In practice, such methods result in the loss function decreases rapidly at the beginning of training but then, after a relatively small number of steps, significantly slow down. The loss may even appear to stagnate over the period of a large number of epochs, only to then suddenly start to decrease fast again for no apparent reason. This so-called plateau phenomenon manifests itself in many learning tasks. The present work aims to identify and quantify the root causes of plateau phenomenon.analysis is carried out in the setting of univariate ReLU networks. No assumptions are made on the number of neurons relative to the number of training data, and our results hold for both the lazy and adaptive regimes. Here, the main findings are: plateaux correspond to periods during which activation patterns remain constant, where activation pattern refers to the number of data points that activate a given neuron; quantification of convergence of the gradient flow dynamics; and, characterization stationary points in terms solutions of local least squares regression lines over subsets of the training data. Based on these conclusions, we propose a new iterative training method, the Active Neuron Least Squares (ANLS), characterised by the explicit adjustment of the activation pattern at each step, which is designed to enable a quick exit from a plateau. Illustrative numerical examples are included throughout.

97 MATHEMATICS AND COMPUTING↗

Correspondence between neuroevolution and gradient descent

Abstract We show analytically that training a neural network by conditioned stochastic mutation or neuroevolution of its weights is equivalent, in the limit of small mutations, to gradient descent on the loss function in the presence of Gaussian white noise. Averaged over independent realizations of the learning process, neuroevolution is equivalent to gradient descent on the loss function. We use numerical simulation to show that this correspondence can be observed for finite mutations, for shallow and deep neural networks. Our results provide a connection between two families of neural-network training methods that are usually considered to be fundamentally different.

97 MATHEMATICS AND COMPUTING↗

Gradient-Based Optimization of the Common Research Model Wing Subject to CFD-Based Gust and Flutter Constraints

The linearized frequency-domain method was recently implemented in the stabilized finite element solver in NASA’s FUN3D code. Previous work by the authors used this method for enforcing flutter constraints during gradient-based optimizations. More recently, the solver was expanded to account for continuous (also known as stochastic) gust responses. This paper expands on recent Common Research Model wing optimization work, which demonstrated gradient-based optimization with flutter and stochastic gust constraints, among others. While that work utilized FUN3D for static aeroelastic solutions but relied on doublet lattice aerodynamics for gust and flutter responses, the present work replaces these unsteady aerodynamic analyses with those of FUN3D’s linearized frequency-domain solver. With analytic derivatives available, gradient-based optimization is performed through the use of the OpenMDAO/MPhys libraries with over 700 shape, structural, and aerodynamic design variables and over 10 nonlinear constraints. Comparisons of analysis results and optimized designs are made between doublet lattice and linearized frequency-domain solutions.

aeroelasticity↗

Employing Sensitivity Derivatives for Robust Optimization under Uncertainty in CFD

A robust optimization is demonstrated on a two-dimensional inviscid airfoil problem in subsonic flow. Given uncertainties in statistically independent, random, normally distributed flow parameters (input variables), an approximate first-order statistical moment method is employed to represent the Computational Fluid Dynamics (CFD) code outputs as expected values with variances. These output quantities are used to form the objective function and constraints. The constraints are cast in probabilistic terms; that is, the probability that a constraint is satisfied is greater than or equal to some desired target probability. Gradient-based robust optimization of this stochastic problem is accomplished through use of both first and second-order sensitivity derivatives. For each robust optimization, the effect of increasing both input standard deviations and target probability of constraint satisfaction are demonstrated. This method provides a means for incorporating uncertainty when considering small deviations from input mean values.

Newman, Perry A.↗

A relation between cosmic-ray fluctuations, gradient, and diffusion coefficient

The motion of charged particles in a stochastic magnetic field is considered via a generalized quasi-linear expansion of Liouville's equation. The result is an equation relating cosmic-ray scintillations to particle gradients and to magnetic-field fluctuations (or diffusion coefficient). The resulting theory may be regarded as an example of a fluctuation-dissipation phenomenon, in which the diffusion coefficient plays the role of the dissipative parameter. The resonant interaction between particles and the random interplanetary magnetic field is considered explicitly, and it is shown that observed scintillations of high-energy (about 1 GeV) cosmic rays may be reasonably explained by the model.

Owens, A. J.↗

An Online Dynamic Amplitude-Correcting Gradient Estimation Technique to Align X-ray Focusing Optics

High-brightness X-ray pulses, as generated at synchrotrons and X-ray free electron lasers (XFELs), are used in a variety of scientific experiments. At these facilities, measurements often require optical equipment, e.g Compound Refractive Lenses (CRLs) to be precisely aligned and focused. The lateral alignment of CRLs to a beamline requires precise positioning along four axes: two translational, and the two rotational. At a synchrotron, alignment is often accomplished manually. However, XFEL beamlines present a beam brightness that fluctuates stochastically, making manual alignment a time-consuming endeavor. Automation using simplex or classic stochastic descent often fails, given the errant gradient estimates. Herein we present a dynamic-amplitude correction to the usual gradient based on the combination of a generalized finite difference stencil and a time-dependent sampling pattern. Intensity is recorded periodically, then used to normalize numerical derivatives against fluctuations. Error expectation is analyzed, and efficacy is demonstrated on classic benchmarks. We provide a proof of concept by laterally aligning optics on a simulated XFEL beamline using data recorded at both synchrotron and XFEL facilities.

97 MATHEMATICS AND COMPUTING↗

Stochastic projective splitting

Here, we present a new, stochastic variant of the projective splitting (PS) family of algorithms for inclusion problems involving the sum of any finite number of maximal monotone operators. This new variant uses a stochastic oracle to evaluate one of the operators, which is assumed to be Lipschitz continuous, and (deterministic) resolvents to process the remaining operators. Our proposal is the first version of PS with such stochastic capabilities. We envision the primary application being machine learning (ML) problems, with the method’s stochastic features facilitating “mini-batch” sampling of datasets. Since it uses a monotone operator formulation, the method can handle not only Lipschitz-smooth loss minimization, but also min–max and noncooperative game formulations, with better convergence properties than the gradient descent-ascent methods commonly applied in such settings. The proposed method can handle any number of constraints and nonsmooth regularizers via projection and proximal operators. We prove almost-sure convergence of the iterates to a solution and a convergence rate result for the expected residual, and close with numerical experiments on a distributionally robust sparse logistic regression problem.

97 MATHEMATICS AND COMPUTING↗

Decadal Spiciness Variability in the Subtropical-Tropical Pacific in the CESM2 Large Ensemble

Tropical Pacific decadal variations impact weather and climate around the world and are also connected to variations in the global warming trend. The mechanisms driving these long-term modulations, particularly the role of subsurface ocean dynamics, are still debated. Here, we investigate the dynamics of spiciness (density-compensated temperature and salinity) anomalies in the tropical and subtropical Pacific, which are hypothesized as a possible driving mechanism of decadal climate variability. Based on the analysis of 100 realizations from the Community Earth System Model Version 2 Large Ensemble (CESM2-LE), we demonstrate a coupling between the subtropics and the equatorial Pacific by propagating spiciness anomalies at decadal time scales. The CESM2-LE simulates spiciness variability along a subduction path from the subtropics to the equator with frequency spectra that show the highest power at low frequencies and a power decay proportional to a −4 slope for frequencies greater than 0.01 cycles per months, corresponding to periods smaller than ∼8.5 years. Signals that originate in the Southern Hemisphere (SH) dominate and arrive with a larger magnitude at the equator compared to spiciness anomalies from the Northern Hemisphere (NH). Spiciness anomalies from the SH have shorter propagation times and are strengthened along their pathway as stochastic wind stress curl forcing generates anomalous baroclinic ocean pressure gradients. These pressure gradients generate spiciness anomalies via anomalous advection across climatological spiciness gradients in the SH. We conclude that the observed spiciness variance at decadal time scales is consistent with a forcing by stochastic wind variations that are low-pass filtered by ocean dynamics.

54 ENVIRONMENTAL SCIENCES↗

Sizing and Topology Design of an Aeroelastic Wingbox Under Uncertainty

The goals of this work are to use a nested optimizer to conduct simultaneous sizing (inner level) and topology (outer level) design of a wingbox, considering uncertainties in the safety factors used to define the aeroelastic constraints. These uncertainties, propagated via sampling-driven polynomial chaos, are explicitly introduced at the inner level of the method, during gradient-based sizing optimization, resulting in a stochastic optimal sizing distribution. Measures of robustness in the total structural mass are then passed to the outer level, where a global optimizer evolves the topology parameters. The results demonstrate design choices needed to improve robustness in the face of uncertain safety factors, and the various physical mechanisms driving this process.

Stanford, Bret K.↗

A New Simple-to-Configure Self-Perturbing Multivariable Extremum-Seeking Controller

This paper presents a new stochastic relay-based extremum-seeking controller (ESC) for multi-input-single-output (MISO) systems. The algorithm was developed with the goal of simplifying configuration to enable easier deployment to real-world problems. A solution is developed first for a static map and then adapted for a general class of dynamic systems. The number of configurable parameters is one per input channel for the static case and only one additional parameter is needed for the dynamic version. The problem of gradient identifiability is solved via the use of stochastic relay gains and a simple stability proof for the static case is presented. Simulation tests demonstrate the performance of the strategy for optimizing both static and dynamic systems.

Salsbury, Timothy [BATTELLE (PACIFIC NW LAB)]↗

An effective description of charge diffusion and energy transport in a charged plasma from holography

We discuss the physics of sound propagation and charge diffusion in a plasma with non-vanishing charge density. Our analysis culminates the program initiated in to construct an open effective field theory of low-lying modes of the stress tensor and charge current in such plasmas. We model the plasma holographically as a Reissner-Nordström-AdS d+1 black hole, and study linearized fluctuations of longitudinally polarized scalar gravitons and photons in this background. We demonstrate that the perturbations can be decoupled and repackaged into the dynamics of two designer scalars, whose gravitational coupling is modulated by a non-trivial dilatonic factor. The holographic analysis allows us to isolate the phonon mode from the charge diffusion mode, and identify the combination of currents that corresponds to each of them. We use these results to obtain the real-time Gaussian effective action, which includes both the retarded response and the associated stochastic (Hawking) fluctuations, accurate to quartic order in gradients.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Parameter estimation for terrain modeling from gradient data

A method is developed for modeling terrain surfaces for use on an unmanned Martian roving vehicle. The modeling procedure employs a two-step process which uses gradient as well as height data in order to improve the accuracy of the model's gradient. Least square approximation is used in order to stochastically determine the parameters which describe the modeled surface. A complete error analysis of the modeling procedure is included which determines the effect of instrumental measurement errors on the model's accuracy. Computer simulation is used as a means of testing the entire modeling process which includes the acquisition of data points, the two-step modeling process and the error analysis. Finally, to illustrate the procedure, a numerical example is included.

Dangelo, K. R.↗