Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 109 records · Page 6

Sequential Learning of Active Subspaces

This study shows that in recent years, active subspace methods (ASMs) have become a popular means of performing subspace sensitivity analysis on black-box functions. Naively applied, however, ASMs require gradient evaluations of the target function. In the event of noisy, expensive, or stochastic simulators, evaluating gradients via finite differencing may be infeasible. In such cases, often a surrogate model is employed, on which finite differencing is performed. When the surrogate model is a Gaussian process (GP), we show that the ASM estimator is available in closed form, rendering the finite-difference approximation unnecessary. We use our closed-form solution to develop acquisition functions focused on sequential learning tailored to sensitivity analysis on top of ASMs. We also show that the traditional ASM estimator may be viewed as a method of moments estimator for a certain class of GPs. We demonstrate how uncertainty on GP hyperparameters may be propagated to uncertainty on the sensitivity analysis, allowing model-based confidence intervals on the active subspace. Our methodological developments are illustrated on several examples.

97 MATHEMATICS AND COMPUTING↗

Bias-Variance Trade-Off in Physics-Informed Neural Networks with Randomized Smoothing for High-Dimensional PDEs

Physics-Informed Neural Networks (PINNs) have triggered a paradigm shift in scientific computing, leveraging mesh-free properties and robust approximation capabilities. While proving effective for low-dimensional partial differential equations (PDEs), the computational cost of PINNs remains a hurdle in high-dimensional scenarios. This is particularly pronounced when computing high-order and high-dimensional derivatives in the physics-informed loss. Randomized Smoothing PINN (RS-PINN) introduces Gaussian noise for stochastic smoothing of the original neural net model, enabling the use of Monte Carlo methods for derivative approximation, which eliminates the need for costly automatic differentiation. Despite its computational efficiency, especially in the approximation of high-dimensional derivatives, RS-PINN introduces biases in both loss and gradients, negatively impacting convergence, especially when coupled with stochastic gradient descent (SGD) algorithms. We present a comprehensive analysis of biases in RS-PINN, attributing them to the nonlinearity of the Mean Squared Error (MSE) loss as well as the intrinsic nonlinearity of the PDE itself. We propose tailored bias correction techniques, delineating their application based on the order of PDE nonlinearity. The derivation of an unbiased RS-PINN allows for a detailed examination of its advantages and disadvantages compared to the biased version. Specifically, the biased version has a lower variance and runs faster than the unbiased version, but it is less accurate due to the bias. To optimize the bias-variance trade-off, we combine the two approaches in a hybrid method that balances the rapid convergence of the biased version with the high accuracy of the unbiased version. In addition to methodological contributions, we present an enhanced implementation of RS-PINN. Extensive experiments on diverse high-dimensional PDEs, including Fokker-Planck, Hamilton-Jacobi-Bellman (HJB), viscous Burgers’, Allen-Cahn, and Sine-Gordon equations, illustrate the bias-variance trade-off and highlight the effectiveness of the hybrid RS-PINN. Empirical guidelines are provided for selecting biased, unbiased, or hybrid versions, depending on the dimensionality and nonlinearity of the specific PDE problem.

97 MATHEMATICS AND COMPUTING↗

Stochastic noise can be helpful for variational quantum algorithms

Saddle points constitute a crucial challenge for first-order gradient descent algorithms. In notions of classical machine learning, they are avoided, for example, by means of stochastic gradient descent methods. In this work, we provide evidence that the saddle-points problem can be naturally avoided in variational quantum algorithms by exploiting the presence of stochasticity. We prove convergence guarantees and present practical examples in numerical simulations and on quantum hardware. We argue that the natural stochasticity of variational algorithms can be beneficial for avoiding strict saddle points, i.e., those saddle points with at least one negative Hessian eigenvalue. This insight that some levels of shot noise could help is expected to add a new perspective to notions of near-term variational quantum algorithms. Published by the American Physical Society 2025

Liu, Junyu↗

AADL: Anderson Accelerated Deep Learning

We propose a stable, distributed approach to perform AA that accelerates the convergence rate of stochastic first-order optimizers to train neural networks. Differently from previous works, we do not alter neither the scheme to perform AA nor the loss function minimized during the training. To improve robustness against stagnation, we customize general guidelines that suggest to relax the frequency of AA corrections by performing AA only at the end of an entire training epoch. To improve robustness of AA against the stochastic oscillations of first-order optimizers, we average the gradients computed on consecutive stochastic optimization updates. The improved regularity of the converging sequence and the reduced amplitude of stochastic oscillations across consecutive optimization steps allows AA to efficiently extrapolate an improved converging sequence, thereby overcoming limitations of existing approaches to perform AA on stochastic optimization.

Lupo Pasini, Massimiliano [Oak Ridge National Lab.↗

DPM: A deep learning PDE augmentation method with application to large-eddy simulation

A framework is introduced that leverages known physics to reduce overfitting in machine learning for scientific applications. The partial differential equation (PDE) that expresses the physics is augmented with a neural network that uses available data to learn a description of the corresponding unknown or unrepresented physics. Training within this combined system corrects for missing, unknown, or erroneously represented physics, including discretization errors associated with the PDE's numerical solution. For optimization of the network within the PDE, an adjoint PDE is solved to provide high-dimensional gradients, and a stochastic adjoint method (SAM) further accelerates training. Additionally, the approach is demonstrated for large-eddy simulation (LES) of turbulence. High-fidelity direct numerical simulations (DNS) of decaying isotropic turbulence provide the training data used to learn sub-filter-scale closures for the filtered Navier–Stokes equations. Out-of-sample comparisons show that the deep learning PDE method outperforms widely-used models, even for filter sizes so large that they become qualitatively incorrect. It also significantly outperforms the same neural network when a priori trained based on simple data mismatch, not accounting for the full PDE. Measures of discretization errors, which are well-known to be consequential in LES, point to the importance of the unified training formulation's design, which without modification corrects for them. For comparable accuracy, simulation runtime is significantly reduced. A relaxation of the typical discrete enforcement of the divergence-free constraint in the solver is also successful, instead allowing the DPM to approximately enforce incompressibility physics. Since the training loss function is not restricted to correspond directly to the closure to be learned, training can incorporate diverse data, including experimental data.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

An adaptive stochastic sequential quadratic programming with differentiable exact augmented lagrangians

In this study, we consider solving nonlinear optimization problems with a stochastic objective and deterministic equality constraints. We assume for the objective that its evaluation, gradient, and Hessian are inaccessible, while one can compute their stochastic estimates by, for example, subsampling. We propose a stochastic algorithm based on sequential quadratic programming (SQP) that uses a differentiable exact augmented Lagrangian as the merit function. To motivate our algorithm design, we first revisit and simplify an old SQP method Lucidi developed for solving deterministic problems, which serves as the skeleton of our stochastic algorithm. Based on the simplified deterministic algorithm, we then propose a non-adaptive SQP for dealing with stochastic objective, where the gradient and Hessian are replaced by stochastic estimates but the stepsizes are deterministic and prespecified. Finally, we incorporate a recent stochastic line search procedure Paquette and Scheinberg into the non-adaptive stochastic SQP to adaptively select the random stepsizes, which leads to an adaptive stochastic SQP. The global "almost sure" convergence for both non-adaptive and adaptive SQP methods is established. Numerical experiments on nonlinear problems in CUTEst test set demonstrate the superiority of the adaptive algorithm.

97 MATHEMATICS AND COMPUTING↗

Quantum optimization algorithms: Energetic implications

Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. Here in this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least 2x to 4x more energy efficient than their classical counterparts; why feedback-based quantum optimization is energy-inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of ≥2 0 x. Finally, we use the NchooseK high-level programming model to run optimization problems on both gate-based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate-based counterparts.

97 MATHEMATICS AND COMPUTING↗

Random coordinate descent: A simple alternative for optimizing parameterized quantum circuits

Variational quantum algorithms rely on the optimization of parameterized quantum circuits in noisy settings. The commonly used back-propagation procedure in classical machine learning is not directly applicable in this setting due to the collapse of quantum states after measurements. Thus, gradient estimations constitute a significant overhead in a gradient-based optimization of such quantum circuits. This paper introduces a random coordinate descent algorithm as a practical and easy-to-implement alternative to the full gradient descent algorithm. This algorithm only requires one partial derivative at each iteration. Motivated by the behavior of measurement noise in the practical optimization of parameterized quantum circuits, this paper presents an optimization problem setting that is amenable to analysis. Under this setting, the random coordinate descent algorithm exhibits the same level of stochastic stability as the full gradient approach, making it as resilient to noise. The complexity of the random coordinate descent method is generally no worse than that of the gradient descent and can be much better for various quantum optimization problems with anisotropic Lipschitz constants. Theoretical analysis and extensive numerical experiments validate our findings. Published by the American Physical Society 2024

Ding, Zhiyan (ORCID:000000018863403X)↗

Nonlinear Matrix Approximation with Radial Basis Function Components

We introduce and investigate matrix approximation by decomposition into a sum of radial basis function (RBF) components. An RBF component is a generalization of the outer product between a pair of vectors, where an RBF function replaces the scalar multiplication between individual vector elements. Even though the RBF functions are positive definite, the summation across components is not restricted to convex combinations and allows us to compute the decomposition for any real matrix that is not necessarily symmetric or positive definite. We formulate the problem of seeking such a decomposition as an optimization problem with a nonlinear and non-convex loss function. Several modern versions of the gradient descent method, including their scalable stochastic counterparts, are used to solve this problem. We provide extensive empirical evidence of the effectiveness of the RBF decomposition and that of the gradient-based fitting algorithm. While being conceptually motivated by singular value decomposition (SVD), our proposed nonlinear counterpart outperforms SVD by drastically reducing the memory required to approximate a data matrix with the same L2 error for a wide range of matrix types. For example, it leads to 2 to 6 times memory save for Gaussian noise, graph adjacency matrices, and kernel matrices. Moreover, this proximity-based decomposition can offer additional interpretability in applications that involve, e.g., capturing the inner low-dimensional structure of the data, retaining graph connectivity structure, and preserving the acutance of images.

Rebrova, Elizaveta↗

Machine learning unifies flexibility and efficiency of spinodal structure generation for stochastic biomaterial design

Abstract Porous biomaterials design for bone repair is still largely limited to regular structures (e.g. rod-based lattices), due to their easy parameterization and high controllability. The capability of designing stochastic structure can redefine the boundary of our explorable structure–property space for synthesizing next-generation biomaterials. We hereby propose a convolutional neural network (CNN) approach for efficient generation and design of spinodal structure—an intriguing structure with stochastic yet interconnected, smooth, and constant pore channel conducive to bio-transport. Our CNN-based approach simultaneously possesses the tremendous flexibility of physics-based model in generating various spinodal structures (e.g. periodic, anisotropic, gradient, and arbitrarily large ones) and comparable computational efficiency to mathematical approximation model. We thus successfully design spinodal bone structures with target anisotropic elasticity via high-throughput screening, and directly generate large spinodal orthopedic implants with desired gradient porosity. This work significantly advances stochastic biomaterials development by offering an optimal solution to spinodal structure generation and design.

59 BASIC BIOLOGICAL SCIENCES↗

Plateau Phenomenon in Gradient Descent Training of RELU Networks: Explanation, Quantification, and Avoidance

The ability of neural networks to provide ‘best in class’ approximation across a wide range of applications is well-documented. Nevertheless, the powerful expressivity of neural networks comes to naught if one is unable to effectively train (choose) the parameters defining the network. In general, neural networks are trained by gradient descent type optimization methods,a stochastic variant thereof. In practice, such methods result in the loss function decreases rapidly at the beginning of training but then, after a relatively small number of steps, significantly slow down. The loss may even appear to stagnate over the period of a large number of epochs, only to then suddenly start to decrease fast again for no apparent reason. This so-called plateau phenomenon manifests itself in many learning tasks. The present work aims to identify and quantify the root causes of plateau phenomenon.analysis is carried out in the setting of univariate ReLU networks. No assumptions are made on the number of neurons relative to the number of training data, and our results hold for both the lazy and adaptive regimes. Here, the main findings are: plateaux correspond to periods during which activation patterns remain constant, where activation pattern refers to the number of data points that activate a given neuron; quantification of convergence of the gradient flow dynamics; and, characterization stationary points in terms solutions of local least squares regression lines over subsets of the training data. Based on these conclusions, we propose a new iterative training method, the Active Neuron Least Squares (ANLS), characterised by the explicit adjustment of the activation pattern at each step, which is designed to enable a quick exit from a plateau. Illustrative numerical examples are included throughout.

97 MATHEMATICS AND COMPUTING↗

Correspondence between neuroevolution and gradient descent

Abstract We show analytically that training a neural network by conditioned stochastic mutation or neuroevolution of its weights is equivalent, in the limit of small mutations, to gradient descent on the loss function in the presence of Gaussian white noise. Averaged over independent realizations of the learning process, neuroevolution is equivalent to gradient descent on the loss function. We use numerical simulation to show that this correspondence can be observed for finite mutations, for shallow and deep neural networks. Our results provide a connection between two families of neural-network training methods that are usually considered to be fundamentally different.

97 MATHEMATICS AND COMPUTING↗

An Online Dynamic Amplitude-Correcting Gradient Estimation Technique to Align X-ray Focusing Optics

High-brightness X-ray pulses, as generated at synchrotrons and X-ray free electron lasers (XFELs), are used in a variety of scientific experiments. At these facilities, measurements often require optical equipment, e.g Compound Refractive Lenses (CRLs) to be precisely aligned and focused. The lateral alignment of CRLs to a beamline requires precise positioning along four axes: two translational, and the two rotational. At a synchrotron, alignment is often accomplished manually. However, XFEL beamlines present a beam brightness that fluctuates stochastically, making manual alignment a time-consuming endeavor. Automation using simplex or classic stochastic descent often fails, given the errant gradient estimates. Herein we present a dynamic-amplitude correction to the usual gradient based on the combination of a generalized finite difference stencil and a time-dependent sampling pattern. Intensity is recorded periodically, then used to normalize numerical derivatives against fluctuations. Error expectation is analyzed, and efficacy is demonstrated on classic benchmarks. We provide a proof of concept by laterally aligning optics on a simulated XFEL beamline using data recorded at both synchrotron and XFEL facilities.

97 MATHEMATICS AND COMPUTING↗

Stochastic projective splitting

Here, we present a new, stochastic variant of the projective splitting (PS) family of algorithms for inclusion problems involving the sum of any finite number of maximal monotone operators. This new variant uses a stochastic oracle to evaluate one of the operators, which is assumed to be Lipschitz continuous, and (deterministic) resolvents to process the remaining operators. Our proposal is the first version of PS with such stochastic capabilities. We envision the primary application being machine learning (ML) problems, with the method’s stochastic features facilitating “mini-batch” sampling of datasets. Since it uses a monotone operator formulation, the method can handle not only Lipschitz-smooth loss minimization, but also min–max and noncooperative game formulations, with better convergence properties than the gradient descent-ascent methods commonly applied in such settings. The proposed method can handle any number of constraints and nonsmooth regularizers via projection and proximal operators. We prove almost-sure convergence of the iterates to a solution and a convergence rate result for the expected residual, and close with numerical experiments on a distributionally robust sparse logistic regression problem.

97 MATHEMATICS AND COMPUTING↗

Decadal Spiciness Variability in the Subtropical-Tropical Pacific in the CESM2 Large Ensemble

Tropical Pacific decadal variations impact weather and climate around the world and are also connected to variations in the global warming trend. The mechanisms driving these long-term modulations, particularly the role of subsurface ocean dynamics, are still debated. Here, we investigate the dynamics of spiciness (density-compensated temperature and salinity) anomalies in the tropical and subtropical Pacific, which are hypothesized as a possible driving mechanism of decadal climate variability. Based on the analysis of 100 realizations from the Community Earth System Model Version 2 Large Ensemble (CESM2-LE), we demonstrate a coupling between the subtropics and the equatorial Pacific by propagating spiciness anomalies at decadal time scales. The CESM2-LE simulates spiciness variability along a subduction path from the subtropics to the equator with frequency spectra that show the highest power at low frequencies and a power decay proportional to a −4 slope for frequencies greater than 0.01 cycles per months, corresponding to periods smaller than ∼8.5 years. Signals that originate in the Southern Hemisphere (SH) dominate and arrive with a larger magnitude at the equator compared to spiciness anomalies from the Northern Hemisphere (NH). Spiciness anomalies from the SH have shorter propagation times and are strengthened along their pathway as stochastic wind stress curl forcing generates anomalous baroclinic ocean pressure gradients. These pressure gradients generate spiciness anomalies via anomalous advection across climatological spiciness gradients in the SH. We conclude that the observed spiciness variance at decadal time scales is consistent with a forcing by stochastic wind variations that are low-pass filtered by ocean dynamics.

54 ENVIRONMENTAL SCIENCES↗

A New Simple-to-Configure Self-Perturbing Multivariable Extremum-Seeking Controller

This paper presents a new stochastic relay-based extremum-seeking controller (ESC) for multi-input-single-output (MISO) systems. The algorithm was developed with the goal of simplifying configuration to enable easier deployment to real-world problems. A solution is developed first for a static map and then adapted for a general class of dynamic systems. The number of configurable parameters is one per input channel for the static case and only one additional parameter is needed for the dynamic version. The problem of gradient identifiability is solved via the use of stochastic relay gains and a simple stability proof for the static case is presented. Simulation tests demonstrate the performance of the strategy for optimizing both static and dynamic systems.

Salsbury, Timothy [BATTELLE (PACIFIC NW LAB)]↗

An effective description of charge diffusion and energy transport in a charged plasma from holography

We discuss the physics of sound propagation and charge diffusion in a plasma with non-vanishing charge density. Our analysis culminates the program initiated in to construct an open effective field theory of low-lying modes of the stress tensor and charge current in such plasmas. We model the plasma holographically as a Reissner-Nordström-AdS d+1 black hole, and study linearized fluctuations of longitudinally polarized scalar gravitons and photons in this background. We demonstrate that the perturbations can be decoupled and repackaged into the dynamics of two designer scalars, whose gravitational coupling is modulated by a non-trivial dilatonic factor. The holographic analysis allows us to isolate the phonon mode from the charge diffusion mode, and identify the combination of currents that corresponds to each of them. We use these results to obtain the real-time Gaussian effective action, which includes both the retarded response and the associated stochastic (Hawking) fluctuations, accurate to quartic order in gradients.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Capturing the softening in T91 steel annealed with molten LBE through interfacial gradient plasticity

In order to study the effect of annealing with molten lead–bismuth eutectic alloy (mLBE) on the mechanical properties of T91 steel, micropillar compression tests were analyzed through the consideration of mechanical interface energy terms within gradient plasticity. The stress–strain curves of the micropillars showed significant stochastic effects, characterized by a different elastic and plastic behavior response for each pillar. Among the various pillars, some exhibited a similar trilinear behavior. Scanning electron microscopy (SEM) images of these specimens demonstrated localized slip deformation and slip planes after compression, while transmission electron microscopy (TEM) images showed the presence of a severe slip zone along the random grain boundaries (RGBs), indicating that the GBs played a dominant role in the deformation/slip. By employing interfacial gradient plasticity that can explicitly account for the presence of GBs, the trilinear response was captured, allowing for the determination of the mechanical interface parameter. As anticipated the pillars which underwent severe slip at the GBs had a lower value for the mechanical interface parameter. Finally, annealing with mLBE can, therefore, in some cases result in softening in the overall stress–strain, as it lowers the mechanical interface energy of GBs.

36 MATERIALS SCIENCE↗