Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Stochastic Gradient Descent”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

GentenMPI: Distributed Memory Sparse Tensor Decomposition

GentenMPl is a toolkit of sparse canonical polyadic (CP) tensor decomposition algorithms that is designed to run effectively on distributed-memory high-performance computers. Its use of distributed-memory parallelism enables it to efficiently decompose tensors that are too large for a single compute node's memory. GentenMPl leverages Sandia's decades-long investment in the Trilinos solver framework for much of its parallel-computation capability. Trilinos contains numerical algorithms and linear algebra classes that have been optimized for parallel simulation of complex physical phenomena. This work applies these tools to the data science problem of sparse tensor decomposition. In this report, we describe the use of Trilinos in GentenMPl, extensions needed for sparse tensor decomposition, and implementations of the CP-ALS (CP via alternating least squares) and GCP-SGD (generalized CP via stochastic gradient descent) sparse tensor decomposition algorithms. We show that GentenMPl can decompose sparse tensors of extreme size, e.g., a 12.6-terabyte tensor on 8192 computer cores. We demonstrate that the Trilinos backbone provides good strong and weak scaling of the tensor decomposition algorithms.

97 MATHEMATICS AND COMPUTING↗

End-to-End Differentiable Modeling and Management of the Environment

Focal Area: (2) Data acquisition and assimilation enabled by machine learning, AI, and advanced methods including experimental/network design/optimization. We emphasize the importance of leveraging optimization techniques from AI/machine learning (ML) to solve challenging problems in Earth system modeling. Science Challenge: Automatic differentiation has had a transformative effect on ML by allowing the calculation of gradients of arbitrary functions in an incredibly large class of models. We can potentially realize similar improvements in parameter estimation and control for Earth system models (ESMs) by reimplementing them in computational frameworks from ML. Practitioners working with large (>10 7 parameters) models in ML and AI can obtain good predictive performance in a range of spatiotemporal tasks by making use of optimization via stochastic gradient descent and incorporating prior knowledge at multiple levels. We propose writing ESMs in open-source computational frameworks such as Torch, Tensorflow, and JAX to greatly expand the scope of environmental forecasting and management challenges, which can be addressed by leveraging automatic differentiation and gradient descent-like algorithms. We do not call for a wholesale replacement of physical models with data-driven surrogates, but rather advocate for interleaving physical and empirical equations in a manner that is most faithful to the extent of our scientific knowledge and observational data. Central to this topic is the merging of differentiable physical simulations with differentiable optimization layers, which are now both beginning to come to the forefront.

54 ENVIRONMENTAL SCIENCES↗

Randomized Algorithms for Scientific Computing (RASC)

Randomized algorithms have propelled advances in artificial intelligence (AI) and represent a foundational research area in advancing AI for Science. Future advancements in DOE Office of Science priority areas such as climate science, astrophysics, fusion, advanced materials, combustion, and quantum computing all require randomized algorithms for surmounting challenges of complexity, robustness, and scalability. Advances in data collection and numerical simulation have changed the dynamics of scientific research and motivate the need for randomized algorithms. For instance, advances in imaging technologies such as X-ray ptychography, electron microscopy, electron energy loss spectroscopy, or adaptive optics lattice light-sheet microscopy collect hyperspectral imaging and scattering data in terabytes, at breakneck speed enabled by state-of-the-art detectors. The data collection is exceptionally fast compared with its analysis. Likewise, advances in high-performance architectures have made exascale computing a reality and changed the economies of scientific computing in the process. Floating-point operations that create data are essentially free in comparison with data movement. Thus far, most approaches have focused on creating faster hardware. Ironically, this faster hardware has exacerbated the problem by making data still easier to create. Under such an onslaught, scientists often resort to heuristic deterministic sampling schemes (e.g., low-precision arithmetic, sampling every nth element) and sacrifice potentially valuable accuracy. Dramatically better results can be achieved via randomized algorithms, reducing the data size as much as or more than naive deterministic subsampling can achieve, while retaining the high accuracy of computing on the full data set. By randomized algorithms we mean those algorithms that employ some form of randomness in internal algorithmic decisions to accelerate time to solution, increase scalability, or improve reliability. Examples include matrix sketching for solving large-scale least-squares problems (see Figure 1) and stochastic gradient descent for training machine learning models. We are not recommending heuristic methods but rather randomized algorithms that have certificates of correctness and probabilistic guarantees of optimality and near-optimality. Such approaches can be useful beyond acceleration, for example, in understanding how to avoid measure zero worst-case scenarios that plague methods such as QR matrix factorization.

97 MATHEMATICS AND COMPUTING↗

Streaming Generalized Canonical Polyadic Tensor Decompositions

In this paper, we develop a method which we call OnlineGCP for computing the Generalized Canonical Polyadic (GCP) tensor decomposition of streaming data. GCP differs from traditional canonical polyadic (CP) tensor decompositions as it allows for arbitrary objective functions which the CP model attempts to minimize. This approach can provide better fits and more interpretable models when the observed tensor data is strongly non-Gaussian. In the streaming case, tensor data is gradually observed over time and the algorithm must incrementally update a GCP factorization with limited access to prior data. In this work, we extend the GCP formalism to the streaming context by deriving a GCP optimization problem to be solved as new tensor data is observed, formulate a tunable history term to balance reconstruction of recently observed data with data observed in the past, develop a scalable solution strategy based on segregated solves using stochastic gradient descent methods, describe a software implementation that provides performance and portability to contemporary CPU and GPU architectures and integrates with Matlab for enhanced usability, and demonstrate the utility and performance of the approach and software on several synthetic and real tensor data sets.

97 MATHEMATICS AND COMPUTING↗

Scalable and Energy-Efficient Methods for Interactive Exploration of Scientific Data

The main scientific contributions of this project are the following novel concepts for multidimensional arrays: shape-based similarity join (SIGMOD 2016), incremental view maintenance (SIGMOD 2017), user-defined stencil functions (HPDC 2017), and distributed caching for in-situ processing (SSDBM 2018). Building on our collaboration with the astrophysics group at LBNL, we applied these techniques to the data generated in the Palomar Transient Factory (PTF) astronomical survey. They played a pivotal role in the first-ever observation of a neutron star merger, which produces gravitational waves and turns out to be the origin of heavy elements, including gold. This has lead to a Science magazine article that has received extensive media coverage on ACM TechNews, Slashdot, FiveThirtyEight, and Quanta Magazine, among others. Additionally, two other articles detailing related aspects of the same discovery have been published in the Astrophysical Journal Letters journal. These publications have more than 3,000 citations according to Google Scholar (as of February 2022). This cross-disciplinary collaboration provided very good opportunities to apply database techniques to real-life scientific problems. The fact that they facilitated major discoveries in astrophysics proves the importance of our research. In addition to the work on multidimensional array databases, this project has also developed stochastic gradient descent (SGD) optimization algorithms for training large scale machine learning models, methods for querying in-situ data, and a database query optimizer based on sketch synopses.

79 ASTRONOMY AND ASTROPHYSICS↗

Scalable Second Order Optimization for Machine Learning

Many machine learning (ML) training tasks are essentially optimization processes that would at first glance appear eminently parallelizable and scalable. However, effective acceleration of these tasks with scalable parallel hardware has proven to be elusive. While standard methods for machine learning, e.g., stochastic gradient descent (SGD) for DNNs, tend to be resource efficient, they appear to be fundamentally sequential in nature.

97 MATHEMATICS AND COMPUTING↗

An Adaptive Optimizer for Measurement-Frugal Variational Algorithms

Variational hybrid quantum-classical algorithms (VHQCAs) have the potential to be useful in the era of near-term quantum computing. However, recently there has been concern regarding the number of measurements needed for convergence of VHQCAs. Here, we address this concern by investigating the classical optimizer in VHQCAs. We introduce a novel optimizer called individual Coupled Adaptive Number of Shots (iCANS). This adaptive optimizer frugally selects the number of measurements (i.e., number of shots) both for a given iteration and for a given partial derivative in a stochastic gradient descent. We numerically simulate the performance of iCANS for the variational quantum eigensolver and for variational quantum compiling, with and without noise. In all cases, and especially in the noisy case, iCANS tends to out-perform state-of-the-art optimizers for VHQCAs. We therefore believe this adaptive optimizer will be useful for realistic VHQCA implementations, where the number of measurements is limited.

97 MATHEMATICS AND COMPUTING↗

Reinforcement Learning-based Output Structured Feedback for Distributed Multi-Area Power System Frequency Control

Load frequency control (LFC) is a key factor to maintain the stable frequency in multi-area power systems. As the modern power systems evolve from centralized to decentralized paradigm, LFC needs to consider the decentralized scheme that considers limited information from the information-exchange graph for the generator control of each interconnected area. This paper aims to solve a data-driven constrained LQR problem with mean-variance risk constraints and output structured feedback, and applies this framework to solve the LFC problem in multi-area power systems. By reformulating the constrained optimization problem into a minimax problem, the stochastic gradient descent max-oracle (SGDmax) algorithm with zero-order policy gradient (ZOPG) is adopted to find the optimal feedback gain from the learning, while guaranteeing the convergence. In addition, to improve the adaptation of the proposed learning method to new or varying models, we construct an emulator grid that approximates the dynamics of a physical grid and performs training based on this model. Once the feedback gain is obtained from the emulator grid, it is applied to the physical grid with a robustness test to check whether the controller from the approximated emulator applies to the actual system. Numerical tests show that the obtained feedback controller can successfully control the frequency of each area, while mitigating the uncertainty from the loads, with reliable robustness that ensures the adaptability of the obtained feedback gain to the actual physical grid.

Kwon, Kyung-bin↗

Hutchinson Trace Estimation for high-dimensional and high-order Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) have proven effective in solving partial differential equations (PDEs), especially when some data are available by seamlessly blending data and physics. However, extending PINNs to high-dimensional and even high-order PDEs encounters significant challenges due to the computational cost associated with automatic differentiation in the residual loss function calculation. Herein, we address the limitations of PINNs in handling high-dimensional and high-order PDEs by introducing the Hutchinson Trace Estimation (HTE) method. Starting with the second-order high-dimensional PDEs, which are ubiquitous in scientific computing, HTE is applied to transform the calculation of the entire Hessian matrix into a Hessian vector product (HVP). This approach not only alleviates the computational bottleneck via Taylor-mode automatic differentiation but also significantly reduces memory consumption from the Hessian matrix to an HVP’s scalar output. We further showcase HTE’s convergence to the original PINN loss and its unbiased behavior under specific conditions. Comparisons with the Stochastic Dimension Gradient Descent (SDGD) highlight the distinct advantages of HTE, particularly in scenarios with significant variability and variance among dimensions. We further extend the application of HTE to higher-order and higher-dimensional PDEs, specifically addressing the biharmonic equation. By employing tensor-vector products (TVP), HTE efficiently computes the colossal tensor associated with the fourth-order high-dimensional biharmonic equation, saving memory and enabling rapid computation. The effectiveness of HTE is illustrated through experimental setups, demonstrating comparable convergence rates with SDGD under memory and speed constraints. Additionally, HTE proves valuable in accelerating the Gradient-Enhanced PINN (gPINN) version as well as the Biharmonic equation. Overall, HTE opens up a new capability in scientific machine learning for tackling high-order and high-dimensional PDEs.

Curse of dimensionality↗

Formulation and solution approach for calibrating activity-based travel demand model-system via microsimulation

This study addresses the problem of calibrating utility-maximizing nested logit activity-based travel demand model-systems. After estimation, it is common practice to use aggregate measurements to calibrate the estimated model-system’s parameters prior to their application in transportation planning, policy making, and operations. However, calibration of activity-based model-systems has received much less attention. Existing calibration approaches are myopic heuristics in the sense that they do not consider the fundamental inter-dependencies among choice-models and do not have a systematic way to adjust model parameters. Also, other purely simulation-based approaches do not perform well in large-scale applications. In this study, we focus on utility-maximizing nested logit activity-based model-systems and calibrating aggregate statistics such as activity shares, mode shares, time-dependent & mode-specific OD flows, and time-dependent & mode-specific sensor counts. We formulate the calibration problem as a simulation-based optimization problem and propose a stochastic gradient-based solution procedure to solve it. The solution procedure relies on microsimulation to calculate expectations of the aggregate statistics of interest to the calibration problem. Additionally, we derive approximate analytical expressions for the gradient of the objective function —that are evaluated through microsimulation on mini-batches of the population. The proposed solution procedure is sensitive to the fundamental structure of the activity-based model-system and is non-myopic in considering the dependencies across its model components. The formulated optimization problem is non-convex, highly nonlinear, and potentially has multiple-minima. Lastly, we show —through a real-world application— that the proposed solution procedure outperforms other state-of-the-art purely simulation-based optimization approaches in terms of computational efficiency, stability, and convergence. We also compare various gradient-based solution algorithms to determine the best algorithm to update the parameters. This work has the potential to facilitate wider and easier application of activity-based model-systems.

97 MATHEMATICS AND COMPUTING↗

Sentiment Analysis based Error Detection for Large-Scale Systems

Today's large-scale systems such as High Performance Computing (HPC) Systems are designed/utilized towards exascale computing, inevitably decreasing its reliability due to the increasing design complexity. HPC systems conduct extensive logging of their execution behaviour. In this paper, we leverage the inherent meaning behind the log messages and propose a novel sentiment analysis-based approach for the error detection in large-scale systems, by automatically mining the sentiments in the log messages. Our contributions are four-fold. (1) We develop a machine learning (ML) based approach to automatically build a sentiment lexicon, based on the system log message templates. (2) Using the sentiment lexicon, we develop an algorithm to detect system errors. (3) We develop an algorithm to identify the nodes and components with erroneous behaviors, based on sentiment polarity scores. (4) We evaluate our solution vs. other state-of-the-art machine/deep learning algorithms based on three representative supercomputers' system logs. Experiments show that our error detection algorithm can identify error messages with an average MCC score and f -score of 91% and 96% respectively, while state of the art ML/deep learning model (LSTM) obtains only 67% and 84%. To the best of our knowledge, this is the first work leveraging the sentiments embedded in log entries of large-scale systems for system health analysis.

error detection↗

Risk-Constrained Reinforcement Learning for Inverter-Dominated Power System Controls

Here, this paper develops a risk-aware controller for grid-forming inverters (GFMs) to minimize large frequency oscillations in GFM inverter-dominated power systems. To tackle the high variability from loads/renewables, we incorporate a mean-variance risk constraint into the classical linear quadratic regulator (LQR) formulation for this problem. The risk constraint aims to bound the time-averaged cost of state variability and thus can improve the worst-case performance for large disturbances. The resulting risk-constrained LQR problem is solved through the dual reformulation to a minimax problem, by using a reinforcement learning (RL) method termed as stochastic gradient-descent with max-oracle (SGDmax). In particular, the zero-order policy gradient (ZOPG) approach is used to simplify the gradient estimation using simulated system trajectories. Numerical tests conducted on the IEEE 68-bus system have validated the convergence of our proposed SGDmax for GFM model and corroborate the effectiveness of the risk constraint in improving the worst-case performance while reducing the variability of the overall control cost.

Frequency control↗

Wavefront shaping with a Hadamard basis for scattering soil imaging

Here, soil is a scattering medium that inhibits imaging of plant-microbial-mineral interactions that are essential to plant health and soil carbon sequestration. However, optical imaging in the complex medium of soil has been stymied by the seemingly intractable problems of scattering and contrast. Here, we develop a wavefront shaping method based on adaptive stochastic parallel gradient descent optimization with a Hadamard basis to focus light through soil mineral samples. Our approach allows a sparse representation of the wavefront with reduced dimensionality for the optimization. We further divide the used Hadamard basis set into subsets and optimize a certain subset at once. Simulation and experimental optimization results demonstrate our method has an approximately seven times higher convergence rate and overall better performance compared to that with optimizing all pixels at once. The proposed method can benefit other high-dimensional optimization problems in adaptive optics and wavefront shaping.

47 OTHER INSTRUMENTATION↗

Stabilization of the 81-channel coherent beam combination using machine learning

We develop a rapidly converging algorithm for stabilizing a large channel-count diffractive optical coherent beam combination. An 81-beam combiner is controlled by a novel, machine-learning based, iterative method to correct the optical phases, operating on an experimentally calibrated numerical model. A neural-network is trained to detect phase errors based on interference pattern recognition of uncombined beams adjacent to the combined one. Due to the non-uniqueness of solutions in the full space of possible phases, the network is trained within a limited phase perturbation/error range. This also reduces the number of samples needed for training. Simulations have proven that the network can converge in one step for small phase perturbations. When the trained neural-network is applied to a realistic case of 360 degree full range, an iterative scheme exploits random walking at the beginning, with the accuracy of prediction on phase feedback direction, to allow the neural-network to step into the training range for fast convergence. This neural-network-based iterative method of phase detection works tens of times faster than the commonly used stochastic parallel gradient descent approach (SPGD) using a single-detector and random dither when both are tested with random phase perturbations.

Wang, Dan↗

Stochastic noise can be helpful for variational quantum algorithms

Saddle points constitute a crucial challenge for first-order gradient descent algorithms. In notions of classical machine learning, they are avoided, for example, by means of stochastic gradient descent methods. In this work, we provide evidence that the saddle-points problem can be naturally avoided in variational quantum algorithms by exploiting the presence of stochasticity. We prove convergence guarantees and present practical examples in numerical simulations and on quantum hardware. We argue that the natural stochasticity of variational algorithms can be beneficial for avoiding strict saddle points, i.e., those saddle points with at least one negative Hessian eigenvalue. This insight that some levels of shot noise could help is expected to add a new perspective to notions of near-term variational quantum algorithms. Published by the American Physical Society 2025

Liu, Junyu↗

Quantum optimization algorithms: Energetic implications

Since the dawn of quantum computing (QC), theoretical developments like Shor's algorithm proved the conceptual superiority of QC over traditional computing. However, such quantum supremacy claims are difficult to achieve in practice because of the technical challenges of realizing noiseless qubits. In the near future, QC applications will need to rely on noisy quantum devices that offload part of their work to classical devices. One way to achieve this is by using parameterized quantum circuits in optimization or even in machine learning tasks. The energy requirements of quantum algorithms have not yet been studied extensively. Here in this article, we explore several optimization algorithms using both theoretical insights and numerical experiments to understand their impact on energy consumption. Specifically, we highlight why and how algorithms like quantum natural gradient descent, simultaneous perturbation stochastic approximations or circuit learning methods, are at least 2x to 4x more energy efficient than their classical counterparts; why feedback-based quantum optimization is energy-inefficient; and how techniques like Rosalin can improve the energy efficiency of other algorithms by a factor of ≥2 0 x. Finally, we use the NchooseK high-level programming model to run optimization problems on both gate-based quantum computers and quantum annealers. Empirical data indicate that these optimization problems run faster, have better success rates, and consume less energy on quantum annealers than on their gate-based counterparts.

97 MATHEMATICS AND COMPUTING↗

Correspondence between neuroevolution and gradient descent

Abstract We show analytically that training a neural network by conditioned stochastic mutation or neuroevolution of its weights is equivalent, in the limit of small mutations, to gradient descent on the loss function in the presence of Gaussian white noise. Averaged over independent realizations of the learning process, neuroevolution is equivalent to gradient descent on the loss function. We use numerical simulation to show that this correspondence can be observed for finite mutations, for shallow and deep neural networks. Our results provide a connection between two families of neural-network training methods that are usually considered to be fundamentally different.

97 MATHEMATICS AND COMPUTING↗