Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “stochastic gradient”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Adaptive sampling quasi-Newton methods for zeroth-order stochastic optimization

Here, we consider unconstrained stochastic optimization problems with no available gradient information. Such problems arise in settings from derivative-free simulation optimization to reinforcement learning. We propose an adaptive sampling quasi-Newton method where we estimate the gradients using finite differences of stochastic function evaluations within a common random number framework. We develop modified versions of a norm test and an inner product quasi-Newton test to control the sample sizes used in the stochastic approximations and provide global convergence results to the neighborhood of a locally optimal solution. We present numerical experiments on simulation optimization problems to illustrate the performance of the proposed algorithm. When compared with classical zeroth-order stochastic gradient methods, we observe that our strategies of adapting the sample sizes significantly improve performance in terms of the number of stochastic function evaluations required.

97 MATHEMATICS AND COMPUTING↗

Stochastic properties of ultralight scalar field gradients

Ultralight axion-like particles are well-motivated dark matter candidates that are the target of numerous direct detection efforts. In the vicinity of the Solar System, such particles can be treated as oscillating scalar fields. The velocity dispersion of the Milky Way determines a coherence time of about 10 6 oscillations, beyond which the amplitude of the axion field fluctuates stochastically. Any analysis of data from an axion direct detection experiment must carefully account for this stochastic behavior to properly interpret the results. This is especially true for experiments sensitive to the gradient of the axion field that are unable to collect data for many coherence times. Indeed, the direction, in addition to the amplitude, of the axion field gradient fluctuates stochastically. We present the first complete stochastic treatment for the gradient of the axion field, including multiple computationally efficient methods for performing likelihood-based data analysis, which can be applied to any axion signal, regardless of coherence time. Additionally, we demonstrate that ignoring the stochastic behavior of the gradient of the axion field can potentially result in failure to discover a true axion signal

79 ASTRONOMY AND ASTROPHYSICS↗

Derivative-free stochastic optimization via adaptive sampling strategies

In this paper, we present a novel derivative-free framework for solving unconstrained stochastic optimization problems. Many problems in fields ranging from simulation optimization to reinforcement learning to quantum computing involve settings where only stochastic function values are obtained via a zeroth-order oracle, which has no available gradient information and necessitates the usage of derivative-free optimization methodologies. Our approach includes estimating gradients using stochastic function evaluations and integrating adaptive sampling techniques to control the accuracy in these stochastic approximations. Our framework encapsulates several gradient estimation techniques, including standard finite-difference, Gaussian smoothing, sphere smoothing, randomized coordinate finite-difference, and randomized subspace finite-difference methods. We provide theoretical convergence guarantees for our framework and analyze the worst-case iteration and sample complexities associated with each gradient estimation method. Finally, we demonstrate the empirical performance of the methods on logistic regression and nonlinear least squares problems.

Adaptive sampling↗

A Scalable Gradient Free Method for Bayesian Experimental Design with Implicit Models

Bayesian experimental design (BED) is to answer the question that how to choose designs that maximize the information gathering. For implicit models, where the likelihood is intractable but sampling is possible, conventional BED methods have difficulties in efficiently estimating the posterior distribution and maximizing the mutual information (MI) between data and parameters. Recent work proposed the use of gradient ascent to maximize a lower bound on MI to deal with these issues. However, the approach requires a sampling path to compute the pathwise gradient of the MI lower bound with respect to the design variables, and such a pathwise gradient is usually inaccessible for implicit models. In this paper, we propose a novel approach that leverages recent advances in stochastic approximate gradient ascent incorporated with a smoothed variational MI estimator for efficient and robust BED. Without the necessity of pathwise gradients, our approach allows the design process to be achieved through a unified procedure with an approximate gradient for implicit models. Several experiments show that our approach outperforms baseline methods, and significantly improves the scalability of BED in high-dimensional problems.

Zhang, Jiaxin↗

Balanced stochastic versus deterministic assembly processes benefit diverse yet uneven ecosystem functions in representative agroecosystems

Ecological assembly processes, by influencing community composition, determine ecosystem functions of microbiomes. However, debate remains on how stochastic versus deterministic assembly processes influence ecosystem functions such as carbon and nutrient cycling. Towards a better understanding, we investigated three types of agroecosystems (the upland, paddy, and flooded) that represent a gradient of stochastic versus deterministic assembly processes. Carbon and nutrient cycling multifunctionality, characterized by nine enzymes associated with soil carbon, nitrogen, phosphorous and sulfur cycling, was evaluated and then associated with microbial assembly processes and cooccurrence patterns of vital ecological groups. Our results suggest that strong deterministic processes favour microorganisms with convergent functions (as in the upland agroecosystem), while stochasticity-dominated processes lead to divergent functions (as in the flooded agroecosystem). To benefit agroecosystems services, we speculate that it is critical for a system to maintain balance between its stochastic and deterministic assembly processes (as in the paddy agroecosystem). By doing so, the system can preserve a diverse array of functional traits and also allow for particular traits to flourish. To further confirm this speculation, it is necessary to develop a systematic knowledge beyond merely characterizing general patterns towards the associations among community assembly, composition, and ecosystem functions.

Liu, Wenjing↗

Stochastic average model methods

We consider the solution of finite-sum minimization problems, such as those appearing in nonlinear least-squares or general empirical risk minimization problems. We are motivated by problems in which the summand functions are computationally expensive and evaluating all summands on every iteration of an optimization method may be undesirable. Here we present the idea of stochastic average model (SAM) methods, inspired by stochastic average gradient methods. SAM methods sample component functions on each iteration of a trust-region method according to a discrete probability distribution on component functions; the distribution is designed to minimize an upper bound on the variance of the resulting stochastic model. We present promising numerical results concerning an implemented variant extending the derivative-free model-based trust-region solver POUNDERS, which we name SAM-POUNDERS.

97 MATHEMATICS AND COMPUTING↗

Global stochastic optimization of stellarator coil configurations

In the construction of a stellarator, the manufacturing and assembling of the coil system is a dominant cost. These coils need to satisfy strict engineering tolerances, and if those are not met the project could be cancelled as in the case of the National Compact Stellarator Experiment (NCSX) project. Therefore, our goal is to find coil configurations that increase construction tolerances without compromising the performance of the magnetic field. In this paper, we develop a gradient-based stochastic optimization model which seeks robust stellarator coil configurations in high dimensions. In particular, we design a two-step method: first, we perform an approximate global search by a sample efficient trust-region Bayesian optimization; second, we refine the minima found in step one with a stochastic local optimizer. To this end, we introduce two stochastic local optimizers: BFGS applied to the sample average approximation; and Adam, equipped with a control variate for variance reduction. Numerical simulations performed on a W7-X-like coil configuration demonstrate that our global optimization approach finds a variety of promising local solutions at less than 0.1% of the cost of previous work, which considered solely local stochastic optimization.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

The onset distribution of rotating m , n = 2 , 1 tearing modes and its consequences on the stability of high-confinement-mode plasmas in DIII-D

Abstract Analysis of a multi-scenario database of over 13 000 DIII-D H-mode discharges shows that the m , n = 2 , 1 magnetic islands are dominantly pressure gradient driven, stochastically triggered non-linear instabilities at all edge safety factor ( q 95 ) values. The instability onset time closely follows the exponential distribution in intermediate and high q 95 scenarios and is characterized by near constant onset rate ( λ ), in accordance with Poisson-point processes. This implies that the plasmas are operated in marginally stable conditions, characterized by a small threshold for instability growth and variations in the trigger amplitude and/or the stabilizing mechanisms with temporally uniform random distribution in this database. While the majority of the tearing modes occur in the first current-profile relaxation time of the β N flattop, constant λ throughout the β N flattop shows that the tearing onset is insensitive to the evolution of the equilibrium current profile. In low q 95 scenarios, where a large fraction of the plasmas are operated at low torque, λ increases over the course of the β N flattop, showing that these plasmas evolve toward more unstable conditions. The onset rate rapidly increases with β N , while it does not show a clear dependence on the current gradient at the mode rational surface. Overall, these observations support that the majority of the analyzed 2,1 tearing modes are non-linear, neoclassically driven instabilities and classical stability does not play a dominant role in their onset.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Hutchinson Trace Estimation for high-dimensional and high-order Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) have proven effective in solving partial differential equations (PDEs), especially when some data are available by seamlessly blending data and physics. However, extending PINNs to high-dimensional and even high-order PDEs encounters significant challenges due to the computational cost associated with automatic differentiation in the residual loss function calculation. Herein, we address the limitations of PINNs in handling high-dimensional and high-order PDEs by introducing the Hutchinson Trace Estimation (HTE) method. Starting with the second-order high-dimensional PDEs, which are ubiquitous in scientific computing, HTE is applied to transform the calculation of the entire Hessian matrix into a Hessian vector product (HVP). This approach not only alleviates the computational bottleneck via Taylor-mode automatic differentiation but also significantly reduces memory consumption from the Hessian matrix to an HVP’s scalar output. We further showcase HTE’s convergence to the original PINN loss and its unbiased behavior under specific conditions. Comparisons with the Stochastic Dimension Gradient Descent (SDGD) highlight the distinct advantages of HTE, particularly in scenarios with significant variability and variance among dimensions. We further extend the application of HTE to higher-order and higher-dimensional PDEs, specifically addressing the biharmonic equation. By employing tensor-vector products (TVP), HTE efficiently computes the colossal tensor associated with the fourth-order high-dimensional biharmonic equation, saving memory and enabling rapid computation. The effectiveness of HTE is illustrated through experimental setups, demonstrating comparable convergence rates with SDGD under memory and speed constraints. Additionally, HTE proves valuable in accelerating the Gradient-Enhanced PINN (gPINN) version as well as the Biharmonic equation. Overall, HTE opens up a new capability in scientific machine learning for tackling high-order and high-dimensional PDEs.

Curse of dimensionality↗

Learning-based demand-supply-coupled charging station location problem for electric vehicle demand management

We present a learning-based, demand-supply-coupled optimization model for the charging station location problem (CSLP), aiming to integrate the concept of electric vehicle (EV) charging demand management into the planning of charging infrastructures. In stage one, a gradient boosting-based learning model is developed to predict the charging demand of a charging station based on 15 defined features. Next, in stage two, a demand–supply-coupled CSLP model is developed to optimize the total charging usage rates of both existing and newly selected charging stations. We design a gradient-based stochastic spatial search algorithm to solve the proposed model. A case study with 6-year charging event data from Kansas City Missouri is performed. Results show that the proposed method can generate satisfactory charging demand predictions, and can increase charging usage rates by 14%, outperforming two benchmark approaches. Furthermore, the results of this research are poised to guide agencies in identifying optimal locations for new charging stations.

33 ADVANCED PROPULSION SYSTEMS↗

Sentiment Analysis based Error Detection for Large-Scale Systems

Today's large-scale systems such as High Performance Computing (HPC) Systems are designed/utilized towards exascale computing, inevitably decreasing its reliability due to the increasing design complexity. HPC systems conduct extensive logging of their execution behaviour. In this paper, we leverage the inherent meaning behind the log messages and propose a novel sentiment analysis-based approach for the error detection in large-scale systems, by automatically mining the sentiments in the log messages. Our contributions are four-fold. (1) We develop a machine learning (ML) based approach to automatically build a sentiment lexicon, based on the system log message templates. (2) Using the sentiment lexicon, we develop an algorithm to detect system errors. (3) We develop an algorithm to identify the nodes and components with erroneous behaviors, based on sentiment polarity scores. (4) We evaluate our solution vs. other state-of-the-art machine/deep learning algorithms based on three representative supercomputers' system logs. Experiments show that our error detection algorithm can identify error messages with an average MCC score and f -score of 91% and 96% respectively, while state of the art ML/deep learning model (LSTM) obtains only 67% and 84%. To the best of our knowledge, this is the first work leveraging the sentiments embedded in log entries of large-scale systems for system health analysis.

error detection↗

Risk-Constrained Reinforcement Learning for Inverter-Dominated Power System Controls

Here, this paper develops a risk-aware controller for grid-forming inverters (GFMs) to minimize large frequency oscillations in GFM inverter-dominated power systems. To tackle the high variability from loads/renewables, we incorporate a mean-variance risk constraint into the classical linear quadratic regulator (LQR) formulation for this problem. The risk constraint aims to bound the time-averaged cost of state variability and thus can improve the worst-case performance for large disturbances. The resulting risk-constrained LQR problem is solved through the dual reformulation to a minimax problem, by using a reinforcement learning (RL) method termed as stochastic gradient-descent with max-oracle (SGDmax). In particular, the zero-order policy gradient (ZOPG) approach is used to simplify the gradient estimation using simulated system trajectories. Numerical tests conducted on the IEEE 68-bus system have validated the convergence of our proposed SGDmax for GFM model and corroborate the effectiveness of the risk constraint in improving the worst-case performance while reducing the variability of the overall control cost.

Frequency control↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Post-Disaster Microgrid Formation for Enhanced Distribution System Resilience

This paper proposes a deep reinforcement learning (DRL) based approach for post-disaster critical load restoration in active distribution systems to form microgrids through network reconfiguration to minimize critical load curtailments. Distribution networks are represented as graph networks, and optimal network configurations with microgrids are obtained by searching for the optimal spanning forest. The constraints to the research question being explored are the radial topology and power balance. Unlike existing analytical and population-based approaches, which necessitate the repetition of entire analyses and computation for each outage scenario to find the optimal spanning forest, the proposed approach, once properly trained, can quickly determine the optimal, or near-optimal, spanning forest even when outage scenarios change. When multiple lines fail in the system, the proposed approach forms microgrids with distributed energy resources in active distribution systems to reduce critical load curtailment. The proposed DRL-based model learns the action-value function using the REINFORCE algorithm, which is a model-free reinforcement learning technique based on stochastic policy gradients. A case study was conducted on a 33-node distribution test system, demonstrating the effectiveness of the proposed approach for post-disaster critical load restoration.

active distribution systems↗

Wavefront shaping with a Hadamard basis for scattering soil imaging

Here, soil is a scattering medium that inhibits imaging of plant-microbial-mineral interactions that are essential to plant health and soil carbon sequestration. However, optical imaging in the complex medium of soil has been stymied by the seemingly intractable problems of scattering and contrast. Here, we develop a wavefront shaping method based on adaptive stochastic parallel gradient descent optimization with a Hadamard basis to focus light through soil mineral samples. Our approach allows a sparse representation of the wavefront with reduced dimensionality for the optimization. We further divide the used Hadamard basis set into subsets and optimize a certain subset at once. Simulation and experimental optimization results demonstrate our method has an approximately seven times higher convergence rate and overall better performance compared to that with optimizing all pixels at once. The proposed method can benefit other high-dimensional optimization problems in adaptive optics and wavefront shaping.

47 OTHER INSTRUMENTATION↗

Stabilization of the 81-channel coherent beam combination using machine learning

We develop a rapidly converging algorithm for stabilizing a large channel-count diffractive optical coherent beam combination. An 81-beam combiner is controlled by a novel, machine-learning based, iterative method to correct the optical phases, operating on an experimentally calibrated numerical model. A neural-network is trained to detect phase errors based on interference pattern recognition of uncombined beams adjacent to the combined one. Due to the non-uniqueness of solutions in the full space of possible phases, the network is trained within a limited phase perturbation/error range. This also reduces the number of samples needed for training. Simulations have proven that the network can converge in one step for small phase perturbations. When the trained neural-network is applied to a realistic case of 360 degree full range, an iterative scheme exploits random walking at the beginning, with the accuracy of prediction on phase feedback direction, to allow the neural-network to step into the training range for fast convergence. This neural-network-based iterative method of phase detection works tens of times faster than the commonly used stochastic parallel gradient descent approach (SPGD) using a single-detector and random dither when both are tested with random phase perturbations.

Wang, Dan↗

Adaptive, Active Learning, and Multifidelity Monte Carlo Methods in the MOOSE Stochastic Tools Module

MOOSE is an open-source computational platform for constructing multi-physics models and executing them in a massively parallel fashion. It has a stochastic tools module (STM) for forward/inverse uncertainty quantification (UQ) and surrogate modeling. This presentation details some recent developments to the STM with respect to the implementation of adaptive, active learning, and multifidelity Monte Carlo methods for forward UQ of computational models. Specifically, the adaptive Monte Carlo methods include Markov Chain Monte Carlo (MCMC)-driven algorithms like adaptive importance sampling and parallelized subset simulation for statistical QoI estimation, rare events analysis, and stochastic gradient-free optimization. The active learning methods include Gaussian Process (GP) surrogates and their training via Adam optimization, design of acquisition functions, and integration with samplers like Monte Carlo, adaptive importance, and parallelized subset simulation. These active learning methods are also designed to work in a batch mode, wherein, the required calls to the full computational model are executed in parallel whenever a user-specified batch size is met. The multifidelity methods in STM are broadly divided into two categories: hierarchical, where a defined hierarchy exists among the low-fidelity models, and peer, where all the low-fidelity models are treated equally. A GP surrogate is used to learn the differences between the low- and high-fidelity models in both multifidelity categories, and acquisition functions from the active learning classes are used to decide whether to rely on a low-fidelity model or call the expensive high-fidelity model. Alongside the software description and usage, applications are also presented to nuclear engineering computational models including a TRISO nuclear fuel particle, a reactor pressure vessel, and a heat-pipe microreactor.

97 MATHEMATICS AND COMPUTING↗