Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

CONCURRENT, CONDENSED STEIN VARIATIONAL GRADIENT DESCENT FOR UNCERTAINTY QUANTIFICATION OF NEURAL NETWORKS

In this work, we propose a Stein variational gradient descent (SVGD) method to concurrently sparsify, train, and provide uncertainty quantification (UQ) of a complexly parameterized model, such as a neural network (NN). It employs a graph reconciliation and condensation process to reduce complexity and increase similarity in the Stein ensemble of parameterizations. Therefore, the proposed concurrent, condensed SVGD (ccSVGD) method can provide UQ on parameters, not just outputs. Furthermore, the parameter reduction speeds up the convergence of the Stein gradient descent as it reduces the combinatorial complexity by aligning and differentiating the sensitivity to parameters. These properties are demonstrated with an illustrative example and an application to a mechanical response representation problem in solid mechanics.

42 ENGINEERING↗

An exploration of machine learning models for the determination of reaction coordinates associated with conformational transitions

Determining collective variables (CVs) for conformational transitions is crucial to understanding their dynamics and targeting them in enhanced sampling simulations. Often, CVs are proposed based on intuition or prior knowledge of a system. However, the problem of systematically determining a proper reaction coordinate (RC) for a specific process in terms of a set of putative CVs can be achieved using committor analysis (CA). Identifying essential degrees of freedom that govern such transitions using CA remains elusive because of the high dimensionality of the conformational space. Various schemes exist to leverage the power of machine learning (ML) to extract an RC from CA. Here, we extend these studies and compare the ability of 17 different ML schemes to identify accurate RCs associated with conformational transitions. We tested these methods on an alanine dipeptide in vacuum and on a sarcosine dipeptoid in an implicit solvent. Our comparison revealed that the light gradient boosting machine method outperforms other methods. In order to extract key features from the models, we employed Shapley Additive exPlanations analysis and compared its interpretation with the “feature importance” approach. For the alanine dipeptide, our methodology identifies ϕ and θ dihedrals as essential degrees of freedom in the C7ax to C7eq transition. For the sarcosine dipeptoid system, the dihedrals ψ and ω are the most important for the cisαD to transαD transition. We further argue that analysis of the full dynamical pathway, and not just endpoint states, is essential for identifying key degrees of freedom governing transitions.

Chemistry↗

Single-shot in-line x-ray phase-contrast imaging of void-shockwave interactions in fusion energy materials

Recent breakthroughs in nuclear fusion, specifically the report of reactions exceeding scientific breakeven at the National Ignition Facility (NIF), highlight the potential of inertial fusion energy (IFE) as a sustainable and virtually limitless energy source. However, further progress in IFE requires characterization of defects in ablator materials and how they affect fuel capsule compression. Voids within the ablator can degrade energy yield, but their impact on the density distribution has primarily been studied through simulations, with limited high-resolution experimental validation. To address this, we used the x-ray free-electron laser (XFEL) at the matter in extreme conditions (MECs) instrument at the Linac coherent light source (LCLS) to capture 2D x-ray phase-contrast (XPC) images of a void-bearing sample with a composition similar to inertial confinement fusion (ICF) ablators. By driving a compressive shockwave through the sample using MEC's long-pulse laser system, we analyzed how voids influence shockwave propagation and density distribution during compression. To quantify this impact, we extracted phase information using two phase retrieval algorithms. First, we applied the contrast transfer function (CTF) method, paired with Tikhonov regularization and a fast optimization approach to generate an initial phase estimate. We then refined the result using a projected gradient descent (PGD) method that works directly with the sample's refractive index. Comparing these results with radiation adaptive grid Eulerian (xRAGE) radiation hydrodynamic simulations enables identification of model validation needs or improvements. By calculating phase maps in situ, it becomes possible to reconstruct areal density maps, improving understanding of laser-capsule interactions and advancing IFE research.

Hodge, D. S. [Colorado State Univ., Fort Collins, ↗

Accelerating iterative ptychography with an integrated neural network

Electron ptychography is a powerful and versatile tool for high-resolution and dose-efficient imaging. Iterative reconstruction algorithms are powerful but also computationally expensive due to their relative complexity and the many hyperparameters that must be optimised. Gradient descent-based iterative ptychography is a popular method, but it may converge slowly when reconstructing low spatial frequencies. Here, in this work, we present a method for accelerating a gradient descent-based iterative reconstruction algorithm by training a neural network (NN) that is applied in the reconstruction loop. The NN works in Fourier space and selectively boosts low spatial frequencies, thus enabling faster convergence in a manner similar to accelerated gradient descent algorithms. We discuss the difficulties that arise when incorporating a NN into an iterative reconstruction algorithm and show how they can be overcome with iterative training. We apply our method to simulated and experimental data of gold nanoparticles on amorphous carbon and show that we can significantly speed up ptychographic reconstruction of the nanoparticles.

4DSTEM↗

An adjoint method for determining the sensitivity of island size to magnetic field variations

An adjoint method to calculate the gradient of island width in stellarators is presented and applied to a set of magnetic field configurations. The underlying method for calculation of the island width is that of Cary & Hanson ( Phys. Fluids B, vol. 3, issue 4, 1991, pp. 1006–1014) (with a minor modification), and requires that the residue of the island centre be small. Therefore, the gradient of the residue is calculated in addition. Both the island width and the gradient calculations are verified using an analytical magnetic field configuration introduced by Reiman & Greenside ( Comput. Phys. Commun. , vol. 43, issue 1, 1986, pp. 157–167). The method is also applied to the calculation of the shape gradient of the width of a magnetic island in a National Compact Stellarator Experiment (NCSX) vacuum configuration with respect to positions on a coil. A gradient-based optimization is applied to a magnetic field configuration studied by Hanson & Cary ( Phys. Fluids , vol. 27, issue 4, 1984, pp. 767–769) to minimize stochasticity by adding perturbations to a pair of helical coils. Although only vacuum magnetic fields and an analytical magnetic field model are considered in this work, the adjoint calculation of the island width gradient could also be applied to a magnetohydrodynamic (MHD) equilibrium if the derivative of the magnetic field, with respect to the equilibrium parameters, is known. Using the island width gradient calculation presented here, more general gradient-based optimization methods can be applied to design stellarators with small magnetic islands. Moreover, the sensitivity of the island size may itself be optimized to ensure that coil tolerances, with respect to island size, are kept as high as possible.

Physics↗

Implementation and Exploration of Parameterizations of Large-Scale Dynamics in NCAR's Single Column Atmosphere Model SCAM6

A single column model with parameterized large-scale (LS) dynamics is used to better understand the response of steady-state tropical precipitation to relative sea surface temperature under various representations of radiation, convection, and circulation. The large-scale dynamics are parametrized via the weak temperature gradient (WTG), damped gravity wave (DGW), and spectral weak temperature gradient (Spectral WTG) method in NCAR's Single Column Atmosphere Model (SCAM6). Radiative cooling is either specified or interactive, and the convective parameterization is run using two different values of a parameter that controls the degree of convective inhibition. Results are interpreted in the context of the Global Atmospheric System Studies -Weak Temperature Gradient (GASS-WTG) Intercomparison project. Using the same parameter settings and simulation configuration as in the GASS-WTG Intercomparison project, SCAM6 under the WTG and DGW methods produces erratic results, suggestive of numerical instability. However, when key parameters are changed to weaken the large-scale circulation's damping of tropospheric temperature variations, SCAM6 performs comparably to single column models in the GASS-WTG Intercomparison project. The Spectral WTG method is less sensitive to changes in convection and radiation than are the other two methods, performing qualitatively similarly across all configurations considered. Under all three methods, circulation strength, represented in 1D by grid-scale vertical velocity, is decreased when barriers to convection are reduced. This effect is most extreme under specified radiative cooling, and is shown to come from increased static stability in the column's reference radiative-convective equilibrium profile. This argument can be extended to interactive radiation cases as well, though perhaps less conclusively.

54 ENVIRONMENTAL SCIENCES↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems: Preprint

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different sets of preprocess data for the prediction of the frequency nadir in the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. The training and testing data sets are collected by extensive generation scheduling simulations on the Multi-timescale Integrated Dynamic and Scheduling (MIDAS) toolbox. Numerical results show that all five machine learning methods can achieve high performance accuracy for power system nadir frequency estimation. Among them, the gradient boosting and the XGBoost are clear winners by providing the best prediction accuracy.

data driven↗

A Comparison of Machine Learning Methods for Frequency Nadir Estimation in Power Systems

An increasing penetration level of inverter-based renewable energy resources changes the inertia of power systems, posing challenges for maintaining the desired system frequency stability. An accurate frequency nadir estimation is crucial for power system operators to prepare preventive actions against large frequency excursions. In this paper, five machine learning methods - linear regression, gradient boosting, support vector regression, an artificial neural network, and XGBoost - are applied to two different datasets, i.e., 1) the unit generation dataset and 2) the system total inertia and headroom dataset, for the prediction of the frequency nadir. The training and testing datasets are generated through extensive generation scheduling simulations using Multi-timescale Integrated Dynamic and Scheduling (MI-DAS) toolbox on the Western Electricity Coordinating Council 240-bus system with high renewable penetration levels. Numerical results show that all five machine learning methods perform well in predicting the nadir frequency of the system. Among them, the gradient boosting and the XGBoost are clear winners yielding the best prediction accuracy in terms of four evaluation metrics.

data driven↗

Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient Descent

Despite the great achievements of the modern deep neural networks (DNNs), the vulnerability/robustness of state-of-the-art DNNs raises security concerns in many application domains requiring high reliability. Various adversarial attacks are proposed to sabotage the learning performance of DNN models. Among those, the black-box adversarial attack methods have received special attentions owing to their practicality and simplicity. Black-box attacks usually prefer less queries in order to maintain stealthy and low costs. However, most of the current black-box attack methods adopt the first-order gradient descent method, which may come with certain deficiencies such as relatively slow convergence and high sensitivity to hyper-parameter settings. In this paper, we propose a zeroth-order natural gradient descent (ZO-NGD) method to design the adversarial attacks, which incorporates the zeroth-order gradient estimation technique catering to the black-box attack scenario and the second-order natural gradient descent to achieve higher query efficiency. The empirical evaluations on image classification datasets demonstrate that ZO-NGD can obtain significantly lower model query complexities compared with state-of-the-art attack methods.

Zhao, Pu↗

Apparatus and methods for sample analysis with multi-gradient microfluidics

A device for analyzing biological samples comprises first, second, third, and fourth layers. The first layer comprises a sample chamber in which a sample is positioned. The second layer comprises first, second, and third channels. A third, porous layer is positioned between the first layer and the second layer. A fourth layer composed of a substantially liquid-impermeable material is positioned between the second layer and the third layer. The fourth layer includes first and second pass-through channels that are aligned with the first and second channel, respectively. Fluids that flow in the first and second channels pass through the pass-through channels and diffuse into the sample chamber, establishing a chemical concentration gradient therein. A gas in the sample chamber can diffuse through the third and fourth layers and interact with a fluid flowing in the third channel, establishing a gas concentration gradient in the sample chamber.

Kim, Peter Wonhee↗

Journey Over Destination: Differentiable Sensor Placement Enhances Generalization [Poster]

The challenge of reconstructing spatial fields that change over time from limited sensor data has been a focal point for many research studies. Various machine learning methods have been used in attempts to address this complex issue, including convolutional neural networks. All of the proposed methods share a common requirement that the user needs to manually determine the sensor positions. This requirement remains a limiting factor in the ongoing quest for efficient learning and accurate field reconstruction. This study aims to present a method that enables a model to optimize sensor positions via backpropagation, thereby facilitating the model’s exploration of the spatial domain and enhancing sensor positioning effectively. Indexing naturally incorporates discrete decisions. This operation is nondifferentiable which is a requirement for the application of gradient-based optimization methods. We showcased its effectiveness by training an attention-based neural network, which achieved top-tier performance on two separate datasets. To our knowledge, this represents the first fully end-to-end differentiable workflow for enhancing sensor placement within a neural network model.

58 GEOSCIENCES↗

Ensemble methods for quantification of potassium oxide in ChemCam Mars and laboratory spectra

In this paper we test new approaches for predicting the amount of element oxides in rock samples from the ChemCam instrument suite onboard the NASA Curiosity rover by focusing on K 2 O. Using the expanded dataset compiled by Gasda et al. (2021) with and without the Earth to Mars (E2M and NoE2M) transformation discussed in Clegg et al. (2017) we trained blended submodels using the “double blending” technique and compared these to ensemble methods (Random Forest, ExtraTrees, and Gradient Boosting Regression). We found that ensemble methods performed similar to blended submodels when looking at RMSE-P on the laboratory spectra and provided significant advantages when looking at spectra coming from Mars. For the full model, blended submodels achieved an RMSE-P of 0.62 and 0.60 (E2M and NoE2M respectively) while Gradient Boosting Regression resulted in a slightly improved RMSE-P of 0.59 and 0.60. More importantly, by employing a local RMSE-P estimation technique where model performance is evaluated based on nearby test samples we found that using ensemble methods can lower the quantification limit for K 2 O from the current value of ≈0.6 wt% to ≈0.08 wt% using Extra Trees and Random Forest. This would allow for a much larger range of K 2 O values to be quantified on Mars with greater certainty given that most targets seen on Mars tend to have <1 wt% K2O. Finally, we used both Mean Decrease in Impurity (MDI) and permutation importance techniques to investigate the wavelengths used by the ensemble methods and found that they correspond to known potassium emission lines. This suggests that ensemble methods can provide an easier to train and improved alternative to blended submodels for predicting potassium compositions from Laser Induced Breakdown Spectroscopy (LIBS) data.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

A comparison of eight optimization methods applied to a wind farm layout optimization problem

Abstract. Selecting a wind farm layout optimization method is difficult. Comparisons between optimization methods in different papers can be uncertain due to the difficulty of exactly reproducing the objective function. Comparisons by just a few authors in one paper can be uncertain if the authors do not have experience using each algorithm. In this work we provide an algorithm comparison for a wind farm layout optimization case study between eight optimization methods applied, or directed, by researchers who developed those algorithms or who had other experience using them. We provided the objective function to each researcher to avoid ambiguity about relative performance due to a difference in objective function. While these comparisons are not perfect, we try to treat each algorithm more fairly by having researchers with experience using each algorithm apply each algorithm and by having a common objective function provided for analysis. The case study is from the International Energy Association (IEA) Wind Task 37, based on the Borssele III and IV wind farms with 81 turbines. Of particular interest in this case study is the presence of disconnected boundary regions and concave boundary features. The optimization methods studied represent a wide range of approaches, including gradient-free, gradient-based, and hybrid methods; discrete and continuous problem formulations; single-run and multi-start approaches; and mathematical and heuristic algorithms. We provide descriptions and references (where applicable) for each optimization method, as well as lists of pros and cons, to help readers determine an appropriate method for their use case. All the optimization methods perform similarly, with optimized wake loss values between 15.48 % and 15.70 % as compared to 17.28 % for the unoptimized provided layout. Each of the layouts found were different, but all layouts exhibited similar characteristics. Strong similarities across all the layouts include tightly packing wind turbines along the outer borders, loosely spacing turbines in the internal regions, and allocating similar numbers of turbines to each discrete boundary region. The best layout by annual energy production (AEP) was found using a new sequential allocation method, discrete exploration-based optimization (DEBO). Based on the results in this study, it appears that using an optimization algorithm can significantly improve wind farm performance, but there are many optimization methods that can perform well on the wind farm layout optimization problem, given that they are applied correctly.

17 WIND ENERGY↗

Fault location in High Voltage Multi-terminal dc Networks Using Ensemble Learning

Precise location of faults for large distance power transmission networks is essential for faster repair and restoration process. High Voltage direct current (HVdc) networks using modular multi-level converter (MMC) technology has found its prominence for interconnected multi-terminal networks. This allows for large distance bulk power transmission at lower costs. However, they cope with the challenge of dc faults. Fast and efficient methods to isolate the network under dc faults have been widely studied and investigated. After successful isolation, it is essential to precisely locate the fault. The post-fault voltage and current signatures are a function of multiple factors and thus accurately locating faults on a multi-terminal network is challenging. In this paper, we discuss a novel data-driven ensemble learning based approach for accurate fault location. Here we utilize the eXtreme Gradient Boosting (XGB) method for accurate fault location. The sensitivity of the proposed algorithm to measurement noise, fault location, resistance and current limiting inductance are performed on a radial three-terminal MTdc network designed in Power System Computer Aided Design (PSCAD)/Electromagnetic Transients including dc (EMTdc).

24 POWER TRANSMISSION AND DISTRIBUTION↗

Exploring gauge-fixing conditions with gradient-based optimization

Lattice gauge fixing is required to compute gauge-variant quantities, for example those used in RI-MOM renormalization schemes or as objects of comparison for model calculations. Recently, gauge-variant quantities have also been found to be more amenable to signal-to-noise optimization using contour deformations. These applications motivate systematic parameterization and exploration of gauge-fixing schemes. This work introduces a differentiable parameterization of gauge fixing which is broad enough to cover Landau gauge, Coulomb gauge, and maximal tree gauges. The adjoint state method allows gradient-based optimization to select gauge-fixing schemes that minimize an arbitrary target loss function.

Detmold, William↗

Advancing spatiotemporal forecasts of CO 2 plume migration using deep learning networks with transfer learning and interpretation analysis

Accurate and timely forecasts of CO 2 plume distribution throughout the injection and post-injection phases are crucial for detecting plume migration, assessing leakage risks, and supporting operational decisions in geologic carbon storage (GCS). Current convolutional neural network-based approaches primarily focus on spatial information and overlook temporal dependencies in plume distributions, thus limiting their ability to capture dynamic movement effects and provide accurate predictions of plume migration. In this work, we propose two deep learning models, Auto-Encoder (AE)-LSTM and Encoder-Decoder (ED)-ConvLSTM, each uniquely designed to capture both spatial and temporal features. We apply the proposed methods to forecast the dynamic distribution of CO 2 plumes based on 108 reservoir simulations over a 30-year injection and a 30-year post-injection period. The results indicate that the ED-ConvLSTM model outperforms the AE-LSTM model in accurately predicting the spatiotemporal dynamics of CO 2 plume migration, achieving R 2 values above 0.99. To provide a deeper understanding of these model predictions, we employ a gradient-based explanation method on the trained models. This approach provides insights into the influence of input variables on plume migration forecasts and uncovers the underlying prediction mechanisms of the proposed models. Furthermore, we introduce a transfer learning technique, enabling fast and accurate plume migration forecasting in the post-injection phase by leveraging the trained model during the injection phase. This reduces the necessity for extensive data collection or re-training. In conclusion, the methods proposed in our work enhances the performance and interpretability of CO 2 plume migration forecasts, thereby facilitating informed decision-making throughout the entire lifecycle of GCS applications.

58 GEOSCIENCES↗

Analytical derivatives of the individual state energies in ensemble density functional theory. II. Implementation on graphical processing units (GPUs)

Conical intersections control excited state reactivity, and thus, elucidating and predicting their geometric and energetic characteristics are crucial for understanding photochemistry. Locating these intersections requires accurate and efficient electronic structure methods. Unfortunately, the most accurate methods (e.g., multireference perturbation theories such as XMS-CASPT2) are computationally challenging for large molecules. The state-interaction state-averaged restricted ensemble referenced Kohn–Sham (SI-SA-REKS) method is a computationally efficient alternative. The application of SI-SA-REKS to photochemistry was previously hampered by a lack of analytical nuclear gradients and nonadiabatic coupling matrix elements. We have recently derived analytical energy derivatives for the SI-SA-REKS method and implemented the method effectively on graphical processing units. We demonstrate that our implementation gives the correct conical intersection topography and energetics for several examples. Furthermore, our implementation of SI-SA-REKS is computationally efficient, with observed sub-quadratic scaling as a function of molecular size. This demonstrates the promise of SI-SA-REKS for excited state dynamics of large molecular systems.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Dakota A Multilevel Parallel Object-Oriented Framework for Design Optimization Parameter Estimation Uncertainty Quantification and Sensitivity Analysis: Version 6.12 Theory Manual

The Dakota toolkit provides a flexible and extensible interface between simulation codes and iterative analysis methods. Dakota contains algorithms for optimization with gradient and nongradient-based methods; uncertainty quantification with sampling, reliability, and stochastic expansion methods; parameter estimation with nonlinear least squares methods; and sensitivity/variance analysis with design of experiments and parameter study methods. These capabilities may be used on their own or as components within advanced strategies such as surrogate-based optimization, mixed integer nonlinear programming, or optimization under uncertainty. By employing object-oriented design to implement abstractions of the key components required for iterative systems analyses, the Dakota toolkit provides a flexible and extensible problem-solving environment for design and performance analysis of computational models on high performance computers. This report serves as a theoretical manual for selected algorithms implemented within the Dakota software. It is not intended as a comprehensive theoretical treatment, since a number of existing texts cover general optimization theory, statistical analysis, and other introductory topics. Rather, this manual is intended to summarize a set of Dakota-related research publications in the areas of surrogate-based optimization, uncertainty quantification, and optimization under uncertainty that provide the foundation for many of Dakota's iterative analysis capabilities.

97 MATHEMATICS AND COMPUTING↗