Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “gradient methods”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Analytical gradient-based optimization of CALPHAD model parameters

The calibration of CALPHAD (CALculation of PHAse Diagrams) models involves the solution of a very challenging high-dimensional multiobjective optimization problem. Traditional approaches to parameter fitting predominantly rely on gradient-free methods, which while robust, are computationally inefficient and often scale poorly with model complexity. In this work, we introduce and demonstrate a generalizable framework for analytic gradient-based optimization of the parameters of the CALPHAD model enabled by the recently formalized Jansson derivative technique. This method allows for efficient evaluation of gradients of thermodynamic properties at equilibrium with respect to model parameters, even in the presence of arbitrarily complex internal degrees of freedom. Leveraging these semi-analytic gradients, we employ the conjugate gradient (CG) method to optimize thermodynamic model parameters for four binary alloy systems: Cu-Mg, Fe-Ni, Cr-Ni, and Cr-Fe. Across all systems, CG achieves comparable or superior optimality relative to Bayesian ensemble Markov Chain Monte Carlo (MCMC) with improvements in computational efficiency ranging from one to three orders of magnitude. Furthermore, our results establish a new paradigm for CALPHAD assessments in which high fidelity data-rich model calibration becomes tractable using deterministic gradient-informed algorithms.

CALPHAD↗

SwitchX : Gmin-Gmax Switching for Energy-efficient and Robust Implementation of Binarized Neural Networks on ReRAM Xbars

Memristive crossbars can efficiently implement Binarized Neural Networks (BNNs) wherein the weights are stored in high-resistance states (HRS) and low-resistance states (LRS) of the synapses. We propose SwitchX mapping of BNN weights onto ReRAM crossbars such that the impact of crossbar non-idealities, that lead to degradation in computational accuracy, are minimized. Essentially, SwitchX maps the binary weights in such a manner that a crossbar instance comprises of more HRS than LRS synapses. We find BNNs mapped onto crossbars with SwitchX to exhibit better robustness against adversarial attacks than the standard crossbar mapped BNNs, the baseline. Finally, we combine SwitchX with state-aware training (that further increases the feasibility of HRS states during weight mapping) to boost the robustness of a BNN on hardware. We find that this approach yields stronger defense against adversarial attacks than adversarial training, a state-of the-art software defense. We perform experiments on a VGG16 BNN with benchmark datasets (CIFAR-10, CIFAR-100 and TinyImagenet) and use Fast Gradient Sign Method (ϵ = 0.05 to 0.3) and Projected Gradient Descent (ϵ = $\frac{2}{255}$ to $\frac{32}{255}$, α = $\frac{2}{255}$) adversarial attacks. We show that SwitchX combined with state-aware training can yield upto ~35% improvements in clean accuracy and ~6–16% in adversarial accuracies against conventional BNNs. Furthermore, an important by-product of SwitchX mapping is increased crossbar power savings, owing to an increased proportion of HRS synapses, which is furthered with state-aware training. We obtain upto ~21–22% savings in crossbar power consumption for state-aware trained BNN mapped via SwitchX on 16 × 16 and 32 × 32 crossbars using the CIFAR-10 and CIFAR-100 datasets.

97 MATHEMATICS AND COMPUTING↗

Locally adaptive activation functions with slope recovery for deep and physics-informed neural networks

Here we propose two approaches of locally adaptive activation functions namely, layer-wise and neuron-wise locally adaptive activation functions, which improve the performance of deep and physics-informed neural networks. The local adaptation of activation function is achieved by introducing a scalable parameter in each layer (layer-wise) and for every neuron (neuron-wise) separately, and then optimizing it using a variant of stochastic gradient descent algorithm. In order to further increase the training speed, an activation slope-based slope recovery term is added in the loss function, which further accelerates convergence, thereby reducing the training cost. On the theoretical side, we prove that in the proposed method, the gradient descent algorithms are not attracted to sub-optimal critical points or local minima under practical conditions on the initialization and learning rate, and that the gradient dynamics of the proposed method is not achievable by base methods with any (adaptive) learning rates. We further show that the adaptive activation methods accelerate the convergence by implicitly multiplying conditioning matrices to the gradient of the base method without any explicit computation of the conditioning matrix and the matrix–vector product. The different adaptive activation functions are shown to induce different implicit conditioning matrices. Furthermore, the proposed methods with the slope recovery are shown to accelerate the training process.

97 MATHEMATICS AND COMPUTING↗

Validating methods for modeling composition gradients in planar shock experiments

An interface is Rayleigh–Taylor (RT) unstable when acceleration pushes a less dense material into a more dense one, and the growth of the instability is governed partly by the Atwood number gradient. Double-shell inertial confinement fusion capsules have a foam spacer layer pushing on an inner capsule composed of a beryllium tamper and high-Z inner shell, and so have RT unstable interfaces that require benchmarking. To this end, the results of a planar shock experiment with beryllium/tungsten targets are presented. One target had the normal bilayer construction of beryllium and tungsten in two distinct layers; the second target had the beryllium grading into tungsten with a quasi-exponential profile, motivated by the potential for reduced RT growth with the gradient profile. Simulations mimic the shock profiles for both targets and match the shock velocity to within 5%. These results validate the ability of our simulations to model double-shell capsules with bilayer or graded layer Be/W inner shells, which are needed to design future experiments at the National Ignition Facility.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Deep Generative Models that Solve PDEs: Distributed Computing for Training Large Data-Free Models

Recent progress in scientific machine learning (SciML) has opened up the possibility of training novel neural network architectures that solve complex partial differential equations (PDEs). Several (nearly data free) approaches have been recently reported that successfully solve PDEs, with examples including deep feed forward networks, generative networks, and deep encoder-decoder networks. However, practical adoption of these approaches is limited by the difficulty in training these models, especially to make predictions at large output resolutions (≥1024×1024). Here we report on a software framework for data parallel distributed deep learning that resolves the twin challenges of training these large SciML models - training in reasonable time as well as distributing the storage requirements. Our framework provides several out of the box functionality including (a) loss integrity independent of number of processes, (b) synchronized batch normalization, and (c) distributed higher-order optimization methods. We show excellent scalability of this framework on both cloud as well as HPC clusters, and report on the interplay between bandwidth, network topology and bare metal vs cloud. We deploy this approach to train generative models of sizes hitherto not possible, showing that neural PDE solvers can be viably trained for practical applications. We also demonstrate that distributed higher-order optimization methods are 2-3× faster than stochastic gradient-based methods and provide minimal convergence drift with higher batch-size.

PDEs↗

Sparsity Applications for Gradient‐Based Optimization of Wind Farms

Optimizing wind farms is essential for designing efficient energy systems, especially as farms grow larger and span multiple sites. However, this optimization becomes increasingly challenging due to the rising computational cost associated with more turbines. Gradient‐based optimization methods scale better than gradient‐free approaches for large problems, but the most computationally expensive component remains the calculation of gradients for the objective function and constraint Jacobians. To address this, we propose leveraging sparsity to accelerate gradient evaluations and reduce the size of the constraint Jacobian. Wind farms naturally exhibit sparsity—many turbines do not influence each other under certain wind directions. However, unlike traditional sparse problems with fixed patterns, wind farm sparsity is dynamic, requiring new strategies to handle changing interactions efficiently. This paper presents a study of sparsity in wind farm optimization and introduces several methods to exploit it. These strategies are tested on multiple farms using the analytic Cumulative Curl model, with gradients computed via automatic differentiation (AD). The same sparsity‐aware techniques are also applicable to finite difference (FD) methods, where they can yield even greater speedups due to the high cost of directional evaluations. Results show that sparse methods achieve up to a 10x speedup with less than ± 5% variance in optimized wake losses compared to traditional methods. These findings suggest that sparsity‐aware optimization not only maintains solution quality but also scales efficiently with farm size, enabling more comprehensive design exploration at reduced computational cost.

17 WIND ENERGY↗

A B-spline based gradient-enhanced micropolar implicit material point method for large localized inelastic deformations

The quasi-brittle response of cohesive-frictional materials in numerical simulations is commonly represented by softening plasticity or continuum damage models, either individually or in combination. However, classical models, particularly when coupled with non-associated plasticity, often suffer from ill-posedness and a lack of objectivity in numerical simulations. Moreover, the performance of the finite element method significantly degrades in simulations involving finite strains when mesh distortion reaches excessive levels. This represents a challenge for modeling cohesive-frictional materials, given their tendency to experience strongly localized deformations, such as those occurring during shear band dominated failure. Hence, accurate modeling of the response of cohesive-frictional solids is a demanding task. To address these challenges, we present an extension of the material point method (MPM) for the unified gradient-enhanced micropolar continuum, aiming at the analysis of finite localized inelastic deformations in cohesive-frictional materials. The generalized gradient-enhanced micropolar continuum formulation is employed to tackle challenges related to localization and softening material behavior, while the MPM addresses issues arising from excessive deformations. The method utilizes a B-spline formulation for the rigid background mesh to mitigate the well-known cell crossing errors of the MPM. To demonstrate the performance of the method, 2D and 3D numerical studies on localized failure in sandstone in plane strain compression and triaxial extension tests are presented. A comparison with finite element results confirms the suitability of the formulation. Moreover, an efficient numerical implementation of the formulation is presented, and it is demonstrated that the additional MPM specific overhead is negligible.

B-spline↗

Reduced scaling formulation of CASPT2 analytical gradients using the supporting subspace method

We present a reduced scaling and exact reformulation of state specific complete active space second-order perturbation (CASPT2) analytical gradients in terms of the MP2 and Fock derivatives using the supporting subspace method. This work follows naturally from the supporting subspace formulation of the CASPT2 energy in terms of the MP2 energy using dressed orbitals and Fock builds. For a given active space configuration, the terms corresponding to the MP2-gradient can be evaluated with O(N5) operations, while the rest of the calculations can be computed with O(N3) operations using Fock builds, Fock gradients, and linear algebra. When tensor-hyper-contraction is applied simultaneously, the computational cost can be further reduced to O(N4) for a fixed active space size. The new formulation enables efficient implementation of CASPT2 analytical gradients by leveraging the existing graphical processing unit (GPU)-based MP2 and Fock routines. We present benchmark results that demonstrate the accuracy and performance of the new method. Example applications of the new method in ab initio molecular dynamics simulation and constrained geometry optimization are given.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

MOOSE ProbML: Parallelizable Probabilistic Machine Learning and Uncertainty Quantification Capabilities

The Multiphysics Object Oriented Simulation Environment (MOOSE) is a widely used open- source finite element software for performing multiphysics multiscale simulations in a massively parallel fashion. Recently, the computational team at Idaho National Laboratory (INL) has implemented Probabilistic Machine Learning (ProbML) capabilities in MOOSE—in a parallelized fashion—and enable active learning with large-scale computational models for tasks such as surrogate model development, scale bridging, forward/inverse uncertainty quantification (UQ), Bayesian optimization, etc. This presentation summarizes these developments in MOOSE along with demonstrations on several real applications relevant to nuclear energy. At the fundamental level, samplers like Monte Carlo/Latin Hypercube, variance reduction, parallelized Markov Chain Monte Carlo (MCMC) support uncertainty propagation in both forward and inverse settings. These samplers can be integrated with the Gaussian processes (GP) suite in MOOSE, which offer several variants like scalar GPs, multi-output GPs, and deep GPs, to enable active learning. These GPs can be tuned using gradient-based optimization methods like Adam and its variants or gradient-free methods like the elliptical slice sampler (a variant of MCMC adept under Gaussian settings) for more complex covariance kernels or likelihoods whose gradient computations can be cumbersome. A variety of batch acquisition functions permit parallelized evaluation of the computational model and support different learning objectives with high efficiency like Bayesian inference, global surrogate development, optimization, etc. Furthermore, libtorch integration supports training, evaluation, and re-training of neural networks and other complex machine learning models in active learning settings. The impacts of these developments are shown on several real applications: (1) nuclear fuel inverse UQ and model inadequacy assessment using the Kennedy O’Hagan framework; (2) uncertainty aware surrogate modeling for additive manufacturing to predict field quantities; (3) nuclear reactor rare events analysis; and (4) complex fluid flow prediction using a global surrogate with quantified prediction uncertainty. Finally, the outlook of MOOSE ProbML is discussed for both outer-loop and inner-loop computations in the broad view to accelerate fuels and materials qualification, address gaps in knowledge and data, and assess new reactor/fuel systems.

11 - NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗

Reduced-dimension Bayesian optimization for model calibration of transient vapor compression cycles

Development and calibration of first-principles dynamic models of vapor compression cycles (VCCs) is of critical importance for applications that include control design and fault detection and diagnostics. Nevertheless, the inherent complexity of models that are represented by large systems of differential–algebraic equations leads to significant challenges for model calibration processes that utilize classical gradient-based methods. Bayesian optimization (BO) is a sample-efficient and gradient-free approach using a probabilistic surrogate model and optimal search over a feasible parameter space. Despite the benefits of BO in reducing computational costs, challenges remain in dealing with a high-dimensional calibration task resulting from a large set of parameters that have significant impacts on system behavior and need to be calibrated simultaneously. This paper presents a reduced-dimension BO framework for calibrating transient VCCs models where the calibration space is projected to a low-dimensional subspace for accelerating convergence of the solution algorithm and consequently reducing the number of transient simulations. The proposed approach was demonstrated via two case studies associated with different VCC applications where 10 parameters were calibrated in each case using laboratory measurements. The reduced-dimension BO framework only required 1 / 8 th of the iterations associated with a standard BO method that deals with high-dimensional calibration parameters for converged solutions and yielded comparable accuracy. Furthermore, both calibrated models revealed significant accuracy improvements compared to uncalibrated models.

Ma, Jiacheng↗

Convergence of Hyperbolic Neural Networks Under Riemannian Stochastic Gradient Descent

Abstract We prove, under mild conditions, the convergence of a Riemannian gradient descent method for a hyperbolic neural network regression model, both in batch gradient descent and stochastic gradient descent. We also discuss a Riemannian version of the Adam algorithm. We show numerical simulations of these algorithms on various benchmarks.

Whiting, Wes (ORCID:0000000247505060)↗

Natural evolutionary strategies for variational quantum computation

Abstract Natural evolutionary strategies (NES) are a family of gradient-free black-box optimization algorithms. This study illustrates their use for the optimization of randomly initialized parameterized quantum circuits (PQCs) in the region of vanishing gradients. We show that using the NES gradient estimator the exponential decrease in variance can be alleviated. We implement two specific approaches, the exponential and separable NES, for parameter optimization of PQCs and compare them against standard gradient descent. We apply them to two different problems of ground state energy estimation using variational quantum eigensolver and state preparation with circuits of varying depth and length. We also introduce batch optimization for circuits with larger depth to extend the use of ES to a larger number of parameters. We achieve accuracy comparable to state-of-the-art optimization techniques in all the above cases with a lower number of circuit evaluations. Our empirical results indicate that one can use NES as a hybrid tool in tandem with other gradient-based methods for optimization of deep quantum circuits in regions with vanishing gradients.

Anand, Abhinav (ORCID:0000000280812310)↗

Explaining machine-learning models for gamma-ray detection and identification

As more complex predictive models are used for gamma-ray spectral analysis, methods are needed to probe and understand their predictions and behavior. Recent work has begun to bring the latest techniques from the field of Explainable Artificial Intelligence (XAI) into the applications of gamma-ray spectroscopy, including the introduction of gradient-based methods like saliency mapping and Gradient-weighted Class Activation Mapping (Grad-CAM), and black box methods like Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP). In addition, new sources of synthetic radiological data are becoming available, and these new data sets present opportunities to train models using more data than ever before. In this work, we use a neural network model trained on synthetic NaI(Tl) urban search data to compare some of these explanation methods and identify modifications that need to be applied to adapt the methods to gamma-ray spectral data. We find that the black box methods LIME and SHAP are especially accurate in their results, and recommend SHAP since it requires little hyperparameter tuning. We also propose and demonstrate a technique for generating counterfactual explanations using orthogonal projections of LIME and SHAP explanations.

73 NUCLEAR PHYSICS AND RADIATION PHYSICS↗

The Effects of Compounded Model Size Reductions on Adversarial Robustness

Recent advances in Edge AI and Tiny Machine Learning (TinyML) have enabled the deployment of machine learning models on resource-constrained environments. However, deploying these models on edge devices, such as micro-controllers, requires significant model footprint reduction through a variety of techniques such as quantization, pruning, and clustering. While these optimization methods offer considerable advantages, they potentially introduce AI-related security vulnerabilities, particularly concerning model robustness with respect to adversarial AI attacks. Prior research has extensively examined the impact of quantization on adversarial robustness; however, the effects of alternative reduction techniques and their combinations remain understudied. This paper investigates the impact of model size reduction techniques on adversarial robustness, when applied individually and combined. We utilized Fast Gradient Sign Method (FGSM) and Projected Gradient Descent (PGD) attacks to generate adversarial perturbations for both training and testing data, and then evaluated the models' accuracy under adversarial training conditions. Our findings revealed that reduction techniques generally diminished robustness; although, combining techniques was not found to make robustness any worse than when applied individually. Moreover, specific techniques can potentially enhance resistance to small size perturbations. This research provides insights into the trade-offs between model size reduction and security, establishing a foundation for future investigations into improving adversarial training techniques and methodologies for maintaining robustness while preserving memory footprint benefits.

Austria, Phillipe [ORNL] (ORCID:0000000236223973)↗

Chemically Amplified, Dry-Develop Poly(aldehyde) Photoresist

The catalytic decomposition of poly(phthalaldehyde) with a photoacid generator can be used as dry-develop photoresist, where the exposed film depolymerizes into small molecules to allow the development of features via controlled vaporization. Higher temperatures enabled shorter dry-development times, but also promoted faster photoacid diffusion that compromised pattern fidelity. Trihexylamine was used as a base quencher to counteract acid diffusion in a phthalaldehyde-propanal co-polymer photoresist. The propanal co-monomer in the polymer improves the vaporization rate because it has a higher vapor pressure than phthalaldehyde. Addition of the base quencher was found to improve the contrast, pattern fidelity, and ease-of-handling of the dry-develop resist in a direct-write UV lithography tool. The dry-development of 4 μ m features was achieved with no appreciable residue. For large area features, a spatially variable exposure method was used to direct the residue away from the exposed area. The gradient exposure method was used to produce 100 μ m features. Plasma etching after dry-development was also used to achieve residue-free dry-developed patterns. These results show the benefits of incorporating base additives into a dry-develop depolymerizable resist system and highlight the need for addressing residue formation.

Materials Science↗

Preliminary Theoretical Analysis of Mixed Precision Krylov Subspace Methods (Q3 Report)

The third quarter of the project was spent performing theoretical finite precision analysis of Krylov subspace method variants that use mixed precision. Our focus here is on the Conjugate Gradient (CG) method and the Lanczos method. We have performed an analysis of maximum attainable accuracy for the classical CG method in which 3 precisions are used: a working precision ε, a precision ε IP for the inner product computations, and a precision ε MV for the matrix-vector products. Our results show that performing inner product computations in lower precision does not affect the attainable accuracy. Further, we have performed a complete error analysis of the s-step Lanczos algorithm. In this case, we show that the numerical behavior of the algorithm can be significantly improved by using extra precision in a small part of the computation. We summarize the main theorems in the remainder of the document. Other activities include attending biweekly xSDK meetings. The subsequent quarter will be spent finalizing these results into technical reports and/or manuscripts for submission to journals, as well as identifying opportunities for future work.

97 MATHEMATICS AND COMPUTING↗