Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 55 records · Page 3

Field Work Proposal ERKJ358: Black-box training for scientific machine learning models (Final Report)

The overarching goal of this project is to develop a scalable black-box training capability for scientific machine learning (SciML) problems that are non-trainable with existing automatic differentiation (AD)-based algorithms. AD assumes that a loss function can be decomposed into a sequence of elementary operations whose derivatives are known. This assumption is violated when the loss function includes a black-box physical model (e.g., a legacy simulator). The current strategy, converting a black-box simulator to an AD-enabled code via differential programming, is inflexible and time-, labor-consuming. Thus, black-box optimization is a main workhorse for training SciML models, e.g., in scientific reinforcement learning, hyper-parameter fine tuning, designing SciML models with adversarial robustness, etc.

97 MATHEMATICS AND COMPUTING↗

Vibrational and rotational cooling of electrons by water vapor

The cooling of electrons by vibrational and rotational excitation of water molecules plays an important role in the thermal balance of electrons in cometary ionospheres. The energy-loss function for rotational excitation and deexcitation of H2O by electron impact is calculated theoretically. The rotational cooling rate is calculated using this loss function for a wide range of electron and neutral temperatures. The vibrational cooling rate is calculated using measured values of electron-impact vibrational excitation cross sections. Analytical formulas are provided for some of the cooling rates. The interaction of ions with H2O molecules is also discussed, and a formula is suggested for the momentum-transfer collision frequency.

Cravens, T. E.↗

Stacking AsFMT overexpression with BdPMT loss of function enhances monolignol ferulate production in Brachypodium distachyon

To what degree can the lignin subunits in a monocot be derived from monolignol ferulate (ML-FA) conjugates? This simple question comes with a complex set of variables. Three potential requirements for optimizing ML-FA production are as follows: (1) The presence of an active FERULOYL-CoA MONOLIGNOL TRANSFERASE (FMT) enzyme throughout monolignol production; (2) Suppression or elimination of enzymatic pathways competing for monolignols and intermediates during lignin biosynthesis; and (3) Exclusion of alternative phenolic compounds that participate in lignification. A 16-fold increase in lignin-bound ML-FA incorporation was observed by introducing an AsFMT gene into Brachypodium distachyon. On its own, knocking out the native p-COUMAROYL-CoA MONOLIGNOL TRANSFERASE (BdPMT) pathway that competes for monolignols and the p-coumaroyl-CoA intermediate did not change ML-FA incorporation, nor did partial loss of CINNAMOYL-CoA REDUCTASE1 (CCR1) function, which reduced metabolic flux to monolignols. However, stacking AsFMT into the Bdpmt-1 mutant resulted in a 32-fold increase in ML-FA incorporation into lignin over the wild-type level.

59 BASIC BIOLOGICAL SCIENCES↗

Dynamic calibration of differential equations using machine learning, with application to turbulence models

We present a methodology for calibration of parametric ordinary and partial differential equation models, using off-the-shelf software for back-propagation in Neural Networks (NN). As a prototypical example, we consider calibration of a Reynolds-averaged Navier-Stokes (RANS) turbulence closure model, against ground truth data from direct numerical simulations (DNS) of two different turbulent flows. Numerical time integration is represented as a custom NN, where only the RANS model parameters are trainable. A loss function is defined to quantify the mismatch between the NN prediction and the ground truth over a predefined, finite time integration window. This loss function is then minimized using a gradient descent method utilizing the back-propagation algorithm. Furthermore, this dynamic approach to training is to be contrasted with a static approach, wherein a least square regression estimate for parameters is obtained in the limit of an infinitesimal time integration window. In a first test of static and dynamic approaches against ground truth data generated by the model, the former proves to be significantly faster and more accurate than the latter at recovering the parameters. When both calibration approaches are tested against DNS data, for which it is known that the model cannot achieve a perfect fit, the static approach yields a good prediction only for short times, while the dynamic approach results in physical and stable predictions over the entire integration window. After optimization of the dynamic approach for time step, spatial resolution, stability, and physics-based constraints, we obtain a 50% improvement of outcomes over those obtained from the existing, manually calibrated set of parameters, demonstrating the merits of this systematic and automated procedure.

97 MATHEMATICS AND COMPUTING↗

Correspondence between neuroevolution and gradient descent

Abstract We show analytically that training a neural network by conditioned stochastic mutation or neuroevolution of its weights is equivalent, in the limit of small mutations, to gradient descent on the loss function in the presence of Gaussian white noise. Averaged over independent realizations of the learning process, neuroevolution is equivalent to gradient descent on the loss function. We use numerical simulation to show that this correspondence can be observed for finite mutations, for shallow and deep neural networks. Our results provide a connection between two families of neural-network training methods that are usually considered to be fundamentally different.

97 MATHEMATICS AND COMPUTING↗

INSURE: An Information Theory iNspired diSentanglement and pURification modEl for Domain Generalization

Domain Generalization (DG) aims to learn a generalizable model on the unseen target domain by only training on the multiple observed source domains. Although a variety of DG methods have focused on extracting domain-invariant features, the domain-specific class-relevant features have attracted attention and been argued to benefit generalization to the unseen target domain. To take into account the class-relevant domain-specific information, in this paper we propose an Information theory iNspired diSentanglement and pURification modEl (INSURE) to explicitly disentangle the latent features to obtain sufficient and compact (necessary) class-relevant feature for generalization to the unseen domain. Specifically, we first propose an information theory inspired loss function to ensure the disentangled class-relevant features contain sufficient class label information and the other disentangled auxiliary feature has sufficient domain information. Additionally, we further propose a paired purification loss function to let the auxiliary feature discard all the class-relevant information and thus the class-relevant feature will contain sufficient and compact (necessary) class-relevant information. Moreover, instead of using multiple encoders, we propose to use a learnable binary mask as our disentangler to make the disentanglement more efficient and make the disentangled features complementary to each other. We conduct extensive experiments on five widely used DG benchmark datasets including PACS, VLCS, OfficeHome, TerraIncognita, and DomainNet. The proposed INSURE achieves state-of-the-art performance. We also empirically show that domain-specific class-relevant features are beneficial for domain generalization. The code is available at https://github.com/yuxi120407/INSURE .

97 MATHEMATICS AND COMPUTING↗

A robust approach to Gaussian process implementation

Abstract. Gaussian process (GP) regression is a flexible modeling technique used to predict outputs and to capture uncertainty in the predictions. However, the GP regression process becomes computationally intensive when the training spatial dataset has a large number of observations. To address this challenge, we introduce a scalable GP algorithm, termed MuyGPs, which incorporates nearest-neighbor and leave-one-out cross-validation during training. This approach enables the evaluation of large spatial datasets with state-of-the-art accuracy and speed in certain spatial problems. Despite these advantages, conventional quadratic loss functions used in the MuyGPs optimization, such as root mean squared error (RMSE), are highly influenced by outliers. We explore the behavior of MuyGPs in cases involving outlying observations and, subsequently, develop a robust approach to handle and mitigate their impact. Specifically, we introduce a novel leave-one-out loss function based on the pseudo-Huber function (LOOPH) that effectively accounts for outliers in large spatial datasets within the MuyGPs framework. Our simulation study shows that the LOOPH loss method maintains accuracy despite outlying observations, establishing MuyGPs as a powerful tool for mitigating unusual observation impacts in the large data regime. In the analysis of US ozone data, MuyGPs provides accurate predictions and uncertainty quantification, demonstrating its utility in managing data anomalies. Through these efforts, we advance the understanding of GP regression in spatial contexts.

Mukangango, Juliette↗

Thermodynamic Consistent Neural Networks for Learning Material Interfacial Mechanics

For multilayer materials in thin substrate systems, interfacial failure is one of the most challenges. The traction-separation relations (TSR) quantitatively describe the mechanical behavior of a material interface undergoing openings, which is critical to understand and predict interfacial failures under complex loadings. However, existing theoretical models have limitations on enough complexity and flexibility to well learn the real-world TSR from experimental observations. A neural network can fit well along with the loading paths but often fails to obey the laws of physics, due to a lack of experimental data and understanding of the hidden physical mechanism. In this paper, we propose a thermodynamic consistent neural network (TCNN) approach to build a data-driven model of the TSR with sparse experimental data. The TCNN leverages recent advances in physics-informed neural networks (PINN) that encode prior physical information into the loss function and efficiently train the neural networks using automatic differentiation. We investigate three thermodynamic consistent principles, i.e., positive energy dissipation, steepest energy dissipation gradient, and energy conservative loading path. All of them are mathematically formulated and embedded into a neural network model with a novel defined loss function. A real-world experiment demonstrates the superior performance of TCNN, and we find that TCNN provides an accurate prediction of the whole TSR surface and significantly reduces the violated prediction against the laws of physics.

Zhang, Jiaxin↗

A deep learning approach for semantic segmentation of unbalanced data in electron tomography of catalytic materials

In computed TEM tomography, image segmentation represents one of the most basic tasks with implications not only for 3D volume visualization, but more importantly for quantitative 3D analysis. In case of large and complex 3D data sets, segmentation can be an extremely difficult and laborious task, and thus has been one of the biggest hurdles for comprehensive 3D analysis. Heterogeneous catalysts have complex surface and bulk structures, and often sparse distribution of catalytic particles with relatively poor intrinsic contrast, which possess a unique challenge for image segmentation, including the current state-of-the-art deep learning methods. To tackle this problem, we apply a deep learning-based approach for the multi-class semantic segmentation of a γ-Alumina/Pt catalytic material in a class imbalance situation. Specifically, we used the weighted focal loss as a loss function and attached it to the U-Net’s fully convolutional network architecture. We assessed the accuracy of our results using Dice similarity coefficient (DSC), recall, precision, and Hausdorff distance (HD) metrics on the overlap between the ground-truth and predicted segmentations. Our adopted U-Net model with the weighted focal loss function achieved an average DSC score of 0.96 ± 0.003 in the γ-Alumina support material and 0.84 ± 0.03 in the Pt NPs segmentation tasks. We report an average boundary-overlap error of less than 2 nm at the 90th percentile of HD for γ-Alumina and Pt NPs segmentations. The complex surface morphology of γ-Alumina and its relation to the Pt NPs were visualized in 3D by the deep learning-assisted automatic segmentation of a large data set of high-angle annular dark-field (HAADF) scanning transmission electron microscopy (STEM) tomography reconstructions.

36 MATERIALS SCIENCE↗

Attitude determination using vector observations: A fast optimal matrix algorithm

The attitude matrix minimizing Wahba's loss function is computed directly by a method that is competitive with the fastest known algorithm for finding this optimal estimate. The method also provides an estimate of the attitude error covariance matrix. Analysis of the special case of two vector observations identifies those cases for which the TRIAD or algebraic method minimizes Wahba's loss function.

Markley, F. Landis↗

Attitude determination using vector observations - A fast optimal matrix algorithm

The attitude matrix minimizing Wahba's loss function is computed directly by a method that is competitive with the fastest known algorithm for finding this optimal estimate. The method also provides an estimate of the attitude error covariance matrix. Analysis of the special case of two vector observations identifies those cases for which the TRIAD or algebraic method minimizes Wahba's loss function.

Markley, F. L.↗

Deep Learning-Based Adaptive Remedial Action Scheme with Security Margin for Renewable-Dominated Power Grids

The Remedial Action Scheme (RAS) is designed to take corrective actions after detecting predetermined conditions to maintain system transient stability in large interconnected power grids. However, since RAS is usually designed based on a few selected typical operating conditions, it is not optimal in operating conditions that are not considered in the offline design, especially under frequently and dramatically varying operating conditions due to the increasing integration of intermittent renewables. The deep learning-based RAS is proposed to enhance the adaptivity of RAS to varying operating conditions. During the training, a customized loss function is developed to penalize the negative loss and suggest corrective actions with a security margin to avoid triggering under-frequency and over-frequency relays. Simulation results of the reduced United States Western Interconnection system model demonstrate that the proposed deep learning–based RAS can provide optimal corrective actions for unseen operating conditions while maintaining a sufficient security margin.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A Lie algebraic theory of barren plateaus for deep parameterized quantum circuits

Variational quantum computing schemes train a loss function by sending an initial state through a parametrized quantum circuit, and measuring the expectation value of some operator. Despite their promise, the trainability of these algorithms is hindered by barren plateaus (BPs) induced by the expressiveness of the circuit, the entanglement of the input data, the locality of the observable, or the presence of noise. Up to this point, these sources of BPs have been regarded as independent. In this work, we present a general Lie algebraic theory that provides an exact expression for the variance of the loss function of sufficiently deep parametrized quantum circuits, even in the presence of certain noise models. Our results allow us to understand under one framework all aforementioned sources of BPs. This theoretical leap resolves a standing conjecture about a connection between loss concentration and the dimension of the Lie algebra of the circuit’s generators.

97 MATHEMATICS AND COMPUTING↗

Using model order tests to determine sensory inputs in a motion study

In the study of motion effects on tracking performance, a problem of interest is the determination of what sensory inputs a human uses in controlling his tracking task. In the approach presented here a simple canonical model (FID or a proportional, integral, derivative structure) is used to model the human's input-output time series. A study of significant changes in reduction of the output error loss functional is conducted as different permutations of parameters are considered. Since this canonical model includes parameters which are related to inputs to the human (such as the error signal, its derivatives and integration), the study of model order is equivalent to the study of which sensory inputs are being used by the tracker. The parameters are obtained which have the greatest effect on reducing the loss function significantly. In this manner the identification procedure converts the problem of testing for model order into the problem of determining sensory inputs.

Repperger, D. W.↗

Gradient-enhanced physics-informed neural networks for forward and inverse PDE problems

Deep learning has been shown to be an effective tool in solving partial differential equations (PDEs) through physics-informed neural networks (PINNs). PINNs embed the PDE residual into the loss function of the neural network, and have been successfully employed to solve diverse forward and inverse PDE problems. However, one disadvantage of the first generation of PINNs is that they usually have limited accuracy even with many training points. Here, we propose a new method, gradient-enhanced physics-informed neural networks (gPINNs), for improving the accuracy of PINNs. gPINNs leverage gradient information of the PDE residual and embed the gradient into the loss function. We tested gPINNs extensively and demonstrated the effectiveness of gPINNs in both forward and inverse PDE problems. Our numerical results show that gPINN performs better than PINN with fewer training points. Additionally, we combined gPINN with the method of residual-based adaptive refinement (RAR), a method for improving the distribution of training points adaptively during training, to further improve the performance of gPINN, especially in PDEs with solutions that have steep gradients.

42 ENGINEERING↗

Deep Learning Approaches to Surrogates for Solving the Diffusion Equation for Mechanistic Real-World Simulations

In many mechanistic medical, biological, physical, and engineered spatiotemporal dynamic models the numerical solution of partial differential equations (PDEs), especially for diffusion, fluid flow and mechanical relaxation, can make simulations impractically slow. Biological models of tissues and organs often require the simultaneous calculation of the spatial variation of concentration of dozens of diffusing chemical species. One clinical example where rapid calculation of a diffusing field is of use is the estimation of oxygen gradients in the retina, based on imaging of the retinal vasculature, to guide surgical interventions in diabetic retinopathy. Furthermore, the ability to predict blood perfusion and oxygenation may one day guide clinical interventions in diverse settings, i.e., from stent placement in treating heart disease to BOLD fMRI interpretation in evaluating cognitive function (Xie et al., 2019; Lee et al., 2020). Since the quasi-steady-state solutions required for fast-diffusing chemical species like oxygen are particularly computationally costly, we consider the use of a neural network to provide an approximate solution to the steady-state diffusion equation. Machine learning surrogates, neural networks trained to provide approximate solutions to such complicated numerical problems, can often provide speed-ups of several orders of magnitude compared to direct calculation. Surrogates of PDEs could enable use of larger and more detailed models than are possible with direct calculation and can make including such simulations in real-time or near-real time workflows practical. Creating a surrogate requires running the direct calculation tens of thousands of times to generate training data and then training the neural network, both of which are computationally expensive. Often the practical applications of such models require thousands to millions of replica simulations, for example for parameter identification and uncertainty quantification, each of which gains speed from surrogate use and rapidly recovers the up-front costs of surrogate generation. We use a Convolutional Neural Network to approximate the stationary solution to the diffusion equation in the case of two equal-diameter, circular, constant-value sources located at random positions in a two-dimensional square domain with absorbing boundary conditions. Such a configuration caricatures the chemical concentration field of a fast-diffusing species like oxygen in a tissue with two parallel blood vessels in a cross section perpendicular to the two blood vessels. To improve convergence during training, we apply a training approach that uses roll-back to reject stochastic changes to the network that increase the loss function. The trained neural network approximation is about 1000 times faster than the direct calculation for individual replicas. Because different applications will have different criteria for acceptable approximation accuracy, we discuss a variety of loss functions and accuracy estimators that can help select the best network for a particular application. We briefly discuss some of the issues we encountered with overfitting, mismapping of the field values and the geometrical conditions that lead to large absolute and relative errors in the approximate solution.

60 APPLIED LIFE SCIENCES↗

Designing accurate emulators for scientific processes using calibration-driven deep models

Abstract Predictive models that accurately emulate complex scientific processes can achieve speed-ups over numerical simulators or experiments and at the same time provide surrogates for improving the subsequent analysis. Consequently, there is a recent surge in utilizing modern machine learning methods to build data-driven emulators. In this work, we study an often overlooked, yet important, problem of choosing loss functions while designing such emulators. Popular choices such as the mean squared error or the mean absolute error are based on a symmetric noise assumption and can be unsuitable for heterogeneous data or asymmetric noise distributions. We propose Learn-by-Calibrating, a novel deep learning approach based on interval calibration for designing emulators that can effectively recover the inherent noise structure without any explicit priors. Using a large suite of use-cases, we demonstrate the efficacy of our approach in providing high-quality emulators, when compared to widely-adopted loss function choices, even in small-data regimes.

97 MATHEMATICS AND COMPUTING↗

Multi-variance replica exchange SGMCMC for inverse and forward problems via Bayesian PINN

Physics-informed neural network (PINN) has been successfully applied in solving a variety of nonlinear non-convex forward and inverse problems. However, the training is challenging because of the non-convex loss functions and the multiple optima in the Bayesian inverse problem. In this work, we propose a multi-variance replica exchange stochastic gradient Langevin dynamics method to tackle the challenge of the multiple local optima in the optimization and the challenge of the multiple modal posterior distribution in the inverse problem. Replica exchange methods are capable of escaping from the local traps and accelerating the convergence; two chains with different temperatures are designed where the low temperature chain aims for the local convergence, and the target of the high temperature chain is to travel globally and explore the whole loss function entropy landscape. However, it may not be efficient to solve mathematical inversion problems by using the vanilla replica method directly since the method doubles the computational cost in evaluating the forward solvers (likelihood functions) in the two chains. To address this issue, we propose to make different assumptions on the energy function estimation and this facilities one to use solvers of different fidelities in the likelihood function evaluation. More precisely, one can use a solver with low fidelity in the high temperature chain while using a solver with high fidelity in the low temperature chain. Our proposed method significantly lowers the computational cost in the high temperature chain, meanwhile preserving the accuracy and converging very fast. Here we give an unbiased estimate of the swapping rate and give an estimation of the discretization error of the scheme. To verify our idea, we design and solve four inverse problems which have multiple modes. The proposed method is also employed to train the Bayesian PINN to solve the forward and inverse problems; faster and more accurate convergence has been observed when compared to the stochastic gradient Langevin dynamics (SGLD) method and vanilla replica exchange methods.

97 MATHEMATICS AND COMPUTING↗