Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Adaptive Activation Functions Accelerate Convergence in Deep and Physics-informed Neural Networks

We employ adaptive activation functions for regression in deep and physics-informed neural networks (PINNs) to approximate smooth and discontinuous functions as well as solutions of linear and nonlinear partial differential equations. In particular, we solve the nonlinear Klein-Gordon equation, which has smooth solutions, the nonlinear Burgers equation, which can admit high gradient solutions, and the Helmholtz equation. We introduce a scalable hyper-parameter in the activation function, which can be optimized to achieve best performance of the network as it changes dynamically the topology of the loss function involved in the optimization process. The adaptive activation function has better learning capabilities than the traditional one (fixed activation) as it improves greatly the convergence rate, especially at early training, as well as the solution accuracy. To better understand the learning process, we plot the neural network solution in the frequency domain to examine how the network captures successively different frequency bands present in the solution. We consider both forward problems, where the approximate solutions are obtained, as well as inverse problems, where parameters involved in the governing equation are identified. Our simulation results show that the proposed method is a very simple and effective approach to increase the efficiency, robustness and accuracy of the neural network approximation of nonlinear functions as well as solutions of partial differential equations, especially for forward problems. We theoretically prove that in the proposed method, gradient descent algorithms are not attracted to suboptimal critical points or local minima.

machine leaning, Bad minima, Inverse problems, Phy↗

Efficient Training of Deep Neural Operator Networks via Randomized Sampling

Neural operators (NOs) employ deep neural networks to learn the mappings between infinitedimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time prediction of complex dynamics across various scientific and engineering applications. In this work, we introduce a random sampling technique to be adopted during the training of DeepONet, aimed at improving the generalization ability of the model, while significantly reducing the computational time. The proposed approach targets the trunk network of the DeepONet model that outputs the basis functions corresponding to the spatiotemporal locations of the bounded domain on which the physical system is defined. While constructing the loss function, DeepONet training traditionally considers a uniform grid of spatiotemporal points at which all the output functions are evaluated for each iteration. This approach leads to a larger batch size, resulting in poor generalization and increased memory demands, due to the limitations of the stochastic gradient descent (SGD) optimizer. The proposed random sampling over the inputs of the trunk net mitigates these challenges, improving generalization and reducing the memory requirements during training, resulting in significant computational gains. We validate our hypothesis through three benchmark examples, demonstrating substantial reductions in training time while achieving comparable or lower overall test errors relative to the traditional training approach. Our results indicate that incorporating randomization in the trunk network inputs during training enhances the efficiency and robustness of DeepONet, offering a promising avenue for improving the framework’s performance in modeling complex physical systems.

Karumuri, Sharmila [Department of Civil & Systems ↗

On the influence of over-parameterization in manifold based surrogates and deep neural operators

Constructing accurate and generalizable approximators (surrogate models) for complex physico-chemical processes exhibiting highly non-smooth dynamics is challenging. The main question is what type of surrogate models we should construct and should these models be under-parameterized or over-parameterized. In this work, we propose new developments and perform comparisons for two promising approaches: manifold-based polynomial chaos expansion (m-PCE) and the deep neural operator (DeepONet), and we examine the effect of over-parameterization on generalization. While m-PCE enables the construction of a mapping by first identifying low-dimensional embeddings of the input functions, parameters, and quantities of interest (QoIs), a neural operator learns the nonlinear mapping via the use of deep neural networks. Here, we demonstrate the performance of these methods in terms of generalization accuracy by solving the 2D time-dependent Brusselator reaction-diffusion system with uncertainty sources, modeling an autocatalytic chemical reaction between two species. We first propose an extension of the m-PCE by constructing a mapping between latent spaces formed by two separate embeddings of the input functions and the output QoIs. To further enhance the accuracy of the DeepONet, we introduce weight self-adaptivity in the loss function. We demonstrate that the performance of m-PCE and DeepONet is comparable for cases of relatively smooth input-output mappings. However, when highly non-smooth dynamics is considered, DeepONet shows higher approximation accuracy. We also find that for m-PCE, modest over-parameterization leads to better generalization, both within and outside of distribution, whereas aggressive over-parameterization leads to over-fitting. In contrast, an even highly over-parameterized DeepONet leads to better generalization for both smooth and non-smooth dynamics. Furthermore, we compare the performance of the above models with another recently proposed operator learning model, the Fourier Neural Operator, and show that its over-parameterization also leads to better generalization. Taken together, our studies show that m-PCE can provide very good accuracy at very low training cost, whereas a highly over-parameterized DeepONet can provide better accuracy and robustness to noise but at higher training cost. In both methods, the inference cost is negligible.

97 MATHEMATICS AND COMPUTING↗

Plateau Phenomenon in Gradient Descent Training of RELU Networks: Explanation, Quantification, and Avoidance

The ability of neural networks to provide ‘best in class’ approximation across a wide range of applications is well-documented. Nevertheless, the powerful expressivity of neural networks comes to naught if one is unable to effectively train (choose) the parameters defining the network. In general, neural networks are trained by gradient descent type optimization methods,a stochastic variant thereof. In practice, such methods result in the loss function decreases rapidly at the beginning of training but then, after a relatively small number of steps, significantly slow down. The loss may even appear to stagnate over the period of a large number of epochs, only to then suddenly start to decrease fast again for no apparent reason. This so-called plateau phenomenon manifests itself in many learning tasks. The present work aims to identify and quantify the root causes of plateau phenomenon.analysis is carried out in the setting of univariate ReLU networks. No assumptions are made on the number of neurons relative to the number of training data, and our results hold for both the lazy and adaptive regimes. Here, the main findings are: plateaux correspond to periods during which activation patterns remain constant, where activation pattern refers to the number of data points that activate a given neuron; quantification of convergence of the gradient flow dynamics; and, characterization stationary points in terms solutions of local least squares regression lines over subsets of the training data. Based on these conclusions, we propose a new iterative training method, the Active Neuron Least Squares (ANLS), characterised by the explicit adjustment of the activation pattern at each step, which is designed to enable a quick exit from a plateau. Illustrative numerical examples are included throughout.

97 MATHEMATICS AND COMPUTING↗

FunDiff: diffusion models over function spaces for physics-informed generative modeling

Recent advances in generative modeling-particularly diffusion models and flow matching-have been widely used for synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. To address this, we introduce FunDiff, an efficient and robust framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generates continuous functions that can be evaluated at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, demonstrating that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We further demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results indicate that our method can generate physically consistent samples with high fidelity to the target distribution, and exhibit robustness to noisy and low-resolution data.

Wang, Sifan [Yale University, New Haven, CT (Unite↗

Physics-informed Karhunen-Loeve and Neural Network Approximations for Solving Inverse Differential Equation Problems

Here we present the PI-CKL-NN method for parameter estimation in differential equation (DE) models given sparse measurements of the parameters and states. In the proposed approach, the space- or time-dependent parameters are approximated by Karhunen-Loeve (KL) expansions that are conditioned on the parameters’ measurements, and the states are approximated by deep neural networks (DNNs). The unknown weights in the KL expansions and DNNs are found my minimizing the cost function that enforces the measurements of the states the DE constraint. Regularization is achieved by adding the l2 norm of the conditional KL coefficients into the loss function. Our approach assumes that the parameter fields are correlated in space or time and enforces the statistical knowledge (the mean and the covariance function) in addition to the DE constraints and measurements as opposed to the physics-informed neural network (PINN) and other similar physics-informed machine learning methods where only DE constraints and data are used for parameter estimation. We use the PI-CKL-NN method for parameter estimation in an ordinary differential equation with an unknown time-dependent parameter and the one- and two-dimensional partial differential diffusion equations with unknown space-dependent diffusion coefficients. We also demonstrate that PI-CKL-NN is more accurate than the PINN method, especially when the observations of the parameters are very sparse

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

New machine learning techniques for simulation-based inference: InferoStatic nets, kernel score estimation, and kernel likelihood ratio estimation

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

Kong, Kyoungchul↗

Codebase release 0.1 for infstat

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential \varphi φ . In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

Kong, Kyoungchul↗

New Machine Learning Techniques for Simulation-Based Inference: InferoStatic Nets, Kernel Score Estimation, and Kernel Likelihood Ratio Estimation

We propose an intuitive, machine-learning approach to multiparameter inference, dubbed the InferoStatic Networks (ISN) method, to model the score and likelihood ratio estimators in cases when the probability density can be sampled but not computed directly. The ISN uses a backend neural network that models a scalar function called the inferostatic potential $\varphi$. In addition, we introduce new strategies, respectively called Kernel Score Estimation (KSE) and Kernel Likelihood Ratio Estimation (KLRE), to learn the score and the likelihood ratio functions from simulated data. We illustrate the new techniques with some toy examples and compare to existing approaches in the literature. We mention en passant some new loss functions that optimally incorporate latent information from simulations into the training procedure.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Transforming the $v$ World: A New Multivariate Transformer Energy Estimator for NOvA

The NOvA Transformer Energy Estimator (Transformer_EE) is a universal machine learning tool currently used to infer the incoming beam neutrino energy and the outgoing lepton energy in both near andfar detectors. It uses a unique, highly flexible framework for simultaneous multivariate prediction that supports many possible loss functions. A spectral reweighting and flattening scheme lessens training bias. A feature noising subroutine enables adversarial-like training, mitigating sensitivities to certain systematic effects at marginal resolution loss at inference time. The state of the Transformer_EE will be reviewed, and its robustness with respect to several NOvA Near and Far Detector systematics highlighted.

Tong, Leon [Minnesota U.] (ORCID:0000000231625965)↗

Triangle Method for Dense ReLU Layers [SWR-25-72]

This software is an implementation of the methods for initializing and training neural networks to be more efficient per parameter, described more fully below and in the related publication: In theory, depth should make a ReLU network EXPONENTIALLY more efficient by enabling it to produce an exponential number of piecewise linear sections in its output. This reasoning is largely based on the work of mathematicians that have hand-constructed networks that make good use of depth. In practice however, even very deep ReLU networks that have been randomly initialized will behave identically to their shallow counterparts - missing an entire exponential dimension of efficiency. The triangle method is a first attempt at realizing the exponential potential of deep networks. Instead of randomly setting weights, we force pairs of neurons in each layer learn to build triangles (i.e. functions from [0,1] -> [0,1] that look like triangles). This is a very efficient pattern for generating lots of linear pieces because composing two triangular functions doubles the number of pieces with each composition. The triangle method is more than just a different initialization, it is a new paradigm of training. Instead of making direct updates to the matrix weights, we do an extra step of backpropagation to collect the derivatives of the loss function with respect to the shapes of the triangles, training them to tilt left or right. This process essentially holds the networks hand throughout the loss landscape and forces it to always use depth effectively by producing triangular shapes internally. This can produce several orders of magnitude of improvement on convex one-dimensional regression problems. Much more theoretical work is needed to realize its full potential beyond this context, but the implementation in this repository will still work in arbitrary numbers of dimensions. The file Triangle_Method.py is a generalized form of the method that will build each neuron its own custom 1-d convex activation function (with exponential efficiency). Example usage on one dimensional problems can be found in Example_Usage.ipynb and an example of using this in a real neural network can be found in Example_VGG16_CIFAR10.ipynb.

Milkert, Max [National Renewable Energy Laboratory↗

Graph Metric Learning Quantifies Morphological Differences between Two Genotypes of Shoot Apical Meristem Cells in Arabidopsis

We present a method for learning “spectrally descriptive” edge weights for graphs. We generalize a previously known distance measure on graphs (Graph Diffusion Distance), thereby allowing it to be tuned to minimize an arbitrary loss function. Because all steps involved in calculating this modified GDD are differentiable, we demonstrate that it is possible for a small neural network model to learn edge weights which minimize loss. We apply this method to discriminate between graphs constructed from shoot apical meristem images of two genotypes of Arabidopsis thaliana specimens: wild-type and trm678 triple mutants with cell division phenotype. Training edge weights and kernel parameters with contrastive loss produces a learned distance metric with large margins between these graph categories. We demonstrate this by showing improved performance of a simple k-nearest-neighbors classifier on the learned distance matrix. We also demonstrate a further application of this method to biological image analysis. Once trained, we use our model to compute the distance between the biological graphs and a set of graphs output by a cell division simulator. Comparing simulated cell division graphs to biological ones allows us to identify simulation parameter regimes which characterize mutant vs. wild-type Arabidopsis cells. We find that trm678 mutant cells are characterized by increased randomness of division planes and decreased ability to avoid previous vertices between cell walls.

59 BASIC BIOLOGICAL SCIENCES↗

Chapter 4: Physically informed deep learning networks for simulating microstructure evolution of 3D polycrystals

As discussed in the previous chapter, high energy diffraction microscopy (HEDM) is used to study the micromechanical evolution of a material during in situ loading. HEDM experiments have been used to verify crystal plasticity (CP) simulations [119, 91, 90, 120], for experimental planning, material design, and to further analyze experimental results. However, Fast Fourier transform-based CP (CP-FFT) or finite element-based CP (CP-FE) methods are often too slow to be used in real-time during an experiment. CP-FFT is faster than CP-FE simulations due to the absence of meshing, but can still take hours to simulate the response of a single volume depending on the size and number of strain steps [127]. Reducing computation time would create a larger exploration space in planning and design, and enable faster analysis of experimental results and real-time feedback during an experiment. This research expands upon previous works to develop a workflow for predicting the full-field evolution of a 3D polycrystal. The workflow is simplified from previous works to predict only orientation and elastic strain tensors (from which stress tensors are calculated). The network is physically informed through loss functions and network architecture for a more robust model. The orientation predictions are informed about the cubic crystal symmetry of the material by incorporating disorientation and misorientation information into the network architecture and loss. The Von Mises stress is used to enforce the correct stress-strain trends in the strain tensor predictions. Additional total strain steps from the elastic and elastoplastic region are included to better capture the stress-strain evolution at smaller total strain steps. Material and hardening parameters are additional inputs into the networks to further inform the network and to study the network’s ability to predict different materials other than those used for training.

36 MATERIALS SCIENCE↗

Encoder–decoder neural network for solving the nonlinear Fokker–Planck–Landau collision operator in XGC

An encoder–decoder neural network has been used to examine the possibility for acceleration of a partial integro-differential equation, the Fokker–Planck–Landau collision operator. This is part of the governing equation in the massively parallel particle-in-cell code XGC, which is used to study turbulence in fusion energy devices. The neural network emphasizes physics-inspired learning, where it is taught to respect physical conservation constraints of the collision operator by including them in the training loss, along with the ℓ 2 loss. In particular, network architectures used for the computer vision task of semantic segmentation have been used for training. A penalization method is used to enforce the ‘soft’ constraints of the system and integrate error in the conservation properties into the loss function. During training, quantities representing the particle density, momentum and energy for all species of the system are calculated at each configuration vertex, mirroring the procedure in XGC. This simple training has produced a median relative loss, across configuration space, of the order of 10 –4 , which is low enough if the error is of random nature, but not if it is of drift nature in time steps. The run time for the current Picard iterative solver of the operator is O(n 2 ), where n is the number of plasma species. As the XGC1 code begins to attack problems including a larger number of species, the collision operator will become expensive computationally, making the neural network solver even more important, especially since its training only scales as O(n). Here, a wide enough range of collisionality has been considered in the training data to ensure the full domain of collision physics is captured. An advanced technique to decrease the losses further will be subject of a subsequent report. Eventual work will include expansion of the network to include multiple plasma species.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Deep transfer operator learning for partial differential equations under conditional shift

Transfer learning enables the transfer of knowledge gained while learning to perform one task (source) to a related but different task (target), hence addressing the expense of data acquisition and labelling, potential computational power limitations and dataset distribution mismatches. Here, we propose a new transfer learning framework for task-specific learning (functional regression in partial differential equations) under conditional shift based on the deep operator network (DeepONet). Task-specific operator learning is accomplished by fine-tuning task-specific layers of the target DeepONet using a hybrid loss function that allows for the matching of individual target samples while also preserving the global properties of the conditional distribution of the target data. Inspired by conditional embedding operator theory, we minimize the statistical distance between labelled target data and the surrogate prediction on unlabelled target data by embedding conditional distributions onto a reproducing kernel Hilbert space. We demonstrate the advantages of our approach for various transfer learning scenarios involving nonlinear partial differential equations under diverse conditions due to shifts in the geometric domain and model dynamics. Our transfer learning framework enables fast and efficient learning of heterogeneous tasks despite considerable differences between the source and target domains.

42 ENGINEERING↗

Shape restricted additive hazards models: Monotone, unimodal, and U‐shape hazard functions

We consider estimation of the semiparametric additive hazards model with an unspecified baseline hazard function where the effect of a continuous covariate has a specific shape but otherwise unspecified. Such estimation is particularly useful for a unimodal hazard function, where the hazard is monotone increasing and monotone decreasing with an unknown mode. A popular approach of the proportional hazards model is limited in such setting due to the complicated structure of the partial likelihood. Our model defines a quadratic loss function, and its simple structure allows a global Hessian matrix that does not involve parameters. Thus, once the global Hessian matrix is computed, a standard quadratic programming method can be applicable by profiling all possible locations of the mode. However, the quadratic programming method may be inefficient to handle a large global Hessian matrix in the profiling algorithm due to a large dimensionality, where the dimension of the global Hessian matrix and number of hypothetical modes are the same order as the sample size. We propose the quadratic pool adjacent violators algorithm to reduce computational costs. The proposed algorithm is extended to the model with a time‐dependent covariate with monotone or U‐shape hazard function. In simulation studies, our proposed method improves computational speed compared to the quadratic programming method, with bias and mean square error reductions. We analyze data from a recent cardiovascular study.

Mathematical & Computational Biology↗

Exact enforcement of temporal continuity in sequential physics-informed neural networks

The use of deep learning methods in scientific computing represents a potential paradigm shift in engineering problem solving. One of the most prominent developments is Physics-Informed Neural Networks (PINNs), in which neural networks are trained to satisfy partial differential equations (PDEs). While this method shows promise, the standard version has been shown to struggle in accurately predicting the dynamic behavior of time-dependent problems. To address this challenge, methods have been proposed that decompose the time domain into multiple segments, employing a distinct neural network in each segment and directly incorporating continuity between them in the loss function of the minimization problem. In this work we introduce a method to exactly enforce continuity between successive time segments via a solution ansatz. This hard constrained sequential PINN (HCS-PINN) method is simple to implement and eliminates the need for any loss terms associated with temporal continuity. The method is tested for a number of benchmark problems involving both linear and non-linear PDEs. Examples include various first order time dependent problems in which traditional PINNs struggle, namely advection, Allen–Cahn, and Korteweg–de Vries equations. Furthermore, second and third order time-dependent problems are demonstrated via wave and Jerky dynamics examples, respectively. Notably, the Jerky dynamics problem is chaotic, making the problem especially sensitive to temporal accuracy. Finally, the numerical experiments conducted with the proposed method demonstrated superior convergence and accuracy over both traditional PINNs and the soft-constrained counterparts.

42 ENGINEERING↗

A novel technique for minimizing energy functional using neural networks

An energy functional describes the equilibrium state of a system. In this work, we present a novel technique, Functional Optimization using Neural Networks (FONN), for minimizing the system’s energy. FONN utilizes neural networks to process information at discrete grid points, considering their interactions with neighboring grid points, to update the state of the system. The training process involves formulating a loss function based on the system’s energy, and with the help of multiple fine-tuning steps, the method employs a progressive energy reduction technique that decreases the energy in multiple steps. FONN’s effectiveness is demonstrated across various problems, including the minimization of the heat and Lyapunov energy. Furthermore, the paper explores the minimization of the elastic bending energy with an area constraint.

97 MATHEMATICS AND COMPUTING↗