Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Entropy-Infused Deep Learning Loss Function for Capturing Extreme Values in Wind Power Forecasting

Extreme scenarios in wind power generation occur with higher frequency and larger magnitude in the recent years due to the ever-increasing extreme meteorological factors. Accurate forecasting of the occurrence of extreme values in wind power generation is of great concern to ensure reliable power system operation. Recently, deep learning models have surged in popularity for wind power forecasting, with the mean squared error (MSE) loss function being commonly used. However, the MSE loss function, being sensitive to extreme values, disproportionately penalizes larger errors, cannot adequately capture the extreme values present in wind energy data, and novel loss functions have seldom been tailored for wind power forecasting. To this end, in this paper, we introduce a novel loss function specifically crafted to capture extreme values in wind power forecasting. The experimental results with four fundamental deep learning methods on open source wind power dataset validate that the new loss function is efficient and superior in all cases compared to MSE in capturing extreme values while maintaining forecasting performance.

17 WIND ENERGY↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Levenberg–Marquardt multi-classification using hinge loss function

Incorporating higher-order optimization functions, such as Levenberg-Marquardt (LM) have revealed better generalizable solutions for deep learning problems. However, these higher-order optimization functions suffer from very large processing time and training complexity especially as training datasets become large, such as in multi-view classification problems, where finding global optima is a very costly problem. To solve this issue, we develop a solution for LM-enabled classification with, to the best of knowledge first-time implementation of hinge loss, for multiview classification. Hinge loss allows the neural network to converge faster and perform better than other loss functions such as logistic or square loss rates. Here we prove our method by experimenting with various multiclass classification challenges of varying complexity and training data size. The empirical results show the training time and accuracy rates achieved, highlighting how our method outperforms in all cases, especially when training time is limited. Our paper presents important results in the relationship between optimization and loss functions and how these can impact deep learning problems.

97 MATHEMATICS AND COMPUTING↗

Precision measurement of the electron energy-loss function in tritium and deuterium gas for the KATRIN experiment

Abstract The KATRIN experiment is designed for a direct and model-independent determination of the effective electron anti-neutrino mass via a high-precision measurement of the tritium $$\upbeta $$ β -decay endpoint region with a sensitivity on $$m_\nu $$ m ν of 0.2 $$\hbox {eV}/\hbox {c}^2$$ eV / c 2 (90% CL). For this purpose, the $$\upbeta $$ β -electrons from a high-luminosity windowless gaseous tritium source traversing an electrostatic retarding spectrometer are counted to obtain an integral spectrum around the endpoint energy of 18.6 keV. A dominant systematic effect of the response of the experimental setup is the energy loss of $$\upbeta $$ β -electrons from elastic and inelastic scattering off tritium molecules within the source. We determined the energy-loss function in-situ with a pulsed angular-selective and monoenergetic photoelectron source at various tritium-source densities. The data was recorded in integral and differential modes; the latter was achieved by using a novel time-of-flight technique. We developed a semi-empirical parametrization for the energy-loss function for the scattering of 18.6-keV electrons from hydrogen isotopologs. This model was fit to measurement data with a 95% $$\hbox {T}_2$$ T 2 gas mixture at 30 K, as used in the first KATRIN neutrino-mass analyses, as well as a $$\hbox {D}_2$$ D 2 gas mixture of 96% purity used in KATRIN commissioning runs. The achieved precision on the energy-loss function has abated the corresponding uncertainty of $$\sigma (m_\nu ^2)< {{10}^{-2}}{\hbox {eV}^{2}}$$ σ ( m ν 2 ) < 10 - 2 eV 2 [1] in the KATRIN neutrino-mass measurement to a subdominant level.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Adaptive Quantum Generative Training using an Unbounded Loss Function

We propose a generative quantum learning algorithm using the Adaptive Derivative-Assembled Problem Tailored ansatz (ADAPT) framework in which the loss function to be minimized is the maximal quantum Rényi divergence of order two, an unbounded function that mitigates barren plateaus which inhibit training variational circuits. We benchmark this method against other state-of-the-art adaptive algorithms by learning random two-local thermal states. We perform numerical experiments of up to 12 qubits comparing our method learning algorithms that use linear objective functions and show that Rényi-ADAPT is capable of constructing shallow quantum circuits competitive with existing methods, while the gradients remain favorable resulting from the maximal Rényi divergence loss function.

quantum algorithms, quantum machine learning, quan↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

CrysFieldExplorer : rapid optimization of the crystal field Hamiltonian

A new approach to the fast optimization of crystal electric field (CEF) parameters to fit experimental data is presented. This approach is implemented in a lightweight Python-based program, CrysFieldExplorer. The main novelty of the method is the development of a unique loss function, referred to as the spectrum characteristic loss (L Spectrum ), which is based on the characteristic polynomial of the Hamiltonian matrix. Particle swarm optimization and a covariance matrix adaptation evolution strategy are used to find the minimum of the total loss function. It is demonstrated that CrysFieldExplorer can perform direct fitting of CEF parameters to any experimental data such as a neutron spectrum, susceptibility or magnetization measurements etc. CrysFieldExplorer can handle a large number of non-zero CEF parameters and reveal multiple local and global minimum solutions. Crystal field theory, the loss function, and the implementation and limitations of the program are discussed within the context of two examples.

42 ENGINEERING↗

Quantum Neural Networks: Issues, Training, and Applications

Our work in the field aims at explaining the limitations and expressive power of Quantum Machine Learning models, as well as finding feasible training algorithms that could be implemented in near-term Quantum Computers. The promise of Quantum Machine Learning is that by incorporating quantum effects, such as entanglement, into machine learning models researchers can improve model performance and understand more complex datasets. This pledge is particularly pronounced in the design of Quantum neural networks (QNNs), a promising framework for creating quantum algorithms, that promise to outperform classical models by combining the speedups of quantum computation with the widespread successes of deep learning. We show that applying this approach alone to quantum deep learning is problematic given that an excess of entanglement between the hidden and visible layers can destroy the predictive power of our QNN models. We address the barren plateau problem by suggesting the use of a generative, unbounded, nonlinear loss function with simple gradients. The loss function quantifies how much the quantum states generated by the QNNs differ from the data and the goal during training is to minimize it. Finally, we showcase how to use generative training to construct a "classical-quantum" neural network to accurately interpolate between the ground states of a Molecular Hamiltonian, a central question in Quantum Chemistry.

97 MATHEMATICS AND COMPUTING↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Applying Physics-Informed Neural Networks to Solve Navier–Stokes Equations for Laminar Flow around a Particle

In recent years, Physics-Informed Neural Networks (PINNs) have drawn great interest among researchers as a tool to solve computational physics problems. Unlike conventional neural networks, which are black-box models that “blindly” establish a correlation between input and output variables using a large quantity of labeled data, PINNs directly embed physical laws (primarily partial differential equations) within the loss function of neural networks. By minimizing the loss function, this approach allows the output variables to automatically satisfy physical equations without the need for labeled data. The Navier–Stokes equation is one of the most classic governing equations in thermal fluid engineering. This study constructs a PINN to solve the Navier–Stokes equations for a 2D incompressible laminar flow problem. Flows passing around a 2D circular particle are chosen as the benchmark case, and an elliptical particle is also examined to enrich the research. The velocity and pressure fields are predicted by the PINNs, and the results are compared with those derived from Computational Fluid Dynamics (CFD). Additionally, the particle drag force coefficient is calculated to quantify the discrepancy in the results of the PINNs as compared to CFD outcomes. The drag coefficient maintained an error within 10% across all test scenarios.

Hu, Beichao (ORCID:0009000163215151)↗

Train, Inform, Borrow, or Combine? Approaches to Process–Guided Deep Learning for Groundwater–Influenced Stream Temperature Prediction

Although groundwater discharge is a critical stream temperature control process, it is not explicitly represented in many stream temperature models, an omission that may reduce predictive accuracy, hinder management of aquatic habitat, and decrease user confidence. We assessed the performance of a previously-described process-guided deep learning model of stream temperature in the Delaware River Basin (USA). We found lower accuracy (root mean square error [RMSE] of 1.71 versus 1.35°C) and stronger seasonal bias (absolute mean monthly bias of 1.06 vs. 0.68°C) for reaches primarily influenced by deep groundwater as compared to atmospheric conditions. We then tested four approaches for improving groundwater process representation: (a) a custom loss function leveraging the unique patterns of air and water temperature coupling characteristic of different temperature drivers, (b) inclusion of additional groundwater-relevant catchment attributes, (c) incorporation of additional process model outputs, and (d) a composite model. The custom loss function and the additional attributes significantly improved the predictive accuracy in groundwater-dominated reaches (RMSE of 1.37 and 1.26°C) and reduced the seasonal bias (absolute mean monthly bias of 0.44 and 0.48°C), but neither approach could identify holdout groundwater reaches. Variable importance analysis indicates the custom loss function nudges the model to use the existing inputs more efficiently, whereas with the added features the model relies on a broader suite of inputs. This analysis is a substantial step toward more accurately representing groundwater discharge processes in stream temperature models and will improve predictive accuracy and inform habitat management.

54 ENVIRONMENTAL SCIENCES↗

Transcription Factor 4 loss-of-function is associated with deficits in progenitor proliferation and cortical neuron content

Transcription Factor 4 ( TCF4) has been associated with autism, schizophrenia, and other neuropsychiatric disorders. However, how pathological TCF4 mutations affect the human neural tissue is poorly understood. Here, we derive neural progenitor cells, neurons, and brain organoids from skin fibroblasts obtained from children with Pitt-Hopkins Syndrome carrying clinically relevant mutations in TCF4 . We show that neural progenitors bearing these mutations have reduced proliferation and impaired capacity to differentiate into neurons. We identify a mechanism through which TCF4 loss-of-function leads to decreased Wnt signaling and then to diminished expression of SOX genes, culminating in reduced progenitor proliferation in vitro. Moreover, we show reduced cortical neuron content and impaired electrical activity in the patient-derived organoids, phenotypes that were rescued after correction of TCF4 expression or by pharmacological modulation of Wnt signaling. This work delineates pathological mechanisms in neural cells harboring TCF4 mutations and provides a potential target for therapeutic strategies for genetic disorders associated with this gene.

59 BASIC BIOLOGICAL SCIENCES↗

LossLens: Diagnostics for Machine Learning Through Loss Landscape Visual Analytics

Modern machine learning often relies on optimizing a neural network's parameters using a loss function to learn complex features. Beyond training, examining the loss function with respect to a network's parameters (i.e., as a loss landscape) can reveal insights into the architecture and learning process. While the local structure of the loss landscape surrounding an individual solution can be characterized using a variety of approaches, the global structure of a loss landscape, which includes potentially many local minima corresponding to different solutions, remains far more difficult to conceptualize and visualize. To address this difficulty, we introduce LossLens, a visual analytics framework that explores loss landscapes at multiple scales. LossLens integrates metrics from global and local scales into a comprehensive visual representation, enhancing model diagnostics. Here we demonstrate LossLens through two case studies: visualizing how residual connections influence a ResNet-20, and visualizing how physical parameters influence a physics-informed neural network (PINN) solving a simple convection problem.

97 MATHEMATICS AND COMPUTING↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Multilabel proportion prediction and out-of-distribution detection on gamma spectra of short-lived fission products

In the machine learning problem of multilabel classification, the objective is to determine for each test instance which classes the instance belongs to. In this work, we consider an extension of multilabel classification, called multilabel proportion prediction, in the context of radioisotope identification (RIID) using gamma spectra data. We aim to not only predict radioisotope proportions, but also identify out-of-distribution (OOD) spectra. We achieve this goal by viewing gamma spectra as discrete probability distributions, and based on this perspective, we develop a custom semi-supervised loss function that combines a traditional supervised loss with an unsupervised reconstruction error function. Our approach was motivated by its application to the analysis of short-lived fission products from spent nuclear fuel. In particular, we demonstrate that a neural network model trained with our loss function can successfully predict the relative proportions of 37 radioisotopes simultaneously. The model trained with synthetic data was then applied to measurements taken by Pacific Northwest National Laboratory (PNNL) to conduct analysis typically done by subject-matter experts. Here, we also extend our approach to successfully identify when measurements are OOD, and thus should not be trusted, whether due to the presence of a novel source or novel proportions.

Anomaly detection↗

HomPINNs: Homotopy physics-informed neural networks for learning multiple solutions of nonlinear elliptic differential equations

Physics-informed neural networks (PINNs) based machine learning is an emerging framework for solving nonlinear differential equations. However, due to the implicit regularity of neural network structure, PINNs can only find the flattest solution in most cases by minimizing the loss functions. In this paper, we combine PINNs with the homotopy continuation method, a classical numerical method to compute isolated roots of polynomial systems, and propose a new deep learning framework, named homotopy physics-informed neural networks (HomPINNs), for solving multiple solutions of nonlinear elliptic differential equations. The implementation of an HomPINN is a homotopy process that is composed of the training of a fully connected neural network, named the starting neural network, and training processes of several PINNs with different tracking parameters. The starting neural network is to approximate a starting function constructed by the trivial solutions, while other PINNs are to minimize the loss functions defined by boundary condition and homotopy functions, varying with different tracking parameters. These training processes are regraded as different steps of a homotopy process, and a PINN is initialized by the well-trained neural network of the previous step, while the first starting neural network is initialized using the default initialization method. Finally, several numerical examples are presented to show the efficiency of our proposed HomPINNs, including reaction-diffusion equations with a heart-shaped domain.

97 MATHEMATICS AND COMPUTING↗

Z-Target Radiography Postprocessing With A Deep Convolution Neural Network

Analyzing X-ray radiographs is crucial for understanding target behavior in Inertial Confinement Fusion (ICF) and High Energy Density (HED) platforms. However, the density of Magneto Raleigh Taylor (MRT) bands and limitations of target materials often obscure relevant spike growth and density information. To address this issue, machine learning postprocessing techniques can be applied to remove darkened regions in radiography images. In this study, a novel method is presented for removing MRT darkened regions from z-target radiographs using a convolutional neural network (CNN). The CNN, consisting of six layers, treats the darkened regions as noise and employs a mixed loss function and end-to-end frameworks to suppress them while preserving sharpness. The six-layer architecture is designed to effectively learn features when provided with a larger volume of learning space. Each layer is optimized using a mixed loss function that combines a standard loss pixel approach with a multi-scaled structural similarity index loss, which considers luminance, contrast, and structure in local neighborhoods. This approach is particularly beneficial for capturing the stochastic structure of MRT limbs. Due to the limited availability of experimental data, training is conducted using synthetic target radiography from 3D Alegra simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗