Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Entropy-Infused Deep Learning Loss Function for Capturing Extreme Values in Wind Power Forecasting

Extreme scenarios in wind power generation occur with higher frequency and larger magnitude in the recent years due to the ever-increasing extreme meteorological factors. Accurate forecasting of the occurrence of extreme values in wind power generation is of great concern to ensure reliable power system operation. Recently, deep learning models have surged in popularity for wind power forecasting, with the mean squared error (MSE) loss function being commonly used. However, the MSE loss function, being sensitive to extreme values, disproportionately penalizes larger errors, cannot adequately capture the extreme values present in wind energy data, and novel loss functions have seldom been tailored for wind power forecasting. To this end, in this paper, we introduce a novel loss function specifically crafted to capture extreme values in wind power forecasting. The experimental results with four fundamental deep learning methods on open source wind power dataset validate that the new loss function is efficient and superior in all cases compared to MSE in capturing extreme values while maintaining forecasting performance.

17 WIND ENERGY↗

Improved loss functions for machine-learned atomic potentials

Machine learning (ML) has become an invaluable tool across a wide array of domains in science as researchers find new ways to leverage its predictive power. This is especially true in chemistry, where ML is used to fit chemical properties or desirable attributes to the local structure of molecules and materials. In the pursuit of greater accuracy, it is relatively simple to increase the size or complexity of such models, although this often requires simultaneously seeking larger datasets in order to both fit and interpret the larger number of parameters. However, it is equally important to assess the quality and relative importance of the data and how these factors impact the training process. We, therefore, investigate the impact of using different loss functions for training neural network potentials (NNPs), as the loss function defines the error and parameter gradients used to train the NNP. In particular, we test the mean-squared error and Huber loss functions and, using insight from these functions, derive a new loss function based on the Asinh function, which yields significant improvement in the accuracy and generality of NNPs. We show that by discounting/minimizing errors and anomalies in the optimization process, both the Huber and Asinh loss functions improve the training of NNPs, leading to a final potential with a greater effective dimensionality.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Levenberg–Marquardt multi-classification using hinge loss function

Incorporating higher-order optimization functions, such as Levenberg-Marquardt (LM) have revealed better generalizable solutions for deep learning problems. However, these higher-order optimization functions suffer from very large processing time and training complexity especially as training datasets become large, such as in multi-view classification problems, where finding global optima is a very costly problem. To solve this issue, we develop a solution for LM-enabled classification with, to the best of knowledge first-time implementation of hinge loss, for multiview classification. Hinge loss allows the neural network to converge faster and perform better than other loss functions such as logistic or square loss rates. Here we prove our method by experimenting with various multiclass classification challenges of varying complexity and training data size. The empirical results show the training time and accuracy rates achieved, highlighting how our method outperforms in all cases, especially when training time is limited. Our paper presents important results in the relationship between optimization and loss functions and how these can impact deep learning problems.

97 MATHEMATICS AND COMPUTING↗

Precision measurement of the electron energy-loss function in tritium and deuterium gas for the KATRIN experiment

Abstract The KATRIN experiment is designed for a direct and model-independent determination of the effective electron anti-neutrino mass via a high-precision measurement of the tritium $$\upbeta $$ β -decay endpoint region with a sensitivity on $$m_\nu $$ m ν of 0.2 $$\hbox {eV}/\hbox {c}^2$$ eV / c 2 (90% CL). For this purpose, the $$\upbeta $$ β -electrons from a high-luminosity windowless gaseous tritium source traversing an electrostatic retarding spectrometer are counted to obtain an integral spectrum around the endpoint energy of 18.6 keV. A dominant systematic effect of the response of the experimental setup is the energy loss of $$\upbeta $$ β -electrons from elastic and inelastic scattering off tritium molecules within the source. We determined the energy-loss function in-situ with a pulsed angular-selective and monoenergetic photoelectron source at various tritium-source densities. The data was recorded in integral and differential modes; the latter was achieved by using a novel time-of-flight technique. We developed a semi-empirical parametrization for the energy-loss function for the scattering of 18.6-keV electrons from hydrogen isotopologs. This model was fit to measurement data with a 95% $$\hbox {T}_2$$ T 2 gas mixture at 30 K, as used in the first KATRIN neutrino-mass analyses, as well as a $$\hbox {D}_2$$ D 2 gas mixture of 96% purity used in KATRIN commissioning runs. The achieved precision on the energy-loss function has abated the corresponding uncertainty of $$\sigma (m_\nu ^2)< {{10}^{-2}}{\hbox {eV}^{2}}$$ σ ( m ν 2 ) < 10 - 2 eV 2 [1] in the KATRIN neutrino-mass measurement to a subdominant level.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Adaptive Quantum Generative Training using an Unbounded Loss Function

We propose a generative quantum learning algorithm using the Adaptive Derivative-Assembled Problem Tailored ansatz (ADAPT) framework in which the loss function to be minimized is the maximal quantum Rényi divergence of order two, an unbounded function that mitigates barren plateaus which inhibit training variational circuits. We benchmark this method against other state-of-the-art adaptive algorithms by learning random two-local thermal states. We perform numerical experiments of up to 12 qubits comparing our method learning algorithms that use linear objective functions and show that Rényi-ADAPT is capable of constructing shallow quantum circuits competitive with existing methods, while the gradients remain favorable resulting from the maximal Rényi divergence loss function.

quantum algorithms, quantum machine learning, quan↗

Effect of coronal elemental abundances on the radiative loss function

The solar photosphere and corona abundances tabulated by Meyer (1985) and the chromospheric abundances given by Murphy (1985) are used here to recalculate radiative loss functions for equilibrium, low-density, optically thin plasmas. Results from a representative standard photospheric abundance set and from coronal and chromospheric abundance sets showing depletions of up to a factor of four in certain elemental abundances are compared. A significant difference is found for both the coronal and chromospheric abundance sets, with the peak of the radiative loss curve shifted closer to 10 to the 6th K than to the standard 2 x 10 to the 5th K found from photospheric abundances. Consequences of these new calculations, in particular for the cool loop model of Antiochos and Noci (1986), are discussed.

Cook, J. W.↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

CrysFieldExplorer : rapid optimization of the crystal field Hamiltonian

A new approach to the fast optimization of crystal electric field (CEF) parameters to fit experimental data is presented. This approach is implemented in a lightweight Python-based program, CrysFieldExplorer. The main novelty of the method is the development of a unique loss function, referred to as the spectrum characteristic loss (L Spectrum ), which is based on the characteristic polynomial of the Hamiltonian matrix. Particle swarm optimization and a covariance matrix adaptation evolution strategy are used to find the minimum of the total loss function. It is demonstrated that CrysFieldExplorer can perform direct fitting of CEF parameters to any experimental data such as a neutron spectrum, susceptibility or magnetization measurements etc. CrysFieldExplorer can handle a large number of non-zero CEF parameters and reveal multiple local and global minimum solutions. Crystal field theory, the loss function, and the implementation and limitations of the program are discussed within the context of two examples.

42 ENGINEERING↗

Quantum Neural Networks: Issues, Training, and Applications

Our work in the field aims at explaining the limitations and expressive power of Quantum Machine Learning models, as well as finding feasible training algorithms that could be implemented in near-term Quantum Computers. The promise of Quantum Machine Learning is that by incorporating quantum effects, such as entanglement, into machine learning models researchers can improve model performance and understand more complex datasets. This pledge is particularly pronounced in the design of Quantum neural networks (QNNs), a promising framework for creating quantum algorithms, that promise to outperform classical models by combining the speedups of quantum computation with the widespread successes of deep learning. We show that applying this approach alone to quantum deep learning is problematic given that an excess of entanglement between the hidden and visible layers can destroy the predictive power of our QNN models. We address the barren plateau problem by suggesting the use of a generative, unbounded, nonlinear loss function with simple gradients. The loss function quantifies how much the quantum states generated by the QNNs differ from the data and the goal during training is to minimize it. Finally, we showcase how to use generative training to construct a "classical-quantum" neural network to accurately interpolate between the ground states of a Molecular Hamiltonian, a central question in Quantum Chemistry.

97 MATHEMATICS AND COMPUTING↗

Attitude-Independent Magnetometer Calibration for Spin-Stabilized Spacecraft

The paper describes a three-step estimator to calibrate a Three-Axis Magnetometer (TAM) using TAM and slit Sun or star sensor measurements. In the first step, the Calibration Utility forms a loss function from the residuals of the magnitude of the geomagnetic field. This loss function is minimized with respect to biases, scale factors, and nonorthogonality corrections. The second step minimizes residuals of the projection of the geomagnetic field onto the spin axis under the assumption that spacecraft nutation has been suppressed by a nutation damper. Minimization is done with respect to various directions of the body spin axis in the TAM frame. The direction of the spin axis in the inertial coordinate system required for the residual computation is assumed to be unchanged with time. It is either determined independently using other sensors or included in the estimation parameters. In both cases all estimation parameters can be found using simple analytical formulas derived in the paper. The last step is to minimize a third loss function formed by residuals of the dot product between the geomagnetic field and Sun or star vector with respect to the misalignment angle about the body spin axis. The method is illustrated by calibrating TAM for the Fast Auroral Snapshot Explorer (FAST) using in-flight TAM and Sun sensor data. The estimated parameters include magnetic biases, scale factors, and misalignment angles of the spin axis in the TAM frame. Estimation of the misalignment angle about the spin axis was inconclusive since (at least for the selected time interval) the Sun vector was about 15 degrees from the direction of the spin axis; as a result residuals of the dot product between the geomagnetic field and Sun vectors were to a large extent minimized as a by-product of the second step.

Natanson, Gregory↗

Toward Physics-informed Neural Networks for 3D Multi-layer Cloud Mask Reconstruction

Three-dimensional (3D) cloud retrievals are critical for understanding their impact on climate and other applications such as aviation safety, weather prediction, and remote sensing. However, obtaining high-resolution and accurate vertical representation of clouds remains unsolved due to the limitations imposed by satellite instrumentation, viewing conditions, and the complexity of cloud dynamics. Cloud masks are essential for comprehending various cloud vertical properties, but deriving accurate 3D cloud masks from 2D satellite imagery data is a challenging task. To tackle these challenges, we introduce a physics-informed loss function for training deep learning models that can extend 2D cloud images into 3D cloud masks. The proposed loss, called CloudMask Loss, is composed of two domain knowledge-informed loss terms: one for evaluating cloud position and thickness, and the other for measuring the number of layers. By combining these loss terms, we improve the trainability of the deep learning models for more accurate and meaningful results. We apply the proposed loss function to different neural networks and demonstrate significant improvements in multi-layer cloud mask reconstruction. Utilizing the same neural network architecture, our proposed loss outperforms standard binary crossentropy loss in terms of multi-layer cloud classification accuracy, number of layers accuracy, and thickness mean absolute error (MAE). The proposed loss function can be readily integrated into various neural network architectures, resulting in substantial performance gains in 3D cloud mask generation.

multi-layer clouds↗

Resource frugal optimizer for quantum machine learning

Quantum-enhanced data science, also known as quantum machine learning (QML), is of growing interest as an application of near-term quantum computers. Variational QML algorithms have the potential to solve practical problems on real hardware, particularly when involving quantum data. However, training these algorithms can be challenging and calls for tailored optimization procedures. Specifically, QML applications can require a large shot-count overhead due to the large datasets involved. In this work, we advocate for simultaneous random sampling over both the dataset as well as the measurement operators that define the loss function. We consider a highly general loss function that encompasses many QML applications, and we show how to construct an unbiased estimator of its gradient. This allows us to propose a shot-frugal gradient descent optimizer called Refoqus (REsource Frugal Optimizer for QUantum Stochastic gradient descent). Our numerics indicate that Refoqus can save several orders of magnitude in shot cost, even relative to optimizers that sample over measurement operators alone.

97 MATHEMATICS AND COMPUTING↗

Applying Physics-Informed Neural Networks to Solve Navier–Stokes Equations for Laminar Flow around a Particle

In recent years, Physics-Informed Neural Networks (PINNs) have drawn great interest among researchers as a tool to solve computational physics problems. Unlike conventional neural networks, which are black-box models that “blindly” establish a correlation between input and output variables using a large quantity of labeled data, PINNs directly embed physical laws (primarily partial differential equations) within the loss function of neural networks. By minimizing the loss function, this approach allows the output variables to automatically satisfy physical equations without the need for labeled data. The Navier–Stokes equation is one of the most classic governing equations in thermal fluid engineering. This study constructs a PINN to solve the Navier–Stokes equations for a 2D incompressible laminar flow problem. Flows passing around a 2D circular particle are chosen as the benchmark case, and an elliptical particle is also examined to enrich the research. The velocity and pressure fields are predicted by the PINNs, and the results are compared with those derived from Computational Fluid Dynamics (CFD). Additionally, the particle drag force coefficient is calculated to quantify the discrepancy in the results of the PINNs as compared to CFD outcomes. The drag coefficient maintained an error within 10% across all test scenarios.

Hu, Beichao (ORCID:0009000163215151)↗

Train, Inform, Borrow, or Combine? Approaches to Process–Guided Deep Learning for Groundwater–Influenced Stream Temperature Prediction

Although groundwater discharge is a critical stream temperature control process, it is not explicitly represented in many stream temperature models, an omission that may reduce predictive accuracy, hinder management of aquatic habitat, and decrease user confidence. We assessed the performance of a previously-described process-guided deep learning model of stream temperature in the Delaware River Basin (USA). We found lower accuracy (root mean square error [RMSE] of 1.71 versus 1.35°C) and stronger seasonal bias (absolute mean monthly bias of 1.06 vs. 0.68°C) for reaches primarily influenced by deep groundwater as compared to atmospheric conditions. We then tested four approaches for improving groundwater process representation: (a) a custom loss function leveraging the unique patterns of air and water temperature coupling characteristic of different temperature drivers, (b) inclusion of additional groundwater-relevant catchment attributes, (c) incorporation of additional process model outputs, and (d) a composite model. The custom loss function and the additional attributes significantly improved the predictive accuracy in groundwater-dominated reaches (RMSE of 1.37 and 1.26°C) and reduced the seasonal bias (absolute mean monthly bias of 0.44 and 0.48°C), but neither approach could identify holdout groundwater reaches. Variable importance analysis indicates the custom loss function nudges the model to use the existing inputs more efficiently, whereas with the added features the model relies on a broader suite of inputs. This analysis is a substantial step toward more accurately representing groundwater discharge processes in stream temperature models and will improve predictive accuracy and inform habitat management.

54 ENVIRONMENTAL SCIENCES↗