Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

Train, Inform, Borrow, or Combine? Approaches to Process–Guided Deep Learning for Groundwater–Influenced Stream Temperature Prediction

Although groundwater discharge is a critical stream temperature control process, it is not explicitly represented in many stream temperature models, an omission that may reduce predictive accuracy, hinder management of aquatic habitat, and decrease user confidence. We assessed the performance of a previously-described process-guided deep learning model of stream temperature in the Delaware River Basin (USA). We found lower accuracy (root mean square error [RMSE] of 1.71 versus 1.35°C) and stronger seasonal bias (absolute mean monthly bias of 1.06 vs. 0.68°C) for reaches primarily influenced by deep groundwater as compared to atmospheric conditions. We then tested four approaches for improving groundwater process representation: (a) a custom loss function leveraging the unique patterns of air and water temperature coupling characteristic of different temperature drivers, (b) inclusion of additional groundwater-relevant catchment attributes, (c) incorporation of additional process model outputs, and (d) a composite model. The custom loss function and the additional attributes significantly improved the predictive accuracy in groundwater-dominated reaches (RMSE of 1.37 and 1.26°C) and reduced the seasonal bias (absolute mean monthly bias of 0.44 and 0.48°C), but neither approach could identify holdout groundwater reaches. Variable importance analysis indicates the custom loss function nudges the model to use the existing inputs more efficiently, whereas with the added features the model relies on a broader suite of inputs. This analysis is a substantial step toward more accurately representing groundwater discharge processes in stream temperature models and will improve predictive accuracy and inform habitat management.

54 ENVIRONMENTAL SCIENCES↗

Transcription Factor 4 loss-of-function is associated with deficits in progenitor proliferation and cortical neuron content

Transcription Factor 4 ( TCF4) has been associated with autism, schizophrenia, and other neuropsychiatric disorders. However, how pathological TCF4 mutations affect the human neural tissue is poorly understood. Here, we derive neural progenitor cells, neurons, and brain organoids from skin fibroblasts obtained from children with Pitt-Hopkins Syndrome carrying clinically relevant mutations in TCF4 . We show that neural progenitors bearing these mutations have reduced proliferation and impaired capacity to differentiate into neurons. We identify a mechanism through which TCF4 loss-of-function leads to decreased Wnt signaling and then to diminished expression of SOX genes, culminating in reduced progenitor proliferation in vitro. Moreover, we show reduced cortical neuron content and impaired electrical activity in the patient-derived organoids, phenotypes that were rescued after correction of TCF4 expression or by pharmacological modulation of Wnt signaling. This work delineates pathological mechanisms in neural cells harboring TCF4 mutations and provides a potential target for therapeutic strategies for genetic disorders associated with this gene.

59 BASIC BIOLOGICAL SCIENCES↗

LossLens: Diagnostics for Machine Learning Through Loss Landscape Visual Analytics

Modern machine learning often relies on optimizing a neural network's parameters using a loss function to learn complex features. Beyond training, examining the loss function with respect to a network's parameters (i.e., as a loss landscape) can reveal insights into the architecture and learning process. While the local structure of the loss landscape surrounding an individual solution can be characterized using a variety of approaches, the global structure of a loss landscape, which includes potentially many local minima corresponding to different solutions, remains far more difficult to conceptualize and visualize. To address this difficulty, we introduce LossLens, a visual analytics framework that explores loss landscapes at multiple scales. LossLens integrates metrics from global and local scales into a comprehensive visual representation, enhancing model diagnostics. Here we demonstrate LossLens through two case studies: visualizing how residual connections influence a ResNet-20, and visualizing how physical parameters influence a physics-informed neural network (PINN) solving a simple convection problem.

97 MATHEMATICS AND COMPUTING↗

A strong loss-of-function mutation in RAN1 results in constitutive activation of the ethylene response pathway as well as a rosette-lethal phenotype

A recessive mutation was identified that constitutively activated the ethylene response pathway in Arabidopsis and resulted in a rosette-lethal phenotype. Positional cloning of the gene corresponding to this mutation revealed that it was allelic to responsive to antagonist1 (ran1), a mutation that causes seedlings to respond in a positive manner to what is normally a competitive inhibitor of ethylene binding. In contrast to the previously identified ran1-1 and ran1-2 alleles that are morphologically indistinguishable from wild-type plants, this ran1-3 allele results in a rosette-lethal phenotype. The predicted protein encoded by the RAN1 gene is similar to the Wilson and Menkes disease proteins and yeast Ccc2 protein, which are integral membrane cation-transporting P-type ATPases involved in copper trafficking. Genetic epistasis analysis indicated that RAN1 acts upstream of mutations in the ethylene receptor gene family. However, the rosette-lethal phenotype of ran1-3 was not suppressed by ethylene-insensitive mutants, suggesting that this mutation also affects a non-ethylene-dependent pathway regulating cell expansion. The phenotype of ran1-3 mutants is similar to loss-of-function ethylene receptor mutants, suggesting that RAN1 may be required to form functional ethylene receptors. Furthermore, these results suggest that copper is required not only for ethylene binding but also for the signaling function of the ethylene receptors.

NASA Discipline Plant Biology↗

Cooling of solar flares plasmas. 1: Theoretical considerations

Theoretical models of the cooling of flare plasma are reexamined. By assuming that the cooling occurs in two separate phase where conduction and radiation, respectively, dominate, a simple analytic formula for the cooling time of a flare plasma is derived. Unlike earlier order-of-magnitude scalings, this result accounts for the effect of the evolution of the loop plasma parameters on the cooling time. When the conductive cooling leads to an 'evaporation' of chromospheric material, the cooling time scales L(exp 5/6)/p(exp 1/6), where the coronal phase (defined as the time maximum temperature). When the conductive cooling is static, the cooling time scales as L(exp 3/4)n(exp 1/4). In deriving these results, use was made of an important scaling law (T proportional to n(exp 2)) during the radiative cooling phase that was forst noted in one-dimensional hydrodynamic numerical simulations (Serio et al. 1991; Jakimiec et al. 1992). Our own simulations show that this result is restricted to approximately the radiative loss function of Rosner, Tucker, & Vaiana (1978). for different radiative loss functions, other scaling result, with T and n scaling almost linearly when the radiative loss falls off as T(exp -2). It is shown that these scaling laws are part of a class of analytic solutions developed by Antiocos (1980).

Cargill, Peter J.↗

Real-Time Attitude Independent Three Axis Magnetometer Calibration

In this paper new real-time approaches for three-axis magnetometer sensor calibration are derived. These approaches rely on a conversion of the magnetometer-body and geomagnetic-reference vectors into an attitude independent observation by using scalar checking. The goal of the full calibration problem involves the determination of the magnetometer bias vector, scale factors and non-orthogonality corrections. Although the actual solution to this full calibration problem involves the minimization of a quartic loss function, the problem can be converted into a quadratic loss function by a centering approximation. This leads to a simple batch linear least squares solution. In this paper we develop alternative real-time algorithms based on both the extended Kalman filter and Unscented filter. With these real-time algorithms, a full magnetometer calibration can now be performed on-orbit during typical spacecraft mission-mode operations. Simulation results indicate that both algorithms provide accurate integer resolution in real time, but the Unscented filter is more robust to large initial condition errors than the extended Kalman filter. The algorithms are also tested using actual data from the Transition Region and Coronal Explorer (TRACE).

Crassidis, John L.↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Multilabel proportion prediction and out-of-distribution detection on gamma spectra of short-lived fission products

In the machine learning problem of multilabel classification, the objective is to determine for each test instance which classes the instance belongs to. In this work, we consider an extension of multilabel classification, called multilabel proportion prediction, in the context of radioisotope identification (RIID) using gamma spectra data. We aim to not only predict radioisotope proportions, but also identify out-of-distribution (OOD) spectra. We achieve this goal by viewing gamma spectra as discrete probability distributions, and based on this perspective, we develop a custom semi-supervised loss function that combines a traditional supervised loss with an unsupervised reconstruction error function. Our approach was motivated by its application to the analysis of short-lived fission products from spent nuclear fuel. In particular, we demonstrate that a neural network model trained with our loss function can successfully predict the relative proportions of 37 radioisotopes simultaneously. The model trained with synthetic data was then applied to measurements taken by Pacific Northwest National Laboratory (PNNL) to conduct analysis typically done by subject-matter experts. Here, we also extend our approach to successfully identify when measurements are OOD, and thus should not be trusted, whether due to the presence of a novel source or novel proportions.

Anomaly detection↗

HomPINNs: Homotopy physics-informed neural networks for learning multiple solutions of nonlinear elliptic differential equations

Physics-informed neural networks (PINNs) based machine learning is an emerging framework for solving nonlinear differential equations. However, due to the implicit regularity of neural network structure, PINNs can only find the flattest solution in most cases by minimizing the loss functions. In this paper, we combine PINNs with the homotopy continuation method, a classical numerical method to compute isolated roots of polynomial systems, and propose a new deep learning framework, named homotopy physics-informed neural networks (HomPINNs), for solving multiple solutions of nonlinear elliptic differential equations. The implementation of an HomPINN is a homotopy process that is composed of the training of a fully connected neural network, named the starting neural network, and training processes of several PINNs with different tracking parameters. The starting neural network is to approximate a starting function constructed by the trivial solutions, while other PINNs are to minimize the loss functions defined by boundary condition and homotopy functions, varying with different tracking parameters. These training processes are regraded as different steps of a homotopy process, and a PINN is initialized by the well-trained neural network of the previous step, while the first starting neural network is initialized using the default initialization method. Finally, several numerical examples are presented to show the efficiency of our proposed HomPINNs, including reaction-diffusion equations with a heart-shaped domain.

97 MATHEMATICS AND COMPUTING↗

Z-Target Radiography Postprocessing With A Deep Convolution Neural Network

Analyzing X-ray radiographs is crucial for understanding target behavior in Inertial Confinement Fusion (ICF) and High Energy Density (HED) platforms. However, the density of Magneto Raleigh Taylor (MRT) bands and limitations of target materials often obscure relevant spike growth and density information. To address this issue, machine learning postprocessing techniques can be applied to remove darkened regions in radiography images. In this study, a novel method is presented for removing MRT darkened regions from z-target radiographs using a convolutional neural network (CNN). The CNN, consisting of six layers, treats the darkened regions as noise and employs a mixed loss function and end-to-end frameworks to suppress them while preserving sharpness. The six-layer architecture is designed to effectively learn features when provided with a larger volume of learning space. Each layer is optimized using a mixed loss function that combines a standard loss pixel approach with a multi-scaled structural similarity index loss, which considers luminance, contrast, and structure in local neighborhoods. This approach is particularly beneficial for capturing the stochastic structure of MRT limbs. Due to the limited availability of experimental data, training is conducted using synthetic target radiography from 3D Alegra simulations.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Physics constrained learning for data-driven inverse modeling from sparse observations

Deep neural networks (DNN) have been used to model nonlinear relations between physical quantities. Those DNNs are embedded in physical systems described by partial differential equations (PDE) and trained by minimizing a loss function that measures the discrepancy between predictions and observations in some chosen norm. This loss function often includes the PDE constraints as a penalty term when only sparse observations are available. As a result, the PDE is only satisfied approximately by the solution. However, the penalty term typically slows down the convergence of the optimizer for stiff problems. We present a new approach that trains the embedded DNNs while numerically satisfying the PDE constraints. We develop an algorithm that enables differentiating both explicit and implicit numerical solvers in reverse-mode automatic differentiation. This allows the gradients of the DNNs and the PDE solvers to be computed in a unified framework. We demonstrate that our approach enjoys faster convergence and better stability in relatively stiff problems compared to the penalty method. Furthermore, our approach allows for the potential to solve and accelerate a wide range of data-driven inverse modeling, where the physical constraints are described by PDEs and need to be satisfied accurately.

97 MATHEMATICS AND COMPUTING↗

Enforcing Self-Consistent Kinematic Constraints in Neutrino Energy Estimators

Machine learning algorithms have long been utilized across many experimental collaborations within the neutrino physics community in applications to ascertain the singular kinematic quantity of initial neutrino energy for use in neutrino oscillation analyses. However, most of these algorithms do not incorporate a coherent physical picture of initial neutrino kinematics, opting to introduce loss functions involving knowledge of only |pν |. Here, we argue for the introduction of composite loss functions utilizing the full kinematic description of the neutrino, pν ≡ (E, px, py , pz ), compiling all relevant energy and angle information consistently. The use of such a fully defined variable can be seen as a usage of Physics Informed Machine Learning.

Richi, R. R.↗

ICDARTS: Improving the Stability of Cyclic DARTS

Cyclic DARTS (CDARTS) is a Differentiable Architecture Search (DARTS)-based approach to neural architecture search (NAS) that uses a cyclic feedback mechanism to train search and evaluation networks concurrently. This training protocol aims to optimize the search process and evaluate the deep evaluation network comprised of discretized candidate operations. However, this approach introduces a loss function for the evaluation network dependent on the search network. The dissimilarity between the evaluation network’s loss function used during the search and retraining phases results in a search network that is a sub-optimal proxy for the final evaluation network accessed during retraining. We present a revised approach that removes the dependency of the evaluation network weights upon those of the search network. In addition, we introduce a modified process for relaxing the search network’s zero operations that allows these operations to be retained in the final evaluation networks.

Herron, Emily↗

Multi-Label Classification with Constraint-Based Learning for Hierarchical Consistency

We explore the limitations of traditional crossentropy loss in a hierarchical multi-label classification setting and introduce a novel loss function. This function is designed to integrate hierarchical constraints directly into the training process. By incorporating such constraints into the loss, our approach slightly improves the logical consistency of predictions in structured domains. We demonstrate the efficacy of our approach through experiments on primary site and histology classification by using electronic pathology reports. These results show that our proposed hierarchical loss function enhances the model's ability to produce predictions that are logically consistent with the natural data hierarchies, and it slightly improves predictive accuracy. Our framework may be extended to other hierarchical domains, however the performance gains are context specific.

Shivanna, Abhishek [ORNL] (ORCID:0009000665228593)↗

Mars Terrain Segmentation with Less Labels

Planetary rover systems need to perform terrain segmentation to identify drivable areas as well as identify specific types of soil for sample collection. The latest Martian terrain segmentation methods rely on supervised learning which is very data hungry and difficult to train where only a small number of labeled samples are available. Moreover, the semantic classes are defined differently for different applications (e.g., rover traversal vs. geological) and as a result the network has to be trained from scratch each time, which is an inefficient use of resources. This research proposes a semi-supervised learning framework for Mars terrain segmentation where a deep segmentation network trained in an unsupervised manner on unlabeled images is transferred to the task of terrain segmentation trained on few labeled images. The network incorporates a backbone module which is trained using a contrastive loss function and an output atrous convolution module which is trained using a pixel-wise cross-entropy loss function. Evaluation results using the metric of segmentation accuracy show that the proposed method with contrastive pre-training outperforms plain supervised learning by 2%-10%. Moreover, the proposed model is able to achieve a segmentation accuracy of 91.1% using only 161 training images (1% of the original dataset) compared to 81.9% with plain supervised learning.

Wilson, Brian D↗

Fourier Neural Networks as Function Approximators and Differential Equation Solvers

We present a Fourier neural network (FNN) that can be mapped directly to the Fourier decomposition. The choice of activation and loss function yields results that replicate a Fourier series expansion closely while preserving a straightforward architecture with a single hidden layer. The simplicity of this network architecture facilitates the integration with any other higher-complexity networks, at a data pre- or postprocessing stage. We validate this FNN on naturally periodic smooth functions and on piecewise continuous periodic functions. We showcase the use of this FNN for modeling or solving partial differential equations with periodic boundary conditions. The main advantages of the current approach are the validity of the solution outside the training region, interpretability of the trained model, and simplicity of use.

Fourier decomposition↗