Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

The expanded LaGrangian system for constrained optimization problems

Smooth penalty functions can be combined with numerical continuation/bifurcation techniques to produce a class of robust and fast algorithms for constrainted optimization problems. The key to the development of these algorithms is the Expanded Lagrangian System which is derived and analyzed in this work. This parameterized system of nonlinear equations contains the penalty path as a solution, provides a smooth homotopy into the first-order necessary conditions, and yields a global optimization technique. Furthermore, the inevitable ill-conditioning present in a sequential optimization algorithm is removed for three penalty methods: the quadratic penalty function for equality constraints, and the logarithmic barrier function (an interior method) and the quadratic loss function (an interior method) for inequality constraints. Although these techniques apply to optimization in general and to linear and nonlinear programming, calculus of variations, optimal control and parameter identification in particular, the development is primarily within the context of nonlinear programming.

Poore, A. B.↗

The expanded Lagrangian system for constrained optimization problems

Smooth penalty functions can be combined with numerical continuation/bifurcation techniques to produce a class of robust and fast algorithms for constrained optimization problems. The key to the development of these algorithms is the Expanded Lagrangian System which is derived and analyzed in this work. This parameterized system of nonlinear equations contains the penalty path as a solution, provides a smooth homotopy into the first-order necessary conditions, and yields a global optimization technique. Furthermore, the inevitable ill-conditioning present in a sequential optimization algorithm is removed for three penalty methods: the quadratic penalty function for equality constraints, and the logarithmic barrier function (an interior method) and the quadratic loss function (an interior method) for inequality constraints. Although these techniques apply to optimization in general and to linear and nonlinear programming, calculus of variations, optimal control and parameter identification in particular, the development is primarily within the context of nonlinear programming.

Poore, A. B.↗

Multidimensional Distributional Neural Network Output Demonstrated in Super‐Resolution of Surface Wind Speed

Accurate quantification of uncertainty in neural network predictions remains a central challenge for scientific applications involving high-dimensional, correlated data. While existing methods capture either aleatoric or epistemic uncertainty, few offer closed-form, multidimensional distributions that preserve spatial correlation while remaining computationally tractable. In this work, we present a framework for training neural networks with a multidimensional Gaussian loss, generating a closed-form predictive distribution over outputs informed by non-identically distributed training data. Our approach captures aleatoric uncertainty by iteratively estimating the means and covariance matrices, and is demonstrated on a super-resolution example out-of-training-sample. We leverage a Fourier representation of the covariance matrix to stabilize network training and preserve spatial correlation. We introduce a novel regularization strategy—referred to as information sharing—that interpolates between image-specific and global covariance estimates, enabling convergence of the super-resolution downscaling network trained on image-specific distributional loss functions. This framework allows for efficient sampling, explicit correlation modeling, and extensions to more complex distribution families all without disrupting prediction performance. We demonstrate the method on a surface wind speed downscaling task and discuss its broader applicability to uncertainty-aware prediction in scientific models.

17 WIND ENERGY↗

Paw-Net: Stacking ensemble deep learning for segmenting scanning electron microscopy images of fine-grained shale samples

Segmentation of scanning electron microscopy (SEM) images is critical yet time-consuming for geological analyses, as it needs to differentiate the boundaries for different mineral objects to facilitate subsequent analyses, such as porosity calculation. Recently, various machine learning methods, especially convolutional neural networks (CNNs), have been explored to segment SEM images of fine-grained shale samples. However, we found that general CNNs do not yield optimal performance due to insufficient training data and imbalanced objects in SEM images. This work has revised the U-Net architecture, a popular approach for biomedical image analyses, by incorporating a loss function that addresses the imbalance issue. Furthermore, we used the ensemble learning method to train multiple models and combined the results to improve the overall performance of segmentation. We prepared 2162 sub-images from raw SEM images in our experiments and divided them into training, validation, and testing datasets. The overall results show that our method improves the average Intersection over Union (IOU) of mineral objects from 0.49 to 0.58, compared to the original U-Net model. Our method can clearly distinguish each object from others with boundaries, even in highly imbalanced images. Training our models takes less than three minutes using a single GPU, while manual labeling can take up to three hours for each image. Furthermore, the method helps geoscientists gain insights quickly and effectively by building neural network models from a small dataset of SEM images.

58 GEOSCIENCES↗

Deep learning symmetries and their Lie groups, algebras, and subalgebras from first principles

Abstract We design a deep-learning algorithm for the discovery and identification of the continuous group of symmetries present in a labeled dataset. We use fully connected neural networks to model the symmetry transformations and the corresponding generators. The constructed loss functions ensure that the applied transformations are symmetries and the corresponding set of generators forms a closed (sub)algebra. Our procedure is validated with several examples illustrating different types of conserved quantities preserved by symmetry. In the process of deriving the full set of symmetries, we analyze the complete subgroup structure of the rotation groups SO (2), SO (3), and SO (4), and of the Lorentz group S O ( 1 , 3 ) . Other examples include squeeze mapping, piecewise discontinuous labels, and SO (10), demonstrating that our method is completely general, with many possible applications in physics and data science. Our study also opens the door for using a machine learning approach in the mathematical study of Lie groups and their properties.

97 MATHEMATICS AND COMPUTING↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Anomaly Detection in Liquid Sodium Cold Trap Operation with Multisensory Data Fusion Using Long Short-Term Memory Autoencoder

Sodium-cooled fast reactors (SFR), which use high temperature fluid near ambient pressure as coolant, are one of the most promising types of GEN IV reactors. One of the unique challenges of SFR operation is purification of high temperature liquid sodium with a cold trap to prevent corrosion and obstructing small orifices. We have developed a deep learning long short-term memory (LSTM) autoencoder for continuous monitoring of a cold trap and detection of operational anomaly. Transient data were obtained from the Mechanisms Engineering Test Loop (METL) liquid sodium facility at Argonne National Laboratory. The cold trap purification at METL is monitored with 31 variables, which are sensors measuring fluid temperatures, pressures and flow rates, and controller signals. Loss-of-coolant type anomaly in the cold trap operation was generated by temporarily choking one of the blowers, which resulted in temperature and flow rate spikes. The input layer of the autoencoder consisted of all the variables involved in monitoring the cold trap. The LSTM autoencoder was trained on the data corresponding to cold trap startup and normal operation regime, with the loss function calculated as the mean absolute error (MAE). The loss during training was determined to follow log-normal density distribution. During monitoring, we investigated a performance of the LSTM autoencoder for different loss threshold values, set at a progressively increasing number of standard deviations from the mean. The anomaly signal in the data was gradually attenuated, while preserving the noise of the original time series, so that the signal-to-noise ratio (SNR) averaged across all sensors decreased below unity. Results demonstrate detection of anomalies with sensor-averaged SNR < 1.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack

Existing black-box adversarial attacks on image classifiers update the perturbation at each iteration from only a small number of queries of the loss function. Since the queries contain very limited information about the loss, black-box methods usually require much more queries than white-box methods. We propose to improve the query efficiency of black-box methods by exploiting the smoothness of the local loss landscape. However, many adversarial losses are not locally smooth with respect to pixel perturbations. To resolve this issue, our first contribution is to theoretically and experimentally justify that the adversarial losses of many standard and robust image classifiers behave like parabolas with respect to perturbations in the Fourier domain. Our second contribution is to exploit the parabolic landscape to build a quadratic approximation of the loss around the current state, and use this approximation to interpolate the loss value as well as update the perturbation without additional queries. Since the local region is already informed by the quadratic fitting, we use large perturbation steps to explore far areas. We demonstrate the efficiency of our method on MNIST, CIFAR-10 and ImageNet datasets for various standard and robust models, as well as on Google Cloud Vision. The experimental results show that exploiting the loss landscape can help significantly reduce the number of queries and increase the success rate. Our codes are available at https://github.com/HoangATran/BABIES.

Tran, Hoang↗

Data optimization for large batch distributed training of deep neural networks

Distributed training in deep learning (DL) is common practice as data and models grow. The current practice for distributed training of deep neural networks faces the challenges of communication bottlenecks when operating at scale, and model accuracy deterioration with an increase in global batch size. Present solutions focus on improving message exchange efficiency as well as implementing techniques to tweak batch sizes and models in the training process. The loss of training accuracy typically happens because the loss function gets trapped in a local minima. We observe that the loss landscape minimization is shaped by both the model and training data and propose a data optimization approach that utilizes machine learning to implicitly smooth out the loss landscape resulting in fewer local minima. Our approach filters out data points which are less important to feature learning, enabling us to speed up the training of models on larger batch sizes to improved accuracy.

Gahlot, Shubhankar↗

On the Stability Analysis of Astrophysical Cooling Functions

To model the temperature evolution of optically thin astrophysical environments at MHD scales, radiative and collisional cooling rates are typically either pretabulated or fit into a functional form and then input into MHD codes as a radiative loss function. Thermal balance requires estimates of the analogous heating rates, which are harder to calculate, and due to uncertainties in the underlying dissipative heating processes these rates are often simply parameterized. The resulting net cooling function defines an equilibrium curve that varies with density and temperature. Such cooling functions can make the gas prone to thermal instability (TI), which will cause departures from equilibrium. There has been no systematic study of thermally unstable parameter space for nonequilibrium states. Motivated by our recent finding that there is a related linear instability, catastrophic cooling instability, that can dominate over TI, here we carry out such a study. We show that Balbus instability criteria for TI can be used to define a critical cooling rate, Λ c , that permits a nonequilibrium analysis of cooling functions through the mapping of TI zones. We furthermore illustrate how thermal conduction modifies the shape of TI zones. Upon applying a Λ c -based stability analysis to coronal loop simulations, we find that loops undergoing periodic episodes of coronal rain formation are linearly unstable to catastrophic cooling instability, while TI is stabilized by thermal conduction.

79 ASTRONOMY AND ASTROPHYSICS↗

On the Stability Analysis of Astrophysical Cooling Functions

To model the temperature evolution of optically thin astrophysical environments at MHD scales, radiative and collisional cooling rates are typically either pretabulated or fit into a functional form and then input into MHD codes as a radiative loss function. Thermal balance requires estimates of the analogous heating rates, which are harder to calculate, and due to uncertainties in the underlying dissipative heating processes these rates are often simply parameterized. The resulting net cooling function defines an equilibrium curve that varies with density and temperature. Such cooling functions can make the gas prone to thermal instability (TI), which will cause departures from equilibrium. There has been no systematic study of thermally unstable parameter space for nonequilibrium states. Motivated by our recent finding that there is a related linear instability, catastrophic cooling instability, that can dominate over TI, here we carry out such a study. We show that Balbus instability criteria for TI can be used to define a critical cooling rate, Λc, that permits a nonequilibrium analysis of cooling functions through the mapping of TI zones. We furthermore illustrate how thermal conduction modifies the shape of TI zones. Upon applying a Λc-based stability analysis to coronal loop simulations, we find that loops undergoing periodic episodes of coronal rain formation are linearly unstable to catastrophic cooling instability, while TI is stabilized by thermal conduction.

Amanda Stricklan↗

Multi-resolution partial differential equations preserved learning framework for spatiotemporal dynamics

Traditional data-driven deep learning models often struggle with high training costs, error accumulation, and poor generalizability in complex physical processes. Physics-informed deep learning (PiDL) addresses these challenges by incorporating physical principles into the model. Most PiDL approaches regularize training by embedding governing equations into the loss function, yet this depends heavily on extensive hyperparameter tuning to weigh each loss term. To this end, we propose to leverage physics prior knowledge by “baking” the discretized governing equations into the neural network architecture via the connection between the partial differential equations (PDE) operators and network structures, resulting in a PDE-preserved neural network (PPNN). This method, embedding discretized PDEs through convolutional residual networks in a multi-resolution setting, largely improves the generalizability and long-term prediction accuracy, outperforming conventional black-box models. The effectiveness and merit of the proposed methods have been demonstrated across various spatiotemporal dynamical systems governed by spatiotemporal PDEs, including reaction-diffusion, Burgers’, and Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

Prediction of laser beam spatial profiles in a high-energy laser facility by use of deep learning

We adapt the significant advances achieved recently in the field of generative artificial intelligence/machine-learning to laser performance modeling in multipass, high-energy laser systems with application to high-shot-rate facilities relevant to inertial fusion energy. Advantages of neural-network architectures include rapid prediction capability, data-driven processing, and the possibility to implement such architectures within future low-latency, low-power consumption photonic networks. Four models were investigated that differed in their generator loss functions and utilized the U-Net encoder/decoder architecture with either a reconstruction loss alone or combined with an adversarial network loss. We achieved inference times of 1.3 ms for a 256 × 256 pixel near-field beam with errors in predicted energy of the order of 1% over most of the energy range. It is shown that prediction errors are significantly reduced by ensemble averaging the models with different weight initializations. These results suggest that including the temporal dimension in such models may provide accurate, real-time spatiotemporal predictions of laser performance in high-shot-rate laser systems.

47 OTHER INSTRUMENTATION↗

Electron deposition in water vapor, with atmospheric applications.

Examination of the consequences of electron impact on water vapor in terms of the microscopic details of excitation, dissociation, ionization, and combinations of these processes. Basic electron-impact cross-section data are assembled in many forms and are incorporated into semianalytic functions suitable for analysis with digital computers. Energy deposition in water vapor is discussed, and the energy loss function is presented, along with the 'electron volts per ion pair' and the efficiencies of energy loss in various processes. Several applications of electron and water-vapor interactions in the atmospheric sciences are considered, in particular, H2O comets, aurora and airglow, and lightning.

Olivero, J. J.↗

Data-driven Vulnerability Analysis of Networked Pipeline System

This paper introduces an attack generation framework for evaluating the vulnerability of nonlinear networked pipeline systems. The vulnerability analysis is formulated as determining the presence of feasible attack sets, defined by boundary functions representing the effectiveness and stealthiness of attack signals with respect to the objective and attack detection module. The framework utilizes three data-driven models, including two discriminative models that learn the boundary functions and a generative model that produces elements of the feasible attack set. A new loss function ensures successful attack generation with high probability.

03 NATURAL GAS↗

Electron energy deposition in CO2.

Semiempirical cross sections of Strickland and Green (1969) are compared with and supplemented by more recent data. The composite set is used in the energy degradation calculation of electrons with emphasis on energies below 100 ev. Good agreements are obtained with accepted values of the loss function and the average electron volts per ion pair at high energies. Efficiencies associated with various loss channels are obtained. The relationship with the recent Mariner UV data is discussed.

Sawada, T.↗

The inverse problem of the optimal regulator.

The inverse problem of the optimal regulator is considered for a general class of multi-input systems with integral-type performance indices. A new phase variable canonical form is shown to be convenient for this analysis. The advantage of the canonical form is to separate the state variables into subvectors of directly controlled, indirectly controlled, and uncontrollable components. Necessary and sufficient conditions for optimized performance indices are given. With the nonlinearities of the system restricted to functions of the directly controlled state variables, additional results are developed about the nonnegative property of optimized loss functions.

Yokoyama, R.↗