Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 91 records · Page 5

Paw-Net: Stacking ensemble deep learning for segmenting scanning electron microscopy images of fine-grained shale samples

Segmentation of scanning electron microscopy (SEM) images is critical yet time-consuming for geological analyses, as it needs to differentiate the boundaries for different mineral objects to facilitate subsequent analyses, such as porosity calculation. Recently, various machine learning methods, especially convolutional neural networks (CNNs), have been explored to segment SEM images of fine-grained shale samples. However, we found that general CNNs do not yield optimal performance due to insufficient training data and imbalanced objects in SEM images. This work has revised the U-Net architecture, a popular approach for biomedical image analyses, by incorporating a loss function that addresses the imbalance issue. Furthermore, we used the ensemble learning method to train multiple models and combined the results to improve the overall performance of segmentation. We prepared 2162 sub-images from raw SEM images in our experiments and divided them into training, validation, and testing datasets. The overall results show that our method improves the average Intersection over Union (IOU) of mineral objects from 0.49 to 0.58, compared to the original U-Net model. Our method can clearly distinguish each object from others with boundaries, even in highly imbalanced images. Training our models takes less than three minutes using a single GPU, while manual labeling can take up to three hours for each image. Furthermore, the method helps geoscientists gain insights quickly and effectively by building neural network models from a small dataset of SEM images.

58 GEOSCIENCES↗

Deep learning symmetries and their Lie groups, algebras, and subalgebras from first principles

Abstract We design a deep-learning algorithm for the discovery and identification of the continuous group of symmetries present in a labeled dataset. We use fully connected neural networks to model the symmetry transformations and the corresponding generators. The constructed loss functions ensure that the applied transformations are symmetries and the corresponding set of generators forms a closed (sub)algebra. Our procedure is validated with several examples illustrating different types of conserved quantities preserved by symmetry. In the process of deriving the full set of symmetries, we analyze the complete subgroup structure of the rotation groups SO (2), SO (3), and SO (4), and of the Lorentz group S O ( 1 , 3 ) . Other examples include squeeze mapping, piecewise discontinuous labels, and SO (10), demonstrating that our method is completely general, with many possible applications in physics and data science. Our study also opens the door for using a machine learning approach in the mathematical study of Lie groups and their properties.

97 MATHEMATICS AND COMPUTING↗

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Anomaly Detection in Liquid Sodium Cold Trap Operation with Multisensory Data Fusion Using Long Short-Term Memory Autoencoder

Sodium-cooled fast reactors (SFR), which use high temperature fluid near ambient pressure as coolant, are one of the most promising types of GEN IV reactors. One of the unique challenges of SFR operation is purification of high temperature liquid sodium with a cold trap to prevent corrosion and obstructing small orifices. We have developed a deep learning long short-term memory (LSTM) autoencoder for continuous monitoring of a cold trap and detection of operational anomaly. Transient data were obtained from the Mechanisms Engineering Test Loop (METL) liquid sodium facility at Argonne National Laboratory. The cold trap purification at METL is monitored with 31 variables, which are sensors measuring fluid temperatures, pressures and flow rates, and controller signals. Loss-of-coolant type anomaly in the cold trap operation was generated by temporarily choking one of the blowers, which resulted in temperature and flow rate spikes. The input layer of the autoencoder consisted of all the variables involved in monitoring the cold trap. The LSTM autoencoder was trained on the data corresponding to cold trap startup and normal operation regime, with the loss function calculated as the mean absolute error (MAE). The loss during training was determined to follow log-normal density distribution. During monitoring, we investigated a performance of the LSTM autoencoder for different loss threshold values, set at a progressively increasing number of standard deviations from the mean. The anomaly signal in the data was gradually attenuated, while preserving the noise of the original time series, so that the signal-to-noise ratio (SNR) averaged across all sensors decreased below unity. Results demonstrate detection of anomalies with sensor-averaged SNR < 1.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack

Existing black-box adversarial attacks on image classifiers update the perturbation at each iteration from only a small number of queries of the loss function. Since the queries contain very limited information about the loss, black-box methods usually require much more queries than white-box methods. We propose to improve the query efficiency of black-box methods by exploiting the smoothness of the local loss landscape. However, many adversarial losses are not locally smooth with respect to pixel perturbations. To resolve this issue, our first contribution is to theoretically and experimentally justify that the adversarial losses of many standard and robust image classifiers behave like parabolas with respect to perturbations in the Fourier domain. Our second contribution is to exploit the parabolic landscape to build a quadratic approximation of the loss around the current state, and use this approximation to interpolate the loss value as well as update the perturbation without additional queries. Since the local region is already informed by the quadratic fitting, we use large perturbation steps to explore far areas. We demonstrate the efficiency of our method on MNIST, CIFAR-10 and ImageNet datasets for various standard and robust models, as well as on Google Cloud Vision. The experimental results show that exploiting the loss landscape can help significantly reduce the number of queries and increase the success rate. Our codes are available at https://github.com/HoangATran/BABIES.

Tran, Hoang↗

Data optimization for large batch distributed training of deep neural networks

Distributed training in deep learning (DL) is common practice as data and models grow. The current practice for distributed training of deep neural networks faces the challenges of communication bottlenecks when operating at scale, and model accuracy deterioration with an increase in global batch size. Present solutions focus on improving message exchange efficiency as well as implementing techniques to tweak batch sizes and models in the training process. The loss of training accuracy typically happens because the loss function gets trapped in a local minima. We observe that the loss landscape minimization is shaped by both the model and training data and propose a data optimization approach that utilizes machine learning to implicitly smooth out the loss landscape resulting in fewer local minima. Our approach filters out data points which are less important to feature learning, enabling us to speed up the training of models on larger batch sizes to improved accuracy.

Gahlot, Shubhankar↗

On the Stability Analysis of Astrophysical Cooling Functions

To model the temperature evolution of optically thin astrophysical environments at MHD scales, radiative and collisional cooling rates are typically either pretabulated or fit into a functional form and then input into MHD codes as a radiative loss function. Thermal balance requires estimates of the analogous heating rates, which are harder to calculate, and due to uncertainties in the underlying dissipative heating processes these rates are often simply parameterized. The resulting net cooling function defines an equilibrium curve that varies with density and temperature. Such cooling functions can make the gas prone to thermal instability (TI), which will cause departures from equilibrium. There has been no systematic study of thermally unstable parameter space for nonequilibrium states. Motivated by our recent finding that there is a related linear instability, catastrophic cooling instability, that can dominate over TI, here we carry out such a study. We show that Balbus instability criteria for TI can be used to define a critical cooling rate, Λ c , that permits a nonequilibrium analysis of cooling functions through the mapping of TI zones. We furthermore illustrate how thermal conduction modifies the shape of TI zones. Upon applying a Λ c -based stability analysis to coronal loop simulations, we find that loops undergoing periodic episodes of coronal rain formation are linearly unstable to catastrophic cooling instability, while TI is stabilized by thermal conduction.

79 ASTRONOMY AND ASTROPHYSICS↗

On the Stability Analysis of Astrophysical Cooling Functions

To model the temperature evolution of optically thin astrophysical environments at MHD scales, radiative and collisional cooling rates are typically either pretabulated or fit into a functional form and then input into MHD codes as a radiative loss function. Thermal balance requires estimates of the analogous heating rates, which are harder to calculate, and due to uncertainties in the underlying dissipative heating processes these rates are often simply parameterized. The resulting net cooling function defines an equilibrium curve that varies with density and temperature. Such cooling functions can make the gas prone to thermal instability (TI), which will cause departures from equilibrium. There has been no systematic study of thermally unstable parameter space for nonequilibrium states. Motivated by our recent finding that there is a related linear instability, catastrophic cooling instability, that can dominate over TI, here we carry out such a study. We show that Balbus instability criteria for TI can be used to define a critical cooling rate, Λc, that permits a nonequilibrium analysis of cooling functions through the mapping of TI zones. We furthermore illustrate how thermal conduction modifies the shape of TI zones. Upon applying a Λc-based stability analysis to coronal loop simulations, we find that loops undergoing periodic episodes of coronal rain formation are linearly unstable to catastrophic cooling instability, while TI is stabilized by thermal conduction.

Amanda Stricklan↗

Multi-resolution partial differential equations preserved learning framework for spatiotemporal dynamics

Traditional data-driven deep learning models often struggle with high training costs, error accumulation, and poor generalizability in complex physical processes. Physics-informed deep learning (PiDL) addresses these challenges by incorporating physical principles into the model. Most PiDL approaches regularize training by embedding governing equations into the loss function, yet this depends heavily on extensive hyperparameter tuning to weigh each loss term. To this end, we propose to leverage physics prior knowledge by “baking” the discretized governing equations into the neural network architecture via the connection between the partial differential equations (PDE) operators and network structures, resulting in a PDE-preserved neural network (PPNN). This method, embedding discretized PDEs through convolutional residual networks in a multi-resolution setting, largely improves the generalizability and long-term prediction accuracy, outperforming conventional black-box models. The effectiveness and merit of the proposed methods have been demonstrated across various spatiotemporal dynamical systems governed by spatiotemporal PDEs, including reaction-diffusion, Burgers’, and Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

Prediction of laser beam spatial profiles in a high-energy laser facility by use of deep learning

We adapt the significant advances achieved recently in the field of generative artificial intelligence/machine-learning to laser performance modeling in multipass, high-energy laser systems with application to high-shot-rate facilities relevant to inertial fusion energy. Advantages of neural-network architectures include rapid prediction capability, data-driven processing, and the possibility to implement such architectures within future low-latency, low-power consumption photonic networks. Four models were investigated that differed in their generator loss functions and utilized the U-Net encoder/decoder architecture with either a reconstruction loss alone or combined with an adversarial network loss. We achieved inference times of 1.3 ms for a 256 × 256 pixel near-field beam with errors in predicted energy of the order of 1% over most of the energy range. It is shown that prediction errors are significantly reduced by ensemble averaging the models with different weight initializations. These results suggest that including the temporal dimension in such models may provide accurate, real-time spatiotemporal predictions of laser performance in high-shot-rate laser systems.

47 OTHER INSTRUMENTATION↗

Electron deposition in water vapor, with atmospheric applications.

Examination of the consequences of electron impact on water vapor in terms of the microscopic details of excitation, dissociation, ionization, and combinations of these processes. Basic electron-impact cross-section data are assembled in many forms and are incorporated into semianalytic functions suitable for analysis with digital computers. Energy deposition in water vapor is discussed, and the energy loss function is presented, along with the 'electron volts per ion pair' and the efficiencies of energy loss in various processes. Several applications of electron and water-vapor interactions in the atmospheric sciences are considered, in particular, H2O comets, aurora and airglow, and lightning.

Olivero, J. J.↗

Data-driven Vulnerability Analysis of Networked Pipeline System

This paper introduces an attack generation framework for evaluating the vulnerability of nonlinear networked pipeline systems. The vulnerability analysis is formulated as determining the presence of feasible attack sets, defined by boundary functions representing the effectiveness and stealthiness of attack signals with respect to the objective and attack detection module. The framework utilizes three data-driven models, including two discriminative models that learn the boundary functions and a generative model that produces elements of the feasible attack set. A new loss function ensures successful attack generation with high probability.

03 NATURAL GAS↗

Electron energy deposition in CO2.

Semiempirical cross sections of Strickland and Green (1969) are compared with and supplemented by more recent data. The composite set is used in the energy degradation calculation of electrons with emphasis on energies below 100 ev. Good agreements are obtained with accepted values of the loss function and the average electron volts per ion pair at high energies. Efficiencies associated with various loss channels are obtained. The relationship with the recent Mariner UV data is discussed.

Sawada, T.↗

The inverse problem of the optimal regulator.

The inverse problem of the optimal regulator is considered for a general class of multi-input systems with integral-type performance indices. A new phase variable canonical form is shown to be convenient for this analysis. The advantage of the canonical form is to separate the state variables into subvectors of directly controlled, indirectly controlled, and uncontrollable components. Necessary and sufficient conditions for optimized performance indices are given. With the nonlinearities of the system restricted to functions of the directly controlled state variables, additional results are developed about the nonnegative property of optimized loss functions.

Yokoyama, R.↗

Learning effective stochastic differential equations from microscopic simulations: Linking stochastic numerics to deep learning

We identify effective stochastic differential equations (SDEs) for coarse observables of fine-grained particle- or agent-based simulations; these SDEs then provide useful coarse surrogate models of the fine scale dynamics. We approximate the drift and diffusivity functions in these effective SDEs through neural networks, which can be thought of as effective stochastic ResNets. The loss function is inspired by, and embodies, the structure of established stochastic numerical integrators (here, Euler–Maruyama and Milstein); our approximations can thus benefit from backward error analysis of these underlying numerical schemes. They also lend themselves naturally to “physics-informed” gray-box identification when approximate coarse models, such as mean field equations, are available. Existing numerical integration schemes for Langevin-type equations and for stochastic partial differential equations can also be used for training; we demonstrate this on a stochastically forced oscillator and the stochastic wave equation. Our approach does not require long trajectories, works on scattered snapshot data, and is designed to naturally handle different time steps per snapshot. We consider both the case where the coarse collective observables are known in advance, as well as the case where they must be found in a data-driven manner.

97 MATHEMATICS AND COMPUTING↗

Imaging nanoscale carrier, thermal, and structural dynamics with time-resolved and ultrafast electron energy-loss spectroscopy

Time-resolved and ultrafast electron energy-loss spectroscopy (EELS) is an emerging technique for measuring photoexcited carriers, lattice dynamics, and near-fields across femtosecond to microsecond timescales. When performed in either a specialized scanning transmission electron microscope or ultrafast electron microscope (UEM), time-resolved and ultrafast EELS can directly image charge carriers, lattice vibrations, and heat dissipation following photoexcitation or applied bias. Yet, recent advances in theoretical calculations and electron optics are often required to realize the full potential of ultrafast EEL spectrum imaging. Here, in this review, we present a comprehensive overview of the recent progress in the theory and instrumentation of time-resolved and ultrafast EELS. We begin with an introduction to the technique, followed by a physical description of the loss function. We outline approaches for calculating and interpreting ground-state and transient EEL spectra spanning low-loss plasmons to core-level excitations analogous to x-ray absorption. We then survey the current state of time-resolved and ultrafast EELS techniques beyond photon-induced near-field electron microscopy, highlighting abilities to image carrier and thermal dynamics. Finally, we examine future directions enabled by emerging technologies, including electron beam monochromation, in situ and operando cells, laser-free UEM, and high-speed direct electron detectors. These advances position time-resolved and ultrafast EELS as a critical tool for uncovering nanoscale dynamic processes in quantum materials and solar energy conversion devices.

Computational methods↗

Deep-learning methods for contrast enhancement and artifact reduction in cryo-electron tomography: a systematic analysis of the state of the art and proposed improvements

Cryo-electron tomography (cryo-ET) has emerged as the preferred technique for visualizing the organization of macromolecular complexes in situ and resolving their structures at subnanometre resolution [Tegunov et al. (2021)View full citation, Nat. Methods, 18, 186–193]. Despite improvements in data quality as a result of advances in detector technology, microscope stability and stage precision, the analysis and interpretation of tomograms remains challenging due to a low signal-to-noise ratio and reconstruction artifacts stemming from experimental constraints in specimen tilt during data collection resulting in a missing wedge in the Fourier space. Recently, self-supervised deep-learning methods have been proposed for contrast enhancement and reduction of resolution anisotropy in reconstructed tomograms. Here, we evaluate several state-of-the-art deep-learning methods which aim to improve the interpretability of cryo-ET reconstructions, with a focus on their performance on downstream tasks of template matching, sub­tomogram averaging and segmentation. We propose new training architectures and a loss function based on Fourier shell correlation that show improved performance over the standard U-Net with L1/L2 losses. We demonstrate our analysis on four diverse experimental datasets: purified 80S ribosomes, in situ Chlamydomonas reinhardtii, immature HIV-1 virus-like particles and INS-1E cells.

contrast enhancement↗