Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Automated RF Phase Adjustment for Beam Stabilization in the Fermilab Linac

The Fermilab Linac experiences longitudinal beam phase drift, leading to increased particle loss, conventionally cor- rected through labor-intensive manual RF adjustments. This project explores machine learning-based automation for drift correction, employing a prototype-based classification approach. Our model utilizes a 34-dimensional feature set (RF settings and BPM readings) and leverages a 7x27 response matrix for system modeling. To overcome limited real-world data, we generate synthetic data, enhancing model training and generalizability. Custom loss functions, including a sur- rogate energy-consistent loss and a temporal smoothness constraint, ensure physically plausible drift predictions. The goal is a robust system for autonomous phase adjustments, ensuring stable beam acceleration and reduced manual intervention.

Chichili, R. R. [U. Illinois, Chicago]↗

Anomaly Detection in Liquid Sodium Cold Trap Operation with Multisensory Data Fusion Using Long Short-Term Memory Autoencoder

Sodium-cooled fast reactors (SFR), which use high temperature fluid near ambient pressure as coolant, are one of the most promising types of GEN IV reactors. One of the unique challenges of SFR operation is purification of high temperature liquid sodium with a cold trap to prevent corrosion and obstructing small orifices. We have developed a deep learning long short-term memory (LSTM) autoencoder for continuous monitoring of a cold trap and detection of operational anomaly. Transient data were obtained from the Mechanisms Engineering Test Loop (METL) liquid sodium facility at Argonne National Laboratory. The cold trap purification at METL is monitored with 31 variables, which are sensors measuring fluid temperatures, pressures and flow rates, and controller signals. Loss-of-coolant type anomaly in the cold trap operation was generated by temporarily choking one of the blowers, which resulted in temperature and flow rate spikes. The input layer of the autoencoder consisted of all the variables involved in monitoring the cold trap. The LSTM autoencoder was trained on the data corresponding to cold trap startup and normal operation regime, with the loss function calculated as the mean absolute error (MAE). The loss during training was determined to follow log-normal density distribution. During monitoring, we investigated a performance of the LSTM autoencoder for different loss threshold values, set at a progressively increasing number of standard deviations from the mean. The anomaly signal in the data was gradually attenuated, while preserving the noise of the original time series, so that the signal-to-noise ratio (SNR) averaged across all sensors decreased below unity. Results demonstrate detection of anomalies with sensor-averaged SNR < 1.

21 SPECIFIC NUCLEAR REACTORS AND ASSOCIATED PLANTS↗

Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack

Existing black-box adversarial attacks on image classifiers update the perturbation at each iteration from only a small number of queries of the loss function. Since the queries contain very limited information about the loss, black-box methods usually require much more queries than white-box methods. We propose to improve the query efficiency of black-box methods by exploiting the smoothness of the local loss landscape. However, many adversarial losses are not locally smooth with respect to pixel perturbations. To resolve this issue, our first contribution is to theoretically and experimentally justify that the adversarial losses of many standard and robust image classifiers behave like parabolas with respect to perturbations in the Fourier domain. Our second contribution is to exploit the parabolic landscape to build a quadratic approximation of the loss around the current state, and use this approximation to interpolate the loss value as well as update the perturbation without additional queries. Since the local region is already informed by the quadratic fitting, we use large perturbation steps to explore far areas. We demonstrate the efficiency of our method on MNIST, CIFAR-10 and ImageNet datasets for various standard and robust models, as well as on Google Cloud Vision. The experimental results show that exploiting the loss landscape can help significantly reduce the number of queries and increase the success rate. Our codes are available at https://github.com/HoangATran/BABIES.

Tran, Hoang↗

Data optimization for large batch distributed training of deep neural networks

Distributed training in deep learning (DL) is common practice as data and models grow. The current practice for distributed training of deep neural networks faces the challenges of communication bottlenecks when operating at scale, and model accuracy deterioration with an increase in global batch size. Present solutions focus on improving message exchange efficiency as well as implementing techniques to tweak batch sizes and models in the training process. The loss of training accuracy typically happens because the loss function gets trapped in a local minima. We observe that the loss landscape minimization is shaped by both the model and training data and propose a data optimization approach that utilizes machine learning to implicitly smooth out the loss landscape resulting in fewer local minima. Our approach filters out data points which are less important to feature learning, enabling us to speed up the training of models on larger batch sizes to improved accuracy.

Gahlot, Shubhankar↗

On the Stability Analysis of Astrophysical Cooling Functions

To model the temperature evolution of optically thin astrophysical environments at MHD scales, radiative and collisional cooling rates are typically either pretabulated or fit into a functional form and then input into MHD codes as a radiative loss function. Thermal balance requires estimates of the analogous heating rates, which are harder to calculate, and due to uncertainties in the underlying dissipative heating processes these rates are often simply parameterized. The resulting net cooling function defines an equilibrium curve that varies with density and temperature. Such cooling functions can make the gas prone to thermal instability (TI), which will cause departures from equilibrium. There has been no systematic study of thermally unstable parameter space for nonequilibrium states. Motivated by our recent finding that there is a related linear instability, catastrophic cooling instability, that can dominate over TI, here we carry out such a study. We show that Balbus instability criteria for TI can be used to define a critical cooling rate, Λ c , that permits a nonequilibrium analysis of cooling functions through the mapping of TI zones. We furthermore illustrate how thermal conduction modifies the shape of TI zones. Upon applying a Λ c -based stability analysis to coronal loop simulations, we find that loops undergoing periodic episodes of coronal rain formation are linearly unstable to catastrophic cooling instability, while TI is stabilized by thermal conduction.

79 ASTRONOMY AND ASTROPHYSICS↗

Multi-resolution partial differential equations preserved learning framework for spatiotemporal dynamics

Traditional data-driven deep learning models often struggle with high training costs, error accumulation, and poor generalizability in complex physical processes. Physics-informed deep learning (PiDL) addresses these challenges by incorporating physical principles into the model. Most PiDL approaches regularize training by embedding governing equations into the loss function, yet this depends heavily on extensive hyperparameter tuning to weigh each loss term. To this end, we propose to leverage physics prior knowledge by “baking” the discretized governing equations into the neural network architecture via the connection between the partial differential equations (PDE) operators and network structures, resulting in a PDE-preserved neural network (PPNN). This method, embedding discretized PDEs through convolutional residual networks in a multi-resolution setting, largely improves the generalizability and long-term prediction accuracy, outperforming conventional black-box models. The effectiveness and merit of the proposed methods have been demonstrated across various spatiotemporal dynamical systems governed by spatiotemporal PDEs, including reaction-diffusion, Burgers’, and Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

Prediction of laser beam spatial profiles in a high-energy laser facility by use of deep learning

We adapt the significant advances achieved recently in the field of generative artificial intelligence/machine-learning to laser performance modeling in multipass, high-energy laser systems with application to high-shot-rate facilities relevant to inertial fusion energy. Advantages of neural-network architectures include rapid prediction capability, data-driven processing, and the possibility to implement such architectures within future low-latency, low-power consumption photonic networks. Four models were investigated that differed in their generator loss functions and utilized the U-Net encoder/decoder architecture with either a reconstruction loss alone or combined with an adversarial network loss. We achieved inference times of 1.3 ms for a 256 × 256 pixel near-field beam with errors in predicted energy of the order of 1% over most of the energy range. It is shown that prediction errors are significantly reduced by ensemble averaging the models with different weight initializations. These results suggest that including the temporal dimension in such models may provide accurate, real-time spatiotemporal predictions of laser performance in high-shot-rate laser systems.

47 OTHER INSTRUMENTATION↗

Data-driven Vulnerability Analysis of Networked Pipeline System

This paper introduces an attack generation framework for evaluating the vulnerability of nonlinear networked pipeline systems. The vulnerability analysis is formulated as determining the presence of feasible attack sets, defined by boundary functions representing the effectiveness and stealthiness of attack signals with respect to the objective and attack detection module. The framework utilizes three data-driven models, including two discriminative models that learn the boundary functions and a generative model that produces elements of the feasible attack set. A new loss function ensures successful attack generation with high probability.

03 NATURAL GAS↗

Learning effective stochastic differential equations from microscopic simulations: Linking stochastic numerics to deep learning

We identify effective stochastic differential equations (SDEs) for coarse observables of fine-grained particle- or agent-based simulations; these SDEs then provide useful coarse surrogate models of the fine scale dynamics. We approximate the drift and diffusivity functions in these effective SDEs through neural networks, which can be thought of as effective stochastic ResNets. The loss function is inspired by, and embodies, the structure of established stochastic numerical integrators (here, Euler–Maruyama and Milstein); our approximations can thus benefit from backward error analysis of these underlying numerical schemes. They also lend themselves naturally to “physics-informed” gray-box identification when approximate coarse models, such as mean field equations, are available. Existing numerical integration schemes for Langevin-type equations and for stochastic partial differential equations can also be used for training; we demonstrate this on a stochastically forced oscillator and the stochastic wave equation. Our approach does not require long trajectories, works on scattered snapshot data, and is designed to naturally handle different time steps per snapshot. We consider both the case where the coarse collective observables are known in advance, as well as the case where they must be found in a data-driven manner.

97 MATHEMATICS AND COMPUTING↗

Imaging nanoscale carrier, thermal, and structural dynamics with time-resolved and ultrafast electron energy-loss spectroscopy

Time-resolved and ultrafast electron energy-loss spectroscopy (EELS) is an emerging technique for measuring photoexcited carriers, lattice dynamics, and near-fields across femtosecond to microsecond timescales. When performed in either a specialized scanning transmission electron microscope or ultrafast electron microscope (UEM), time-resolved and ultrafast EELS can directly image charge carriers, lattice vibrations, and heat dissipation following photoexcitation or applied bias. Yet, recent advances in theoretical calculations and electron optics are often required to realize the full potential of ultrafast EEL spectrum imaging. Here, in this review, we present a comprehensive overview of the recent progress in the theory and instrumentation of time-resolved and ultrafast EELS. We begin with an introduction to the technique, followed by a physical description of the loss function. We outline approaches for calculating and interpreting ground-state and transient EEL spectra spanning low-loss plasmons to core-level excitations analogous to x-ray absorption. We then survey the current state of time-resolved and ultrafast EELS techniques beyond photon-induced near-field electron microscopy, highlighting abilities to image carrier and thermal dynamics. Finally, we examine future directions enabled by emerging technologies, including electron beam monochromation, in situ and operando cells, laser-free UEM, and high-speed direct electron detectors. These advances position time-resolved and ultrafast EELS as a critical tool for uncovering nanoscale dynamic processes in quantum materials and solar energy conversion devices.

Computational methods↗

Deep-learning methods for contrast enhancement and artifact reduction in cryo-electron tomography: a systematic analysis of the state of the art and proposed improvements

Cryo-electron tomography (cryo-ET) has emerged as the preferred technique for visualizing the organization of macromolecular complexes in situ and resolving their structures at subnanometre resolution [Tegunov et al. (2021)View full citation, Nat. Methods, 18, 186–193]. Despite improvements in data quality as a result of advances in detector technology, microscope stability and stage precision, the analysis and interpretation of tomograms remains challenging due to a low signal-to-noise ratio and reconstruction artifacts stemming from experimental constraints in specimen tilt during data collection resulting in a missing wedge in the Fourier space. Recently, self-supervised deep-learning methods have been proposed for contrast enhancement and reduction of resolution anisotropy in reconstructed tomograms. Here, we evaluate several state-of-the-art deep-learning methods which aim to improve the interpretability of cryo-ET reconstructions, with a focus on their performance on downstream tasks of template matching, sub­tomogram averaging and segmentation. We propose new training architectures and a loss function based on Fourier shell correlation that show improved performance over the standard U-Net with L1/L2 losses. We demonstrate our analysis on four diverse experimental datasets: purified 80S ribosomes, in situ Chlamydomonas reinhardtii, immature HIV-1 virus-like particles and INS-1E cells.

contrast enhancement↗

Generative Vulnerability Assessment for Cyber-Physical Systems

Cyber-physical systems (CPS) are highly susceptible to malicious attacks due to their complex dynamics and interconnectivity. A comprehensive understanding of their vulnerabilities is essential for designing effective resilience measures. This paper presents a data-driven attack generative system for evaluating the vulnerability of CPS. The proposed approach formulates the vulnerability assessment problem as determining the feasibility of a specific attack set based on two boundary functions that represent the effectiveness and stealthiness of attacks. The attack generative model is trained using a custom loss function, with two universal approximators designed to learn the effectiveness and stealthiness functions simultaneously. Theoretical results for successful generation and asymptotic convergence of the resulting training algorithm are given. As a result, the proposed approach is evaluated via numerical simulation of an IEEE 14-bus system and gas pipeline systems, demonstrating its viability in learning how to attack nonlinear CPS and identify potential vulnerabilities.

Computer systems organization↗

Collective excitations and low-energy ionization signatures of relativistic particles in silicon detectors

Abstract Solid-state detectors with a low energy threshold have several applications, including searches of non-relativistic halo dark-matter particles with sub-GeV masses. When searching for relativistic, beyond-the-Standard-Model particles with enhanced cross sections for small energy transfers, a small detector with a low energy threshold may have better sensitivity than a larger detector with a higher energy threshold. In this paper, we calculate the low-energy ionization spectrum from high-velocity particles scattering in a dielectric material. We consider the full material response including the excitation of bulk plasmons. We generalize the energy-loss function to relativistic kinematics, and benchmark existing tools used for halo dark-matter scattering against electron energy-loss spectroscopy data. Compared to calculations commonly used in the literature, such as the Photo-Absorption-Ionization model or the free-electron model, including collective effects shifts the recoil ionization spectrum towards higher energies, typically peaking around 4–6 electron-hole pairs. We apply our results to the three benchmark examples: millicharged particles produced in a beam, neutrinos with a magnetic dipole moment produced in a reactor, and upscattered dark-matter particles. Our results show that the proper inclusion of collective effects typically enhances a detector’s sensitivity to these particles, since detector backgrounds, such as dark counts, peak at lower energies.

Physics↗

Reducing the Parameter Dependency of Phase-Picking Neural Networks with Dice Loss

Training a neural network for picking seismic phase arrivals has been commonly posed as a segmentation problem. It is a highly imbalanced segmentation problem in the sense that the background vastly dominates the foreground because we are trying to pick the optimal single sample point that represents the arrival of a seismic phase in a many seconds long time window. Here, we test the Dice loss, which is a preferred loss function for highly imbalanced image segmentation problems. We show that phase-picking neural networks trained on the Dice loss behave in a binary fashion for which the prediction output is almost always either nearly 1 or nearly 0. This feature removes the strong dependence of data processing workflows on the prediction score threshold, which is an otherwise critical parameter to determine when using neural networks trained on the cross-entropy loss. When strategically used, models trained on the Dice loss can reduce the parameter dependency of machine learning-based seismic monitoring.

58 GEOSCIENCES↗

Efficient Training of Deep Neural Operator Networks via Randomized Sampling

Neural operators (NOs) employ deep neural networks to learn the mappings between infinitedimensional function spaces. Deep operator network (DeepONet), a popular NO architecture, has demonstrated success in the real-time prediction of complex dynamics across various scientific and engineering applications. In this work, we introduce a random sampling technique to be adopted during the training of DeepONet, aimed at improving the generalization ability of the model, while significantly reducing the computational time. The proposed approach targets the trunk network of the DeepONet model that outputs the basis functions corresponding to the spatiotemporal locations of the bounded domain on which the physical system is defined. While constructing the loss function, DeepONet training traditionally considers a uniform grid of spatiotemporal points at which all the output functions are evaluated for each iteration. This approach leads to a larger batch size, resulting in poor generalization and increased memory demands, due to the limitations of the stochastic gradient descent (SGD) optimizer. The proposed random sampling over the inputs of the trunk net mitigates these challenges, improving generalization and reducing the memory requirements during training, resulting in significant computational gains. We validate our hypothesis through three benchmark examples, demonstrating substantial reductions in training time while achieving comparable or lower overall test errors relative to the traditional training approach. Our results indicate that incorporating randomization in the trunk network inputs during training enhances the efficiency and robustness of DeepONet, offering a promising avenue for improving the framework’s performance in modeling complex physical systems.

Karumuri, Sharmila [Department of Civil & Systems ↗

On the influence of over-parameterization in manifold based surrogates and deep neural operators

Constructing accurate and generalizable approximators (surrogate models) for complex physico-chemical processes exhibiting highly non-smooth dynamics is challenging. The main question is what type of surrogate models we should construct and should these models be under-parameterized or over-parameterized. In this work, we propose new developments and perform comparisons for two promising approaches: manifold-based polynomial chaos expansion (m-PCE) and the deep neural operator (DeepONet), and we examine the effect of over-parameterization on generalization. While m-PCE enables the construction of a mapping by first identifying low-dimensional embeddings of the input functions, parameters, and quantities of interest (QoIs), a neural operator learns the nonlinear mapping via the use of deep neural networks. Here, we demonstrate the performance of these methods in terms of generalization accuracy by solving the 2D time-dependent Brusselator reaction-diffusion system with uncertainty sources, modeling an autocatalytic chemical reaction between two species. We first propose an extension of the m-PCE by constructing a mapping between latent spaces formed by two separate embeddings of the input functions and the output QoIs. To further enhance the accuracy of the DeepONet, we introduce weight self-adaptivity in the loss function. We demonstrate that the performance of m-PCE and DeepONet is comparable for cases of relatively smooth input-output mappings. However, when highly non-smooth dynamics is considered, DeepONet shows higher approximation accuracy. We also find that for m-PCE, modest over-parameterization leads to better generalization, both within and outside of distribution, whereas aggressive over-parameterization leads to over-fitting. In contrast, an even highly over-parameterized DeepONet leads to better generalization for both smooth and non-smooth dynamics. Furthermore, we compare the performance of the above models with another recently proposed operator learning model, the Fourier Neural Operator, and show that its over-parameterization also leads to better generalization. Taken together, our studies show that m-PCE can provide very good accuracy at very low training cost, whereas a highly over-parameterized DeepONet can provide better accuracy and robustness to noise but at higher training cost. In both methods, the inference cost is negligible.

97 MATHEMATICS AND COMPUTING↗

Plateau Phenomenon in Gradient Descent Training of RELU Networks: Explanation, Quantification, and Avoidance

The ability of neural networks to provide ‘best in class’ approximation across a wide range of applications is well-documented. Nevertheless, the powerful expressivity of neural networks comes to naught if one is unable to effectively train (choose) the parameters defining the network. In general, neural networks are trained by gradient descent type optimization methods,a stochastic variant thereof. In practice, such methods result in the loss function decreases rapidly at the beginning of training but then, after a relatively small number of steps, significantly slow down. The loss may even appear to stagnate over the period of a large number of epochs, only to then suddenly start to decrease fast again for no apparent reason. This so-called plateau phenomenon manifests itself in many learning tasks. The present work aims to identify and quantify the root causes of plateau phenomenon.analysis is carried out in the setting of univariate ReLU networks. No assumptions are made on the number of neurons relative to the number of training data, and our results hold for both the lazy and adaptive regimes. Here, the main findings are: plateaux correspond to periods during which activation patterns remain constant, where activation pattern refers to the number of data points that activate a given neuron; quantification of convergence of the gradient flow dynamics; and, characterization stationary points in terms solutions of local least squares regression lines over subsets of the training data. Based on these conclusions, we propose a new iterative training method, the Active Neuron Least Squares (ANLS), characterised by the explicit adjustment of the activation pattern at each step, which is designed to enable a quick exit from a plateau. Illustrative numerical examples are included throughout.

97 MATHEMATICS AND COMPUTING↗

FunDiff: diffusion models over function spaces for physics-informed generative modeling

Recent advances in generative modeling-particularly diffusion models and flow matching-have been widely used for synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. To address this, we introduce FunDiff, an efficient and robust framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generates continuous functions that can be evaluated at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, demonstrating that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We further demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results indicate that our method can generate physically consistent samples with high fidelity to the target distribution, and exhibit robustness to noisy and low-resolution data.

Wang, Sifan [Yale University, New Haven, CT (Unite↗