Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

MAL33 drives natural variation in maltose metabolism in Saccharomyces eubayanus

Maltose is one of the most abundant sugars in brewer’s wort, and its efficient utilization is critical for successful fermentation. However, maltose consumption varies naturally among Saccharomyces eubayanus strains isolated from different host trees, such as Quercus and Nothofagus. To identify the genetic determinants underlying these phenotypic differences, we performed bulk segregant analysis (BSA) and quantitative trait loci (QTL) mapping using an F 2 offspring derived from QC18 (Quercus-associated) and CL467.1 (Nothofagus-associated) strains. QTL mapping identified two significant genomic regions on subtelomeric loci of chromosomes V-R and XVI-L, each containing complete MAL loci composed of MAL32 (encoding maltase), MAL31 (transporter), and MAL33 (transcriptional activator) genes. Comparative polymorphism analyses identified mutations in MAL32 and MAL33 of QC18, including frameshift mutations resulting in premature stop codons. Functional validation demonstrated that the heterologous expression of MAL33 ChrV from CL467.1 fully restored maltose utilization in QC18, indicating the functional presence of MAL33 cis-regulatory sequences and MAL32 and MAL31 genes in QC18. While structural protein predictions identified truncation and impaired functionality in the maltose-responsive activation domain of Mal33p from QC18, overexpression of QC18’s own MAL33 ChrV allele also improved maltose metabolism, suggesting dosage-dependent transcriptional limitations rather than complete functional loss. These results indicate that allelic variations in the maltose-responsive activation domain of Mal33p result in differences in maltose consumption between strains. Here, we hypothesized that reduced maltose metabolism in QC18 is an adaptive response to the distinct sugar composition in Quercus robur bark, contrasting with the starch-rich environment of Nothofagus pumilio. These findings highlight subtelomeric MAL gene diversity as a reservoir of genetic variation, representing a key evolutionary mechanism that influences maltose adaptation among natural Saccharomyces isolates.

evolutionary plasticity↗

CFM-ID 4.0 – a web server for accurate MS-based metabolite identification

The CFM-ID 4.0 web server (https://cfmid.wishartlab.com) is an online tool for predicting, annotating and interpreting tandem mass (MS/MS) spectra of small molecules. It is specifically designed to assist researchers pursuing studies in metabolomics, exposomics and analytical chemistry. More specifically, CFM-ID 4.0 supports the: 1) prediction of electrospray ionization quadrupole time-of-flight tandem mass spectra (ESI-QTOF-MS/MS) for small molecules over multiple collision energies (10 eV, 20 eV, and 40 eV); 2) annotation of ESI-QTOF-MS/MS spectra given the structure of the compound; and 3) identification of a small molecule that generated a given ESI-QTOF-MS/MS spectrum at one or more collision energies. The CFM-ID 4.0 web server makes use of a substantially improved MS fragmentation algorithm, a much larger database of experimental and in silico predicted MS/MS spectra and improved scoring methods to offer more accurate MS/MS spectral prediction and MS/MS-based compound identification. Compared to earlier versions of CFM-ID, this new version has an MS/MS spectral prediction performance that is ~22% better and a compound identification accuracy that is ~35% better on a standard (CASMI 2016) testing dataset. CFM-ID 4.0 also features a neutral loss function that allows users to identify similar or substituent compounds where no match can be found using CFM-ID’s regular MS/MS-to-compound identification utility. Finally, the CFM-ID 4.0 web server now offers a much more refined user interface that is easier to use, supports molecular formula identification (from MS/MS data), provides more interactively viewable data (including proposed fragment ion structures) and displays MS mirror plots for comparing predicted with observed MS/MS spectra. These improvements should make CFM-ID 4.0 much more useful to the community and should make small molecule identification much easier, faster, and more accurate.

59 BASIC BIOLOGICAL SCIENCES↗

Acoustic plasmons and conducting carriers in hole-doped cuprate superconductors

The layered crystal structures of cuprates enable collective charge excitations fundamentally different from those of three-dimensional metals. Acoustic plasmons have been observed in electron-doped cuprates by resonant inelastic X-ray scattering (RIXS); in contrast, whether acoustic plasmons exist in hole-doped cuprates is under debate, despite extensive measurements. This contrast led us to investigate the charge excitations of hole-doped cuprate La 2-x Sr x CuO 4 . Here we present incidentenergy-dependent RIXS measurements and calculations of collective charge response via the loss function to reconcile the aforementioned issues. Our results provide evidence for the acoustic plasmons of Zhang-Rice singlet (ZRS), which has a character of the Cu 3d x 2 -y 2 strongly hybridised with the O 2p orbitals; the metallic behaviour is implied to result from the movement of ZRS rather than the simple hopping of O 2p holes.

36 MATERIALS SCIENCE↗

Learning from many collider events at once

There have been a number of recent proposals to enhance the performance of machine learning strategies for collider physics by combining many distinct events into a single ensemble feature. To evaluate the efficacy of these proposals, we study the connection between single-event classifiers and multievent classifiers under the assumption that collider events are independent and identically distributed. We show how one can build optimal multievent classifiers from single-event classifiers, and we also show how to construct multievent classifiers such that they produce optimal single-event classifiers. This is illustrated for a Gaussian example as well as for classification tasks relevant for searches and measurements at the Large Hadron Collider. We extend our discussion to regression tasks by showing how they can be phrased in terms of parametrized classifiers. Empirically, we find that training a single-event (per-instance) classifier is more effective than training a multievent (per-ensemble) classifier, as least for the cases we studied, and we relate this fact to properties of the loss function gradient in the two cases. While we did not identify a clear benefit from using multievent classifiers in the collider context, we speculate on the potential value of these methods in cases involving only approximate independence, as relevant for jet substructure studies.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

Analytic Theory for the Dynamics of Wide Quantum Neural Networks

Here, parametrized quantum circuits can be used as quantum neural networks and have the potential to outperform their classical counterparts when trained for addressing learning problems. To date, much of the results on their performance on practical problems are heuristic in nature. In particular, the convergence rate for the training of quantum neural networks is not fully understood. Here, we analyze the dynamics of gradient descent for the training error of a class of variational quantum machine learning models. We define wide quantum neural networks as parametrized quantum circuits in the limit of a large number of qubits and variational parameters. Then, we find a simple analytic formula that captures the average behavior of their loss function and discuss the consequences of our findings. For example, for random quantum circuits, we predict and characterize an exponential decay of the residual training error as a function of the parameters of the system. Finally, we validate our analytic results with numerical experiments.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Representation Learning via Quantum Neural Tangent Kernels

Variational quantum circuits are used in quantum machine learning and variational quantum simulation tasks. Designing good variational circuits or predicting how well they perform for given learning or optimization tasks is still unclear. Here we discuss these problems, analyzing variational quantum circuits using the theory of neural tangent kernels. We define quantum neural tangent kernels, and derive dynamical equations for their associated loss function in optimization and learning tasks. We analytically solve the dynamics in the frozen limit, or lazy training regime, where variational angles change slowly and a linear perturbation is good enough. We extend the analysis to a dynamical setting, including quadratic corrections in the variational angles. We then consider a hybrid quantum classical architecture and define a large-width limit for hybrid kernels, showing that a hybrid quantum classical neural network can be approximately Gaussian. The results presented here show limits for which analytical understandings of the training dynamics for variational quantum circuits, used for quantum machine learning and optimization problems, are possible. These analytical results are supported by numerical simulations of quantum machine-learning experiments.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Wilson loops with neural networks

Wilson loops are essential objects in QCD and have been pivotal in scale setting and demonstrating confinement. Various generalizations are crucial for computations needed in effective field theories. In lattice gauge theory, Wilson loop calculations face challenges, including excited-state contamination at short times and the signal-to-noise ratio issue at longer times. To address these problems, we develop a new method by using neural networks to parametrize interpolators for the static quark-antiquark pair. We construct gauge-equivariant layers for the network and train it to find the ground state of the system. The trained network itself is then treated as our new observable for the inference. Our results demonstrate a significant improvement in the signal compared to traditional Wilson loops, performing as well as Coulomb-gauge Wilson-line correlators while maintaining gauge invariance. Additionally, we present an example where the optimized ground state is used to measure the static force directly, as well as another example combining this method with the multilevel algorithm. Finally, we extend the formalism to find excited-state interpolators for static quark-antiquark systems. To our knowledge, this work is the first study of neural networks with a physically motivated loss function for Wilson loops.

Bellscheidt, Verena [Massachusetts Inst. of Techno↗

Introducing a multiscale feature integration network for inpainting with applications to enhanced CMB map reconstruction

We introduce a novel neural network, SkyReconNet, which combines the expanded receptive fields of dilated convolutional layers along with standard convolutions, to capture both the global and local features for reconstructing the missing information in an image. We implement our network to inpaint the masked regions in a full-sky cosmic microwave background (CMB) map. Inpainting CMB maps is a particularly formidable challenge when dealing with extensive and irregular masks, such as galactic masks which can obscure substantial fractions of the sky. The hybrid design of SkyReconNet leverages the strengths of standard and dilated convolutions to accurately predict CMB fluctuations in the masked regions by effectively utilizing the information from surrounding unmasked areas. During training, the network optimizes its weights by minimizing a composite loss function that combines the structural similarity index measure (SSIM) and mean squared error (MSE). SSIM preserves the essential structural features of the CMB, ensuring an accurate and coherent reconstruction of the missing CMB fluctuations, while MSE minimizes the pixelwise deviations, thus enhancing the overall accuracy of the predictions. The predicted CMB maps and their corresponding angular power spectra align closely with the targets, achieving the performance limited only by the fundamental uncertainty of cosmic variance. The network’s generic architecture enables application to other physics-based challenges involving data with missing or defective pixels, systematic artifacts, etc. In conclusion, our results demonstrate its effectiveness in addressing the challenges posed by large irregular masks, offering a significant inpainting tool not only for CMB analyses but also for image-based experiments across disciplines where such data imperfections are prevalent.

Cosmic microwave background↗

Optimizers for stabilizing likelihood-free inference

A growing number of applications in particle physics and beyond use neural networks as unbinned likelihood ratio estimators applied to real or simulated data. Precision requirements on the inference tasks demand a high-level of stability from these networks, which are affected by the stochastic nature of training. We show how physics concepts can be used to stabilize network training through a physics-inspired optimizer. In particular, the energy conserving descent (ECD) optimization framework uses classical Hamiltonian dynamics on the space of network parameters to reduce the dependence on the initial conditions while also stabilizing the result near the minimum of the loss function. We develop a version of this optimizer known as , which has few free hyperparameters with limited ranges guided by physical reasoning. We apply to representative likelihood-ratio estimation tasks in particle physics and find on average that it out-performs the widely used Adam optimizer. We expect that ECD will be a useful tool for wide array of data-limited problems, where it is computationally expensive to exhaustively optimize hyperparameters and mitigate fluctuations with ensembling.

Monte Carlo methods↗

End-to-end orientation estimation from 2D cryo-EM1images

Cryo-electron microscopy (cryo-EM) is a Nobel Prize-winning technique for deter-mining high-resolution 3D structures of biological macromolecules. A 3D structure is reconstructed from hundreds of thousands of noisy 2D projection images. However, existing 3D reconstruction methods are still time-consuming, and one of the major computational bottlenecks is to recover the unknown orientation of the particle in16each 2D image. The dominant methods typically exploit expensive global search on each image to estimate the missing orientations. Here, a novel end-to-end supervised learning method is introduced to directly recover the missing orientations from 2D cryo-EM images. A neural network is used to approximate the mapping from images to orientations. Furthermore, a robust loss function is proposed for optimizing the parameters of the network, which can handle both asymmetric and symmetric 3D structures. Experiments on synthetic datasets with various symmetry types confirm that the neural network is capable of recovering orientations from 2D cryo-EM images, and the results on one real cryo-EM dataset further demonstrate its potential in more challenging imaging conditions.

3D reconstruction↗

DENSECL: Haze Mitigation Using Dense Blocks and Contrastive Loss Regularization

Haze, which occurs as a result of the scattering of light in the atmosphere by small particles, diminishes the visibility of scene objects, inflicting important image applications such as object detection. To address the problem, this paper introduces a new physics-based end-to-end deep learning approach to haze mitigation in outdoor scenes, including those in airborne images. The proposed model named DenseCL is designed with dense blocks and adopts a contrastive loss function as an additional regularization. The model also maintains the cycle consistency by remapping the dehazed outputs into a hazy image using the physics-based light scattering function. DenseCL has been trained with publicly available outdoor images and demonstrates outstanding performance on outdoor, indoor, and remotely sensed nonhomogeneous haze satellite images.

Mitra, Somosmita↗

Mixture-of-Experts for Multi-Domain Defect Identification in Non-Destructive Inspection

Composite materials are widely used in aircraft structures because of their superior mechanical properties. However, their complex failure modes require sophisticated inspection methods to ensure structural integrity. Ultrasonic testing (UT) is a common non-destructive inspection (NDI) technique for aircraft composites that can detect internal and external defects with high resolution and accuracy. Despite their effectiveness, traditional UT methods rely on the manual interpretation of ultrasonic signals, which is time-consuming, labor-intensive, and subjective. Furthermore, processing such large-scale data, particularly across materials of varying thicknesses, significantly increases the computational demands of deep learning model optimization. To overcome these challenges, we propose an efficient sparse mixture-of-experts (MoE) model with a multi-level loss function and introduce four novel training objectives to improve computational efficiency and accuracy in identifying surface defects in composite aircraft materials. Here, we evaluated our approach on material with multiple thicknesses or domains comprising various defects. Our experimental results demonstrate higher accuracy and F1-Score, with only 10% training epochs compared to baseline MoE.

composite materials↗

AutoAtlas: Neural Network for 3D Unsupervised Partitioning and Representation Learning

Here we present a novel neural network architecture called AutoAtlas for fully unsupervised partitioning and representation learning of 3D brain Magnetic Resonance Imaging (MRI) volumes. AutoAtlas consists of two neural network components: one neural network to perform multi-label partitioning based on local texture in the volume, and a second neural network to compress the information contained within each partition. We train both of these components simultaneously by optimizing a loss function that is designed to promote accurate reconstruction of each partition, while encouraging spatially smooth and contiguous partitioning, and discouraging relatively small partitions. We show that the partitions adapt to the subject specific structural variations of brain tissue while consistently appearing at similar spatial locations across subjects. AutoAtlas also produces very low dimensional features that represent local texture of each partition. We demonstrate prediction of metadata associated with each subject using the derived feature representations and compare the results to prediction using features derived from FreeSurfer anatomical parcellation. Since our features are intrinsically linked to distinct partitions, we can then map values of interest, such as partition-specific feature importance scores onto the brain for visualization.

42 ENGINEERING↗

Dynamic Parameter Estimation with Physics-based Neural Ordinary Differential Equations

Accurate estimation of dynamic parameters of gen-erators is crucial to building a reliable model for dynamical studies and reliable operation of the power system. This paper develops a physics-based neural ordinary differential equations (ODE) approach to learn the parameters of generator dynamic model using phasor measurement units (PMU) data. We design a physics-based neural network to represent the swing equations of the power system dynamics. A loss function is defined as the difference between dynamic simulation results from the physics-based neural networks and pseudo PMU measurements. The parameters of generator dynamic model are iteratively updated using the neural ODEs and the adjoint method. By exploiting the mini-batch scheme in neural ODE training, the parameter estimation performance is significantly improved. Numerical study results on a 3-machine 9-bus system show that the proposed algorithm outperforms state-of-the-art baseline method in both computation time and dynamic parameter estimation accuracy.

24 POWER TRANSMISSION AND DISTRIBUTION↗

Deep Learning-Based Weather-Related Power Outage Prediction with Socio-Economic and Power Infrastructure Data

This paper presents a deep learning-based approach for hourly power outage probability prediction within census tracts encompassing a utility company's service territory. Two distinct deep learning models, conditional Multi-Layer Perceptron (MLP) and unconditional MLP, were developed to forecast power outage probabilities, leveraging a rich array of input features gathered from publicly available sources including weather data, weather station locations, power infrastructure maps, socio-economic and demographic statistics, and power outage records. Given a one-hour-ahead weather forecast, the models predict the power outage probability for each census tract, taking into account both the weather prediction and the location's characteristics. The deep learning models employed different loss functions to optimize prediction performance. Our experimental results underscore the significance of socio-economic factors in enhancing the accuracy of power outage predictions at the census tract level.

24 POWER TRANSMISSION AND DISTRIBUTION↗

High-Resolution Synthetic Solar Irradiance Sequence Generation: An LSTM-Based Generative Adversarial Network

The rapid growth of renewable energy resources penetration is bringing more challenges to power system planning and operation. Relevant renewable energy integration studies, such as the capability and dynamic performance of inverter-based resources' primary frequency response and fast frequency response, require high-resolution renewable generation output data that are representative of renewable energy resources. This paper focuses on creating synthetic but realistic solar irradiance data and proposes a long short-term memory-based generative adversarial network to generate high-resolution (second-level) solar irradiance sequences from low-resolution (minute-level) measurements. Combined with a classifier to recognize the solar irradiance patterns, the proposed model is trained using multi-loss functions to accurately capture the temporal correlations among both high-resolution and low-resolution sequences. Verification of the proposed approach is performed on the data set of the Oahu Solar Measurement Grid collected through the National Renewable Energy Laboratory. The results of the case studies demonstrate the proposed approach's capability to capture the statistical characteristics of different solar irradiance patterns and to generate high-quality synthetic solar irradiance sequences in high resolution.

dynamic scheduling↗

A Weakly Supervised Machine Learning Procedure for Magnet Quench Diagnostics

Voltage taps remain the standard and reliable diagnostic tool for detecting quenches in superconducting magnets. However, they identify a quench only at the time of voltage rise and do not provide information on earlier physical precursors. In this work, we investigate whether acoustic emission data can reveal precursor activity that occurs before conventional voltage detection using machine learning techniques. We introduce an event selection method and a weakly supervised machine learning procedure to learn data-driven criteria for identifying potential acoustic precursors to quenches. Two Convolutional Neural Network (CNN) architectures are trained: one on acoustic sensor events from our selection procedure and one on the Fast Fourier Transforms (FFTs) of these events. Both networks are trained iteratively using confidence-weighted loss functions to associate certain subsets of training data with a precursor label. We evaluate the performance of these models by examining the time distribution of events classified as potential precursors relative to the quench onset. Results indicate that the proposed approach can possibly distinguish acoustic emission events occurring closer to the quench from earlier acoustic activity during ramping, suggesting the potential for flagging quench precursors in acoustic data.

Khan, Maira [Fermilab] (ORCID:0009000891602387)↗