Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss function regularization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 37 records · Page 2

A Variational Autoencoder Model Toward Molecular Structure Representation Learning of Fuels

Here, in this work, a Variational Autoencoder (VAE)-based data-driven modeling framework is developed with the overarching goal of enabling fuel design. The VAE model is trained on a large dataset with several chemical species to learn a compressed latent space molecular representation. Chemical structure in the form of Simplified Molecular Input Line Entry System (SMILES) string is fed as input, encoded into the VAE latent space, and decoded back to the SMILES string using Long Short-Term Memory (LSTM) networks. Complexities of the VAE training loss function are thoroughly examined by varying the weightage (beta (𝜷) parameter) of the latent space regularization term, thereby assessing the balance between reconstruction accuracy and validity, and focusing on both accurate molecular structure reconstruction and latent space consistency. Two different strategies for 𝜷 variation are evaluated: linear annealing and cyclic annealing. In addition, the impact of total correlation adjustment and hierarchical priors is also studied with regard to the balance between reconstruction fidelity and latent space regularization, and potential issues such as posterior collapse, over-regularization, and poor disentanglement of latent variables. Overall, the best performance of the model is achieved with hierarchical priors and incrementally increasing 𝜷 from 0 to a threshold value of 0.25 over 75 epochs. The generative VAE model can be readily coupled with Quantitative Structure–Property Relationship (QSPR) analysis to develop an integrated end-to-end framework for fuel-property prediction and molecular design of novel promising fuels.

fuel design↗

HomPINNs: Homotopy physics-informed neural networks for learning multiple solutions of nonlinear elliptic differential equations

Physics-informed neural networks (PINNs) based machine learning is an emerging framework for solving nonlinear differential equations. However, due to the implicit regularity of neural network structure, PINNs can only find the flattest solution in most cases by minimizing the loss functions. In this paper, we combine PINNs with the homotopy continuation method, a classical numerical method to compute isolated roots of polynomial systems, and propose a new deep learning framework, named homotopy physics-informed neural networks (HomPINNs), for solving multiple solutions of nonlinear elliptic differential equations. The implementation of an HomPINN is a homotopy process that is composed of the training of a fully connected neural network, named the starting neural network, and training processes of several PINNs with different tracking parameters. The starting neural network is to approximate a starting function constructed by the trivial solutions, while other PINNs are to minimize the loss functions defined by boundary condition and homotopy functions, varying with different tracking parameters. These training processes are regraded as different steps of a homotopy process, and a PINN is initialized by the well-trained neural network of the previous step, while the first starting neural network is initialized using the default initialization method. Finally, several numerical examples are presented to show the efficiency of our proposed HomPINNs, including reaction-diffusion equations with a heart-shaped domain.

97 MATHEMATICS AND COMPUTING↗

Loss Landscape Analysis for Reliable Quantized ML Models for Scientific Sensing

In this paper, we propose a method to perform empirical analysis of the loss landscape of machine learning (ML) models. The method is applied to two ML models for scientific sensing, which necessitates quantization to be deployed and are subject to noise and perturbations due to experimental conditions. Our method allows assessing the robustness of ML models to such effects as a function of quantization precision and under different regularization techniques -- two crucial concerns that remained underexplored so far. By investigating the interplay between performance, efficiency, and robustness by means of loss landscape analysis, we both established a strong correlation between gently-shaped landscapes and robustness to input and weight perturbations and observed other intriguing and non-obvious phenomena. Our method allows a systematic exploration of such trade-offs a priori, i.e., without training and testing multiple models, leading to more efficient development workflows. This work also highlights the importance of incorporating robustness into the Pareto optimization of ML models, enabling more reliable and adaptive scientific sensing systems.

Baldi, Tommaso [Pisa, Scuola Normale Superiore]↗

Decreases in polyunsaturated fatty acid content improve heat stress tolerance during flowering and silicle development in pennycress (Thlaspi arvense L.)

Introduction: Pennycress (Thlaspi arvense L.) is an emerging intermediate oilseed crop grown in the offseason between primary summer crops to produce three cash crops in two years. Previous efforts to improve seed oil quality produced Clustered Regularly Interspaced Short Palindromic Repeats (CRISPR) genome-edited lines with decreased polyunsaturated fatty acids (PUFAs) levels, through loss of function of the FATTY ACID DESATURASE 2 (FAD2), REDUCED OLEATE DESATURATION1 (ROD1), and FATTY ACID ELONGASE1 (FAE1) genes. While seed oil compositions were previously characterized, it remains unknown how vegetative and reproductive tissue compositions might differ and affect tolerance to high temperature (HT) conditions.Methods: In four growth chamber experiments, we explored HT tolerance during flowering and silicle development. Plants were subjected to a 34 °C day/28 °C night regime and compared to control plants maintained at 20 °C. Pollen grain viability at a range of temperatures, lipid peroxidation and proline content in leaves and silicles following HT, and seed yield were measured.Results: Both fad2 and rod1 mutant lines had relatively higher pollen viability (71% and 54% respectively) under moderately elevated temperature (28 °C) compared to wild-type controls (37%). They also showed smaller decreases in seed yield (0% and 40% for fad2 and rod1 respectively, compared to 61% for wild type), following HT exposure during late flowering and early silicle development. Silicles of fad2 plants experienced 65% less lipid peroxidation under HT and 55% less buildup of proline, signifying less stress.Discussion: The differential results of fad2 and rod1 are likely due to the role of FAD2 in membranes in all tissues, whereas ROD1 predominantly affects triacylglycerol (TAG) composition in oil-accumulating tissues including pollen. Our results indicate that decreasing PUFAs, through gene editing, can increase heat tolerance in reproductive tissues as an auxiliary benefit accompanying improved seed oil quality.

59 BASIC BIOLOGICAL SCIENCES↗

CFM-ID 4.0 – a web server for accurate MS-based metabolite identification

The CFM-ID 4.0 web server (https://cfmid.wishartlab.com) is an online tool for predicting, annotating and interpreting tandem mass (MS/MS) spectra of small molecules. It is specifically designed to assist researchers pursuing studies in metabolomics, exposomics and analytical chemistry. More specifically, CFM-ID 4.0 supports the: 1) prediction of electrospray ionization quadrupole time-of-flight tandem mass spectra (ESI-QTOF-MS/MS) for small molecules over multiple collision energies (10 eV, 20 eV, and 40 eV); 2) annotation of ESI-QTOF-MS/MS spectra given the structure of the compound; and 3) identification of a small molecule that generated a given ESI-QTOF-MS/MS spectrum at one or more collision energies. The CFM-ID 4.0 web server makes use of a substantially improved MS fragmentation algorithm, a much larger database of experimental and in silico predicted MS/MS spectra and improved scoring methods to offer more accurate MS/MS spectral prediction and MS/MS-based compound identification. Compared to earlier versions of CFM-ID, this new version has an MS/MS spectral prediction performance that is ~22% better and a compound identification accuracy that is ~35% better on a standard (CASMI 2016) testing dataset. CFM-ID 4.0 also features a neutral loss function that allows users to identify similar or substituent compounds where no match can be found using CFM-ID’s regular MS/MS-to-compound identification utility. Finally, the CFM-ID 4.0 web server now offers a much more refined user interface that is easier to use, supports molecular formula identification (from MS/MS data), provides more interactively viewable data (including proposed fragment ion structures) and displays MS mirror plots for comparing predicted with observed MS/MS spectra. These improvements should make CFM-ID 4.0 much more useful to the community and should make small molecule identification much easier, faster, and more accurate.

59 BASIC BIOLOGICAL SCIENCES↗

Evolutionary Design of Controlled Structures

Basic physical concepts of structural delay and transmissibility are provided for simple rod and beam structures. Investigations show the sensitivity of these concepts to differing controlled-structures variables, and to rational system modeling effects. An evolutionary controls/structures design method is developed. The basis of the method is an accurate model formulation for dynamic compensator optimization and Genetic Algorithm based updating of sensor/actuator placement and structural attributes. One and three dimensional examples from the literature are used to validate the method. Frequency domain interpretation of these controlled structure systems provide physical insight as to how the objective is optimized and consequently what is important in the objective. Several disturbance rejection type controls-structures systems are optimized for a stellar interferometer spacecraft application. The interferometric designs include closed loop tracking optics. Designs are generated for differing structural aspect ratios, differing disturbance attributes, and differing sensor selections. Physical limitations in achieving performance are given in terms of average system transfer function gains and system phase loss. A spacecraft-like optical interferometry system is investigated experimentally over several different optimized controlled structures configurations. Configurations represent common and not-so-common approaches to mitigating pathlength errors induced by disturbances of two different spectra. Results show that an optimized controlled structure for low frequency broadband disturbances achieves modest performance gains over a mass equivalent regular structure, while an optimized structure for high frequency narrow band disturbances is four times better in terms of root-mean-square pathlength. These results are predictable given the nature of the physical system and the optimization design variables. Fundamental limits on controlled performance are discussed based on the measured and fit average system transfer function gains and system phase loss.

Masters, Brett P.↗

Multidimensional Distributional Neural Network Output Demonstrated in Super‐Resolution of Surface Wind Speed

Accurate quantification of uncertainty in neural network predictions remains a central challenge for scientific applications involving high-dimensional, correlated data. While existing methods capture either aleatoric or epistemic uncertainty, few offer closed-form, multidimensional distributions that preserve spatial correlation while remaining computationally tractable. In this work, we present a framework for training neural networks with a multidimensional Gaussian loss, generating a closed-form predictive distribution over outputs informed by non-identically distributed training data. Our approach captures aleatoric uncertainty by iteratively estimating the means and covariance matrices, and is demonstrated on a super-resolution example out-of-training-sample. We leverage a Fourier representation of the covariance matrix to stabilize network training and preserve spatial correlation. We introduce a novel regularization strategy—referred to as information sharing—that interpolates between image-specific and global covariance estimates, enabling convergence of the super-resolution downscaling network trained on image-specific distributional loss functions. This framework allows for efficient sampling, explicit correlation modeling, and extensions to more complex distribution families all without disrupting prediction performance. We demonstrate the method on a surface wind speed downscaling task and discuss its broader applicability to uncertainty-aware prediction in scientific models.

17 WIND ENERGY↗

FunDiff: diffusion models over function spaces for physics-informed generative modeling

Recent advances in generative modeling-particularly diffusion models and flow matching-have been widely used for synthesizing discrete data such as images and videos. However, adapting these models to physical applications remains challenging, as the quantities of interest are continuous functions governed by complex physical laws. To address this, we introduce FunDiff, an efficient and robust framework for generative modeling in function spaces. FunDiff combines a latent diffusion process with a function autoencoder architecture to handle input functions with varying discretizations, generates continuous functions that can be evaluated at arbitrary locations, and seamlessly incorporate physical priors. These priors are enforced through architectural constraints or physics-informed loss functions, ensuring that generated samples satisfy fundamental physical laws. We theoretically establish minimax optimality guarantees for density estimation in function spaces, demonstrating that diffusion-based estimators achieve optimal convergence rates under suitable regularity conditions. We further demonstrate the practical effectiveness of FunDiff across diverse applications in fluid dynamics and solid mechanics. Empirical results indicate that our method can generate physically consistent samples with high fidelity to the target distribution, and exhibit robustness to noisy and low-resolution data.

Wang, Sifan [Yale University, New Haven, CT (Unite↗

Recipes for when physics fails: recovering robust learning of physics informed neural networks

Abstract Physics-informed neural networks (PINNs) have been shown to be effective in solving partial differential equations by capturing the physics induced constraints as a part of the training loss function. This paper shows that a PINN can be sensitive to errors in training data and overfit itself in dynamically propagating these errors over the domain of the solution of the PDE. It also shows how physical regularizations based on continuity criteria and conservation laws fail to address this issue and rather introduce problems of their own causing the deep network to converge to a physics-obeying local minimum instead of the global minimum. We introduce Gaussian process (GP) based smoothing that recovers the performance of a PINN and promises a robust architecture against noise/errors in measurements. Additionally, we illustrate an inexpensive method of quantifying the evolution of uncertainty based on the variance estimation of GPs on boundary data. Robust PINN performance is also shown to be achievable by choice of sparse sets of inducing points based on sparsely induced GPs. We demonstrate the performance of our proposed methods and compare the results from existing benchmark models in literature for time-dependent Schrödinger and Burgers’ equations.

97 MATHEMATICS AND COMPUTING↗

A comparison between bright points in a coronal hole and a quiet-sun region

A comparison is made of the morphological structure and temporal behavior of the emission from coronal bright points in a coronal hole and a quiet region, using data from the Harvard EUV experiment on Skylab. It is found that, in both regions, coronal bright points are located at network boundaries and cover a range of sizes from 10 to 40 in in linear extent. In a given bright pint, the peaks of emission in the six different lines, measured simultaneously through the same instrument slit, are not always cospatial, implying that bright points consist of a complex of small-scale loops at different temperatures. The intensity of bright points in both regions is also characterized by a significant temporal variability in all the wavelengths measured. This variability exhibits no regular periodicity. Yet the ratio of the varying (ac) to the constant (dc) components of the emission, in all the bright points studied, has a local maximum at 1-2 x 10 to the 5th k which coincides with the peak of the radiative loss function, and another local maximum at Mg x (1.4 x 10 to the 6th K). It is found that coronal bright points in a coronal hole or a quiet region are indistinguishable structures, and, therefore, conclude that they are independent of the overlying background corona.

Habbal, Shadia Rifai↗

How Well Does Kohn–Sham Regularizer Work for Weakly Correlated Systems?

Kohn–Sham regularizer (KSR) is a differentiable machine learning approach to finding the exchange-correlation functional in Kohn–Sham density functional theory that works for strongly correlated systems. Here we test KSR for a weak correlation. We propose spin-adapted KSR (sKSR) with trainable local, semilocal, and nonlocal approximations found by minimizing density and total energy loss. We assess the atoms-to-molecules generalizability by training on one-dimensional (1D) H, He, Li, Be, and Be 2+ and testing on 1D hydrogen chains, LiH, BeH 2 , and helium hydride complexes. The generalization error from our semilocal approximation is comparable to other differentiable approaches, but our nonlocal functional outperforms any existing machine learning functionals, predicting ground-state energies of test systems with a mean absolute error of 2.7 mH.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

AADL: Anderson Accelerated Deep Learning

We propose a stable, distributed approach to perform AA that accelerates the convergence rate of stochastic first-order optimizers to train neural networks. Differently from previous works, we do not alter neither the scheme to perform AA nor the loss function minimized during the training. To improve robustness against stagnation, we customize general guidelines that suggest to relax the frequency of AA corrections by performing AA only at the end of an entire training epoch. To improve robustness of AA against the stochastic oscillations of first-order optimizers, we average the gradients computed on consecutive stochastic optimization updates. The improved regularity of the converging sequence and the reduced amplitude of stochastic oscillations across consecutive optimization steps allows AA to efficiently extrapolate an improved converging sequence, thereby overcoming limitations of existing approaches to perform AA on stochastic optimization.

Lupo Pasini, Massimiliano [Oak Ridge National Lab.↗

An end-to-end deep learning method for solving nonlocal Allen–Cahn and Cahn–Hilliard phase-field models

Here, we propose an efficient end-to-end deep learning method for solving nonlocal Allen–Cahn (AC) and Cahn–Hilliard (CH) phase-field models. One motivation for this effort emanates from the fact that discretized partial differential equation-based AC or CH phase-field models result in diffuse interfaces between phases, with the only recourse for remediation is to severely refine the spatial grids in the vicinity of the true moving sharp interface whose width is determined by a grid-independent parameter that is substantially larger than the local grid size. In this work, we introduce non-mass conserving nonlocal AC or CH phase-field models with regular, logarithmic, or obstacle double-well potentials. Because of non-locality, some of these models feature totally sharp interfaces separating phases. The discretization of such models can lead to a transition between phases whose width is only a single grid cell wide. Another motivation is to use deep learning approaches to ameliorate the otherwise high cost of solving discretized nonlocal phase-field models. To this end, loss functions of the customized neural networks are defined using the residual of the fully discrete approximations of the AC or CH models, which results from applying a Fourier collocation method and a temporal semi-implicit approximation. To address the long-range interactions in the models, we tailor the architecture of the neural network by incorporating a nonlocal kernel as an input channel to the neural network model. We then provide the results of extensive computational experiments to illustrate the accuracy, predictive capabilities, and cost reductions of the proposed method.

42 ENGINEERING↗

ARPIST: Provably accurate and stable numerical integration over spherical triangles

Numerical integration on spherical triangles, including the computation of their areas, is a core computation in geomathematics. The commonly used techniques sometimes suffer from instabilities and significant loss of accuracy. We describe a new algorithm, called ARPIST, for accurate and stable integration of functions on spherical triangles. ARPIST is based on an easy-to-implement transformation to the spherical triangle from its corresponding linear triangle via radial projection to achieve high accuracy and efficiency. More importantly, ARPIST overcomes potential instabilities in computing the Jacobian of the transformation, even for poorly shaped triangles that may occur at poles in regular longitude-latitude meshes, by avoiding potential catastrophic rounding errors. We compare our proposed technique with L’Huilier’s Theorem for computing the area of spherical triangles, and also compare it with the recently developed LSQST method (Beckmann et al., 2014) and a radial-basis-function-based technique (Reeger and Fornberg, 2016) for integration of smooth functions on spherical triangulations. In conclusion, our results show that ARPIST enables better or comparable accuracy over previous methods while being easier to implement, significantly faster, and more tolerant of poor element shapes.

97 MATHEMATICS AND COMPUTING↗

CRAGE-CRISPR facilitates rapid activation of secondary metabolite biosynthetic gene clusters in bacteria

With the advent of genome sequencing and mining technologies, secondary metabolite biosynthetic gene clusters (BGCs) within bacterial genomes are becoming easier to predict. For subsequent BGC characterization, clustered regularly interspaced short palindromic repeats (CRISPR) has contributed to knocking out target genes and/or modulating their expression; however, CRISPR is limited to strains for which robust genetic tools are available. Here we present a strategy that combines CRISPR with chassis-independent recombinase-assisted genome engineering (CRAGE), which enables CRISPR systems in diverse bacteria. To demonstrate CRAGE-CRISPR, we select 10 polyketide/non-ribosomal peptide BGCs in Photorhabdus luminescens as models and create their deletion and activation mutants. Subsequent loss- and gain-of-function studies confirm 22 secondary metabolites associated with the BGCs, including a metabolite from a previously uncharacterized BGC. These results demonstrate that the CRAGE-CRISPR system is a simple yet powerful approach to rapidly perturb expression of defined BGCs and to profile genotype-phenotype relationships in bacteria.

59 BASIC BIOLOGICAL SCIENCES↗

CaloTrilogy: Toward a Breakthrough in One-Step, End-to-End, Physics-Guided Shower Generation for Modern Calorimeters

High-precision calorimeter simulation at current and future colliders imposes rapidly growing computational demands, motivating the development of machine-learning surrogates for traditional Monte Carlo tools such as Geant4. Flow matching and diffusion-based generative models have become leading approaches for high-dimensional fast simulation because of their sample quality, but typically require ${\cal O}(100)$ function evaluations at inference and often rely on auxiliary networks to constrain global observables, compromising streamlined end-to-end generation. We introduce a unified framework that improves the balance between speed, shower quality, and physics fidelity. The method combines: (i) an average velocity field integrator that enables sampling in one or a few evaluations; (ii) a learned generative prior in shower space, constructed from data rather than random noise; and (iii) physics-guided loss terms that impose inductive biases on key observables during training. These elements are training time regularizers, preserving end-to-end inference with no additional cost. With only one or a few evaluation steps, the model achieves shower quality competitive with state-of-the-art flow and diffusion approaches, tested on several public high granularity calorimeter datasets. The results demonstrate inter-layer shower structure consistent with the underlying physics, providing a strong candidate for future fast simulation workflows.

Jiang, Cheng [Edinburgh U.]↗

Physics-guided dual implicit neural representations for source separation

Significant challenges exist in efficient data analysis of most advanced experimental and observational techniques because the collected signals often include unwanted contributions, such as background and signal distortions, that can obscure the physically relevant information of interest. To address this, we have developed a self-supervised machine-learning approach for source separation using a dual implicit neural representation framework that jointly trains two neural networks: one for approximating distortions of the physical signal of interest and the other for learning the effective background contribution. Our method learns directly from the raw data by minimizing a reconstruction-based loss function without requiring labeled data or pre-defined dictionaries. We demonstrate the effectiveness of our framework by considering a challenging case study involving large-scale simulated, as well as experimental, momentum-energy-dependent inelastic neutron scattering data in a four-dimensional parameter space, characterized by heterogeneous background contributions and unknown distortions to the target signal. The method is found to successfully separate physically meaningful signals from a complex or structured background even when the signal characteristics vary across all four dimensions of the parameter space. An analytical approach that informs the choice of the regularization parameter is presented. Our method offers a versatile framework for addressing source separation problems across diverse domains, ranging from superimposed signals in astronomical measurements to structural features in biomedical image reconstructions.

47 OTHER INSTRUMENTATION↗