Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss functions”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Enhancing high-fidelity neural network potentials through low-fidelity sampling

The efficacy of neural network potentials (NNPs) critically depends on the quality of the configurational datasets used for training. Prior research using empirical potentials has shown that well-selected liquid–solid transitional configurations of a metallic system can be translated to other metallic systems. This study demonstrates that such validated configurations can be relabeled using density functional theory (DFT) calculations, thereby enhancing the development of high-fidelity NNPs. Training strategies and sampling approaches are efficiently assessed using empirical potentials and subsequently relabeled via DFT in a highly parallelized fashion for high-fidelity NNP training. Our results reveal that relying solely on energy and force for NNP training is inadequate to prevent overfitting, highlighting the necessity of incorporating stress terms into the loss functions. To optimize training involving force and stress terms, we propose employing transfer learning to fine-tune the weights, ensuring that the potential surface is smooth for these quantities composed of energy derivatives. This approach markedly improves the accuracy of elastic constants derived from simulations in both empirical potential-based NNPs and relabeled DFT-based NNPs. Overall, this study offers significant insights into leveraging empirical potentials to expedite the development of reliable and robust NNPs at the DFT level.

97 MATHEMATICS AND COMPUTING↗

4D Multimodal Co-attention Fusion Network with Latent Contrastive Alignment for Alzheimer’s Diagnosis

Multimodal neuroimaging provides complementary structural and functional insights into both human brain organization and disease-related dynamics. Recent studies demonstrate enhanced diagnostic sensitivity for Alzheimer’s disease (AD) through synergistic integration of neuroimaging data (e.g., sMRI, fMRI) with tabular data (e.g., behavioral and cognitive tests). However, the intrinsic heterogeneity across modalities (e.g., 4D spatiotemporal fMRI dynamics vs. 3D anatomical sMRI structure) presents critical challenges for discriminative feature fusion, often leading to information loss or biased fusion. To bridge this gap, we propose M2M-AlignNet: a multimodal co-attention network with latent alignment for early AD diagnosis using sMRI and fMRI. At the core of our approach is a multi-patch-to-multi-patch (M2M) contrastive loss function that quantifies and reduces representational discrepancies via weighted patch correspondence, explicitly aligning fMRI components across brain regions with their sMRI structural substrates without one-to-one constraints. Additionally, we propose a latent-as-query co-attention module to autonomously discover fusion patterns, circumventing modality prioritization biases while minimizing feature redundancy. We conduct extensive experiments to confirm the effectiveness of our method and highlight the correspondence between fMRI and sMRI as AD biomarkers.

Wei, Yuxiang [Georgia Institute of Technology]↗

Any Two Learning Algorithms Are (Almost) Exactly Identical

This paper shows that if one is provided with a loss function, it can be used in a natural way to specify a distance measure quantifying the similarity of any two supervised learning algorithms, even non-parametric algorithms. Intuitively, this measure gives the fraction of targets and training sets for which the expected performance of the two algorithms differs significantly. Bounds on the value of this distance are calculated for the case of binary outputs and 0-1 loss, indicating that any two learning algorithms are almost exactly identical for such scenarios. As an example, for any two algorithms A and B, even for small input spaces and training sets, for less than 2e(-50) of all targets will the difference between A's and B's generalization performance of exceed 1%. In particular, this is true if B is bagging applied to A, or boosting applied to A. These bounds can be viewed alternatively as telling us, for example, that the simple English phrase 'I expect that algorithm A will generalize from the training set with an accuracy of at least 75% on the rest of the target' conveys 20,000 bytes of information concerning the target. The paper ends by discussing some of the subtleties of extending the distance measure to give a full (non-parametric) differential geometry of the manifold of learning algorithms.

Wolpert, David H.↗

Physics-informed neural networks for identification of material properties using standing waves

A metallic structure in its initial stage of failure involves plastic deformation or environmental degradation that changes the elastic modulus and density. This work presents the detection of change in wave velocity (a function of elastic modulus and density) as a system identification problem. A physics-informed neural network (PINN) is proposed to solve the system identification problem. The PINN takes the spatial coordinates of scanning locations and time as inputs and provides the displacement and wave velocity as outputs. The governing partial differential equation of standing waves in a rod is incorporated into the neural network as physics in the form of a loss function. The wave velocity vector is randomly initiated. During the training of the network, physics is used to determine and update the wave velocity target vector from the network’s displacement predictions. The measured data, comprising sparse displacement response on the rod structure, are used to train the PINN. The wave velocity at the sparse locations on the rod is learned from the predicted displacements during the training. Using the predictions of the trained network, the response of free vibration or material property variation can be reconstructed at unscanned locations on the structure to obtain high-resolution maps for full-field imaging to detect and localize the changes caused by plastic deformation. The PINN’s sparse scanning and simultaneous prediction capability during training can lead to high scanning and data-processing speeds. This capability yields a nondestructive evaluation system that can predict the presence of degraded material locations as the structural vibrations are scanned and processed in real time.

Rathod, Vivek↗

Challenges in Training PINNs: A Loss Landscape Perspective

This paper explores challenges in training Physics Informed Neural Networks (PINNs), emphasizing the role of the loss landscape in the training process. We examine difficulties in minimizing the PINN loss function, particularly due to ill conditioning caused by differential operators in the residual term. We compare gradient-based optimizers Adam, L-BFGS, and their combination Adam+L-FGS, showing the superiority of Adam+L-BFGS, and introduce a novel secondorder optimizer, NysNewton-CG (NNCG), which significantly improves PINN performance. Theoretically, our work elucidates the connection between ill-conditioned differential operators and ill-conditioning in the PINN loss and shows the benefits of combining first- and second-order optimization methods. Our work presents valuable insights and more powerful optimization strategies for training PINNs, which could improve the utility of PINNs for solving difficult partial differential equations.

Rathore, Pratik↗

Neural conditional reweighting

There is a growing use of neural network classifiers as unbinned, high-dimensional (and variable-dimensional) reweighting functions. To date, the focus has been on marginal reweighting, where a subset of features are used for reweighting while all other features are integrated over. There are some situations, though, where it is preferable to condition on auxiliary features instead of marginalizing over them. Here, we introduce neural conditional reweighting, which extends neural marginal reweighting to the conditional case. This approach is particularly relevant in high-energy physics experiments for reweighting detector effects conditioned on particle-level truth information. Furthermore we leverage a custom loss function that not only allows us to achieve neural conditional reweighting through a single training procedure, but also yields sensible interpolation even in the presence of phase space holes. As a specific example, we apply neural conditional reweighting to the energy response of high-energy jets, which could be used to improve the modeling of physics objects in parametrized fast simulation packages.

72 PHYSICS OF ELEMENTARY PARTICLES AND FIELDS↗

Spatial Statistical Data Fusion (SSDF)

As remote sensing for scientific purposes has transitioned from an experimental technology to an operational one, the selection of instruments has become more coordinated, so that the scientific community can exploit complementary measurements. However, tech nological and scientific heterogeneity across devices means that the statistical characteristics of the data they collect are different. The challenge addressed here is how to combine heterogeneous remote sensing data sets in a way that yields optimal statistical estimates of the underlying geophysical field, and provides rigorous uncertainty measures for those estimates. Different remote sensing data sets may have different spatial resolutions, different measurement error biases and variances, and other disparate characteristics. A state-of-the-art spatial statistical model was used to relate the true, but not directly observed, geophysical field to noisy, spatial aggregates observed by remote sensing instruments. The spatial covariances of the true field and the covariances of the true field with the observations were modeled. The observations are spatial averages of the true field values, over pixels, with different measurement noise superimposed. A kriging framework is used to infer optimal (minimum mean squared error and unbiased) estimates of the true field at point locations from pixel-level, noisy observations. A key feature of the spatial statistical model is the spatial mixed effects model that underlies it. The approach models the spatial covariance function of the underlying field using linear combinations of basis functions of fixed size. Approaches based on kriging require the inversion of very large spatial covariance matrices, and this is usually done by making simplifying assumptions about spatial covariance structure that simply do not hold for geophysical variables. In contrast, this method does not require these assumptions, and is also computationally much faster. This method is fundamentally different than other approaches to data fusion for remote sensing data because it is inferential rather than merely descriptive. All approaches combine data in a way that minimizes some specified loss function. Most of these are more or less ad hoc criteria based on what looks good to the eye, or some criteria that relate only to the data at hand.

Braverman, Amy J.↗

Electronic and optical properties of plutonium metal and oxides from Reflection Electron Energy Loss Spectroscopy

Reflection Electron Energy Loss Spectra (REELS) have been acquired on plutonium δ-Pu-Ga alloys, α-Pu 2 O 3 and PuO 2 . The position of the peaks attributable to the Pu 6p excitation can be used as ‘fingerprint’ for these species. X-ray Photoelectron Spectroscopy has been used to aid interpretation of the REELS spectra. The band gap of PuO 2 was determined from the onset of the energy loss in the REELS spectrum. The REELS data have been evaluated within the dielectric response theory by means of the QUEELS-ε(k, w)-REELS software. In this work, the Energy Loss Function, dielectric function, refractive index and extinction coefficient for these materials have been determined in the ultraviolet and extreme ultraviolet range and are compared to previously reported results.

36 MATERIALS SCIENCE↗

Gearbox bearing crack growth prognostics and uncertainty quantification with physics-informed machine learning

This paper introduces the extreme theory of functional connections (X-TFC), a physics-informed machine learning algorithm, and tailors it to estimate the remaining useful life (RUL) of wind turbine gearbox bearings experiencing fatigue crack growth. Unlike purely data-driven methods, X-TFC embeds a physics model, based on Head's theory in this work, into its training objective. The core of X-TFC is a random-projection single-layer neural network trained via an extreme learning machine, which requires only limited damage progression data and solves for output weights with a least-squares optimization algorithm. A composite loss function balances the network's fit to observed degradation data against the residuals of the governing crack growth differential equation, ensuring the learned damage trajectory remains physically plausible. When applied to a vibration-based health-index (HI) dataset measured during the growth of a crack on the inner ring of a high-speed bearing in a wind turbine gearbox (Bechhoefer and Dubé, 2020), X-TFC achieves near-zero prediction bias. Even when trained on only the first 10 %–20 % of the damage progression data, with sufficient physics weighting its predictions remain monotonic and smooth, delivering high prognosability and trendability. To quantify the epistemic uncertainty, we employ a Monte Carlo ensemble of independently initialized X-TFC models trained on noise-perturbed data, which yields confidence intervals around each RUL estimate and captures both model-parameter and epistemic uncertainty. In addition to a vibration-based HI, we demonstrate that the proposed framework can be directly applied to a supervisory control and data acquisition (SCADA) data-based HI (Eftekhari Milani et al., 2026) measured during similar wind turbine gearbox bearing crack faults, preserving its accuracy and interpretability. This extension shows the versatility of our approach, which is applicable to bearings of multiple gearbox manufacturers, models, and ratings using only SCADA data. By integrating domain knowledge with machine learning, X-TFC offers a rapid, reliable tool for crack prognostics. Its adaptability to other bearing failure modes, such as pitch bearing ring cracks, positions X-TFC as a powerful enabler of data-driven, physics-informed asset management in the wind energy sector and beyond.

17 WIND ENERGY↗

Hutchinson Trace Estimation for high-dimensional and high-order Physics-Informed Neural Networks

Physics-Informed Neural Networks (PINNs) have proven effective in solving partial differential equations (PDEs), especially when some data are available by seamlessly blending data and physics. However, extending PINNs to high-dimensional and even high-order PDEs encounters significant challenges due to the computational cost associated with automatic differentiation in the residual loss function calculation. Herein, we address the limitations of PINNs in handling high-dimensional and high-order PDEs by introducing the Hutchinson Trace Estimation (HTE) method. Starting with the second-order high-dimensional PDEs, which are ubiquitous in scientific computing, HTE is applied to transform the calculation of the entire Hessian matrix into a Hessian vector product (HVP). This approach not only alleviates the computational bottleneck via Taylor-mode automatic differentiation but also significantly reduces memory consumption from the Hessian matrix to an HVP’s scalar output. We further showcase HTE’s convergence to the original PINN loss and its unbiased behavior under specific conditions. Comparisons with the Stochastic Dimension Gradient Descent (SDGD) highlight the distinct advantages of HTE, particularly in scenarios with significant variability and variance among dimensions. We further extend the application of HTE to higher-order and higher-dimensional PDEs, specifically addressing the biharmonic equation. By employing tensor-vector products (TVP), HTE efficiently computes the colossal tensor associated with the fourth-order high-dimensional biharmonic equation, saving memory and enabling rapid computation. The effectiveness of HTE is illustrated through experimental setups, demonstrating comparable convergence rates with SDGD under memory and speed constraints. Additionally, HTE proves valuable in accelerating the Gradient-Enhanced PINN (gPINN) version as well as the Biharmonic equation. Overall, HTE opens up a new capability in scientific machine learning for tackling high-order and high-dimensional PDEs.

Curse of dimensionality↗

Differentiable programming for online training of a neural artificial viscosity function within a staggered grid Lagrangian hydrodynamics scheme

Lagrangian methods to solve the inviscid Euler equations produce numerical oscillations near shock waves. A common approach to reducing these oscillations is to add artificial viscosity (AV) to the discrete equations. The AV term acts as a dissipative mechanism that attenuates oscillations by smearing the shock across a finite number of computational cells. However, AV introduces several control parameters that are not determined by the underlying physical model, and hence, in practice are tuned to the characteristics of a given problem. We seek to improve the standard quadratic-linear AV form by replacing it with a learned neural function that reduces oscillations relative to exact solutions of the Euler equations, resulting in a hybrid numerical-neural hydrodynamic solver. Because AV is an artificial construct that exists solely to improve the numerical properties of a hydrodynamic code, there is no offline ‘viscosity data’ against which a neural network can be trained before inserting into a numerical simulation, thus requiring online training. We achieve this via differentiable programming, i.e. end-to-end backpropagation or adjoint solution through both the neural and differential equation code, using automatic differentiation of the hybrid code in the Julia programming language to calculate the necessary loss function gradients. A novel offline pre-training step accelerates training by initializing the neural network to the default numerical AV scheme, which can be learned rapidly by space-filling sampling over the AV input space. We find that online training over early time steps of simulation is sufficient to learn a neural AV function that reduces numerical oscillations in long-term hydrodynamic shock simulations. These results offer an early proof-of-principle that online differentiable training of hybrid numerical schemes with novel neural network components can improve certain performance aspects existing in purely numerical schemes.

97 MATHEMATICS AND COMPUTING↗

Solving Inverse Stochastic Problems from Discrete Particle Observations Using the Fokker--Planck Equation and Physics-Informed Neural Networks

The Fokker--Planck (FP) equation governing the evolution of the probability density function (PDF) is applicable to many disciplines, but it requires specification of the coefficients for each case, which can be functions of space-time and not just constants and hence require the development of a data-driven modeling approach. When the data available is directly on the PDF, there exist methods for inverse problems that can be employed to infer the coefficients and thus determine the FP equation and subsequently obtain its solution. Herein, we address a more realistic scenario, where only sparse data are given on the particles' positions at a few time instants, which are not sufficient to accurately construct directly the PDF even at those times from existing methods, e.g., kernel estimation algorithms. To this end, we develop a general framework based on physics-informed neural networks (PINNs) that introduces a new loss function using the Kullback--Leibler divergence to connect the stochastic samples with the FP equation to simultaneously learn the equation and infer the multidimensional PDF at all times. In particular, we consider two types of inverse problems, type I, where the FP equation is known but the initial PDF is unknown, and type II, in which, in addition to the unknown initial PDF, the drift and diffusion terms are also unknown. In both cases, we investigate problems with either Brownian or Lévy noise or a combination of both. Here, we demonstrate the new PINN framework in detail in the one-dimensional (1D) case, but we also provide results for up to five dimensions demonstrating that we can infer both the FP equation and dynamics simultaneously at all times with high accuracy using only very few discrete observations of the particles.

97 MATHEMATICS AND COMPUTING↗

Spacecraft automated operations

Trends in automation of planetary spacecraft are examined using data from missions as far back as Mariner '67 and up to the highly sophisticated Galileo. Nine design considerations which influence the degree of automation such as protection against catastrophic failures, highly repetitive functions, loss of spacecraft communications, and the need for near-real-time adaptivity are discussed. Rapid growth of automation is shown in terms of on-board hardware by plots of number of processors on board, the average speed of processors, and total core memory. The number of commands transmitted from the ground has grown to 5 million bits in Voyager, so that increases in mission complexity have increased both in spacecraft automation and ground operations. Achieving greater automation by transferring ground operations to the spacecraft with the current means of controlling missions, are considered noting proposed changes. For the future, improved computer technology, more microprocessors and increased core storage will be used, and the number of automated functions and their complexity will grow. It is concluded that using the growing computational capability of spacecraft will achieve more autonomy thus reversing the trend of increased mission complexity and cost.

Bird, T. H.↗

SO(3)-invariance of informed-graph-based deep neural network for anisotropic elastoplastic materials

This work examines the frame-invariance (and the lack thereof) exhibited in simulated anisotropic elasto-plastic responses generated from supervised machine learning of classical multi-layer and informed-graph-based neural networks, and proposes different remedies to fix this drawback. The inherent hierarchical relations among physical quantities and state variables in an elasto-plasticity model are first represented as informed, directed graphs, where three variations of the graph are tested. While feed-forward neural networks are used to train path-independent constitutive relations (e.g., elasticity), recurrent neural networks are used to replicate responses that depends on the deformation history, i.e. or path dependent. In dealing with the objectivity deficiency, we use the spectral form to represent tensors and, subsequently, three metrics, the Euclidean distance between the Euler Angles, the distance from the identity matrix, and geodesic on the unit sphere in Lie algebra, can be employed to constitute objective functions for the supervised machine learning. In this, the aim is to minimize the measured distance between the true and the predicted 3D rotation entities. Following this, we conduct numerical experiments on how these metrics, which are theoretically equivalent, may lead to differences in the efficiency of the supervised machine learning as well as the accuracy and robustness of the resultant models. Neural network models trained with tensors represented in component form for a given Cartesian coordinate system are used as a benchmark. Our numerical tests show that, even given the same amount of information and data, the quality of the anisotropic elasto-plasticity model is highly sensitive to the way tensors are represented and measured. The results reveal that using a loss function based on geodesic on the unit sphere in Lie algebra together with an informed, directed graph yield significantly more accurate rotation prediction than the other tested approaches.

42 ENGINEERING↗

Physiology of a microgravity environment invited review: microgravity and skeletal muscle

Spaceflight (SF) has been shown to cause skeletal muscle atrophy; a loss in force and power; and, in the first few weeks, a preferential atrophy of extensors over flexors. The atrophy primarily results from a reduced protein synthesis that is likely triggered by the removal of the antigravity load. Contractile proteins are lost out of proportion to other cellular proteins, and the actin thin filament is lost disproportionately to the myosin thick filament. The decline in contractile protein explains the decrease in force per cross-sectional area, whereas the thin-filament loss may explain the observed postflight increase in the maximal velocity of shortening in the type I and IIa fiber types. Importantly, the microgravity-induced decline in peak power is partially offset by the increased fiber velocity. Muscle velocity is further increased by the microgravity-induced expression of fast-type myosin isozymes in slow fibers (hybrid I/II fibers) and by the increased expression of fast type II fiber types. SF increases the susceptibility of skeletal muscle to damage, with the actual damage elicited during postflight reloading. Evidence in rats indicates that SF increases fatigability and reduces the capacity for fat oxidation in skeletal muscles. Future studies will be required to establish the cellular and molecular mechanisms of the SF-induced muscle atrophy and functional loss and to develop effective exercise countermeasures.

short duration↗

Robust scalable initialization for Bayesian variational inference with multi-modal Laplace approximations

Predictive modeling typically relies on Bayesian model calibration to provide uncertainty quantification. Variational inference utilizing fully independent (“mean-field”) Gaussian distributions are often used as approximate probability density functions. This simplification is attractive since the number of variational parameters grows only linearly with the number of unknown model parameters. However, the resulting diagonal covariance structure and unimodal behavior can be too restrictive to provide useful approximations of intractable Bayesian posteriors that exhibit highly non-Gaussian behavior, including multimodality. High-fidelity surrogate posteriors for these problems can be obtained by considering the family of Gaussian mixtures. Gaussian mixtures are capable of capturing multiple modes and approximating any distribution to an arbitrary degree of accuracy, while maintaining some analytical tractability. Unfortunately, variational inference using Gaussian mixtures with full-covariance structures suffers from a quadratic growth in variational parameters with the number of model parameters. The existence of multiple local minima due to strong nonconvex trends in the loss functions often associated with variational inference present additional complications, These challenges motivate the need for robust initialization procedures to improve the performance and computational scalability of variational inference with mixture models. In this work, we propose a method for constructing an initial Gaussian mixture model approximation that can be used to warm-start the iterative solvers for variational inference. The procedure begins with a global optimization stage in model parameter space. In this step, local gradient-based optimization, globalized through multistart, is used to determine a set of local maxima, which we take to approximate the mixture component centers. Around each mode, a local Gaussian approximation is constructed via the Laplace approximation. Finally, the mixture weights are determined through constrained least squares regression. The robustness and scalability of the proposed methodology is demonstrated through application to an ensemble of synthetic tests using high-dimensional, multimodal probability density functions. Here, the practical aspects of the approach are demonstrated with inversion problems in structural dynamics.

97 MATHEMATICS AND COMPUTING↗

Staffing implications of software productivity models

The attributes of software project staffing and productivity implied by equating the effects of two popular software models in a small neighborhood of a given effort-duration point are investigated. The first model presupposes that organizational productivity decreases as a function of the project staff size due to interfacing and intercommunication. The second, the so-called software equation, relates the product size to effort and duration through a power law tradeoff formula. The conclusions that may be reached by assuming that both of these describe project behavior, the former as a global phenomenon and the latter as a localized effect in a small neighborhood of a given effort duration point, are that (1) there is a calculable maximum effective staff level, which, if exceeded, reduces the project production rate, (2) there is a calculable maximum extent to which effort and time may be traded effectively, (3) it becomes ineffective in a practical sense to expend more than an additional 25 to 50% of resources in order to reduce delivery time, and (4) the team production efficiency can be computed directly from the staff level, the slope of the intercommunication loss function, and the ratio of exponents in the software equation.

Tausworthe, R. C.↗

Bandgap engineering of SrZrS 3 chalcogenide perovskite via substitutional doping for photovoltaic applications: a first-principles DFT study

Strontium zirconium sulfide (SrZrS 3 ) has garnered significant attention for photovoltaic (PV) applications due to its excellent optoelectronic properties, high chemical and moisture stability, and non-toxicity. However, the bandgaps of both the α- and β-phases lie outside the optimum ranges for both single-junction solar cells (SJSCs) and tandem solar cells (TSCs), thereby limiting their applications in PV technologies. In this study, we employed hybrid density functional theory to engineer the band gaps of α- and β-SrZrS 3 through substitutional doping. We considered three dopants (Hf, Sn, and Ti) at the Zr site of SrZrS 3 with varying doping concentrations and investigated their effects on the structural, electronic, and optical properties of the materials. We found that Sn and Ti doping effectively lowers the band gaps of both α- and β-SrZrS 3 , whereas Hf doping increases them. For x values up to 0.25, the band gaps of the α-SrZr 1−x Sn x S 3 and α-SrZr 1−x Ti x S 3 are within the optimum range for SJSCs, and those of β-SrZr 1−x Sn x S 3 and β-SrZr 1−x Ti x S 3 lie within the optimum range for Si/perovskite as well as perovskite/perovskite TSCs. The three dopants exhibited significant effects on the optical properties of both α- and β-SrZrS 3 , including the absorption coefficient, energy-loss functions, reflectivity, and refractivity spectra. Thermodynamic stability analysis revealed that for both phases, SrZr 1−x Hf x S 3 can be synthesized via exothermic processes, whereas the formation of SrZr 1−x Ti x S 3 and SrZr1−xSn x S 3 is endothermic and hence, not thermodynamically favorable. Further analysis showed that SrZr 1−x Ti x S 3 in both α- and β-phases are stable under thermodynamic equilibrium conditions, whereas SrZr 1−x Sn x S 3 is prone to dissociation into ternary phases (SrZrS 3 and SrSnS 3 ), especially at higher doping concentrations. These results show that Ti doping is effective in tuning the band gaps of α- and β-SrZrS 3 toward the optimal values for PV applications.

14 SOLAR ENERGY↗