Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “loss function regularization”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 19 records

Leveraging large language models to address data scarcity in machine learning for graphene synthesis

Machine learning in experimental materials science faces significant challenges due to the scarcity of data, which are costly and time-consuming to generate, particularly when relying on in-house experiments. Literature data mining offers a potential solution but introduces issues like mixed data quality, inconsistent formats, and non-uniform reporting of synthesis parameters, resulting in partially missing and heterogeneous features across the dataset. Here, we propose data imputation and feature engineering methods that employ pre-trained large language models (LLMs) to enhance machine learning performance on scarce, heterogeneous datasets, demonstrated on graphene CVD synthesis data and the ML-HydPARK hydrogen storage dataset. GPT models perform data imputation via tailored prompting and semantic normalization of inconsistently reported features through embeddings, for example, to harmonize the complex nomenclature of CVD substrates. Beyond yielding more diverse and richer feature representations than traditional methods such as K-nearest neighbors (KNN) and Multivariate Imputation by Chained Equations (MICE), LLM-based data imputation is evaluated against dataset characteristics and prompting strategies. We vary the level of autonomy granted to the LLM, from generic prompting that leverages pre-trained knowledge for autonomous data generation to data-informed prompting that constrains outputs using target-specific information, and demonstrate which level of autonomy yields superior imputation performance across datasets and feature types. The proposed data engineering methods markedly improve downstream performance; for example, in graphene layer number classification using a support vector machine (SVM), binary accuracy increases from 39% to 65% and ternary accuracy from 52% to 72%. Fine-tuning experiments on both datasets show that combining our proposed LLM-based data imputation and feature encoding methods with numerical machine learning predictors outperforms standalone fine-tuned LLM predictors in data-scarce settings. The proposed strategies emphasize data enhancement techniques rather than refining learning architectures or regularizing loss functions, offering a broadly applicable framework for improving machine learning performance on scarce, inhomogeneous datasets.

Chemical vapor deposition↗

Connect the Dots: In Situ 4-D Seismic Monitoring of CO 2 Storage With Spatio-Temporal CNNs

4-D seismic imaging has been widely used in CO 2 sequestration projects to monitor the fluid flow in the volumetric subsurface region that is not sampled by wells. Ideally, real-time monitoring and near-future forecasting would provide site operators with great insights to understand the dynamics of the subsurface reservoir and assess any potential risks. However, due to obstacles such as high deployment cost, availability of acquisition equipment, exclusion zones around surface structures, only very sparse seismic imaging data can be obtained during monitoring. That leads to an unavoidable and growing knowledge gap over time. The operator needs to understand the fluid flow throughout the project lifetime and the seismic data are only available at a limited number of times. This is insufficient for understanding reservoir behavior. To overcome those challenges, we have developed spatio-temporal neural-network-based models that can produce high-fidelity interpolated or extrapolated images effectively and efficiently. Specifically, our models are built on an autoencoder, and incorporate the long short-term memory (LSTM) structure with a new loss function regularized by optical flow. We validate the performance of our models using real 4-D post-stack seismic imaging data acquired at the Sleipner CO 2 sequestration field. We employ two different strategies in evaluating our models. Numerically, we compare our models with different baseline approaches using classic pixel-based metrics. We also conduct a blind survey and collect a total of 20 responses from domain experts to evaluate the quality of data generated by our models. Finally, via both numerical and expert evaluation, we conclude that our models can produce high-quality 2-D/3-D seismic imaging data at a reasonable cost, offering the possibility of real-time monitoring or even near-future forecasting of the CO 2 storage reservoir.

4-D seismic imaging↗

Inter-well connectivity detection in CO 2 WAG projects using statistical recurrent unit models

Routine well-wise injection and production measurements contain significant information on subsurface structure and properties. Data-driven technology that interprets surface data into subsurface structure or properties can assist operators in making informed decisions by providing a better understanding of field assets. Our machine-learning framework is built on the statistical recurrent unit (SRU) model and interprets well-based injection/production data into inter-well connectivity without relying on a geologic model. We test it on synthetic and field-scale CO 2 EOR projects utilizing the water-alternating-gas (WAG) process. SRU is a special type of recurrent neural network (RNN) that allows for better characterization of temporal trends, by learning various statistics of the input at different time scales. In our application, the complete states (injection rate, pressure and cumulative injection) at injectors and pressure states at producers are fed to SRU as the input and the phase rates at producers are treated as the output. Once the SRU is trained and validated, it is then used to assess the connectivity of each injector to any producer using permutation variable importance method, wherein inputs corresponding to an injector are shuffled and the increase in prediction error at a given producer is recorded as the importance (connectivity metric) of the injector to the producer. This method is tested in both synthetic and field-scale cases. The validation of the proposed data-driven inter-well connectivity assessment is performed using synthetic data from simulation models where inter-well connectivity can be easily measured using the streamline-based flux allocation. The SRU model is shown to offer excellent prediction performance on the synthetic case. Despite significant measurement noise and frequent well shut-ins imposed in the field-scale case, the SRU model offers good prediction accuracy, the overall relative error of the phase production rates at most producers ranges from 10% to 30%. It is shown that the dominant connections identified by the data-driven method and streamline method are in close agreement. This significantly improves confidence in our data-driven procedure. The novelty of this work is that it is purely data-driven method and can directly interpret routine surface measurements to intuitive subsurface knowledge. Furthermore, the streamline-based validation procedure provides physics-based backing to the results obtained from data analytics. This study results in a reliable and efficient data analytics framework that is well-suited for large field applications.

42 ENGINEERING↗

Using spatio-temporal graph neural networks to estimate fleet-wide photovoltaic performance degradation patterns

Accurate estimation of photovoltaic (PV) system performance is crucial for determining its feasibility as a power generation technology and financial asset. PV-based energy solutions offer a viable alternative to traditional energy resources due to their superior Levelized Cost of Energy (LCOE). A significant challenge in assessing the LCOE of PV systems lies in understanding the Performance Loss Rate (PLR) for large fleets of PV systems. Estimating the PLR of PV systems becomes increasingly important in the rapidly growing PV industry. Precise PLR estimation benefits PV users by providing real-time monitoring of PV module performance, while explainable PLR estimation assists PV manufacturers in studying and enhancing the performance of their products. However, traditional PLR estimation methods based on statistical models have notable drawbacks. Firstly, they require user knowledge and decision-making. Secondly, they fail to leverage spatial coherence for fleet-level analysis. Additionally, these methods inherently assume the linearity of degradation, which is not representative of real world degradation. To overcome these challenges, we propose a novel graph deep learning-based decomposition method called the Spatio-Temporal Graph Neural Network for fleet-level PLR estimation (PV-stGNN-PLR). PV-stGNN-PLR decomposes the power timeseries data into aging and fluctuation components, utilizing the aging component to estimate PLR. PV-stGNN-PLR exploits spatial and temporal coherence to derive PLR estimation for all systems in a fleet and imposes flatness and smoothness regularization in loss function to ensure the successful disentanglement between aging and fluctuation. We have evaluated PV-stGNN-PLR on three simulated PV datasets consisting of 100 inverters from 5 sites. Experimental results show that PV-stGNN-PLR obtains a reduction of 33.9% and 35.1% on average in Mean Absolute Percent Error (MAPE) and Euclidean Distance (ED) in PLR degradation pattern estimation compared to the state-of-the-art PLR estimation methods.

14 SOLAR ENERGY↗

Surrogate-based Analysis and Optimization

A major challenge to the successful full-scale development of modem aerospace systems is to address competing objectives such as improved performance, reduced costs, and enhanced safety. Accurate, high-fidelity models are typically time consuming and computationally expensive. Furthermore, informed decisions should be made with an understanding of the impact (global sensitivity) of the design variables on the different objectives. In this context, the so-called surrogate-based approach for analysis and optimization can play a very valuable role. The surrogates are constructed using data drawn from high-fidelity models, and provide fast approximations of the objectives and constraints at new design points, thereby making sensitivity and optimization studies feasible. This paper provides a comprehensive discussion of the fundamental issues that arise in surrogate-based analysis and optimization (SBAO), highlighting concepts, methods, techniques, as well as practical implications. The issues addressed include the selection of the loss function and regularization criteria for constructing the surrogates, design of experiments, surrogate selection and construction, sensitivity analysis, convergence, and optimization. The multi-objective optimal design of a liquid rocket injector is presented to highlight the state of the art and to help guide future efforts.

Queipo, Nestor V.↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Data imbalance in drug response prediction: multi-objective optimization approach in deep learning setting

Abstract Drug response prediction (DRP) methods tackle the complex task of associating the effectiveness of small molecules with the specific genetic makeup of the patient. Anti-cancer DRP is a particularly challenging task requiring costly experiments as underlying pathogenic mechanisms are broad and associated with multiple genomic pathways. The scientific community has exerted significant efforts to generate public drug screening datasets, giving a path to various machine learning models that attempt to reason over complex data space of small compounds and biological characteristics of tumors. However, the data depth is still lacking compared to application domains like computer vision or natural language processing domains, limiting current learning capabilities. To combat this issue and improves the generalizability of the DRP models, we are exploring strategies that explicitly address the imbalance in the DRP datasets. We reframe the problem as a multi-objective optimization across multiple drugs to maximize deep learning model performance. We implement this approach by constructing Multi-Objective Optimization Regularized by Loss Entropy loss function and plugging it into a Deep Learning model. We demonstrate the utility of proposed drug discovery methods and make suggestions for further potential application of the work to achieve desirable outcomes in the healthcare field.

Biochemistry & Molecular Biology↗

Extension of the PINN diffusion model to k-eigenvalue problems

This paper extends our recent work on the Physics-Informed Neural Networks (PINN) approach for the fixed source diffusion models and applies it to the diffusion theory based k-eigenvalue problems. To make the PINN equitable for the eigenvalue problems, we introduce a novel integral regularization term to the loss function in the framework, and allow the direct inference of the principal eigenvalue and the associated eigenfunction. The regularization term enforces a pre-defined value on the integration of the model predictions, and this value can be directly related to a physical property of the system. We also introduce an additional learnable parameter to approximate the principal eigenvalue. As a proof of principle, we solve the one-group two-dimensional k-eigenvalue neutron diffusion equation in this work. We then provide two numerical examples to demonstrate the applicability of the PINN approach. In each example, we solve the k-eigenvalue diffusion equation in a multi-region configuration constrained with a set of Robin boundary conditions for generality. We use a FEM solution based on the power-iteration method to verify the results of the PINN solution. The results showed relative percentage error in the predicted eigenvalue of about 0.77% and about 1.2% for example 1 and example 2, respectively. The mean absolute error in the predicted flux for example 1 is ∼ 0.002 and for example 2 is ∼ 0.0024. These results indicate some preliminary successes of the PINN application to k-eigenvalue problems. (authors)

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

DENSECL: Haze Mitigation Using Dense Blocks and Contrastive Loss Regularization

Haze, which occurs as a result of the scattering of light in the atmosphere by small particles, diminishes the visibility of scene objects, inflicting important image applications such as object detection. To address the problem, this paper introduces a new physics-based end-to-end deep learning approach to haze mitigation in outdoor scenes, including those in airborne images. The proposed model named DenseCL is designed with dense blocks and adopts a contrastive loss function as an additional regularization. The model also maintains the cycle consistency by remapping the dehazed outputs into a hazy image using the physics-based light scattering function. DenseCL has been trained with publicly available outdoor images and demonstrates outstanding performance on outdoor, indoor, and remotely sensed nonhomogeneous haze satellite images.

Mitra, Somosmita↗

A physics-informed and hierarchically regularized data-driven model for predicting fluid flow through porous media

This paper presents a new deep learning data-driven model for predicting structure dependent pore-fluid velocity fields in rock. The model is based on a Convolutional Auto-Encoder (CAE) artificial neural network capable of learning from image data generated by direct numerical simulations of fluid flow through pore-structures, such as by Lattice Boltzmann or molecular dynamics methods. The main novelty of the model in comparison to previous CAE-based data-driven approaches consists of three parts. The first is a methodology for decomposing the full-domain of the porous media into sub-regions, or “sub-domains”, in order to reduce the overall size of the CAE, batch process the sub-domains in parallel, and enable the CAE to learn local and generalizable nonlinear mappings of pore-fluid velocities. The second consists of embedding the finite difference solutions of the incompressible Navier-Stokes and continuity equations into convolutional layers prior to the CAE in order to provide the CAE with knowledge of fluid dynamics physics (PhyFlow). The third main novelty is that the training of the CAE is regularized with a hierarchical loss function that encourages the learning of fluid flow patterns (in a way similar to ranked modes in principal component analysis), ranking from most to least important. This is shown to increase the stability in learning, reduce over-fitting, and promote interpretability of the CAE neural network layers (HierCAE). The comprehensive new data-driven model, which we call the PhyFlow-HierCAE model, is shown to exhibit improved accuracy and generalizability of flow field predictions over conventional CAE models, attributable to the embedded physical knowledge and the hierarchical regularization, as well as realize orders of magnitude speed-ups in computation times as a surrogate for the direct numerical simulations. Examples of training and forward predictions on unseen pore-structures are provided and evaluated for data from Lattice Boltzmann and molecular dynamics simulations of pore-fluid flow. The model is shown to be a fast and accurate emulator (or “surrogate”) for predicting effective permeability of unseen pore-structures based on learning from relatively small direct numerical simulation datasets.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Group-equivariant autoencoder for identifying spontaneously broken symmetries

We introduce the group-equivariant autoencoder (GE autoencoder), a deep neural network (DNN) method that locates phase boundaries by determining which symmetries of the Hamiltonian have spontaneously broken at each temperature. We use group theory to deduce which symmetries of the system remain intact in all phases, and then use this information to constrain the parameters of the GE autoencoder such that the encoder learns an order parameter invariant to these “never-broken” symmetries. This procedure produces a dramatic reduction in the number of free parameters such that the GE-autoencoder size is independent of the system size. We include symmetry regularization terms in the loss function of the GE autoencoder so that the learned order parameter is also equivariant to the remaining symmetries of the system. By examining the group representation by which the learned order parameter transforms, we are then able to extract information about the associated spontaneous symmetry breaking. We test the GE autoencoder on the 2D classical ferromagnetic and antiferromagnetic Ising models, finding that the GE autoencoder (1) accurately determines which symmetries have spontaneously broken at each temperature; (2) estimates the critical temperature in the thermodynamic limit with greater accuracy, robustness, and time efficiency than a symmetry-agnostic baseline autoencoder; and (3) detects the presence of an external symmetry-breaking magnetic field with greater sensitivity than the baseline method. Lastly, we describe various key implementation details, including a quadratic-programming-based method for extracting the critical temperature estimate from trained autoencoders and calculations of the DNN initialization and learning rate settings required for fair model comparisons.

42 ENGINEERING↗

Keeping a Beat on the Heart

Feel the relief of a patient suffering from heart arrhythmia, who is able to return home while having her heart monitored by health professionals 24 hours a day, without the fear that she will miss an important indicator and suffer a fatal heart attack - using technology originally developed to conduct experiments on the Space Shuttle. Approximately 400,000 Americans die every year from sudden heart attacks . Medical research revealed that patterns of electrical activity in the heart can act as predictors of these lethal cardiac events known as arrhythmias. Fortunately, certain arrhythmias such as ventricular fibrillation (loss of regular heartbeat and subsequent loss of function) and ventricular tachycardia (rapid heartbeats), can be detected and appropriately treated. Today, patients at moderate risk of arrhythmias can benefit from technology that would permit long- term continuous monitoring of electrical cardiac rhythms outside the hospital environment in the comfort of their own homes. Medical telemetry systems, also known as telemedicine, are evolving rapidly as wireless communication technology advances, evidenced by the commercial products and research prototypes for remote health monitoring that have appeared in recent years. Wireless systems allow patients to move freely in their home and work environment while being monitored remotely by health care professionals.

Liszka, Kathy J.↗

Magnetic field mapping of inaccessible regions using physics-informed neural networks

A difficult problem concerns the determination of magnetic field components within an experimentally inaccessible region when direct field measurements are not feasible. In this paper, we propose a new method of accessing magnetic field components using non-disruptive magnetic field measurements on a surface enclosing the experimental region. Magnetic field components in the experimental region are predicted by solving a set of partial differential equations (Ampere’s law and Gauss’ law for magnetism) numerically with the aid of physics-informed neural networks (PINNs). Prediction errors due to noisy magnetic field measurements and small number of magnetic field measurements are regularized by the physics information term in the loss function. We benchmark our model by comparing it with an older method. The new method we present will be of broad interest to experiments requiring precise determination of magnetic field components, such as searches for the neutron electric dipole moment.

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Multi-resolution partial differential equations preserved learning framework for spatiotemporal dynamics

Traditional data-driven deep learning models often struggle with high training costs, error accumulation, and poor generalizability in complex physical processes. Physics-informed deep learning (PiDL) addresses these challenges by incorporating physical principles into the model. Most PiDL approaches regularize training by embedding governing equations into the loss function, yet this depends heavily on extensive hyperparameter tuning to weigh each loss term. To this end, we propose to leverage physics prior knowledge by “baking” the discretized governing equations into the neural network architecture via the connection between the partial differential equations (PDE) operators and network structures, resulting in a PDE-preserved neural network (PPNN). This method, embedding discretized PDEs through convolutional residual networks in a multi-resolution setting, largely improves the generalizability and long-term prediction accuracy, outperforming conventional black-box models. The effectiveness and merit of the proposed methods have been demonstrated across various spatiotemporal dynamical systems governed by spatiotemporal PDEs, including reaction-diffusion, Burgers’, and Navier-Stokes equations.

97 MATHEMATICS AND COMPUTING↗

Jensen–Shannon divergence based novel loss functions for Bayesian neural networks

Bayesian neural networks (BNNs) are state-of-the-art machine learning methods that can naturally regularize and systematically quantify uncertainties using their stochastic parameters. Kullback–Leibler (KL) divergence-based variational inference used in BNNs suffer from unstable optimization and challenges in approximating light-tailed posteriors due to the unbounded nature of the KL divergence. To resolve these issues, we formulate a novel loss function for BNNs based on a new modification to the generalized Jensen–Shannon (JS) divergence, which is bounded. In addition, we propose a Geometric JS divergence-based loss, which is computationally efficient since it can be evaluated analytically. We found that the JS divergence-based variational inference is intractable, and hence employed a constrained optimization framework to formulate these losses. Our theoretical analysis and empirical experiments on multiple regression and classification data sets suggest that the proposed losses perform better than the KL divergence-based loss, especially when the data sets are noisy or biased. Specifically, there are approximately 5% and 8% improvements in accuracy for a noise-added CIFAR-10 dataset and a regression dataset, respectively. There is about 13% reduction in false negative predictions of a biased histopathology dataset. Additionally, we quantify and compare the uncertainty metrics for the regression and classification tasks.

97 MATHEMATICS AND COMPUTING↗

Short term hearing loss in general aviation operations, phase 1, part 1

The effects of light aircraft noise on six subjects during flight operations were investigated. The noise environment in the Piper Apache light aircraft was found to be capable of producing hearing threshold shifts. The following are the principal findings and conclusions: (1) Through most of the frequency range for which measurements were taken (500 to 6000 Hz), there was a regular progression showing increased loss of auditory acuity as a function of increased exposure time. (2) Extensive variability was found in the results among subjects, and in the measured loss at discrete frequencies for each subject. (3) The principal loss of hearing occurred at the low frequencies, around 500 Hz.

Parker, J. F., Jr.↗

Physics-informed Karhunen-Loeve and Neural Network Approximations for Solving Inverse Differential Equation Problems

Here we present the PI-CKL-NN method for parameter estimation in differential equation (DE) models given sparse measurements of the parameters and states. In the proposed approach, the space- or time-dependent parameters are approximated by Karhunen-Loeve (KL) expansions that are conditioned on the parameters’ measurements, and the states are approximated by deep neural networks (DNNs). The unknown weights in the KL expansions and DNNs are found my minimizing the cost function that enforces the measurements of the states the DE constraint. Regularization is achieved by adding the l2 norm of the conditional KL coefficients into the loss function. Our approach assumes that the parameter fields are correlated in space or time and enforces the statistical knowledge (the mean and the covariance function) in addition to the DE constraints and measurements as opposed to the physics-informed neural network (PINN) and other similar physics-informed machine learning methods where only DE constraints and data are used for parameter estimation. We use the PI-CKL-NN method for parameter estimation in an ordinary differential equation with an unknown time-dependent parameter and the one- and two-dimensional partial differential diffusion equations with unknown space-dependent diffusion coefficients. We also demonstrate that PI-CKL-NN is more accurate than the PINN method, especially when the observations of the parameters are very sparse

71 CLASSICAL AND QUANTUM MECHANICS, GENERAL PHYSIC↗

Predictive Modeling for Differential Diagnosis and Mortality Risk Assessment

The prevalence of electronic health record (EHR) systems has brought prodigious biomedical informatics opportunity. Automated machine learning methods can effectively utilize such data and have become common tools for healthcare predictive modeling. Researches in medical informatics have explored the potential of deep learning and classical models in emergent care scenarios. In particular, predicting differential diagnoses for admissions have proven useful in decreasing unnecessary lab tests and improving inpatient triage decision-making. Moreover, identification of high-risk patients for in-hospital mortality is vitally important to maximize allocation of medical resources.The Medical Information Mart for Intensive Care (MIMIC-III) database, containing de-identified critical care inpatient was used in our study. This data set captures hospital patient laboratory measurements, pharmacologic prescriptions, diagnostic data and procedure event recordings. When considering adult patients and discounting admissions with ICU length of stay less than 24 hours, there were 37,787 unique admissions and 30,414 total patients. We examined the top 25 most prevalent ICD-9 group-level disease specificities in MIMIC-III using a multi-label classification model. In-hospital mortality was modeled as binary classification with 4,155 (13%) adult patients that expired, of which 3,138 (75.5%) were in the ICU setting. The metrics AUC, F1 score, sensitivity and specificity values calculated for each disease label measured prediction performance.The usage of ICD-9 group codes reduced feature dimension from 14,567 to 942 and greatly improved distribution of patient diagnostic categories. Disease temporal patterns were captured by considering the most frequently sampled 6 vital signs and 13 laboratory values. Missing data were imputed at each time-stamp. Time-series raw hourly average values were converted into 5 summary features (mean, standard deviation, number of observations, min & max values). Patient demographic variables such as age, gender, marital status and ethnicity were also factored into the modeling. Choi et al showed that contextual embedding of medical data, diagnostic and procedural codes alone can predict future diagnoses with sensitivity as high as 0.79. We utilized an embedding technique called word2vec which allowed sparse representations of medical history to be transformed into dense word vectors. The mappings captured contextual information by treating each admission as a sentence and learning the most likely neighboring words in a sliding window fashion. Binary and multi-label classification was achieved via collapse models, which do not consider temporal information, as well as recurrent neural networks with regularization, Softmax output layer activation together with categorical cross-entropy as the loss function.

US Army collaboration↗