Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Transfer Learning Model”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 325 records · Page 18

High dimensional predictions of suicide risk in 4.2 million US Veterans using ensemble transfer learning

We present an ensemble transfer learning method to predict suicide from Veterans Affairs (VA) electronic medical records (EMR). A diverse set of base models was trained to predict a binary outcome constructed from reported suicide, suicide attempt, and overdose diagnoses with varying choices of study design and prediction methodology. Each model used twenty cross-sectional and 190 longitudinal variables observed in eight time intervals covering 7.5 years prior to the time of prediction. Ensembles of seven base models were created and fine-tuned with ten variables expected to change with study design and outcome definition in order to predict suicide and combined outcome in a prospective cohort. The ensemble models achieved c-statistics of 0.73 on 2-year suicide risk and 0.83 on the combined outcome when predicting on a prospective cohort of ~4.2 M veterans. The ensembles rely on nonlinear base models trained using a matched retrospective nested case-control (Rcc) study cohort and show good calibration across a diversity of subgroups, including risk strata, age, sex, race, and level of healthcare utilization. In addition, a linear Rcc base model provided a rich set of biological predictors, including indicators of suicide, substance use disorder, mental health diagnoses and treatments, hypoxia and vascular damage, and demographics. Similar content being viewed by others

60 APPLIED LIFE SCIENCES↗

Multiscale CFD simulation of biomass fast pyrolysis with a machine learning derived intra-particle model and detailed pyrolysis kinetics

Coupling particle and reactor scale models is as essential as reactor fluid dynamics and particle motion for accurate Computational Fluid Dynamic (CFD) simulations of biomass fast pyrolysis reactors due to intraparticle heat transfer and chemical reactions controlling conversion time and product distributions. Direct online coupling of a particle model with a reactor model is computationally expensive, while offline coupling is case-dependent. In this research, solutions from a series of particle pyrolysis simulations were regressed with Artificial Neural Network (ANN). Furthermoer, this machine learning-derived model predicted the same temperature and conversion profiles compared with particle resolved simulation while the isothermal approach overpredicted the temperature by 130 K and underpredicted the conversion time by 30 s. The ANN model was then integrated into CFD simulations of fluidized bed biomass fast pyrolysis with varied feedstocks via coupling PyTorch and MFiX. The averaged error of simulation predicted bio-oil yields with four feedstocks is 6.4%. This multi-scale approach provides an efficient tool for the coupled particle and reactor scale simulations of biomass pyrolysis.

09 BIOMASS FUELS↗

Multitask Recommender Systems for Cancer Drug Response

The problem we are currently trying to address is that there are many types of cancer drugs and many types of cancers and there is not always experimental data for a specific cancer type and cancer drug interaction. While there is a large possible set of feasible drug and cancer combinations, testing each pair is not realistic due to the high monetary cost of cell-based assays. Thus, this leaves researchers with a difficult choice of what drugs they should test on specific cancer types. This issue is known as the cold-start problem. Our focus is on developing recommender systems capable of addressing the cold-start problem as it relates to interaction between cancer types and cancer drugs. One of the most effective ways to address the cold-start problem is through large data analysis, however due to the cost prohibitive nature of cancer research the largest available data set size is the Genomics of Drug Sensitivity in Cancer with 494,973 genomic associations. To achieve optimal model performance on the cold-start problem, it is advantageous to employ multitask algorithms that are capable of transferring information between cancer datasets. The aim of this report is to draw from adaptations and state of the art developments in both algorithms for recommender systems and multitask learning to model the interaction between cancer cell lines and cancer drugs. Cancer cell lines are defined by the US National Cancer Institute as "cancer cells that keep dividing and growing over time, under certain conditions in a laboratory". This paper will focus on evaluating the performance of Neural Collaborative Filtering and Gaussian Processes, as well as their multitask adaptations, on cancer datasets from CCLE, NCI60, GDSC and CTRP. These methods will be evaluated on model performance in regression prediction but also in interpretability.

60 APPLIED LIFE SCIENCES↗

Reinforcement Learning for Load-balanced Parallel Particle Tracing

We explore an online reinforcement learning (RL) paradigm to dynamically optimize parallel particle tracing performance in distributed-memory systems. Our method combines three novel components: (1) a work donation algorithm, (2) a high-order workload estimation model, and (3) a communication cost model. First, we design an RL-based work donation algorithm. Our algorithm monitors workloads of processes and creates RL agents to donate data blocks and particles from high-workload processes to low-workload processes to minimize program execution time. The agents learn the donation strategy on the fly based on reward and cost functions designed to consider processes' workload changes and data transfer costs of donation actions. Second, we propose a workload estimation model, helping RL agents estimate the workload distribution of processes in future computations. Third, we design a communication cost model that considers both block and particle data exchange costs, helping RL agents make effective decisions with minimized communication costs. We demonstrate that our algorithm adapts to different flow behaviors in large-scale fluid dynamics, ocean, and weather simulation data. Our algorithm improves parallel particle tracing performance in terms of parallel efficiency, load balance, and costs of I/O and communication for evaluations with up to 16,384 processors.

Distributed and parallel particle tracing↗

Deconvoluting thermomechanical effects in X-ray diffraction data using machine learning

X-ray diffraction is ideal for probing the sub-surface state during complex or rapid thermomechanical loading of crystalline materials. However, challenges arise as the size of diffraction volumes increases due to spatial broadening and because of the inability to deconvolute the effects of different lattice deformation mechanisms. Here, we present a novel approach that uses combinations of physics-based modeling and machine learning to deconvolve thermal and mechanical elastic strains for diffraction data analysis. The method builds on a previous effort to extract thermal strain distribution information from diffraction data. The new approach is applied to extract the evolution of the thermomechanical state during laser melting of an Inconel 625 wall specimen which produces significant residual stress upon cooling. A combination of heat transfer and fluid flow, elasto-plasticity and X-ray diffraction simulations is used to generate training data for machine-learning (Gaussian process regression, GPR) models that map diffracted intensity distributions to underlying thermomechanical strain fields. First-principles density functional theory is used to determine accurate temperature-dependent thermal expansion and elastic stiffness used for elasto-plasticity modeling. The trained GPR models are found to be capable of deconvoluting the effects of thermal and mechanical strains, in addition to providing information about underlying strain distributions, even from complex diffraction patterns with irregularly shaped peaks.

36 MATERIALS SCIENCE↗

Active learning of a crystal plasticity flow rule from discrete dislocation dynamics simulations

Continuum-scale material deformation models, such as crystal plasticity (CP), can significantly enhance their predictive accuracy by incorporating input from lower-scale (i.e. mesoscale) models. The procedure to generate and extract the relevant information is however typically complex and ad hoc, involving decision and intervention by domain experts, leading to long development times. In this study, we develop a principled approach for calibration of continuum-scale models using lower scale information by representing a CP flow rule as a Gaussian process model. This representation allows for efficient parameter space exploration, guided by the uncertainty embedded in the model through a process known as Bayesian optimization (BO). We demonstrate a semi-autonomous BO loop which instantiates discrete dislocation dynamics simulations whose initial conditions are automatically chosen to optimize the uncertainty of a model CP flow rule. Our self-guided computational pipeline efficiently generated a dataset and corresponding model whose error, uncertainty, and physical feature sensitivities were validated with comparison to an independent dataset four times larger, demonstrating a valuable and efficient active learning implementation readily transferable to similar material systems.

36 MATERIALS SCIENCE↗

AP-Net: An atomic-pairwise neural network for smooth and transferable interaction potentials

Intermolecular interactions are critical to many chemical phenomena, but their accurate computation using ab initio methods is often limited by computational cost. The recent emergence of machine learning (ML) potentials may be a promising alternative. Useful ML models should not only estimate accurate interaction energies but also predict smooth and asymptotically correct potential energy surfaces. However, existing ML models are not guaranteed to obey these constraints. Indeed, systemic deficiencies are apparent in the predictions of our previous hydrogen-bond model as well as the popular ANI-1X model, which we attribute to the use of an atomic energy partition. As a solution, we propose an alternative atomic-pairwise framework specifically for intermolecular ML potentials, and we introduce AP-Net—a neural network model for interaction energies. The AP-Net model is developed using this physically motivated atomic-pairwise paradigm and also exploits the interpretability of symmetry adapted perturbation theory (SAPT). We show that in contrast to other models, AP-Net produces smooth, physically meaningful intermolecular potentials exhibiting correct asymptotic behavior. Initially trained on only a limited number of mostly hydrogen-bonded dimers, AP-Net makes accurate predictions across the chemically diverse S66x8 dataset, demonstrating significant transferability. On a test set including experimental hydrogen-bonded dimers, AP-Net predicts total interaction energies with a mean absolute error of 0.37 kcal mol−1, reducing errors by a factor of 2–5 across SAPT components from previous neural network potentials. The pairwise interaction energies of the model are physically interpretable, and an investigation of predicted electrostatic energies suggests that the model “learns” the physics of hydrogen-bonded interactions.

Glick, Zachary L. (ORCID:0000000309002849)↗

Machine-learning modeling of magnetization dynamics in quasi-equilibrium and driven metallic spin systems

Here, we present a perspective on recent progress in machine-learning (ML) force-field approaches for large-scale Landau–Lifshitz–Gilbert (LLG) simulations of metallic spin systems. Building on a generalization of the Behler–Parrinello (BP) architecture originally developed for quantum molecular dynamics, we develop scalable and transferable ML models that faithfully capture the complex, environment-dependent electron-mediated exchange fields characteristic of itinerant magnets. A central ingredient of this framework is the implementation of symmetry-aware magnetic descriptors based on group-theoretical bispectrum formalisms. Leveraging these ML force fields, LLG simulations faithfully reproduce hallmark non-collinear magnetic orders—such as the 120° and tetrahedral states—on the triangular lattice, and successfully capture the complex spin textures emerging in the mixed-phase states of a square-lattice double-exchange model under thermal quench. We further discuss a generalized potential theory that extends the BP formalism to incorporate both conservative and nonconservative electronic torques, thereby enabling ML models to learn nonequilibrium exchange fields from computationally demanding microscopic approaches such as nonequilibrium Green’s-function techniques. This extension yields quantitatively accurate predictions of voltage-driven domain-wall motion and establishes a foundation for quantum-accurate, multiscale modeling of nonequilibrium spin dynamics and spintronic functionalities.

Descriptors↗

Analytical Modeling of Exoplanet Transit Spectroscopy with Dimensional Analysis and Symbolic Regression

Abstract The physical characteristics and atmospheric chemical composition of newly discovered exoplanets are often inferred from their transit spectra, which are obtained from complex numerical models of radiative transfer. Alternatively, simple analytical expressions provide insightful physical intuition into the relevant atmospheric processes. The deep-learning revolution has opened the door for deriving such analytical results directly with a computer algorithm fitting to the data. As a proof of concept, we successfully demonstrate the use of symbolic regression on synthetic data for the transit radii of generic hot-Jupiter exoplanets to derive a corresponding analytical formula. As a preprocessing step, we use dimensional analysis to identify the relevant dimensionless combinations of variables and reduce the number of independent inputs, which improves the performance of the symbolic regression. The dimensional analysis also allowed us to mathematically derive and properly parameterize the most general family of degeneracies among the input atmospheric parameters that affect the characterization of an exoplanet atmosphere through transit spectroscopy.

79 ASTRONOMY AND ASTROPHYSICS↗

A leaf-level spectral library to support high-throughput plant phenotyping: predictive accuracy and model transfer

Abstract Leaf-level hyperspectral reflectance has become an effective tool for high-throughput phenotyping of plant leaf traits due to its rapid, low-cost, multi-sensing, and non-destructive nature. However, collecting samples for model calibration can still be expensive, and models show poor transferability among different datasets. This study had three specific objectives: first, to assemble a large library of leaf hyperspectral data (n=2460) from maize and sorghum; second, to evaluate two machine-learning approaches to estimate nine leaf properties (chlorophyll, thickness, water content, nitrogen, phosphorus, potassium, calcium, magnesium, and sulfur); and third, to investigate the usefulness of this spectral library for predicting external datasets (n=445) including soybean and camelina using extra-weighted spiking. Internal cross-validation showed satisfactory performance of the spectral library to estimate all nine traits (mean R2=0.688), with partial least-squares regression outperforming deep neural network models. Models calibrated solely using the spectral library showed degraded performance on external datasets (mean R2=0.159 for camelina, 0.337 for soybean). Models improved significantly when a small portion of external samples (n=20) was added to the library via extra-weighted spiking (mean R2=0.574 for camelina, 0.536 for soybean). The leaf-level spectral library greatly benefits plant physiological and biochemical phenotyping, whilst extra-weight spiking improves model transferability and extends its utility.

59 BASIC BIOLOGICAL SCIENCES↗

Parametric and Sensitivity Analysis of a Steam Generator Model Using Python and Machine-Learning Tools

For this study, we used Python and machine-learning tools to perform a comprehensive parametric and sensitivity analysis on a steam generator (SG) model. (The Python model was based on a previously completed MATLAB framework for the Holtec SMR-160 SG.) We investigated the influence of various input parameters (e.g., heat transfer coefficient [HTC], Nusselt number, and heat exchanger effectiveness) on the system’s output. With machine-learning tools such as the Risk Analysis Virtual Environment (RAVEN), which was developed at Idaho National Laboratory, we were then able to perform an automated analysis of the SG inputs’ effect on the HTC. The analysis results give valuable insights into the performance and optimization of SG systems. We found the inlet mass flow rate (MFR) to have the greatest impact on the HTC, followed closely by the inlet temperature, and then pressure. Shifting of the input parameters causes the location of the maximum HTC along the SG length to change incrementally. The cold leg (CL) MFR was also found to impact the HTC magnitude as well as the location of the maximum HTC. At between 0.4–0.9 of the total SG length, the input parameters experience maximum impact on the HTC, leading us to suggest that sensors be efficiently placed on the SG so as to closely and effectively monitor thermal-hydraulic properties during reactor operation. We also found that the sensitivity data calculated manually agrees with the RAVEN – based data, confirming the same range of maximum sensitivity. However, the RAVEN-based analysis showed that cold leg pressure and hot leg temperature have a greater impact on the heat transfer coefficient than the mass flow rate, implying that a manual sensitivity study taking only two samples is not accurate.

20 FOSSIL-FUELED POWER PLANTS↗

Transfer Learning-Based Independent Component Analysis

Understanding the underlying component structure is crucial for multivariate signal analysis. Among all the techniques that try to learn the latent structure, independent component analysis (ICA) is one of the most important and popular methods, which aims to extract independent components from multivariate signals and enables further analysis. For example, in electroencephalogram (EEG) analysis, artifacts filtering and disease detection are conducted based on the independent components of the signals. One critical challenge in existing ICA approaches is that the component extraction accuracy may degrade when the available data of a unit are limited. To address this issue, this paper proposes a transfer learning-based ICA method by innovatively transferring component distribution from a source domain, so that accurate component extraction results can be achieved even when only limited data are available in the target domain. To the best of our knowledge, this is the first work that leverages transfer learning to improve ICA accuracy with limited available data. In particular, we first extract all the independent components from the source domain by maximizing the log-likelihood function with a Newton-like method on a smooth manifold. Then for the target domain, the component with the largest negentropy is extracted in each round. To effectively leverage the knowledge from the source domain and to prevent the negative transfer, we try to find a component in the source domain that matches the component we are extracting. The probability density function of the matched component will then be used to improve the component extraction accuracy if such matched component can be found; otherwise, no knowledge will be transferred. Finally, numerical simulations and a case study with electrocardiogram (ECG) data are conducted, showing the effectiveness of the proposed method in transferring knowledge and reducing negative transfer.

42 ENGINEERING↗

Using ensembles and distillation to optimize the deployment of deep learning models for the classification of electronic cancer pathology reports

One of the goals of the Surveillance, Epidemiology, and End Results (SEER) program is to estimate incidence, prevalence, and mortality of all cancers. To that end, cancer registries across the country maintain a massive database of cancer pathology reports which contain rich information to understand cancer trends. However, these reports are stored in the form of unstructured text, and human annotators are required to read and extract relevant information. In this article, we show that existing deep learning models for automating information extraction from cancer pathology reports can be significantly improved by using ensemble model distillation. We found that by training multiple predictive models and transferring their knowledge to a single, low-resource model, we can reduce the number of highly confident wrong predictions. Our results show that our implemented methods could save 1000s of manual annotation hours.

60 APPLIED LIFE SCIENCES↗

Influence of extreme temperature conditions on CO 2 direct air capture using amino-acid solutions

Geological features play a pivotal role in determining the feasibility of deploying CO₂ direct air capture (DAC) technologies, primarily because they influence the availability of cost-effective energy sources, such as natural gas and geothermal energy, and also due to the potential for CO₂ sequestration. Many regions face challenges due to variable weather conditions including seasonal temperature fluctuations, high or low humidity, and sub-ambient temperatures. These extremes can reduce DAC performance or even lead to catastrophic events. Aqueous solvents considered for DAC systems are particularly vulnerable to seasonal variations in colder climates, where the solvent may underperform or freeze. It is therefore essential to investigate the CO₂ capture efficiency of aqueous solvents across a broad range of environmental temperatures, spanning sub-zero to hot conditions (>30 °C). In this study, DAC operation is examined using a high-flux solvent–air crossflow contactor under two major weather scenarios: (i) cold conditions below 0 °C and (ii) hot conditions above 30 °C. A parametric study is conducted to investigate the contactor performance regarding CO₂ removal efficiency, uptake capacity, and reaction kinetics versus temperature when the air velocity through the contactor exceeds 1 m/s. The efficacy of the contactor is systematically investigated using various anti-freeze amino-acid solvent formulations. A mass-transfer mechanistic model is developed to assess the process performance over a wide temperature range and propose scalable design guidelines. Machine learning is also employed to identify key parameters affecting the CO₂ capture efficiency. It is shown that air velocity and temperature are the primary factors influencing CO₂ uptake. Based on performance data obtained under subfreezing temperatures, a technoeconomic analysis is conducted to evaluate the feasibility of using aqueous solvents in seasonal cold regions. In conclusion, the findings of this study provide valuable insights into siting considerations for deploying solvent-based DAC, thereby contributing to the advancement of sustainable carbon removal solutions.

Air–liquid contactor↗

Machine learning-accelerated path integral molecular dynamics simulations of reactive organic electrolytes

Hydrogen bonded electrolytes that exhibit accelerated proton transport via sequential reactive hops have drawn interest for their promise in clean energy applications. Molecular dynamics simulations of these electrolytes offer the opportunity to uncover microscopic mechanistic details that could be used to design and tune the properties of candidate electrolyte technologies. However, accurately modeling the proton transfer reactions and transport properties that give rise to high charge conductivites in these electrolytes proves computationally challenging because of the need to perform lengthy condensed phase simulations, treating both the electronic and nuclear degrees of freedom quantum mechanically. In this paper, we demonstrate that such a modeling task can be efficiently achieved with the use of density functional theory (DFT)-trained machine learning potentials (MLP) to accelerate path integral molecular dynamics (PIMD) simulations. We highlight the practical utility of this approach by using it to benchmark how closely PIMD simulations employing different DFT exchange–correlation functionals reproduce the composition-dependent densities, diffusion coefficients, and electrical conductivities of mixtures consisting of imidazole and levulinic acid. Even with the speedup afforded by our MLPs, PIMD simulations remain quite expensive. Furthermore, in order to render PIMD more computationally tractable, we introduce and benchmark the accuracy of a ring polymer contraction approach that leverages a computationally efficient short-range MLP to accelerate our PIMD simulations by an additional factor of four.

Chemical bonding↗

Self-supervised and multi-fidelity learning for extended predictive soil spectroscopy

Infrared spectroscopy is a cost-effective, non-destructive, and environmentally benign technology that is increasingly recognized as an important solution for meeting the global demand for soil data. While both near-infrared (NIR) and mid-infrared (MIR) diffuse reflectance spectroscopy enable rapid estimation of soil properties, they present a significant trade-off: NIR offers superior scalability and lower operational costs, whereas MIR provides higher analytical fidelity by capturing fundamental molecular vibrations. In this study, we propose a self-supervised, multi-fidelity learning framework designed to bridge this gap. Our approach leverages large-scale MIR spectral libraries to learn a compact, transferable latent representation, into which NIR spectra are subsequently aligned for downstream prediction. The workflow consists of pretraining a latent model on a large MIR library, adapting the representation using a smaller paired NIR–MIR dataset, and evaluating generalization on an independent external test set. Across a range of chemical and physical soil properties, we found that MIR-derived embeddings improved prediction accuracy relative to baseline models that used raw MIR inputs. Predictions derived from the spectrum conversion (NIR to MIR) task did not match the performance of the original MIR spectra but were similar or superior to predictive performance of NIR-only models, suggesting the unified spectral latent space can effectively leverage the larger and more diverse MIR dataset for prediction of soil properties not well represented in current NIR libraries.

54 ENVIRONMENTAL SCIENCES↗

Transformer quantum state: A multipurpose model for quantum many-body problems

Here, inspired by the advancements in large language models based on transformers, we introduce the transformer quantum state (TQS): a versatile machine learning model for quantum many-body problems. In sharp contrast to Hamiltonian/task specific models, TQS can generate the entire phase diagram, predict field strengths with experimental measurements, and transfer such a knowledge to new systems it has never been trained on before, all within a single model. With specific tasks, fine-tuning the TQS produces accurate results with small computational cost. Versatile by design, TQS can be easily adapted to new tasks, thereby pointing towards a general-purpose model for various challenging quantum problems.

75 CONDENSED MATTER PHYSICS, SUPERCONDUCTIVITY AND↗