Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Evaluating county-level lung cancer incidence from environmental radiation exposure, PM 2.5 , and other exposures with regression and machine learning models

Characterizing the interplay between exposures shaping the human exposome is vital for uncovering the etiology of complex diseases. For example, cancer risk is modified by a range of multifactorial external environmental exposures. Environmental, socioeconomic, and lifestyle factors all shape lung cancer risk. However, epidemiological studies of radon aimed at identifying populations at high risk for lung cancer often fail to consider multiple exposures simultaneously. For example, moderating factors, such as PM 2.5 , may affect the transport of radon progeny to lung tissue. This ecological analysis leveraged a population-level dataset from the National Cancer Institute’s Surveillance, Epidemiology, and End-Results data (2013–17) to simultaneously investigate the effect of multiple sources of low-dose radiation (gross γ activity and indoor radon) and PM 2.5 on lung cancer incidence rates in the USA. County-level factors (environmental, sociodemographic, lifestyle) were controlled for, and Poisson regression and random forest models were used to assess the association between radon exposure and lung and bronchus cancer incidence rates. Tree-based machine learning (ML) method perform better than traditional regression: Poisson regression: 6.29/7.13 (mean absolute percentage error, MAPE), 12.70/12.77 (root mean square error, RMSE); Poisson random forest regression: 1.22/1.16 (MAPE), 8.01/8.15 (RMSE). The effect of PM 2.5 increased with the concentration of environmental radon, thereby confirming findings from previous studies that investigated the possible synergistic effect of radon and PM 2.5 on health outcomes. In summary, the results demonstrated (1) a need to consider multiple environmental exposures when assessing radon exposure’s association with lung cancer risk, thereby highlighting (1) the importance of an exposomics framework and (2) that employing ML models may capture the complex interplay between environmental exposures and health, as in the case of indoor radon exposure and lung cancer incidence.

63 RADIATION, THERMAL, AND OTHER ENVIRON. POLLUTAN↗

Machine learning models inaccurately predict current and future high-latitude C balances

The high-latitude carbon (C) cycle is a key feedback to the global climate system, yet because of system complexity and data limitations, there is currently disagreement over whether the region is a source or sink of C. Recent advances in big data analytics and computing power have popularized the use of machine learning (ML) algorithms to upscale site measurements of ecosystem processes, and in some cases forecast the response of these processes to climate change. Due to data limitations, however, ML model predictions of these processes are almost never validated with independent datasets. To better understand and characterize the limitations of these methods, we develop an approach to independently evaluate ML upscaling and forecasting. We mimic data-driven upscaling and forecasting efforts by applying ML algorithms to different subsets of regional process-model simulation gridcells, and then test ML performance using the remaining gridcells. In this study, we simulate C fluxes and environmental data across Alaska using ecosys, a process-rich terrestrial ecosystem model, and then apply boosted regression tree ML algorithms to training data configurations that mirror and expand upon existing AmeriFLUX eddy-covariance data availability. We first show that a ML model trained using ecosys outputs from currently-available Alaska AmeriFLUX sites incorrectly predicts that Alaska is presently a modeled net C source. Increased spatial coverage of the training dataset improves ML predictions, halving the bias when 240 modeled sites are used instead of 15. However, even this more accurate ML model incorrectly predicts Alaska C fluxes under 21st century climate change because of changes in atmospheric CO 2 , litter inputs, and vegetation composition that have impacts on C fluxes which cannot be inferred from the training data. Our results provide key insights to future C flux upscaling efforts and expose the potential for inaccurate ML upscaling and forecasting of high-latitude C cycle dynamics.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Models for Predicting Molecular UV–Vis Spectra with Quantum Mechanical Properties

Accurate understanding of Ultraviolet–visible (UV–Vis) spectra is critical for highthroughput design of compounds for drug discovery. Experimentally determining UV–Vis spectra can become expensive when dealing with a large quantity of novel molecules. This provides us an opportunity to drive computational advances in molecular property predictions using quantum mechanics and machine learning. In this work, we use both Quantum Mechanically (QM) predicted and measured UV–Vis spectra as input to modify four different machine learning architectures: UVvis-SchNet, UVvis- DTNN, UVvis-Transformer, and UVvis-MPNN. Here we find that the UVvis-MPNN model outperforms the other models when using optimized 3D coordinates and QM predicted spectra as input features. This model has the highest performance for predicting UVVisible spectra with a training RMSE of 0.06 and validation RMSE of 0.08.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Hierarchical Testing of a Hybrid Machine Learning‐Physics Global Atmosphere Model

Machine learning (ML)-based models have demonstrated high skill and computational efficiency, often outperforming conventional physics-based models in weather and subseasonal predictions. While prior studies have assessed their fidelity in capturing synoptic-scale atmospheric dynamics, their performance across timescales and under out-of-distribution forcing, such as +3K or +4K uniform-warming forcings, and the sources of biases remain elusive, to establish the model's reliability for Earth science. Here, we design three sets of experiments targeting synoptic-scale phenomena, interannual variability, and out-of-distribution uniform-warming forcings. We evaluate the Neural General Circulation Model (NeuralGCM), a hybrid model integrating a dynamical core with ML-based component, against observations and physics-based Earth system models (ESMs). At the synoptic scale, NeuralGCM captures the evolution and propagation of extratropical cyclones with performance comparable to ESMs. At the interannual scale, when forced by El Niño-Southern Oscillation sea surface temperature (SST) anomalies, NeuralGCM successfully reproduces associated teleconnection patterns but exhibits deficiencies in capturing nonlinear response. Under out-of-distribution uniform-warming forcings, NeuralGCM simulates similar responses in global-average temperature and precipitation and reproduces large-scale tropospheric circulation features similar to those in ESMs. Notable weaknesses include overestimating the tracks and spatial extent of extratropical cyclones, biases in the teleconnected wave train triggered by tropical SST anomalies, and differences in upper-level warming and stratospheric circulation responses to SST warming compared to physics-based ESMs. The causes of these weaknesses were explored. Despite the noted weaknesses, NeuralGCM reproduces responses across experiments reasonably and performs comparably to ESMs. By integrating a dynamical core with ML, NeuralGCM shows potential for developing ML-based ESMs.

global warming↗

Hierarchical transfer learning: an agile and equitable strategy for machine-learning interatomic models

Machine-learned interatomic models are growing in popularity due to their ability to afford near quantum-accurate predictions for complex phenomena with orders-of-magnitude greater computational efficiency. However, these models struggle when applied to systems of many element types due to the approximately exponential increase in number of parameters that must be determined. To mitigate this challenge, we present a new hierarchical transfer learning approach that allows the fitting problem to be decomposed into smaller independent and reusable parameter blocks that enable development of explicitly chemically extensible ML-IAM. Application of this strategy is demonstrated for C and N mixtures under conditions ranging from nominally ambient to ~10,000 K and 200 GPa for compositions from 0 to 100% N. Ultimately, this strategy makes model generation for chemically complex systems more tractable and efficient, facilitates comprehensive model validation, and makes ML-IAM development for problems of this nature more accessible to users with limited access to extreme computing infrastructure.

Lindsey, Rebecca K. [Univ. of Michigan, Ann Arbor,↗

Evaluation of Machine Learning Models for Automated Data Analysis in In-Service Nuclear Power Plant Inspections

The commercial nuclear power industry is facing a potential shortage of certified nondestructive evaluation (NDE) analysts to meet future in-service inspection demands. Automated data analysis (ADA) currently supports human inspectors in tasks such as eddy current evaluations for steam generator examinations. Machine learning (ML) systems are nearing the capability to pass performance demonstration tests for ultrasonic testing (UT) inspections of reactor pressure vessel upper head penetrations in nuclear power plants (NPPs). Current research and development is focused on assisted analysis (AA) of ADA versus fully automated examinations. This presentation will cover assessment of ML flaw detection on dissimilar metal weld (DMW) piping joints.

36 MATERIALS SCIENCE↗

Prognostic analysis of high-flow nasal cannula therapy and non-invasive ventilation in mild to moderate hypoxemia patients and construction of a machine learning model for 48-h intubation prediction—a retrospective analysis of the MIMIC database

Background This study aims to investigate the clinical outcome between high-flow nasal cannula (HFNC) and non-invasive ventilation (NIV) therapy in mild to moderate hypoxemic patients on the first ICU day and to develop a predictive model of 48-h intubation. Methods The study included adult patients from the MIMIC III and IV databases who first initiated HFNC or NIV therapy due to mild to moderate hypoxemia (100 < PaO2/FiO2 ≤ 300). The 48-h and 30-day intubation rates were compared using cross-sectional and survival analysis. Nine machine learning and six ensemble algorithms were deployed to construct the 48-h intubation predictive models, of which the optimal model was determined by its prediction accuracy. The top 10 risk and protective factors were identified using the Shapley interpretation algorithm. Result A total of 123,042 patients were screened, of which, 673 were from the MIMIC IV database for ventilation therapy comparison (HFNC n = 363, NIV n = 310) and 48-h intubation predictive model construction (training dataset n = 471, internal validation set n = 202) and 408 were from the MIMIC III database for external validation. The NIV group had a lower intubation rate (23.1% vs. 16.1%, p = 0.001), ICU 28-day mortality (18.5% vs. 11.6%, p = 0.014), and in-hospital mortality (19.6% vs. 11.9%, p = 0.007) compared to the HFNC group. Survival analysis showed that the total and 48-h intubation rates were not significantly different. The ensemble AdaBoost decision tree model (internal and external validation set AUROC 0.878, 0.726) had the best predictive accuracy performance. The model Shapley algorithm showed Sequential Organ Failure Assessment (SOFA), acute physiology scores (APSIII), the minimum and maximum lactate value as risk factors for early failure and age, the maximum PaCO 2 and PH value, Glasgow Coma Scale (GCS), the minimum PaO 2 /FiO 2 ratio, and PaO 2 value as protective factors. Conclusion NIV was associated with lower intubation rate and ICU 28-day and in-hospital mortality. Further survival analysis reinforced that the effect of NIV on the intubation rate might partly be attributed to the other impact factors. The ensemble AdaBoost decision tree model may assist clinicians in making clinical decisions, and early organ function support to improve patients’ SOFA, APSIII, GCS, PaCO 2 , PaO 2 , PH, PaO 2 /FiO 2 ratio, and lactate values can reduce the early failure rate and improve patient prognosis.

Fu, Wei↗

Performance Comparison of Machine Learning Models for Ultrasonic Nondestructive Evaluation of Alkali-Silica Reaction in Concrete

Alkali-silica reaction (ASR) causes concrete degradation, leading to cracking, rebar corrosion, and reduced structural integrity, which raises safety concerns. Ultrasonic nondestructive evaluation (NDE) effectively assesses concrete properties and monitors ASR progression. However, its deployment and analysis require specialized expertise and subjective interpretation. As computational power increases, artificial intelligence (AI) and machine learning (ML) algorithms are increasingly being used to automate NDE data analysis across various industries for AI-assisted automation. Regulatory agencies are adapting to this technological shift, prompting a need to evaluate current ML technologies’ capabilities and limitations in assessing concrete material properties and damage. This report presents a comparative analysis of four ML regression models for predicting concrete material damage induced by ASR expansion using long-term ultrasonic data monitoring. The models investigated include linear regression (LR), support vector regression (SVR), shallow neural networks (NN), and deep neural networks (DNN). LR, SVR, and shallow NN models use features extracted from ultrasonic signals, whereas the DNN model processes time-domain ultrasonic signals and frequency spectra directly. The study systematically compared the models’ performance from various perspectives, including model input, prediction performance, and generalization ability. The findings indicate significant variability in model performance, with some ML algorithms achieving very high or very low prediction accuracy depending on the preprocessing and feature engineering (extraction and selection) applied. Key insights include the observation that shallow ML models (LR, SVR, and shallow NNs) require meticulous preprocessing and feature extraction to achieve high accuracy. In contrast, the DNN model, although it bypasses the need for feature engineering, necessitates extensive preprocessing to mitigate noise and computational demands. The SVR model emerged as the top performer among the shallow models, and the DNN model exhibited superior performance on specific datasets but struggled with generalization across specimens from different batches. Additionally, the SVR model is sensitive to temperature variations, whereas the DNN model is robust in this regard. Using recurrent neural networks is recommended for future ASR expansion prediction studies. Recurrent neural networks’ inherent ability to capture temporal dependencies and long-term patterns makes them well suited for analyzing sequential ultrasonic monitoring data. Overall, the results and conclusions of this study could provide insights into the capabilities and effectiveness of ML when applied to ultrasonic NDE data and help identify best practices for using ML for ultrasonic NDE of concrete material properties.

36 MATERIALS SCIENCE↗

Using diverse potentials and scoring functions for the development of improved machine-learned models for protein–ligand affinity and docking pose prediction

The advent of computational drug discovery holds the promise of significantly reducing the effort of experimentalists, along with monetary cost. More generally, predicting the binding of small organic molecules to biological macromolecules has far-reaching implications for a range of problems, including metabolomics. However, problems such as predicting the bound structure of a protein–ligand complex along with its affinity have proven to be an enormous challenge. In recent years, machine learning-based methods have proven to be more accurate than older methods, many based on simple linear regression. Nonetheless, there remains room for improvement, as these methods are often trained on a small set of features, with a single functional form for any given physical effect, and often with little mention of the rationale behind choosing one functional form over another. Moreover, it is not entirely clear why one machine learning method is favored over another. Here, we endeavor to undertake a comprehensive effort towards developing high-accuracy, machine-learned scoring functions, systematically investigating the effects of machine learning method and choice of features, and, when possible, providing insights into the relevant physics using methods that assess feature importance. Here, we show synergism among disparate features, yielding adjusted R 2 with experimental binding affinities of up to 0.871 on an independent test set and enrichment for native bound structures of up to 0.913. When purely physical terms that model enthalpic and entropic effects are used in the training, we use feature importance assessments to probe the relevant physics and hopefully guide future investigators working on this and other computational chemistry problems.

59 BASIC BIOLOGICAL SCIENCES↗

Interpretable Machine Learning Models for Autonomous Characterization of Analogue Ocean World Seawater Chemistry and Biosignature Potential Using Isotope Ratio Data

Background: Future missions to ocean worlds, such as Enceladus and Europa, will attempt to characterize the subsurface seawater chemistry and assess the potential for life. Such missions will be equipped with capabilities to precisely measure volatile isotopes in plumes, atmospheres, and exospheres. Motivation: While large isotopic fractionations can indicate a biological source, there are signatures resulting from abiotic geochemical processes that mimic isotopic biosignatures. While machine learning (ML) has the potential to disentangle competing effects and biotic mimicry, high-dimensional isotope ratio mass spectrometry (IRMS) data is likely to contain noise/irrelevant features and involve complex statistical interactions that make human inference and interpretation difficult. Further, ML predictions with as far-reaching implications as an extraterrestrial biosignature on an ocean world requires the use of interpretable models (i.e., not “black box” models) with physically and mathematically meaningful feature spaces along with false positive diagnostics. Methods: We use volatile CO2 IRMS data of analogue ocean world seawaters to validate an ML approach to provide biogeochemical context for biosignature detection. We employ a feature selection method called nearest-neighbor projected distance regression (NPDR) that detects statistical interactions and helps elucidate the mechanisms of the Random Forest classification models. Results: We train and validate predictive ML models on volatile CO2 IRMS data of analogue ocean world seawaters to predict major salt components (e.g., MgSO4, NaHCO3), pH, ionic strength, and the presence of biosignatures. Features derived from IRMS measurements are augmented with extracted time-series features. Our results show high test accuracy and interpretability, which is increased by interaction network visualization, sample-wise variable importance scores, and single-sample class probability estimates. We demonstrate an ML mission software solution that triggers autonomous data transmission and biogeochemical sample prediction.

geochemistry↗

Machine Learning Modeling Pipeline for Extracting Nuclear Proliferation Events of Interest from Open Data Sources (U)

In FY2020, the Savannah River National Laboratory (SRNL) and the Sanghani Center for Artificial Intelligence and Data Analytics at Virginia Polytechnic Institute and State University entered a collaboration funded by Department of Energy’s (DOE) Office of Defense Nuclear Nonproliferation Research and Development. The project’s mission was to take the first steps toward developing a demonstration prototype system that uses multiple machine learning and data analytics methods on largescale open data sources to identify new, developing, and/or undeclared nuclear programs. Given the SRNL team’s on-site perspective of events culminating in the DOE’s decision to pursue the Savannah River Plutonium Processing Facility (SRPPF), the team targeted the identification of events and indicators in retrospective datasets that pointed to the activity of “fissile core fabrication at the Savannah River Site” prior to the official announcement in May of 2018. A preliminary modeling pipeline was developed in FY20 that showed the datasets contained adequate signal for continuation of efforts. In FY21, a modular demonstration prototype modeling pipeline has continued in development for two text-based data sources: a broad internet archive (Webhose Ltd.) and a decahose Twitter database (i.e., a global sampling of one in every ten Tweets). The techniques that have been developed rely on graph theory and anomaly detection to identify contextual shifts in key words and phrases at various points in time such that indicators of events of interest could be identified and subsequently, events could be extracted from the corpuses. The foundational concept behind the approaches is that contextual shifts in key words and phrases can act as indicators of events of interest. Both datasets have proven successful in extracting events of interest related to pit production at the Savannah River Site prior to the official announcement. In addition, the pipelines have generated a wide range of events broadly summarized as: the awarding of DOE contracts at major sites, DOE investments in various programs, accidents at DOE national laboratories, speculations about the fate of pit production in the DOE complex, domestic and international shipments and receipts of nuclear materials at DOE sites, termination of non-proliferation agreements with Russia, termination of MOX, new weapons development approvals/testing, nuclear posture reviews, major DOE cleanup/production milestones, political opinions, and nuclear watch groups’ opinions, among many others.

97 MATHEMATICS AND COMPUTING↗

Causality guided machine learning model on wetland CH 4 emissions across global wetlands

Wetland CH 4 emissions are among the most uncertain components of the global CH 4 budget. The complex nature of wetland CH 4 processes makes it challenging to identify causal relationships for improving our understanding and predictability of CH 4 emissions. In this study, we used the flux measurements of CH 4 from eddy covariance towers (30 sites from 4 wetlands types: bog, fen, marsh, and wet tundra) to construct a causality-constrained machine learning (ML) framework to explain the regulative factors and to capture CH 4 emissions at sub-seasonal scale. We found that soil temperature is the dominant factor for CH 4 emissions in all studied wetland types. Ecosystem respiration (CO 2 ) and gross primary productivity exert controls at bog, fen, and marsh sites with lagged responses of days to weeks. Integrating these asynchronous environmental and biological causal relationships in predictive models significantly improved model performance. More importantly, modeled CH 4 emissions differed by up to a factor of 4 under a +1°C warming scenario when causality constraints were considered. These results highlight the significant role of causality in modeling wetland CH 4 emissions especially under future warming conditions, while traditional data-driven ML models may reproduce observations for the wrong reasons. Our proposed causality-guided model could benefit predictive modeling, large-scale upscaling, data gap-filling, and surrogate modeling of wetland CH 4 emissions within earth system land models.

54 ENVIRONMENTAL SCIENCES↗

An exploration of machine learning models for the determination of reaction coordinates associated with conformational transitions

Determining collective variables (CVs) for conformational transitions is crucial to understanding their dynamics and targeting them in enhanced sampling simulations. Often, CVs are proposed based on intuition or prior knowledge of a system. However, the problem of systematically determining a proper reaction coordinate (RC) for a specific process in terms of a set of putative CVs can be achieved using committor analysis (CA). Identifying essential degrees of freedom that govern such transitions using CA remains elusive because of the high dimensionality of the conformational space. Various schemes exist to leverage the power of machine learning (ML) to extract an RC from CA. Here, we extend these studies and compare the ability of 17 different ML schemes to identify accurate RCs associated with conformational transitions. We tested these methods on an alanine dipeptide in vacuum and on a sarcosine dipeptoid in an implicit solvent. Our comparison revealed that the light gradient boosting machine method outperforms other methods. In order to extract key features from the models, we employed Shapley Additive exPlanations analysis and compared its interpretation with the “feature importance” approach. For the alanine dipeptide, our methodology identifies ϕ and θ dihedrals as essential degrees of freedom in the C7ax to C7eq transition. For the sarcosine dipeptoid system, the dihedrals ψ and ω are the most important for the cisαD to transαD transition. We further argue that analysis of the full dynamical pathway, and not just endpoint states, is essential for identifying key degrees of freedom governing transitions.

Chemistry↗

Accuracy of predictions made by machine learned models for biocrude yields obtained from hydrothermal liquefaction of organic wastes

Hydrothermal liquefaction (HTL) has potential for converting abundant wet organic wastes into renewable fuels. Because HTL consists of a complex reaction network, deterministic, physics-based prediction of its biocrude yield is prohibitively difficult. Data-driven methods provide an alternative to the physics-based approach; however, rigorous testing must be performed to ensure the accuracy of predictions made by data-driven methods. To this end, a data set was assembled consisting of 570 data points appearing in the open literature. The data set was divided into training, validation, and test sub-sets and used for evaluating different machine learning regression approaches to predict biocrude yield. Among the tested algorithms, Random Forest and eXtreme Gradient Boosting (XGBoost) predicted biocrude yields in a test set that had not been used for training with the greatest accuracy, with root mean square errors (RMSE) of 8.34 and 8.57, respectively. Further refinement of the Random Forest model reduced its RMSE to 8.07. In comparison, predictions of a series of literature models resulted in RMSE ranging from 9.16 in the most accurate case to 27.6 in the least accurate; most literature models yielded RMSE values > 10. Using biocrude yield predictions from the most accurate Random Forest model and a probabilistic economic analysis found that the model accuracy is sufficient to prioritize allocation of resources based on projected minimum fuel selling price. In our report the models and analysis represent a major advance in the ability to use readily available data to predict biocrude yields on new feedstocks that have not previously been studied.

42 ENGINEERING↗

Multivariate Machine Learning Models of Nanoscale Porosity from Ultrafast NMR Relaxometry

Abstract Nanoporous materials are of great interest in many applications, such as catalysis, separation, and energy storage. The performance of these materials is closely related to their pore sizes, which are inefficient to determine through the conventional measurement of gas adsorption isotherms. Nuclear magnetic resonance (NMR) relaxometry has emerged as a technique highly sensitive to porosity in such materials. Nonetheless, streamlined methods to estimate pore size from NMR relaxometry remain elusive. Previous attempts have been hindered by inverting a time domain signal to relaxation rate distribution, and dealing with resulting parameters that vary in number, location, and magnitude. Here we invoke well‐established machine learning techniques to directly correlate time domain signals to BET surface areas for a set of metal‐organic frameworks (MOFs) imbibed with solvent at varied concentrations. We employ this series of MOFs to establish a correlation between NMR signal and surface area via partial least squares (PLS), following screening with principal component analysis, and apply the PLS model to predict surface area of various nanoporous materials. This approach offers a high‐throughput, non‐destructive way to assess porosity in c.a. one minute. We anticipate this work will contribute to the development of new materials with optimized pore sizes for various applications.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Multiscale and Machine Learning Modeling for Additive Manufacturing

Additive manufacturing (AM) techniques provide the opportunity to simultaneously design new materials and components with complex structures in less time, enabling faster material developments. Even though compositionally similar, the texture of the materials produced by such techniques is significantly different from conventionally manufactured materials. Additively manufactured materials produces highly heterogeneous microstructure within a single build. Such variations in the microstructure make qualifying AM products challenging for extreme environment applications. Understanding the AM process and its influence on the materials’ microstructures/properties is paramount for evaluating the workability and performance of the manufactured materials. The performance of AM materials for advanced nuclear reactor applications is of interest to the Advanced Materials and Manufacturing Technologies (AMMT) program under the Department of Energy Office of Nuclear Energy. Hence, considering the microstructural variabilities in the AM products and their impact on the performance of the material, it is important to correlate the process conditions to the final product and establish a process-structure-property- performance (PSPP) correlation for AM materials.

11 NUCLEAR FUEL CYCLE AND FUEL MATERIALS↗