Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 127 records · Page 7

High-Temperature Gas Sensor Materials with Properties Predicted via First-Principles Calculations with Machine Learning Modeling and Experimental Corroboration

Understanding the temperature dependence of functional properties of sensing materials is vital for their applications in combustion environments. The electron-phonon coupling that derives the electronic structure change with temperatures is a key property of interest as it affects other sensing responses. Herein, we first assess the temperature dependence of band gap renormalization in sensing materials by employing Allen-Heine-Cardona (AHC) theory with density functional theory (DFT) simulations corroborated with experimental observation. As the AHC calculations are impractical for high-throughput screening of materials, we employ data-driven Gaussian process regression to predict the parameters employed in the O’Donnell empirical model from a set of physical features. To mitigate the reliability issues arising from the small size of the dataset, we apply a Bayesian technique to improve the generalizability of the data-driven models as well as to quantify the uncertainty associated with theoretical predictions. These models capture well the overall trend of the O’Donnell parameters with respect to a reduced feature set obtained by transforming the available physical features. Quantifying the associated uncertainty helps us understand the reliability of the predictions and, therefore, the variation of bandgap as a function of temperature for other novel materials. The predicted candidates from machine learning models are further validated by experiments and DFT calculations.

bandgap renormalization↗

Development of advanced machine learning models for analysis of plutonium surrogate optical emission spectra

This work investigates and applies machine learning paradigms seldom seen in analytical spectroscopy for quantification of gallium in cerium matrices via processing of laser-plasma spectra. Ensemble regressions, support vector machine regressions, Gaussian kernel regressions, and artificial neural network techniques are trained and tested on cerium-gallium pellet spectra. A thorough hyperparameter optimization experiment is conducted initially to determine the best design features for each model. The optimized models are evaluated for sensitivity and precision using the limit of detection (LoD) and root mean-squared error of prediction (RMSEP) metrics, respectively. Gaussian kernel regression yields the superlative predictive model with an RMSEP of 0.33% and an LoD of 0.015% for quantification of Ga in a Ce matrix. This study concludes that these machine learning methods could yield robust prediction models for rapid quality control analysis of plutonium alloys.

Rao, Ashwin P. (ORCID:0000000319312568)↗

Voltage Mining for (De)lithiation-Stabilized Cathodes and a Machine Learning Model for Li-Ion Cathode Voltage

Advances in lithium-metal anodes have inspired interest in discovery of Li-free cathodes, most of which are natively found in their charged state. This is in contrast to today's commercial lithium-ion battery cathodes, which are more stable in their discharged state. In this study, we combine calculated cathode voltage information from both categories of cathode materials, covering 5577 and 2423 total unique structure pairs, respectively. The resulting voltage distributions with respect to the redox pairs and anion types for both classes of compounds emphasize design principles for high-voltage cathodes, which favor later Period 4 transition metals in their higher oxidation states and more electronegative anions like fluorine or polyanion groups. Generally, cathodes that are found in their charged, delithiated state are shown to exhibit voltages lower than those that are most stable in their lithiated state, in agreement with thermodynamic expectations. Deviations from this trend are found to originate from different anion distributions between redox pairs. In addition, a machine learning model for voltage prediction based on chemical formulas is trained and shows state-of-the-art performance when compared to two established composition-based ML models for material properties predictions, Roost and CrabNet.

25 ENERGY STORAGE↗

Pd–Methyl Bond Energy─Property Correlations, Noncorrelations, Machine Learning Models, and Application to Polymerization Catalysis

Metal–carbon bonds are a key intermediate in a variety of homogeneous organometallic transformations and often determine the critical thermodynamics and kinetics of catalytic processes. Surprisingly, the influence of different ligands on metal–carbon bond strengths has been largely overlooked. Here, in this study, we evaluated nearly 700 experimental Pd–methyl complexes by calculating their bond dissociation energies using density functional theory (DFT) and compared these bond strengths to several fundamental molecular properties, and this revealed several surprising correlations and noncorrelations. Most surprising was that several fundamental properties, such as the bond length, bond force constant, and bond electron density, have no correlation with bond strength, despite these correlations often holding for main-group compounds. We were indeed able to identify key ligand-dependent chemical features/descriptors that provided a highly accurate machine learning model and provided insight into the general factors that control the Pd–carbon bond strength, such as radical delocalization and nucleophilicity. Insights gained from the Pd–Me bond energy analysis were then applied to CO migratory insertion steps that are part of copolymerization reactions.

binding energy↗

Best of both worlds: Enforcing detailed balance in machine learning models of transition rates

The slow microstructural evolution of materials often plays a key role in determining material properties. When the unit steps of the evolution process are slow, direct simulation approaches such as molecular dynamics become prohibitive and Kinetic Monte-Carlo (kMC) algorithms, where the state-to-state evolution of the system is represented in terms of a continuous-time Markov chain, are instead frequently relied upon to efficiently predict long-time evolution. The accuracy of kMC simulations however relies on the complete and accurate knowledge of reaction pathways and corresponding kinetics. This requirement becomes extremely stringent in complex systems such as concentrated alloys where the astronomical number of local atomic configurations makes the a priori tabulation of all possible transitions impractical. Machine learning models of transition kinetics have been used to mitigate this problem by enabling the efficient on-the-fly prediction of kinetic parameters. While conventional KMC methods based on transition state theory naturally yield reversible dynamics that exactly obey the detailed balance criterion, providing strong guarantees on the properties of the stationary distribution, many recently-proposed ML-based approaches to barrier predictions provide no such guarantees. In this study, we derive conditions under which physics-informed ML architectures exactly enforce the detailed balance condition by construction, even when relying on non-extensive descriptions of states in terms of local environments around mobile defects. In conclusion, using the diffusion of a vacancy in a concentrated alloy as an example, we show that such ML architectures also exhibit superior performance in terms of prediction accuracy, demonstrating that the imposition of physical constraints can facilitate the accurate learning of barriers at no increase in computational cost.

36 MATERIALS SCIENCE↗

Optimized Machine Learning Model for Predicting Groundwater Contamination

The use of physical models to predict groundwater contaminant movement remains technically challenging due to the complexity of the phenomena, the heterogeneity of key parameters in nature, and the presence of poorly defined interactive and feedback processes. New approaches to address these challenges are needed. In this study, we evaluate various Artificial Intelligence (AI)-based approaches to understand a hexavalent chromium (Cr(VI)) plumes located on the U.S. Department of Energy’s (DOE) Hanford Site in Richland, WA. The groundwater monitoring dataset used in this study included data from the 100 Area along the Columbia River and included data collected between 2010 to 2019. This study investigates the most prominent contaminant, Cr(VI), with the Extreme Gradient Boosting (XGBoost) machine learning model. The XGBoost models were compared with optimized versions using an Empirical Bayes Search Cross-Validation technique for better prediction. The optimized XGBoost model yielded an R^2 value of 0.99 on the training set and 0.85 on the testing set, whereas XGBoost without optimization yielded a value of 0.83 on the training set and 0.85 on the testing set. This paper provides an overview of a computational method for groundwater contamination modeling that shows promise for improving current remediation efforts.

Mazumdar, Hirak↗

Machine learning models for maintenance cost estimation in delivery trucks using diesel and natural gas fuels

The maintenance costs can represent about 15%–60% of the cost of produced goods depending on the type of goods transported. To comply with stringent emissions regulations, diesel engines are incorporated with complex after-treatment systems that demand increased maintenance. The availability of alternative fuels such as natural gas and propane has fostered the natural gas and propane powertrain systems as well as electrification options for heavy- and medium-duty vehicles. A critical barrier to adopting alternative fuel vehicles has been the lack of knowledge on comparative vehicle maintenance/repair costs with conventional diesel. Moreover, the region of operation, the type of vehicle operation, and seasonal temperature changes also affect the duty cycle which impacts the maintenance and repair costs. This study focuses on estimating the cost-per-mile for heavy-duty vehicles using machine learning models such as random forest, xgboost, neural networks, and a super-learner model. The super-learner model achieved an error as low as 0.0068 $/mile for mean absolute error and 0.0086 $/mile for root mean square error with a coefficient of determination/R-Squared of 97.28%. Specifically, the paper investigates the data collected from the maintenance and repair costs associated with delivery trucks using diesel and natural gas fuels. Since the availability of data is the major constraint, we leveraged the data collected by West Virginia University and the partnership with fleet companies. This allows for additional information related to maintenance costs and fleet-specific maintenance practices of alternative fuel vehicles. This study promotes clean fuel technologies and enables fleet management companies to adopt alternative fuel vehicles in case of similar or lower cost of maintenance compared to diesel vehicles resulting in reduced emissions and total cost of ownership.

Katreddi, Sasanka↗

Machine Learning Modeling for Accelerated Battery Materials Design in the Small Data Regime

Abstract Machine learning (ML)‐based approaches to battery design are relatively new but demonstrate significant promise for accelerating the timeline for new materials discovery, process optimization, and cell lifetime prediction. Battery modeling represents an interesting and unconventional application area for ML, as datasets are often small but some degree of physical understanding of the underlying processes may exist. This review article provides discussion and analysis of several important and increasingly common questions: how ML‐based battery modeling works, how much data are required, how to judge model performance, and recommendations for building models in the small data regime. This article begins with an introduction to ML in general, highlighting several important concepts for small data applications. Previous ionic conductivity modeling efforts are discussed in depth as a case study to illustrate these modeling concepts. Finally, an overview of modeling efforts in major areas of battery design is provided and several areas for promising future efforts are identified, within the context of typical small data constraints.

Sendek, Austin D.↗

A machine learning model for predicting the minimum miscibility pressure of CO 2 and crude oil system based on a support vector machine algorithm approach

CO 2 enhanced oil recovery (EOR) is a potential way for carbon capture, utilization and storage (CCUS). Though, the effect of CO 2 injection is greatly influenced by the reservoir conditions. Typically, Minimum miscible pressure (MMP) is selected as one of the key parameters for the screening and evaluation of prospective CO 2 flooding. Conventional slim tube test is both accurate and widely accepted but it is inefficient. Existing empirical formulas for MMPs are easy to be used but have been proved inaccurate and unreliable. Machine learning-based methods have great advantages in predicting MMP. However, only predication accuracy is discussed for most models without the screening of the main control factors and further validation of the model reliability. In this paper, a new prediction model based on support vector machine (SVM) was developed for pure/impure CO 2 and crude oil system. This study was based on 147 sets of MMP data from the literature with full information on reservoir temperature, oil composition and gas composition. The main control factors were screened by several statistical methods. Unlike the conventional prediction models that verified by only prediction accuracy, learning curve and single factor control variable analysis are further validated to obtain the optimum model.

02 PETROLEUM↗

A Universal Machine Learning Model for Elemental Grain Boundary Energies

The grain boundary (GB) energy has a profound influence on the grain growth and properties of polycrystalline metals. Here, we show that the energy of a GB, normalized by the bulk cohesive energy, can be described purely by four geometric features. By machine learning on a large computed database of 361 small Σ (Σ<10) GBs of more than 50 metals, we develop a model that can predict the grain boundary energies to within a mean absolute error of 0.13 J m –2 . More importantly, this universal GB energy model can be extrapolated to the energies of high Σ GBs without loss in accuracy. These results highlight the importance of capturing fundamental scaling physics and domain knowledge in the design of interpretable, extrapolatable machine learning models for materials science.

36 MATERIALS SCIENCE↗

Micromechanical Surrogate Machine Learning Model for Creep Deformation Modeling

Process variability during the manufacture of gas turbine engine hot section components can significantly affect the material’s resulting microstructure. In casting, for instance, geometric variation within a component (thin sections versus thick sections, radial location) influences cooling rates and the resulting grain size. The high temperature creep response is known to be sensitive to grain size owing to a diffusional creep mechanism which occurs more readily along grain boundaries. Microstructural variation correspondingly drives mechanical behavior which propagates into component scale performance uncertainty. These factors are essential when planning inspection, maintenance, and repair strategies within a reliability framework. These benefits provide opportunities to increase overall energy efficiency through refined margins. Critically, there is an opportunity to bolster existing data-driven reliability models using physics-driven process-structure-property relations. Here we present recent work establishing a framework for evaluating the probabilistic creep performance of high-temperature materials. A novel microstructure-sensitive crystal plasticity finite element model is established that captures both grain boundary and crystallographic deformation effects. The computationally expensive physics model is calibrated using a statistical approach and this high-fidelity model is subsequently used to train a computationally efficient machine learning surrogate model. The surrogate model is essential for sampling a large ensemble of simulated structure-property pair results. The ensemble data are then mined to extract salient trends to be incorporated into a microstructure-sensitive reliability model. The proposed approach represents a novel way to capture microstructure-sensitive trends from physics-based models within a modern reliability framework.

Fernandez-Zelaia, Patxi [ORNL]↗

Evolving Multi-hazard Machine Learning Modeling for Advanced Risk-Informed Infrastructure Resilience Assessment

The socioeconomic impacts of pipeline incidents have escalated over the past three decades, revealing the limitation of traditional risk modeling methods when applied to extensive pipeline networks. This research aims to develop machine learning (ML) models that effectively identify, rank, and predict the diverse hazards and socioeconomic consequences associated with pipeline incidents. Utilizing historical data on pipeline incidents alongside weather and oceanographic data from the 1980s onward, the Houston metropolitan area serves as a testbed for the proposed methodologies. The research segments the combined datasets into three consecutive periods, demonstrating the efficacy of the updated model in predicting future events, particularly concerning precipitation rate data. Despite the challenges posed by a relatively limited dataset, local-level ML modeling offers valuable insights into the spatial and temporal dynamics of multiple hazards that contribute to pipeline incidents. These findings hold significant implications for future research, particularly in understanding and mitigating risks in various locations across the Gulf Coast and other coastal regions.

42 ENGINEERING↗

The Development and Deployment of Machine Learning Models for Aircraft Engine Concept Assessment

In today's competitive landscape, the effective development and utilization of machine-learning (ML) applications have become crucial across diverse economic sectors. This study presents an outline of the procedure involved in creating and implementing ML models for conceptualizing and evaluating aircraft engines. These models leverage supervised deep-learning algorithms to analyze patterns within an open-source repository containing data on both production and research conventional turbofan engines. The main areas of focus encompass crucial engine parameters like thrust-specific fuel consumption (TSFC), engine weight, engine diameter, and turbomachinery stage counts. While the creation of ML models is fundamental for their utilization, ensuring their seamless deployment holds equal significance. To address this aspect, a conversational AI chatbot that specifically focuses on propulsion has been developed. Leveraging natural language processing (NLP) techniques, this chatbot simplifies the deployment of machine learning (ML) models. The comprehensive workflow encompasses several key stages: gathering and enhancing engine data, training and cross validating the ML models, testing and evaluating their performance, and finally, deploying, monitoring, and updating the ML models. By following this systematic approach, the aim is to streamline the development and deployment process of ML models tailored for aircraft engine assessment.

AI Chatbot↗

Uncertainty Quantification for Data-Driven Machine Learning Models in Nuclear Engineering Applications: Where We Are and What Do We Need?

Machine learning (ML) has been leveraged to tackle a diverse range of tasks in almost all branches of nuclear engineering. Many of the successes in ML applications can be attributed to the recent performance breakthroughs in deep learning, the growing availability of computational power, data, and easy-to-use ML libraries. However, these empirical successes have often outpaced our formal understanding of the ML algorithms. An important but under-rated area is uncertainty quantification (UQ) of ML. ML-based models are subject to approximation uncertainty when they are used to make predictions, due to sources including but not limited to, data noise, data coverage, extrapolation, imperfect model architecture and the stochastic training process. The goal of this paper is to clearly explain and illustrate the importance of UQ of ML. We will elucidate the differences in the basic concepts of UQ of physics-based models and data-driven ML models. Various sources of uncertainties in physical modeling and data-driven modeling will be discussed, demonstrated, and compared. We will also present and demonstrate a few techniques to quantify the ML prediction uncertainties, including Monte Carlo dropout, deep ensemble, Bayesian neural networks, Gaussian Processes and conformal prediction. Lastly, we will discuss the need for building a verification, validation and UQ framework to establish ML credibility.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

River Dissolved Oxygen Prediction Using Machine Learning Models and Wireless Sensor Measurements

Simultaneous flooding&heat and droughts&heat events can potentially destabilize hydro-meteorological conditions to deteriorate the water quality of Neches River. Machine learning (ML) models utilizing wireless sensor measurements have been applied to predict water quality and optimize various water management strategies. This study aims to develop ML models to predict dissolved oxygen (DO) prediction under various hydro-meteorological conditions and enhance water management decision-making. Wireless sensor measurements of DO, water temperature, sample depth, conductivity, turbidity, and pH, along with discharge from the United States Geological Survey stations, are collected for model inputs at the Pine Island Bayou C749 station (PIB-C749) and Neches River Saltwater Barrier (SWB). Multilayer perceptron neural networks, recurrent neural networks, long short-term memory (LSTM), and bidirectional LSTM (BiLSTM) with and without attention mechanism (AT) are tested to determine the best model, which is applied the rolling forecast method to predict 14-day DO. Traditional and recurrent transfer learning (TL and RTL) methods are adopted to overcome insufficient data at the SWB. The input feature importance analysis using the integrated gradients (IG) algorithm is applied to determine dominant inputs. The results show LSTM-based models are capable handling long sequential data. AT-BiLSTM and RTL-LSTM demonstrate the best performance at the PIB-C749 (RMSE=0.054) and the SWB (RMSE=0.028), respectively. TL and RTL methods significantly improve model performance at the SWB. DO, temperature, and pH show higher importance, consistent with hydrodynamics and water chemistry. Both best models are applied to predict 14-day DO and demonstrate reasonable performance for decision-making. Hydro-meteorological conditions of 2017 flood and 2012 drought events are simulated and reveal that possible hypoxia occurs after flooding due to increasing temperature and turbidity, and DO concentration decreases significantly under heat and drought conditions. In conclusion, LSTM-based models utilizing wireless sensor data can be a timely and effective approach to make appropriate decisions on water resource management.

54 ENVIRONMENTAL SCIENCES↗

Machine Learning Models to Predict Cognitive Impairment of Rodents Subjected to Space Radiation

This research uses machine-learned computational analyses to predict the cognitive performance impairment of rats induced by irradiation. The experimental data in the analyses is from a rodent model exposed to ≤ 15 cGy of individual Galactic Cosmic Radiation (GCR) ions: 4He, 16O, 28Si, 48Ti, or 56Fe, expected for a Lunar or Mars mission. This work investigates rats at a subject-based level and uses performance scores taken before irradiation to predict impairment in Attentional Set-shifting (ATSET) data post-irradiation. Here, the worst performing rats of the control group define the impairment thresholds based on population analyses via cumulative distribution functions, leading to the labeling of impairment for each subject. A significant finding is the exhibition of a dose-dependent increasing probability of impairment for 1 to 10 cGy of 28Si or 56Fe in the Simple Discrimination (SD) stage of the ATSET, and for 1 to 10 cGy of 56Fe in the Compound Discrimination (CD) stage. On a subject-based level, implementing Machine Learning (ML) classifiers such as the Gaussian Naïve Bayes, Support Vector Machine, and Artificial Neural Networks identifies rats that have a higher tendency for impairment after GCR exposure. The algorithms employ the experimental prescreenperformance scores as multidimensional input features to predict each rodent’s susceptibility to cognitive impairment due to space radiation exposure. The receiver operating characteristic and the precision-recall curves of the ML models show a better prediction of impairment when 56Feis the ion in question in both SD and CD stages. They, however, do not depict impairment due to 4Hein SD and 28Siin CD, suggesting no dose-dependent impairment response in these cases. One key finding of our study is that prescreen performance scores can be used to predict the ATSET performance impairments. This result is significant to crewed space missions as it supports the potential of predicting an astronaut’s impairment in a specific task before spaceflight through the implementation of appropriately trained ML tools. Future research can focus on constructing ML ensemble methods to integrate the findings from the methodologies implemented in this study for morerobust predictionsof cognitive decrements due to space radiation exposure.

space radiation↗

Predicting Intensive Care Unit Length of Stay and Mortality Using Patient Vital Signs: Machine Learning Model Development and Validation

Background: Patient monitoring is vital in all stages of care. In particular, intensive care unit (ICU) patient monitoring has the potential to reduce complications and morbidity, and to increase the quality of care by enabling hospitals to deliver higher-quality, cost-effective patient care, and improve the quality of medical services in the ICU. Objective: We here report the development and validation of ICU length of stay and mortality prediction models. The models will be used in an intelligent ICU patient monitoring module of an Intelligent Remote Patient Monitoring (IRPM) framework that monitors the health status of patients, and generates timely alerts, maneuver guidance, or reports when adverse medical conditions are predicted. Methods: We utilized the publicly available Medical Information Mart for Intensive Care (MIMIC) database to extract ICU stay data for adult patients to build two prediction models: one for mortality prediction and another for ICU length of stay. For the mortality model, we applied six commonly used machine learning (ML) binary classification algorithms for predicting the discharge status (survived or not). For the length of stay model, we applied the same six ML algorithms for binary classification using the median patient population ICU stay of 2.64 days. For the regression-based classification, we used two ML algorithms for predicting the number of days. We built two variations of each prediction model: one using 12 baseline demographic and vital sign features, and the other based on our proposed quantiles approach, in which we use 21 extra features engineered from the baseline vital sign features, including their modified means, standard deviations, and quantile percentages. Results: We could perform predictive modeling with minimal features while maintaining reasonable performance using the quantiles approach. The best accuracy achieved in the mortality model was approximately 89% using the random forest algorithm. The highest accuracy achieved in the length of stay model, based on the population median ICU stay (2.64 days), was approximately 65% using the random forest algorithm. Conclusions: The novelty in our approach is that we built models to predict ICU length of stay and mortality with reasonable accuracy based on a combination of ML and the quantiles approach that utilizes only vital signs available from the patient’s profile without the need to use any external features. This approach is based on feature engineering of the vital signs by including their modified means, standard deviations, and quantile percentages of the original features, which provided a richer dataset to achieve better predictive power in our models.

59 BASIC BIOLOGICAL SCIENCES↗