Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 145 records · Page 8

Medium Energy Electron Flux in Earth's Outer Radiation Belt (MERLIN): A Machine Learning Model

The radiation belts of the Earth, filled with energetic electrons, comprise complex and dynamic systems that pose a significant threat to satellite operation. While various models of electron flux both for low and relativistic energies have been developed, the behavior of medium energy (120–600 keV) electrons, especially in the MEO region, remains poorly quantified. At these energies, electrons are driven by both convective and diffusive transport, and their prediction usually requires sophisticated 4D modeling codes. In this paper, we present an alternative approach using the Light Gradient Boosting (LightGBM) machine learning algorithm. The Medium Energy electRon fLux In Earth's outer radiatioN belt (MERLIN) model takes as input the satellite position, a combination of geomagnetic indices and solar wind parameters including the time history of velocity, and does not use persistence. MERLIN is trained on >15 years of the GPS electron flux data and tested on more than 1.5 years of measurements. Tenfold cross validation yields that the model predicts the MEO radiation environment well, both in terms of dynamics and amplitudes o f flux. Evaluation on the test set shows high correlation between the predicted and observed electron flux (0.8) and low values of absolute error. The MERLIN model can have wide space weather applications, providing information for the scientific community in the form of radiation belts reconstructions, as well as industry for satellite mission design, nowcast of the MEO environment, and surface charging analysis.

79 ASTRONOMY AND ASTROPHYSICS↗

A database of ultrastable MOFs reassembled from stable fragments with machine learning models

High-throughput screening of hypothetical metal-organic framework (MOF) databases can uncover new materials, but their stability in real-world applications is often unknown. We leverage community knowledge and machine learning (ML) models to identify MOFs that are thermally stable and stable upon activation. We separate these MOFs into their building blocks and recombine them to make a new hypothetical MOF database of over 50,000 structures with orders of magnitude more (1) connectivity nets and (2) inorganic building blocks than were present in prior databases. Further, this database shows a 10-fold enrichment of ultrastable MOF structures that are stable upon activation and more than 1 standard deviation more thermally stable than the average experimentally characterized MOF. For nearly 10,000 ultrastable MOFs, we compute elastic moduli to confirm that these materials have good mechanical stability, and we report methane deliverable capacities. We identify privileged metal nodes in ultrastable MOFs that optimize gas storage and mechanical stability simultaneously.

36 MATERIALS SCIENCE↗

Machine learning models for estimating contamination across different curbside collection strategies

Contaminated recyclables, which are frequently discarded as waste, pose a significant challenge to the implementation of a circular economy. These contaminated recyclables impede the circulation of resources, resulting in higher processing costs at material recovery facilities (MRFs). Over the past few decades, machine learning (ML) models such as linear regression (LR), support vector machine (SVM), and random forest (RF) have evolved to provide new methods for predicting inbound contamination rates in addition to traditional statistical models. In this study, we applied ML models to predict inbound contamination rates using demographic features from 15 counties in the U.S. with different curbside collection strategies. In general, we found that ML models outperformed linear mixed models. Specifically, SVM models had the highest performance (R 2 = 0.75; mean absolute error (MAE) = 0.06), which may be due to their ability to model nonlinear relationships between features and inbound contamination rates. Further, the key predictor was population, with poverty rate being positively correlated and median age negatively correlated with inbound contamination rates. To improve the management of contamination and enhance the implementation of a circular economy, better models are needed to understand and estimate inbound contamination rates as well as identify critical factors in the present and future.

54 ENVIRONMENTAL SCIENCES↗

A Robust Schema for Storing and Managing Machine Learning Data and Models

- Machine Learning (ML) has enabled models that can improve efficiency and decrease computational cost - ML models are crucial in enabling Integrated Computational Materials Engineering (ICME) - Large data sets require robust means of storing ML data and models

Brandon L. Hearley↗

The Development and Deployment of Machine Learning Models for Aircraft Engine Concept Assessment

In today's competitive landscape, the effective development and utilization of machine-learning (ML) applications have become imperative across various sectors. This study presents an outline of the procedure involved in creating and implementing ML models for conceptualizing and evaluating aircraft engines. These models leverage supervised deep-learning algorithms to analyze patterns within an open-source repository containing data on both production and research conventional turbofan engines. The main areas of focus encompass crucial engine parameters like thrust-specific fuel consumption (TSFC), engine weight, engine diameter, and turbomachinery stage counts. While the creation of ML models is fundamental for their utilization, ensuring their seamless deployment holds equal significance. To address this aspect, a conversational AI chatbot is constructed, utilizing natural language processing (NLP) techniques, to facilitate the deployment of these ML models. The comprehensive workflow encompasses several key stages: gathering and enhancing engine data, training and cross validating the ML models, testing and evaluating their performance, and finally, deploying, monitoring, and updating the ML models. By following this systematic approach, the aim is to streamline the development and deployment process of ML models tailored for aircraft engine assessment.

Development↗

The Development and Deployment of Machine Learning Models for Aircraft Engine Concept Assessment

In today's competitive landscape, the effective development and utilization of machine-learning (ML) applications have become imperative across various sectors. This study presents an outline of the procedure involved in creating and implementing ML models for conceptualizing and evaluating aircraft engines. These models leverage supervised deep-learning algorithms to analyze patterns within an open-source repository containing data on both production and research conventional turbofan engines. The main areas of focus encompass crucial engine parameters like thrust-specific fuel consumption (TSFC), engine weight, engine diameter, and turbomachinery stage counts. While the creation of ML models is fundamental for their utilization, ensuring their seamless deployment holds equal significance. To address this aspect, a conversational AI chatbot is constructed, utilizing natural language processing (NLP) techniques, to facilitate the deployment of these ML models. The comprehensive workflow encompasses several key stages: gathering and enhancing engine data, training and cross validating the ML models, testing and evaluating their performance, and finally, deploying, monitoring, and updating the ML models. By following this systematic approach, the aim is to streamline the development and deployment process of ML models tailored for aircraft engine assessment.

Development↗

Subtleties in the trainability of quantum machine learning models

A new paradigm for data science has emerged, with quantum data, quantum models, and quantum computational devices. This field, called quantum machine learning (QML), aims to achieve a speedup over traditional machine learning for data analysis. However, its success usually hinges on efficiently training the parameters in quantum neural networks, and the field of QML is still lacking theoretical scaling results for their trainability. Some trainability results have been proven for a closely related field called variational quantum algorithms (VQAs). While both fields involve training a parametrized quantum circuit, there are crucial differences that make the results for one setting not readily applicable to the other. In this work, we bridge the two frameworks and show that gradient scaling results for VQAs can also be applied to study the gradient scaling of QML models. Our results indicate that features deemed detrimental for VQA trainability can also lead to issues such as barren plateaus in QML. Consequently, our work has implications for several QML proposals in the literature. In addition, we provide theoretical and numerical evidence that QML models exhibit further trainability issues not present in VQAs, arising from the use of a training dataset. We refer to these as dataset-induced barren plateaus. These results are most relevant when dealing with classical data, as here the choice of embedding scheme (i.e., the map between classical data and quantum states) can greatly affect the gradient scaling.

97 MATHEMATICS AND COMPUTING↗

Building Trustworthy Machine Learning Models for Astronomy

Astronomy is entering an era of data-driven discovery, due in part to modern machine learning (ML) techniques enabling powerful new ways to interpret observations. This shift in our scientific approach requires us to consider whether we can trust the black box. Here, we overview methods for an often-overlooked step in the development of ML models: building community trust in the algorithms. Trust is an essential ingredient not just for creating more robust data analysis techniques, but also for building confidence within the astronomy community to embrace machine learning methods and results.

46 INSTRUMENTATION RELATED TO NUCLEAR SCIENCE AND ↗

Machine learning models for segmentation and classification of cyanobacterial cells

Abstract Timelapse microscopy has recently been employed to study the metabolism and physiology of cyanobacteria at the single-cell level. However, the identification of individual cells in brightfield images remains a significant challenge. Traditional intensity-based segmentation algorithms perform poorly when identifying individual cells in dense colonies due to a lack of contrast between neighboring cells. Here, we describe a newly developed software package called Cypose which uses machine learning (ML) models to solve two specific tasks: segmentation of individual cyanobacterial cells, and classification of cellular phenotypes. The segmentation models are based on the Cellpose framework, while classification is performed using a convolutional neural network named Cyclass. To our knowledge, these are the first developed ML-based models for cyanobacteria segmentation and classification. When compared to other methods, our segmentation models showed improved performance and were able to segment cells with varied morphological phenotypes, as well as differentiate between live and lysed cells. We also found that our models were robust to imaging artifacts, such as dust and cell debris. Additionally, the classification model was able to identify different cellular phenotypes using only images as input. Together, these models improve cell segmentation accuracy and enable high-throughput analysis of dense cyanobacterial colonies and filamentous cyanobacteria.

Huffine, Clair A.↗

The Use of Machine Learning Models for Predicting the Dielectric Strength of Gases

Technological advancements in high voltage systems have pushed sulfur hexafluoride (SF6) to its operational limits. Furthermore, this gas has other drawbacks including a high liquefaction temperature and a high global warming potential. Therefore, there has been an urgent need to find alternative gases with high dielectric strength (DS). In this work, density functional theory (DFT) is used to calculate molecular descriptors that are fed into an artificial neural network (ANN) and a random forest (RF). These machine learning (ML) models are then used to predict the DS for hundreds of molecules. A finite element model (FEM) is also used to calculate the electric field profile of multiple simple electrode geometries as the applied voltage to the system is increased. Results indicate that the random forest model has better generalization to unseen data than the neural network. The highest DS value predicted by the RF was 2.16 relative to the experimental DS of SF6. The results also demonstrate how choosing a gas with a higher DS and a geometry with minimal edges and corners can significantly increase the operating voltage of an electrical system. Due to its superior generalization, the RF represents the most promising path toward an accurate DS predictor once sufficient experimental data are available.

Mileski, Matthew [AFIT]↗

Physics-based Models, Machine Learning, and Experiment: Towards Understanding Complex Electrode Degradation

Degradation phenomena in Li-ion batteries are highly complex, coupled, and sensitive to use history and operating conditions. In this study, we show how tracking model parameters in continuum-level physics-based models, expedited by machine learning, can be useful in testing hypotheses for degradation mechanisms. An exemplary analysis using this approach is presented for a set of lithium trivanadate ( L i x V 3 O 8 ) cathodes cycled over a range of current rates. A simple cell revival process is combined with the parameter estimates over the course of cycling to extract valuable insights into cathode evolution and eliminate hypothesized degradation mechanisms for these cathodes. The presented approach is expected to be broadly applicable for degradation analysis of other electrodes.

25 ENERGY STORAGE↗

A comparative study of machine learning models for predicting the state of reactive mixing

Mixing phenomena are important mechanisms controlling flow, species transport, and reaction processes in fluids and porous media. Accurate predictions of reactive mixing are critical for many Earth and environmental science problems such as contaminant fate and remediation, macroalgae growth, and plankton biomass evolution. Here, to investigate the evolution of mixing dynamics under different scenarios (e.g., anisotropy, fluctuating velocity fields), a finite-element-based numerical model was built to solve the fast, irreversible bimolecular reaction-diffusion equations to simulate a range of reactive-mixing scenarios. A total of 2,315 simulations were performed using different sets of model input parameters comprising various spatial scales of vortex structures in the velocity field, time-scales associated with velocity oscillations, the perturbation parameter for the vortex-based velocity, anisotropic dispersion contrast (i.e., ratio of longitudinal-to-transverse dispersion), and molecular diffusion. The outputs comprised concentration profiles of reactants and products. The inputs to and outputs from these simulations were concatenated into feature and label matrices, respectively, to train 20 different machine learning (ML) models intended to emulate system behavior. These 20 ML emulators, based on linear methods, Bayesian methods, ensemble learning methods, and multilayer perceptrons (MLPs), were trained to classify the state of mixing and predict three quantities of interest (QoIs) characterizing species production, decay (i.e., average concentration, square of average concentration), and degree of mixing (i.e., variances of species concentration). Unsurprisingly, linear classifiers and regressors failed to reproduce the QoIs; however, ensemble methods (classifiers and regressors) and the MLP model accurately classified the state of reactive mixing and the QoIs. Among ensemble methods, random forest and decision-tree-based AdaBoost faithfully predicted the QoIs. At run time, trained ML emulators produced results times faster than the finite-element simulations. Due to their low computational expense and high accuracy, ensemble and MLP models are excellent emulators for these numerical simulations and great utilities in uncertainty quantification exercises, which can require 1,000s of forward model runs.

97 MATHEMATICS AND COMPUTING↗

Classifying thermodynamic cloud phase using machine learning models

Vertically resolved thermodynamic cloud-phase classifications are essential for studies of atmospheric cloud and precipitation processes. The Department of Energy (DOE) Atmospheric Radiation Measurement (ARM) Thermodynamic Cloud Phase (THERMOCLDPHASE) value-added product (VAP) uses a multi-sensor approach to classify the thermodynamic cloud phase by combining lidar backscatter and depolarization, radar reflectivity, Doppler velocity, spectral width, microwave-radiometer-derived liquid water path, and radiosonde temperature measurements. The measured pixels are classified as ice, snow, mixed phase, liquid (cloud water), drizzle, rain, and liq_driz (liquid+drizzle). We use this product as the ground truth to train three machine learning (ML) models to predict the thermodynamic cloud phase from multi-sensor remote sensing measurements taken at the ARM North Slope of Alaska (NSA) observatory: a random forest (RF), a multi-layer perceptron (MLP), and a convolutional neural network (CNN) with a U-Net architecture. Evaluations against the outputs of the THERMOCLDPHASE VAP with 1 year of data show that the CNN outperforms the other two models, achieving the highest test accuracy, F1 score, and mean intersection over union (IOU). Analysis of ML confidence scores shows that ice, rain, and snow have higher confidence scores, followed by liquid, while mixed, drizzle, and liq_driz have lower scores. Feature importance analysis reveals that the mean Doppler velocity and vertically resolved temperature are the most influential data streams for ML thermodynamic cloud-phase predictions. Lidar measurements exhibit lower feature importance due to rapid signal attenuation caused by the frequent presence of persistent low-level clouds at the NSA site. The ML models' generalization capacity is further evaluated by applying them at another Arctic ARM site in Norway using data taken during the ARM Cold-Air Outbreaks in the Marine Boundary Layer Experiment (COMBLE) field campaign. The models demonstrated similar performance to that observed at the NSA site. Finally, we evaluate the ML models' response to simulated instrument outages and signal degradation and show that a CNN U-Net model trained with input channel dropouts performs better when input fields are missing.

ARM Aerial Facility↗

A Machine Learning Model for Predicting Composition of Catalytic Coprocessing Products from Molecular Beam Mass Spectra

Demand for the development of an automated and integrated refining process for biofuels has increased in recent years due to the lack of generalized process inspection tools. In bio-oil upgrading processes, all process variables are maintained based on the offline specification of intermediates and products. A lack of real-time product specifications in batch-wise monitoring can cause process failure and wasted resources. Therefore, there is a need for a fast and accurate intermediates/product specification tool that can be used for real-time specification to reduce waste and mitigate the risk of process failure. Here, to address this gap, we developed a machine learning (ML) model for predicting speciated bio-oil composition, including paraffin, iso-paraffins, olefins, naphthene, and aromatics. The model is trained using the mass spectra from upgraded products collected in the vapor phase before condensation and predicts the composition of the condensed product. Training ML models using raw mass spectra is challenging due to numerous overlapped peaks originating from different parent compounds. With this in mind, we propose a protocol that (i) transforms raw mass spectra to chemistry-inspired predefined features and (ii) trains decision tree-based models using these features. Our results show that the random forest model was robust against overfitting and had the highest accuracy compared to other models. Moreover, a stochastic ablation method determined the eight most significant features while maximizing the accuracy. Our protocol facilitates real-time compositional analysis of upgraded bio-oils and thus real-time process monitoring. Additionally, this protocol enables the rational design of efficient catalysts and the determination of optimal process conditions.

09 BIOMASS FUELS↗

Applied Machine-Learning Models to Identify Spectral Sub-Types of M Dwarfs from Photometric Surveys

M dwarfs are the most abundant stars in the Solar Neighborhood and they are prime targets for searching for rocky planets in habitable zones. Consequently, a detailed characterization of these stars is in demand. The spectral sub-type is one of the parameters that is used for the characterization and it is traditionally derived from the observed spectra. However, obtaining the spectra of M dwarfs is expensive in terms of observation time and resources due to their intrinsic faintness. We study the performance of four machine-learning (ML) models—K-Nearest Neighbor (KNN), Random Forest (RF), Probabilistic Random Forest (PRF), and Multilayer Perceptron (MLP)—in identifying the spectral sub-types of M dwarfs at a grand scale by deploying broadband photometry in the optical and near-infrared. We trained the ML models by using the spectroscopically identified M dwarfs from the Sloan Digital Sky Survey (SDSS) Data Release (DR) 7, together with their photometric colors that were derived from the SDSS, Two-Micron All-Sky Survey, and Wide-field Infrared Survey Explorer. We found that the RF, PRF, and MLP give a comparable prediction accuracy, 74%, while the KNN provides slightly lower accuracy, 71%. We also found that these models can predict the spectral sub-type of M dwarfs with ~99% accuracy within ±1 sub-type. The five most useful features for the prediction are r - z, r - i, r - J, r - H , and g - z, and hence lacking data in all SDSS bands substantially reduces the prediction accuracy. However, we can achieve an accuracy of over 70% when the r and i magnitudes are available. Since the stars in this study are nearby (d ≲ 1300 pc for 95% of the stars), the dust extinction can reduce the prediction accuracy by only 3%. Finally, we used our optimized RF models to predict the spectral sub-types of M dwarfs from the Catalog of Cool Dwarf Targets for the Transiting Exoplanet Survey Satellite, and we provide the optimized RF models for public use.

79 ASTRONOMY AND ASTROPHYSICS↗

Machine Learning Models to Predict Cognitive Impairment of Rodents Subjected to Space Radiation

INTRODUCTION We use artificial neural networks (ANNs) as an example machine learning (ML) tool to predict the cognitive performance impairment of rats induced by irradiation. The experimental data in the analyses is attentional set-shifting (ATSET) test scores from a rodent model exposed to ≤15 cGy of individual galactic cosmic radiation (GCR) ions: 4He, 28Si, or 56Fe, expected for a Lunar or Mars mission [1]. This work investigates rats at a subject-based level and uses applied dose and performance scores taken before irradiation to predict whether a rat will be impaired when irradiated. The results of this study are significant to crewed space missions as they support the potential of predicting an astronaut’s impairment in a specific task before spaceflight through the implementation of appropriately trained ML tools. METHODS Data used in this work are scores from the ATSET, a multi-stage constrained cognitive flexibility test [2]. Our computational model utilizes the number of attempts to reach the criterion to pass a stage as a behavioral performance measure for rats. We use the post-irradiation scores, generate thresholds from cumulative distribution plots of non-irradiated rats, and calculate the percent of irradiated rats whose scores fall below the threshold to infer how each radiation type/dose affects a population. Rats scoring above the threshold are labeled impaired while the others are non-impaired. We then employ ANNs as a typical ML technique, and use each subject’s individual scores taken before radiation along with the applied dose, to predict their personal susceptibility to cognitive impairment due to space radiation exposure. RESULTS AND CONCLUSION A significant finding is the exhibition of a dose-dependent increasing probability of impairment for 1 to 10 cGy of 28Si or 56Fe in the simple discrimination (SD) stage of the ATSET, and for 1 to 10 cGy of 56Fe in the compound discrimination (CD) stage. On a subject-based level, implementing ML classifiers such as ANNs identifies rats that have a higher tendency for impairment after GCR exposure [1]. The receiver operating characteristic (ROC) and the precision-recall (PR) curves of the ML models show a better prediction of impairment when 56Fe is the ion in question in both SD (Figure 1) and CD stages. They, however, do not depict impairment due to 4He in SD (Figure 1) and 28Si in CD, suggesting no dose-dependent impairment response in these cases. In this work, “good” prediction pertains to “better-than-random-chance”, due to the limited sample size and the high inter- and intra-individual variabilities in response to brain stimulation paradigms, as applicable to both animals and humans. More behavioral tests and biomarkers should be investigated on the same subjects, to be fed to the ML models to capture the agents responsible for performance alterations of some individuals versus others.

machine learning↗

Machine-learning modeling of magnetization dynamics in quasi-equilibrium and driven metallic spin systems

Here, we present a perspective on recent progress in machine-learning (ML) force-field approaches for large-scale Landau–Lifshitz–Gilbert (LLG) simulations of metallic spin systems. Building on a generalization of the Behler–Parrinello (BP) architecture originally developed for quantum molecular dynamics, we develop scalable and transferable ML models that faithfully capture the complex, environment-dependent electron-mediated exchange fields characteristic of itinerant magnets. A central ingredient of this framework is the implementation of symmetry-aware magnetic descriptors based on group-theoretical bispectrum formalisms. Leveraging these ML force fields, LLG simulations faithfully reproduce hallmark non-collinear magnetic orders—such as the 120° and tetrahedral states—on the triangular lattice, and successfully capture the complex spin textures emerging in the mixed-phase states of a square-lattice double-exchange model under thermal quench. We further discuss a generalized potential theory that extends the BP formalism to incorporate both conservative and nonconservative electronic torques, thereby enabling ML models to learn nonequilibrium exchange fields from computationally demanding microscopic approaches such as nonequilibrium Green’s-function techniques. This extension yields quantitatively accurate predictions of voltage-driven domain-wall motion and establishes a foundation for quantum-accurate, multiscale modeling of nonequilibrium spin dynamics and spintronic functionalities.

Descriptors↗