Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine Learning Models”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 577 records · Page 32

Insights into co-pyrolysis of polyethylene terephthalate and polyamide 6 mixture through experiments, kinetic modeling and machine learning

The non-isothermal pyrolysis of polyethylene terephthalate (PET), polyamide 6 (PA6), and their mixtures was studied in a thermogravimetric analyzer at different heating rates. Temperature of maximum decomposition (T max ) decreased by 25–45 °C and 35–55 °C for the PET:PA6 mixtures (3:1, 1:1, 1:3) compared to PET and PA6, respectively. The kinetic analysis was initially carried out using isoconversional method. However, the dependency of activation energy on conversion was observed for the co-pyrolysis of PET and PA6 that suggested the occurrence of multi-step reactions in the mixtures. Distributed activation energy model (DAEM) was used in this study to describe the multistep reactions occurring during pyrolysis of PET:PA6 mixtures. Here, in this work, a four-parallel reaction DAEM was developed to describe the pyrolysis kinetics of PET:PA6 mixtures. The apparent mean activation energies (E o ) for PET, PA6, and mixtures varied in the range of 244–255, 140–215, and 138–255 kJ mol –1 , respectively. The mass loss profiles of PET and PA6 mixtures were also modeled using artificial neural network (ANN). Out of 155 ANN models, the best prediction was made by ANN511 with R 2 greater than 0.997 for both test and unseen data. The interaction effects observed through TGA experiments and subsequent kinetic analysis were further assessed in terms of product composition using analytical pyrolysis coupled with gas chromatograph/mass spectrometer (Py-GC/MS). Co-pyrolysis of PET and PA6 resulted in the formation of new aromatic compounds with nitrogen-containing functional groups, which were not detected when PET or PA6 were pyrolyzed individually.

37 INORGANIC, ORGANIC, PHYSICAL, AND ANALYTICAL CH↗

Uncertainty-aware mixed-variable machine learning for materials design

Abstract Data-driven design shows the promise of accelerating materials discovery but is challenging due to the prohibitive cost of searching the vast design space of chemistry, structure, and synthesis methods. Bayesian optimization (BO) employs uncertainty-aware machine learning models to select promising designs to evaluate, hence reducing the cost. However, BO with mixed numerical and categorical variables, which is of particular interest in materials design, has not been well studied. In this work, we survey frequentist and Bayesian approaches to uncertainty quantification of machine learning with mixed variables. We then conduct a systematic comparative study of their performances in BO using a popular representative model from each group, the random forest-based Lolo model (frequentist) and the latent variable Gaussian process model (Bayesian). We examine the efficacy of the two models in the optimization of mathematical functions, as well as properties of structural and functional materials, where we observe performance differences as related to problem dimensionality and complexity. By investigating the machine learning models’ predictive and uncertainty estimation capabilities, we provide interpretations of the observed performance differences. Our results provide practical guidance on choosing between frequentist and Bayesian uncertainty-aware machine learning models for mixed-variable BO in materials design.

36 MATERIALS SCIENCE↗

Modeling the 4D discharge of lithium-ion batteries with a multiscale time-dependent deep learning framework

The lithium-ion battery (LIB) field is moving towards the direction of investigating spatially resolved physical phenomena in the 3D porous microstructure of electrodes. These pore-scale simulations give new insights into the local dynamics of lithiation/de-lithiation and charge transport, Nevertheless, the computational time of these simulations limits the integration of these models in optimization workflows of cycling conditions or electrode manufacturing processes. Machine learning models present a way of assessing in real-time the performance of materials. While several successful techniques for replicating simulations with machine learning have been proposed, this case study presents a more demanding problem, due to the necessity of understanding the behavior of heterogeneous 3D local data, as it evolves in time: this poses both a scientific and a technical challenge. To this end, we propose an autoregressive multiscale convolutional neural network model to predict relevant quantities at the pore-scale in the solid phase: the lithium concentration (in the active material) and potential (in the active material and carbon binder). Here, these are ultimately used to reconstruct the battery discharge curve. 3D images of the electrode microstructures are the input to the network, trained with a dataset of finite element method simulations to predict the discharge behavior of the cathode side in lithium ion batteries. We propose this machine learning model as a proof-of-concept of the applicability of multiscale networks for time-dependent physics problems. The trained model exhibits very high accuracy (with errors lower than 2 %) in forecasting the discharge behavior of new unseen cathodes.

25 ENERGY STORAGE↗

BEAST DB: Grand-Canonical Database of Electrocatalyst Properties

We present BEAST DB, an open-source database comprised of ab initio electrochemical data computed using grand-canonical density functional theory in implicit solvent at consistent calculation parameters. The database contains over 20,000 surface calculations and covers a broad set of heterogeneous catalyst materials and electrochemical reactions. Calculations were performed at self-consistent fixed potential as well as constant charge to facilitate comparisons to the computational hydrogen electrode. This article presents common use cases of the database to rationalize trends in catalyst activity, screen catalyst material spaces, understand elementary mechanistic steps, analyze the electronic structure, and train machine learning models to predict higher fidelity properties. Users can interact graphically with the database by querying for individual calculations to gain a granular understanding of reaction steps or by querying for an entire reaction pathway on a given material using an interactive reaction pathway tool. BEAST DB will be periodically updated, with planned future updates to include advanced electronic structure data, surface speciation studies, and greater reaction coverage.

database↗

Accurate Prediction of Algal Biomass Lipid, Protein, and Carbohydrate Composition with Machine Learning Regression Modelling of Near-IR Spectra

During large scale algal biomass cultivation, it is difficult to reliably control relative composition to target levels. Rapid determination of chemical composition is feasible by using near infrared (NIR) spectral data. We sought to build and improve on reliable high-throughput screening prediction method based on partial least squares regression (PLSR) by the application of artificial neural networks (ANN) and associated optimization strategies. The algal biomass sample set was designed and created in an iterative process of culturing in physiologically diverse conditions at the GAI field site, followed by compositional analyses at NREL. The workflow allowed us to identify gaps in compositional space for informing the subsequent cultivation and sampling efforts and generated a high quality set of 210 unique samples with chemical analysis results, spectral scanning data, and cultivation metadata. We observed a significant improvement in the performance of carbohydrate content predictions using an optimized ANN model compared to PLSR, with > 16% reduction in mean absolute percent error (MAPE) when tested on the same set of reserved data. The optimized ANN models for FAME and protein prediction performed exceptionally well with 5.99% and 5.09% MAPE, respectively. Application of these methods to detection and quantification of minor biomass constituents that are relevant to certain product streams has shown positive preliminary results, opening the possibility for extensions to the outputs of this powerful data type. All models are accompanied by prediction uncertainties and unsupervised spectral outlier detection to alert an operator to unreliable spectral data. These tools can be deployed for rapid determination of algal culture status, and cultivation and biomass quality improvement.

algal biofuels↗

Advancing molecular machine learning representations with stereoelectronics-infused molecular graphs

Molecular representation is a critical element in our understanding of the physical world and the foundation for modern molecular machine learning. Previous molecular machine learning models have used strings, fingerprints, global features and simple molecular graphs that are inherently information-sparse representations. However, as the complexity of prediction tasks increases, the molecular representation needs to encode higher fidelity information. This work introduces a new approach to infusing quantum-chemical-rich information into molecular graphs via stereoelectronic effects, enhancing expressivity and interpretability. Learning to predict the stereoelectronics-infused representation with a tailored double graph neural network workflow enables its application to any downstream molecular machine learning task without expensive quantum-chemical calculations. We show that the explicit addition of stereoelectronic information substantially improves the performance of message-passing two-dimensional machine learning models for molecular property prediction. We show that the learned representations trained on small molecules can accurately extrapolate to much larger molecular structures, yielding chemical insight into orbital interactions for previously intractable systems, such as entire proteins, opening new avenues of molecular design. Finally, we have developed a web application (simg.cheme.cmu.edu) where users can rapidly explore stereoelectronic information for their own molecular systems.

Boiko, Daniil A↗

Automated algorithms to build active galactic nucleus classifiers

ABSTRACT We present a machine learning model to classify active galactic nuclei (AGNs) and galaxies (AGN-galaxy classifier) and a model to identify type 1 (optically unabsorbed) and type 2 (optically absorbed) AGN (type 1/2 classifier). We test tree-based algorithms, using training samples built from the X-ray Multi-Mirror Mission–Newton (XMM–Newton) catalogue and the Sloan Digital Sky Survey (SDSS), with labels derived from the SDSS survey. The performance was tested making use of simulations and of cross-validation techniques. With a set of features including spectroscopic redshifts and X-ray parameters connected to source properties (e.g. fluxes and extension), as well as features related to X-ray instrumental conditions, the precision and recall for AGN identification are 94 and 93 per cent, while the type 1/2 classifier has a precision of 74 per cent and a recall of 80 per cent for type 2 AGNs. The performance obtained with photometric redshifts is very similar to that achieved with spectroscopic redshifts in both test cases, while there is a decrease in performance when excluding redshifts. Our machine learning model trained on X-ray features can accurately identify AGN in extragalactic surveys. The type 1/2 classifier has a valuable performance for type 2 AGNs, but its ability to generalize without redshifts is hampered by the limited census of absorbed AGN at high redshift.

Falocco, S. (ORCID:0000000299841103)↗

Increased Interpretability for Model-Driven Deception: MARS LDRD Project

Machine learning has been proposed as a solution to several cybersecurity solutions and one of the most promising applications is for digital twins for intrusion detection and driving deceptive defense. However, machine learning techniques often result in a black-box function that is difficult for end users to interpret which for deception limits their ability to effectively define decoys. In this report, an approach to validate the equations learned are accurate is provided and demonstrated. Following, begins the process of addressing this issue for a model-driven deception technology that produces equations representing the physical process controlled by operation technology devices. This research was performed by applying subject matter expert context to machine learned models.

97 MATHEMATICS AND COMPUTING↗

Response of U.S. West Coast Mountain Snowpack to Local Sea Surface Temperature Perturbations: Insights from Numerical Modeling and Machine Learning

Sea surface temperature (SST) significantly modulates the precipitation and temperature over land, with important consequences on land surface processes such as snowpack. Compared to the impact of remote SST, the effect of nearshore/local SST is less well understood. In this study, the impact of local SST on the mountain snowpack of the U.S. West Coast is investigated using two 6-km regional climate simulations driven by the same lateral boundary conditions but with time-varying versus time-invariant and warmer local SSTs during 2003–15. Results show that local SST warming leads to warmer winter with more precipitation over the mountains. Meanwhile, the removal of SST temporal variability results in reduced temperature variability but increased precipitation variability. As a result, winter snow accumulation decreases by ~200 mm per season in the Cascade Mountains in the north but increases by ~100 mm per season in the Sierra Nevada in the south. Such a dipole response results from the competing effects of precipitation and temperature change at different elevations and are amplified by the enhanced atmospheric river moisture transport. To further delineate the relative contributions of different meteorological factors to the snowpack response, two neural network models were developed to predict the snow behaviors at seasonal and monthly scales. These models reveal the dominant influence of the total amount and the average temperature of precipitation on the snowpack response. Furthermore, these findings highlight the sensitivity of mountain snowpack to local SST in the western United States and underscore the importance of local SST and atmospheric rives to accurate snowpack estimations for water management.

54 ENVIRONMENTAL SCIENCES↗

Human limits in machine learning: prediction of potato yield and disease using soil microbiome data

Abstract Background The preservation of soil health is a critical challenge in the 21st century due to its significant impact on agriculture, human health, and biodiversity. We provide one of the first comprehensive investigations into the predictive potential of machine learning models for understanding the connections between soil and biological phenotypes. We investigate an integrative framework performing accurate machine learning-based prediction of plant performance from biological, chemical, and physical properties of the soil via two models: random forest and Bayesian neural network. Results Prediction improves when we add environmental features, such as soil properties and microbial density, along with microbiome data. Different preprocessing strategies show that human decisions significantly impact predictive performance. We show that the naive total sum scaling normalization that is commonly used in microbiome research is one of the optimal strategies to maximize predictive power. Also, we find that accurately defined labels are more important than normalization, taxonomic level, or model characteristics. ML performance is limited when humans can’t classify samples accurately. Lastly, we provide domain scientists via a full model selection decision tree to identify the human choices that optimize model prediction power. Conclusions Our study highlights the importance of incorporating diverse environmental features and careful data preprocessing in enhancing the predictive power of machine learning models for soil and biological phenotype connections. This approach can significantly contribute to advancing agricultural practices and soil health management.

Aghdam, Rosa↗

A Machine Learning Correction Model of the Winter Clear-Sky Temperature Bias over the Arctic Sea Ice in Atmospheric Reanalyses

Atmospheric reanalyses are widely used to estimate the past atmospheric near-surface state over sea ice. They provide boundary conditions for sea ice and ocean numerical simulations and relevant information for studying polar variability and anthropogenic climate change. Previous research revealed the existence of large near-surface temperature biases (mostly warm) over the Arctic sea ice in the current generation of atmospheric reanalyses, which is linked to a poor representation of the snow over the sea ice and the stably stratified boundary layer in the forecast models used to produce the reanalyses. These errors can compromise the employment of reanalysis products in support of polar research. Here, we train a fully connected neural network that learns from remote sensing infrared temperature observations to correct the existing generation of uncoupled atmospheric reanalyses (ERA5, JRA-55) based on a set of sea ice and atmospheric predictors, which are themselves reanalysis products. The advantages of the proposed correction scheme over previous calibration attempts are the consideration of the synoptic weather and cloud state, compatibility of the predictors with the mechanism responsible for the bias, and a self-emerging seasonality and multidecadal trend consistent with the declining sea ice state in the Arctic. The correction leads on average to a 27% temperature bias reduction for ERA5 and 7% for JRA-55 if compared to independent in situ observations from the MOSAiC campaign (respectively, 32% and 10% under clear-sky conditions). These improvements can be beneficial for forced sea ice and ocean simulations, which rely on reanalyses surface fields as boundary conditions.

54 ENVIRONMENTAL SCIENCES↗

Statistical Treatment of Convolutional Neural Network Superresolution of Inland Surface Wind for Subgrid-Scale Variability Quantification

Abstract Machine learning models have been employed to perform either physics-free data-driven or hybrid dynamical downscaling of climate data. Most of these implementations operate over relatively small downscaling factors because of the challenge of recovering fine-scale information from coarse data. This limits their compatibility with many global climate model outputs, often available between ∼50- and 100-km resolution, to scales of interest such as cloud resolving or urban scales. This study systematically examines the capability of a type of superresolving convolutional neural network (SR-CNNs) to downscale surface wind speed data over land from different coarse resolutions (25-, 48-, and 100-km resolution) to 3 km. For each downscaling factor, we consider three convolutional neural network (CNN) configurations that generate superresolved predictions of fine-scale wind speed, which take between one and three input fields: coarse wind speed, fine-scale topography, and diurnal cycle. In addition to fine-scale wind speeds, probability density function parameters are generated through which sample wind speeds can be generated, accounting for the intrinsic stochasticity of wind speed. For assessing generalization to new data, CNN models are tested on regions with different topography and climate that are unseen during training. The evaluation of superresolved predictions focuses on subgrid-scale variability and the recovery of extremes. Models with coarse wind and fine topography as inputs exhibit the best performance when compared with other model configurations, operating across the same downscaling factor. Our diurnal cycle encoding results in lower out-of-sample generalizability when compared with other input configurations.

17 WIND ENERGY↗

Optimal control of the electron temperature profile in DIII-D using machine learning surrogate models

The viability of the tokamak as a potential fusion reactor depends on the ability to keep the plasma in a stable regime while achieving temperatures, densities, and confinement times that are as high as possible. Tokamak scenario development attempts to find plasma regimes that achieve all of these conditions and are accessible with a given set of hardware constraints. This requires the ability to control plasma properties such as the normalized beta, the internal inductance, safety factor, rotation, etc. One property that has received less attention than some of the others, but is no less critical to achieving high performance, is the electron temperature (T e ) profile. In this work, Linear Quadratic Integral (LQI) control is used to develop a controller for the electron temperature profile in DIII-D. The controller is based on a linearized model derived from the transport equation that describes the evolution of the electron temperature, and includes contributions from the neural network surrogate models NubeamNet and MMMnet. Furthermore, the controller is tested in simulation using COTSIM, and is proven capable of tracking a target T e profile.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Numerical modeling based machine learning approach for the optimization of falling - film evaporator in thermal desalination application

Scale formation that drastically increases thermal resistance and reduces freshwater production remains a critical challenge in thermal desalination. Novel designs of falling film evaporator and optimal operating condition hold great promise to mitigate scale formation, and increase heat transfer performance and fresh water production. In this work, CFD simulation based machine learning and multi-objective optimization are performed to identify optimal conditions and tube arrangement for evaporator. Non-dominated sorting genetic algorithm is adopted to determine and analyze the optimal pareto front for multiple objectives in desalination criteria. The errors of training, validation, and testing set are computed to identify an optimal hyperparameter set. For performance ratio, fouling resistance, and water production rate, the average relative error is 2.26%, 3.67%, and 3.24%. At pareto front, both performance ratio and water production rate increase at high temperature with fouling resistance (thermal resistance of the fouling layer) increasing as well. Tradeoffs between mitigating scale formation and enhancing desalination performance are evaluated in optimizations for different objectives. Finally, potential optima are identified and can be applied as guidelines to determine evaporator design and system operating conditions.

42 ENGINEERING↗