Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Machine learning prediction”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 307 records · Page 17

A Machine Learning Framework to Predict Images of Edge-on Protoplanetary Disks

The physical structure and properties of protoplanetary disks are typically derived from spatially resolved disk images. Edge-on disks in particular provide an important view point on the vertical structure and degree of settling of disks. Such analyses rely on radiative transfer (RT) calculations that are generally computationally intensive due to the high optical depth of disks. Here we present a machine learning framework that has the potential to dramatically speed up the forward modeling process by approximating the results of RT calculations. This framework, trained on an initial set of RT calculations, utilizes an autoencoder neural network to enable the generation of synthetic scattered light images of edge-on disks directly from a set of physical parameters. We demonstrate that this framework generates synthetic images 2–3 orders of magnitude faster than using RT calculations. These machine learning-generated images appear to approximate the RT images well, in particular preserving their size and shape. We also find a strong correlation between the latent space representations of the generated disk images and several of their associated physical parameters. Finally, we discuss potential changes to the framework, such as methods to further improve the image quality, extending the framework to multiple wavelengths, and inverting the process to infer physical parameters from observed images. Overall, these new tools have the potential to enable a more efficient and uniform analysis of edge-on disk properties and the initial conditions of planet formation.

79 ASTRONOMY AND ASTROPHYSICS↗

Microscopic and Macroscopic Characterization of Grain Boundary Energy and Strength in Silicon Carbide via Machine-Learning Techniques

Predicting the properties of grain boundaries poses a challenge because of the complex relationships between structural and chemical attributes both at the atomic and continuum scales. Grain boundary systems are typically characterized by parameters used to classify local atomic arrangements in order to extract features such as grain boundary energy or grain boundary strength. The present work utilizes a combination of high-throughput atomistic simulations, macroscopic and microscopic descriptors, and machine-learning techniques to characterize the energy and strength of silicon carbide grain boundaries. Additionally, a diverse data set of symmetric tilt and twist grain boundaries are described using macroscopic metrics such as misorientation, the alignment of critical low-index planes, and the Schmid factor, but also in terms of microscopic metrics, by quantifying the local atomic structure and chemistry at the interface. These descriptors are used to create random-forest regression models, allowing for their relative importance to the grain boundary energy and decohesion stress to be better understood. Results show that while the energetics of the grain boundary were best described using the microscopic descriptors, the ability of the macroscopic descriptors to reasonably predict grain boundaries with low energy suggests a link between the crystallographic orientation and the resultant atomic structure that forms at the grain boundary within this regime. For grain boundary strength, neither microscopic nor macroscopic descriptors were able to fully capture the response individually. However, when both descriptor sets were utilized, the decohesion stress of the grain boundary could be accurately predicted. These results highlight the importance of considering both macroscopic and microscopic factors when constructing constitutive models for grain boundary systems, which has significant implications for both understanding the fundamental mechanisms at work and the ability to bridge length scales.

36 MATERIALS SCIENCE↗

Review of machine learning and deep learning models for toxicity prediction

The ever-increasing number of chemicals has raised public concerns due to their adverse effects on human health and the environment. To protect public health and the environment, it is critical to assess the toxicity of these chemicals. Traditional in vitro and in vivo toxicity assays are complicated, costly, and time-consuming and may face ethical issues. These constraints raise the need for alternative methods for assessing the toxicity of chemicals. Recently, due to the advancement of machine learning algorithms and the increase in computational power, many toxicity prediction models have been developed using various machine learning and deep learning algorithms such as support vector machine, random forest, k-nearest neighbors, ensemble learning, and deep neural network. This review summarizes the machine learning- and deep learning-based toxicity prediction models developed in recent years. Support vector machine and random forest are the most popular machine learning algorithms, and hepatotoxicity, cardiotoxicity, and carcinogenicity are the frequently modeled toxicity endpoints in predictive toxicology. It is known that datasets impact model performance. The quality of datasets used in the development of toxicity prediction models using machine learning and deep learning is vital to the performance of the developed models. The different toxicity assignments for the same chemicals among different datasets of the same type of toxicity have been observed, indicating benchmarking datasets is needed for developing reliable toxicity prediction models using machine learning and deep learning algorithms. This review provides insights into current machine learning models in predictive toxicology, which are expected to promote the development and application of toxicity prediction models in the future.

Research & Experimental Medicine↗

Optimized Machine Learning Model for Predicting Groundwater Contamination

The use of physical models to predict groundwater contaminant movement remains technically challenging due to the complexity of the phenomena, the heterogeneity of key parameters in nature, and the presence of poorly defined interactive and feedback processes. New approaches to address these challenges are needed. In this study, we evaluate various Artificial Intelligence (AI)-based approaches to understand a hexavalent chromium (Cr(VI)) plumes located on the U.S. Department of Energy’s (DOE) Hanford Site in Richland, WA. The groundwater monitoring dataset used in this study included data from the 100 Area along the Columbia River and included data collected between 2010 to 2019. This study investigates the most prominent contaminant, Cr(VI), with the Extreme Gradient Boosting (XGBoost) machine learning model. The XGBoost models were compared with optimized versions using an Empirical Bayes Search Cross-Validation technique for better prediction. The optimized XGBoost model yielded an R^2 value of 0.99 on the training set and 0.85 on the testing set, whereas XGBoost without optimization yielded a value of 0.83 on the training set and 0.85 on the testing set. This paper provides an overview of a computational method for groundwater contamination modeling that shows promise for improving current remediation efforts.

Mazumdar, Hirak↗

Construction of Women’s All-Around Speed Skating Event Performance Prediction Model and Competition Strategy Analysis Based on Machine Learning Algorithms

Introduction Accurately predicting the competitive performance of elite athletes is an essential prerequisite for formulating competitive strategies. Women’s all-around speed skating event consists of four individual subevents, and the competition system is complex and challenging to make accurate predictions on their performance. Objective The present study aims to explore the feasibility and effectiveness of machine learning algorithms for predicting the performance of women’s all-around speed skating event and provide effective training and competition strategies. Methods The data, consisting of 16 seasons of world-class women’s all-around speed skating competition results, used in the present study came from the International Skating Union (ISU). According to the competition rules, distinct features are filtered using lasso regression, and a 5,000 m race model and a medal model are built using a fivefold cross-validation method. Results The results showed that the support vector machine model was the most stable among the 5,000 m race and the medal models, with the highest AUC (0.86, 0.81, respectively). Furthermore, 3,000 m points are the main characteristic factors that decide whether an athlete can qualify for the final. The 11th lap of the 5,000 m, the second lap of the 500 m, and the fourth lap of the 1,500 m are the main characteristic factors that affect the athlete’s ability to win medals. Conclusion Compared with logistic regression, random forest, K-nearest neighbor, naive Bayes, neural network, support vector machine is a more viable algorithm to establish the performance prediction model of women’s all-around speed skating event; excellent performance in the 3,000 m event can facilitate athletes to advance to the final, and athletes with outstanding performance in the 500 m event are more likely competitive for medals.

Liu, Meng↗

Complexity-calibrated benchmarks for machine learning reveal when prediction algorithms succeed and mislead

Abstract Recurrent neural networks are used to forecast time series in finance, climate, language, and from many other domains. Reservoir computers are a particularly easily trainable form of recurrent neural network. Recently, a “next-generation” reservoir computer was introduced in which the memory trace involves only a finite number of previous symbols. We explore the inherent limitations of finite-past memory traces in this intriguing proposal. A lower bound from Fano’s inequality shows that, on highly non-Markovian processes generated by large probabilistic state machines, next-generation reservoir computers with reasonably long memory traces have an error probability that is at least $$\sim 60\%$$ ∼ 60 % higher than the minimal attainable error probability in predicting the next observation. More generally, it appears that popular recurrent neural networks fall far short of optimally predicting such complex processes. These results highlight the need for a new generation of optimized recurrent neural network architectures. Alongside this finding, we present concentration-of-measure results for randomly-generated but complex processes. One conclusion is that large probabilistic state machines—specifically, large $$\epsilon$$ ϵ -machines—are key to generating challenging and structurally-unbiased stimuli for ground-truthing recurrent neural network architectures.

97 MATHEMATICS AND COMPUTING↗

Predicting chatter using machine learning and acoustic signals from low-cost microphones

Machining chatter is a phenomenon resulting from self-oscillation between a machining tool and workpiece. This self-oscillation results in variation on the machined product that reduces the ability to meet desired specifications. Chatter is a widely studied topic as it directly relates to the quality of machined products. Here, this study details the application of a Random Forest (RF) classifier with Recursive Feature Elimination (RFE) to machining audio collected by a single microphone during down-milling operations. This approach allows straightforward feature elimination that results in an easily understood set of analyzed dimensions. Stability is predicted solely based on the classification output of the RF classifier. Our approach proves highly predictive with consistent machining setup and a small sample set. We also review transferability between machining setups and present key findings. Our RF approach demonstrates the ability to analyze and classify chatter through a low-cost approach with limited training data required. The motivation for using a single microphone is to enable detection on machines without other sensors, such as accelerometers, present in the machining setup. The value of the in-process sensor and chatter classifier is highlighted because the machining setup included asymmetric dynamics that reduced the accuracy of the traditional analytical stability solution. We see a natural progression to deploying this audio-only methodology with real-time processing and classification using either a laptop or smartphone. This progression will allow visual indicators during the machining process that can alert machinists of progression into unstable machining processes.

42 ENGINEERING↗

Empirical relationships between environmental factors and soil organic carbon produce comparable prediction accuracy as the Machine Learning

Accurate representation of environmental controllers of soil organic carbon (SOC) stocks in Earth System Model (ESM) land models could reduce uncertainties in future carbon-climate feedback projections. Using empirical relationships between environmental factors and SOC stocks to evaluate land models can help modelers understand prediction biases beyond what can be achieved with the observed SOC stocks alone. In this study, we used 31 observed environmental factors, field SOC observations (n = 6,213) from the continental US, and two Machine Learning approaches [Random Forest (RF) and Generalized Additive Modeling (GAM)] to (1) select important environmental predictors of SOC stocks, (2) derive empirical relationships between environmental factors and SOC stocks, and (3) use the derived relationships to predict SOC stocks and compare the prediction accuracy of simpler model developed with the machine learning predictions. Out of the 31 environmental factors we investigated, 12 were identified as important predictors of SOC stocks by the RF approach. In contrast, the GAM approach identified six (of those 12) environmental factors as important controllers of SOC stocks: potential evapotranspiration, normalized difference vegetation index, soil drainage condition, precipitation, elevation, and net primary productivity. The GAM approach showed minimal SOC predictive importance of the remaining six environmental factors identified by the RF approach. Our derived empirical relations produced comparable prediction accuracy as the GAM and RF approach using only a subset of environmental factors. The empirical relationships we derived using the GAM approach can serve as important benchmarks to evaluate environmental control representations of SOC stocks in ESMs, which could reduce uncertainty in predicting future carbon-climate feedbacks.

54 ENVIRONMENTAL SCIENCES↗

Development of Machine-Learned Interatomic Potentials to Predict Structure, Transport, and Reactivity in Platinum-Based Fuel Cells

Machine-learned interatomic potentials (MLIPs) have rapidly progressed in accuracy, speed, and data efficiency in recent years. However, training robust MLIPs in multicomponent systems remains a challenge. In this work, we train an MLIP to describe hydrated Nafion ionomers and platinum catalysts, which are important components of fuel cells, by constructing a diverse training set to describe the bulk polymer and interfacial catalyst–polymer interactions well. We use our trained MLIP to study the properties of the platinum–Nafion system, including polymer structure, proton mobility in a bulk Nafion polymer and near a platinum-Nafion interface, and reactions near and far from the interface, finding excellent results for structure and reactions contained within our training set. Transport seems to be well described, with both vehicular transport and Grotthuss hopping captured, although converged calculations of diffusivities were not computed because they require calculations of tens of nanoseconds that are challenging with current state-of-the-art MLIPs. The combined insights that this model provides can be leveraged to optimize fuel cell performance, and the approach can be applied to other chemical processes and devices where structure, transport, and reactivity all contribute to the overall observed performance.

33 ADVANCED PROPULSION SYSTEMS↗

Machine Learning Travel Time emulator

This notebook makes seismic phase travel time predictions using machine learning. There are 3 different machine learning models corresponding to the three regimes: local, regional, and teleseismic. After reading in the test data and splitting into the 3 regimes, we use precomputed scalers to scale the input features used to make the travel time predictions. We then load the machine learning models and use them to predict travel times for test data.

Anderson, Gemma↗