Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 73 records · Page 4

Random Forest Prediction of Crystal Structure from Electron Diffraction Patterns

Transmission electron microscopy (TEM) diffraction patterns are regularly used to determine the structure of crystalline materials. Electron diffraction is the most common method to solve for unknown or partially known crystal structures, as it provides direct and interpretable feedback on the orientation of crystal grains under the beam [1]. However, it remains a challenge to determine the crystal structure of a new material or even a new phase of an existing material. Analysis of such materials commonly requires manual exploration and comparison with simulated diffraction patterns. This is often a time consuming process with no obvious start point when many similar structures are possible, and this method cannot be used to determine crystal structure or orientation from structures not included in the diffraction libraries. Therefore, we have developed a machine learning model to determine the crystal structure of a material from its electron diffraction pattern.

36 MATERIALS SCIENCE↗

Machine learning models for rat multigeneration reproductive toxicity prediction

Reproductive toxicity is one of the prominent endpoints in the risk assessment of environmental and industrial chemicals. Due to the complexity of the reproductive system, traditional reproductive toxicity testing in animals, especially guideline multigeneration reproductive toxicity studies, take a long time and are expensive. Therefore, machine learning, as a promising alternative approach, should be considered when evaluating the reproductive toxicity of chemicals. We curated rat multigeneration reproductive toxicity testing data of 275 chemicals from ToxRefDB (Toxicity Reference Database) and developed predictive models using seven machine learning algorithms (decision tree, decision forest, random forest, k-nearest neighbors, support vector machine, linear discriminant analysis, and logistic regression). A consensus model was built based on the seven individual models. An external validation set was curated from the COSMOS database and the literature. The performances of individual and consensus models were evaluated using 500 iterations of 5-fold cross-validations and the external validation data set. The balanced accuracy of the models ranged from 58% to 65% in the 5-fold cross-validations and 45%–61% in the external validations. Prediction confidence analysis was conducted to provide additional information for more appropriate applications of the developed models. The impact of our findings is in increasing confidence in machine learning models. We demonstrate the importance of using consensus models for harnessing the benefits of multiple machine learning models (i.e., using redundant systems to check validity of outcomes). While we continue to build upon the models to better characterize weak toxicants, there is current utility in saving resources by being able to screen out strong reproductive toxicants before investing in vivo testing. The modeling approach (machine learning models) is offered for assessing the rat multigeneration reproductive toxicity of chemicals. Our results suggest that machine learning may be a promising alternative approach to evaluate the potential reproductive toxicity of chemicals.

consensus model↗

Southwest Pacific tropical cyclone development classification utilizing machine learning and synoptic composites

This study evaluates the ability of machine learning algorithms to classify tropical depressions (TDs) and tropical storms (TSs) in the western region of the southwest Pacific Ocean (SWPO). Decision rules are generated to predict the environment required for a depression to fully develop into a mature storm, and the most influential predictors in the classification decision are ranked. TD and TS are discriminated based on a maximum sustained wind speed threshold (≥17 ms -1 ). Various aerosol, thermodynamic, and dynamic parameters are extracted closest to the initiation point of each non-developing and developing sample. The covariates associated with each labelled sample are used to train a decision tree and random forest model. Results using a testing dataset suggest the random forest approach more accurately distinguishes between non-developing and developing samples. The classification accuracy of the decision tree and random forest are 72% and 91%, respectively. Random forest outperformed the decision tree by providing higher accuracy in test data. The most important variables for binary classification are sea salt aerosol optical depth (AOD), 1,000 mb relative humidity, and sea surface temperature. AOD is a quantitative estimate of the aerosols presents in the air through the extinction of a ray of light as it passes through the atmosphere. Mean composite maps constructed in an unsupervised manner have been created for the most important variables identified by the random forest classifier during TD and TS events to highlight the difference in geophysical and aerosol variables' climatology during the two different classifications. This work will advance the risk management strategies for northeastern Australia and other SWPO basin islands to control their tropical cyclone related losses through prioritizing forecasting variables that are the strongest predictors of the strengthening of tropical depressions into tropical cyclones.

54 ENVIRONMENTAL SCIENCES↗

Applied Machine-Learning Models to Identify Spectral Sub-Types of M Dwarfs from Photometric Surveys

M dwarfs are the most abundant stars in the Solar Neighborhood and they are prime targets for searching for rocky planets in habitable zones. Consequently, a detailed characterization of these stars is in demand. The spectral sub-type is one of the parameters that is used for the characterization and it is traditionally derived from the observed spectra. However, obtaining the spectra of M dwarfs is expensive in terms of observation time and resources due to their intrinsic faintness. We study the performance of four machine-learning (ML) models—K-Nearest Neighbor (KNN), Random Forest (RF), Probabilistic Random Forest (PRF), and Multilayer Perceptron (MLP)—in identifying the spectral sub-types of M dwarfs at a grand scale by deploying broadband photometry in the optical and near-infrared. We trained the ML models by using the spectroscopically identified M dwarfs from the Sloan Digital Sky Survey (SDSS) Data Release (DR) 7, together with their photometric colors that were derived from the SDSS, Two-Micron All-Sky Survey, and Wide-field Infrared Survey Explorer. We found that the RF, PRF, and MLP give a comparable prediction accuracy, 74%, while the KNN provides slightly lower accuracy, 71%. We also found that these models can predict the spectral sub-type of M dwarfs with ~99% accuracy within ±1 sub-type. The five most useful features for the prediction are r - z, r - i, r - J, r - H , and g - z, and hence lacking data in all SDSS bands substantially reduces the prediction accuracy. However, we can achieve an accuracy of over 70% when the r and i magnitudes are available. Since the stars in this study are nearby (d ≲ 1300 pc for 95% of the stars), the dust extinction can reduce the prediction accuracy by only 3%. Finally, we used our optimized RF models to predict the spectral sub-types of M dwarfs from the Catalog of Cool Dwarf Targets for the Transiting Exoplanet Survey Satellite, and we provide the optimized RF models for public use.

79 ASTRONOMY AND ASTROPHYSICS↗

Label Assist: Personalized Travel Models for Longitudinal Data Collection

Understanding travel behavior is crucial to transportation decarbonization. OpenPATH is an open-source mobility platform which collects and analyzes human travel behavior at the individual level. The mobile application passively senses trips and prompts users to label them. However, users find the labeling process burdensome; less than half the trips are typically labeled, making much of the data unusable in aggregate analyses of mobility patterns. Prior work has addressed the response fatigue challenge through automated mode inference using sensor data, but sensors cannot capture all aspects of travel behavior. We explore an alternative approach in which we leverage prior user input to predict travel choices in novel trips. We first explore trip clustering methods and develop a novel two-step pipeline using DBSCAN and SVMs to extract realistic geospatial clusters. We then propose two strategies to predict trip labels: (i) clustering trips and extrapolating labels for similar trips, and (ii) random forest classification. The random forest approach is able to achieve - $70-80% accuracy (purpose: 72%, mode: 79%, replaced mode: 81%). These novel approaches to trip classification allow us to increase the rate of user labeling by suggesting predicted labels to be verified by the user. Unlabeled trips can also contribute to aggregate analyses, using label predictions and their associated confidences as a substitute. While there exist other travel survey apps with the ability to infer travel choices, to our knowledge, this is the first paper to describe such a supervised system and rigorously evaluate it.

ADVANCED PROPULSION SYSTEMS,ENERGY PLANNING, POLIC↗

Weather and Random Forest-based Load Profiling Approximation Models and its Transferability across Climate Zones

This study is to provide predictive understanding of the associations of various weather attributes with residential and commercial load profiles, for a variety of climate zones and seasons. In this work, machine learning (ML) approaches were used to identify and quantify the impacts of various weather attributes on residential and commercial electricity demand and its components across the western United States. Performance and transferability of the developed ML models were then evaluated across different temperate zones (e.g., southern, middle, and northern US) and across coastal, mid-continent, and wet zones, with inputs of weather condition data from the National Oceanic and Atmospheric Administration (NOAA) at representative weather stations. The predictive models were developed based on the ranked/screened factors using the regression tree (RT) and random forest (RF) approaches, for five different scenarios (seasons).

load composite, Random Forest, regression tree, lo↗

How Should Machine Learning Be Successfully Used for Wind Speed Vertical Extrapolation?

An accurate characterization of the wind resource available at hub-height is required for an efficient and bankable wind farm project. However, direct measurement of wind speed at the constantly increasing height of the hub of commercial wind turbines is oftentimes challenging and expensive, so that it is common practice to vertically extrapolate the wind resource from lower and more easily accessible levels. Conventional techniques for wind speed vertical extrapolation include the use of a power law and a logarithmic profile. While simple, the limits in accuracy of these methods have been shown in various studies. Recently, machine learning has been proposed as a new method to vertically extrapolate winds. All the published studies on the topic assess the performance of machine learning techniques in vertically extrapolating the wind resource at the same location where the algorithm has been trained. However, in real-world applications, the wind resource is measured at the instrument location, but it then needs to be extrapolated at hub height at the location of the wind turbines within the find farm. To be able to fully recommend the use of machine learning techniques over the simple power law and logarithmic law, the spatial variability of the performance improvements of the machine learning approaches needs to be assessed. Here, we propose a round-robin validation of a machine learning-based method for wind speed extrapolation. We use 20 months of observations at four locations spanning a 100 km wide region at the Southern Great Plains (SGP) atmospheric observatory, in north-central Oklahoma. At each location, we train a random forest to predict 30-min average wind speed at 143 m AGL. We use as input features lidar wind speed at 65 m AGL, time of day, sonic anemometer wind speed at 4 m AGL, turbulent kinetic energy, and Obukhov length. First, we perform a same-site comparison of the performance of the proposed random forest against the conventional techniques for wind speed extrapolation (namely power law and logarithmic profile, with widely accepted stability corrections). We find that the random forest outperforms the power law in vertically extrapolating wind speed in all the considered stability regimes, with a 33% reduction in MAE for stable conditions, and a 31% reduction in unstable conditions. Similar results are found when comparing predictions of extrapolated winds from the logarithmic profile and the random forest with the observed values. Next, we propose a round-robin validation, to use the random forest trained at each site to extrapolate wind speed at the remaining three sites. We find that the performance of the random forest approach degrades when the algorithm is tested at a site different than the training one. However, even under those circumstances, the machine learning-based approach still outperforms the conventional techniques for wind speed extrapolation, with, on average, a reduction in mean absolute error between 15 and 20% over the conventional methods, with the largest benefits obtained under stable conditions.

Monte Carlo↗