Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 235 records · Page 13

A Machine Learning–Based Tire Life Prediction Framework for Increasing Life of Commercial Vehicle Tires

In the commercial freight industry, tire retreading decisions are often conservative due to limited knowledge of a tire’s remaining service life. This practice leads to increased costs and material waste. This paper proposes a machine learning–based approach for estimating tire casing life and retreadability, focusing on usage data rather than wear information. This approach could extend the tire’s lifespan and reduce landfill waste. Data integration from diverse tire casing measurement sources presents challenges, including imbalanced removal data. Our methodology addresses these challenges by using historical inspection, telematics, and finite element modeling (FEM) datasets. We introduce “Tire Casing Energy” as a comprehensive usage input and apply a Variance-Reduction Synthetic Minority Oversampling Technique (VR-SMOTE) for data imbalance rectification. A random forest model is used to estimate the state of the tire casing and the casing removal probability, with Bayesian optimization applied for hyperparameter tuning, enhancing model accuracy. Here, the proposed prediction framework is able to differentiate different truck fleets and tire locations based on their usage parameters. With the aid of this machine learning model, the importance and sensitivity of different tire usage parameters can be obtained, which is beneficial to maximize tire life.

Data balancing↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Identification of mechanisms driving heterogeneous void growth in ductile aluminum

Void growth plays a central role in ductile fracture, yet the specific mechanisms that control this remain obscure. Classical models, such as those proposed by Rice and Tracey in 1969, are able to capture average rates of void growth, but cannot capture the heterogeneity of individual void growth. Building on recent work, the present study employs laboratory-based diffraction contrast tomography and in-situ x-ray computed tomography to investigate the effect of grain structure and other microstructural factors on void growth in an Al-2219 alloy. Crystal plasticity finite element (CP-FE) modeling is used alongside experimental data to evaluate the contributions of local mechanical states, grain orientation, grain size, and neighboring microstructural features. No strong linear relationships are found with any of the considered descriptors and void growth rate. Potential complex nonlinear relationships are explored with the use of a random forest regression model, which identifies initial void volume, void aspect ratio, local normal stress state, local shear stress state, and local equivalent plastic strain (EQPS) as features that most improve void growth rate predictions. The combination of these analyses suggests that these features should be prioritized to improve models of void growth.

Diffraction contrast tomography (DCT)↗

Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem

Coastal terrestrial-aquatic interfaces (TAIs) are crucial contributors to global biogeochemical cycles and carbon exchange. A systematic evaluation of the interaction between coastal catchment properties and carbon dioxide (CO2) emission by soil respiration is significant for assessing carbon dynamics and predicting the future trajectory of atmospheric CO2 concentrations in coastal TAIs. The soil CO2 efflux in these transition zones is however poorly understood due to the high spatiotemporal dynamics of TAIs, as various sub-ecosystems in this region are compressed and expanded by complex influences of tides, changes in river levels, climate, and land use. We focus on the Chesapeake Bay region to (i) investigate the spatial heterogeneity of the coastal ecosystem and identify spatial zones with similar environmental characteristics based on the spatial data layers, including vegetation index (kNDVI), climate, landcover, diversity, topography, soil property, and relative tidal elevation; (ii) understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. Specifically, we employed hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, followed by the determination of main driving factors using Random Forest regression and SHapley Additive exPlanations. Maximum and minimum temperature are the main drivers common to all sub-ecosystems, while each region also has additional unique major drivers that differentiate them from one another. Precipitation exerts an influence on vegetated lands, while soil pH value holds importance specifically in forested lands. In croplands characterized by high clay content and low sand content, the significant role is attributed to bulk density. Wetlands demonstrate the importance of both elevation and sand content, with clay content being more relevant in non-inundated wetlands than in inundated wetlands. The topographic wetness index significantly contributes to the mixed vegetation areas, including shrub, grass, pasture, and forest. Additionally, our research reveals that dense vegetation land covers and urban/developed areas exhibit distinct soil property drivers. Overall, there is no one-size-fits-all approach to modeling carbon fluxes in coastal TAIs, and our study highlights the importance of further research and monitoring practices to improve our understanding of carbon dynamics and promote the sustainable management of coastal TAIs.

54 ENVIRONMENTAL SCIENCES↗

Landsat 8 monitoring of multi-depth suspended sediment concentrations in Lake Erie’s Maumee River using machine learning

Satellite remote sensing has been widely used to map suspended sediment concentration (SSC) in waterbodies. However, due to the complexity of sediment-water interactions, it has been difficult to derive linear and non-linear regression equations to reliably predict SSC, especially when trying to estimate depth of integrated sediment. Herein, this study uses Landsat 8 OLI (Operational Land Imager) sensor to map SSC within the Maumee River in Ohio, USA, at multiple depth intervals (15, 61, 91, and 182 cm). Simple linear least squares regression (LLSR), and three common machine learning models: random forest (RF), support vector regression (SVR), and model averaged neural network (MANN) were used to estimate SSC at the depth intervals. All machine learning models significantly outperformed LLSR while RF performed the best. In both RF and MANN, R2 (coefficient of determination) increases with depth with a maximum R2 of 0.89 and 0.83, respectively, at a depth of 0–182 cm. The results show that machine learning models can implement nonlinear relationships that produce better predictions than traditional linear regression methods in estimating depth integrated SSC, especially when samples are limited.

47 OTHER INSTRUMENTATION↗

A Machine Learning Initializer for Newton-Raphson AC Power Flow Convergence

Power flow computations are fundamental to many power system studies. Obtaining a converged power flow case is not a trivial task especially in large power grids due to the non-linear nature of the power flow equations. One key challenge is that the widely used Newton based power flow methods are sensitive to the initial voltage magnitude and angle estimates, and a bad initial estimate would lead to non-convergence. This paper addresses this challenge by developing a random-forest (RF) machine learning model to provide better initial voltage magnitude and angle estimates towards achieving power flow convergence. This method was implemented on a real ERCOT 6102 bus system under various operating conditions. By providing better Newton-Raphson initialization, the RF model precipitated the solution of 2,106 cases out of 3,899 non-converging dispatches. These cases could not be solved from flat start or by initialization with the voltage solution of a reference case. Finally, results obtained from the RF initializer performed better when compared with DC power flow initialization, Linear regression, and Decision Trees.

random forest↗

Data-Driven Security Assessment of Power Grids Based on Machine Learning Approach: Preprint

Data-driven security assessment provides key indicators on power system stability using simulations on scheduling models, as opposed to dynamic simulations that are more time-consuming. This paper investigates data-driven security assessment of power grids based on machine learning. Multivariate random forest regression is used as the machine learning algorithm due to its high robustness to the input data. Three stability issues are analyzed using the proposed machine learning tool, including transient stability, frequency stability and small signal stability. The estimation values from machine learning tool are compared with those from dynamic simulations. Results show that the proposed machine learning tool can effectively predict the stability margins for the three stability metrics.

14 SOLAR ENERGY↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Coronado Ecological Conservation: Assessing Vegetation Change Due to Border Wall Construction and Shifting Social Trails

Species monitoring is essential for mitigating the impacts of plant invasion, such as radical changes in an area’s ecosystem, degraded soil health, increased wildfire severity, landslides, and increased flooding. For this project, NASA DEVELOP partnered with the National Park Service (NPS) to investigate invasive species in disturbed lands: specifically, areas affected by off-trail travel and U.S.-Mexico border construction activities. The team assessed how construction has impacted the distribution of Lehmann’s lovegrass and Russian thistle invasives throughout Coronado National Memorial, AZ from 1986-2022. Using data from Landsat 5 and 8, Sentinel-2, NAIP, and PlanetScope, the team computed NDVI, NDMI, MSAVI2, EVI, and Tasseled Cap Wetness, Brightness, and Greenness transformations as vegetation health indicators to input into various machine learning algorithms. To minimize noise, the team conducted Principal Component Analysis on vegetation indices and spectral bands before running k-means clustering and random forest classification algorithms. Between all datasets, the team found that the median area fully overtaken by invasive plants was 5.37% of the park’s total area in 2022. The NPS will use end products to help increase restoration efforts in disturbed areas with high concentrations of invasive plants, and this project can serve as a jumping off point for future invasive species monitoring. The NPS’s collection of ground data for 2022-2023, in conjunction with future data collection, will notably improve the accuracy of classification models, leading to more precise monitoring of invasive species spread over time.

Coronado National Memorial↗

Development of a metamodelling framework for building energy models with application to fifth-generation district heating and cooling networks

Fully defined physics-based building energy models can accurately represent building systems; however, generating models based on high-level parameters is time consuming and simulation time of complex models can be slow. This article discusses the development of a Metamodelling Framework to create metamodels from a building energy modelling dataset. The framework generates metamodels using either linear regression, random forests, or support vector regressions. A fifth-generation district heating and cooling system analysis use case was used to motivate the development of the framework. The use case required quick and accurate representations of annual building loads reported hourly. Typical annual building modelling approaches can result in a runtime of 10 min. The metamodels runtime was reduced to less than 10 s to load and run an annual simulation with user-defined covariates. The results of the metamodel performance and an abbreviated topology analysis based on the motivating use case will be presented.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Exploring Sensitivity of ICF Outputs to Design Parameters in Experiments Using Machine Learning

We report building a sustainable burn platform in inertial confinement fusion (ICF) requires an understanding of the complex coupling of physical processes and the effects that key experimental design changes have on implosion performance. While simulation codes are used to model ICF implosions, incomplete physics and the need for approximations deteriorate their predictive capability. Identification of relationships between controllable design inputs and measurable outcomes can help guide the future design of experiments and development of simulation codes, which can potentially improve the accuracy of the computational models used to simulate ICF implosions. In this article, we leverage developments in machine learning (ML) and methods for ML feature importance/sensitivity analysis to identify complex relationships in ways that are difficult to process using expert judgment alone. We present work using random forest (RF) regression for prediction of yield, velocity, and other experimental outcomes given a suite of design parameters, along with an assessment of important relationships and uncertainties in the prediction model. We show that RF models are capable of learning and predicting on ICF experimental data with high accuracy, and we extract feature importance metrics that provide insight into the physical significance of different controllable design inputs for various ICF design configurations. These results can be used to augment expert intuition and simulation results for optimal design of future ICF experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Prediction of Aircraft Estimated Time of Arrival Using A Supervised Learning Approach

We present a novel data-driven approach for prediction of the estimated time of arrival (ETA) of aircraft in the terminal area via the implementation of a Random Forest regression model. The model uses data fused from a number of sources (flight track, weather, flight plan information, etc.) and provides predictions for the remaining flight time for aircraft landing at Dallas/Fort Worth (DFW) International Airport. The predictions are made when the aircraft is at a distance of 200-miles from the airport. The results show that the model is able to predict estimated time of arrival to within ± 5 min for 90% of the flights in the test data with the mean absolute error being lower at 145 seconds. This paper covers the entire pipeline of data collection, preprocessing, setup and training of the ML model, and the results obtained for DFW.

Machine learning↗

Use of Machine Learning to Reduce Uncertainties in Particle Number Concentration and Aerosol Indirect Radiative Forcing Predicted by Climate Models

The radiative forcing of anthropogenic aerosols associated with aerosol–cloud interactions (RF(sub aci)) remains the largest source of uncertainty in climate prediction. The calculation of particle number concentration (PNC), one of the critical parameters affecting RF(sub aci), is generally simplified in climate models. Here we employ outputs from long-term (30-years) simulations of a global size-resolved (sectional) aerosol microphysics model and a machine-learning tool to develop a Random Forest Regression Model (RFRM) for PNC. We have implemented the PNC RFRM in GISS-ModelE2.1 with a mass-based One-Moment Aerosol module, which is one of CMIP6 models. Compared to the default setting, the GISS-ModelE2.1 simulation based on RFRM reduces the changes of cloud droplet number concentration associated with anthropogenic emissions, and decreases the RF(sub aci) from −1.46 W⋅m(exp −2) to −1.11 W⋅m(exp −2). This work highlights a promising approach based on machine learning to reduce uncertainties of climate models in predicting PNC and RF(sub aci) without compromising their computing efficiency.

Radiative forcing↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗

Applications of Machine Learning to Predicting Core-collapse Supernova Explosion Outcomes

Most existing criteria derived from progenitor properties of core-collapse supernovae are not very accurate in predicting explosion outcomes. We present a novel look at identifying the explosion outcome of core-collapse supernovae using a machine-learning approach. Informed by a sample of 100 2D axisymmetric supernova simulations evolved with F ornax , we train and evaluate a random forest classifier as an explosion predictor. Furthermore, we examine physics-based feature sets including the compactness parameter, the Ertl condition, and a newly developed set that characterizes the silicon/oxygen interface. With over 1500 supernovae progenitors from 9-27 M ⊙ , we additionally train an autoencoder to extract physics-agnostic features directly from the progenitor density profiles. We find that the density profiles alone contain meaningful information regarding their explodability. Both the silicon/oxygen and autoencoder features predict the explosion outcome with ≈90% accuracy. In anticipation of much larger multidimensional simulation sets, we identify future directions in which machine-learning applications will be useful beyond the explosion outcome prediction.

79 ASTRONOMY AND ASTROPHYSICS↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predicting oxidation damage in ultra high-temperature borides: A machine learning approach

Ultra-high temperature (UHT) borides are ceramics materials with melting points above 3000 °C for structural applications in extreme environments. However, at temperatures exceeding 1600 °C and under oxidizing conditions, the material suffers from detrimental degradation. Optimized design and performance of diboride materials under such extreme conditions requires filling the missing composition-microstructure-oxidation gap. This study proposes a computational data-driven framework to connect the processing and microstructure of Ultra-high temperature borides with the oxidation damage. Random Forest Regressor (RFR) model is adopted to forecast the oxide scale thickness developed after oxidation testing based on processing variables and microstructural features. The model trained on a dataset consisting of 107 samples of experimental data extracted from the literature aims to predict oxidation damage. With proper data manipulation and fine model tuning, the predictor could forecast the oxide scale thickness of UHT diborides with a Mean Absolute Error of 37.45 μm and an R-square of 0.83. This model could be used as a high-throughput scheme to design and test new UHT diborides materials computationally. Furthermore, a model with larger composition capabilities could also be developed in the future as more experimental data become available.

36 MATERIALS SCIENCE↗

Network-Scale Ubiquitous Volume Estimation Using Tree-Based Ensemble Learning Methods

Currently ubiquitous volume data for roadway networks remains the key missing dimension in traffic operations. Most volume data are average annual daily traffic (AADT) measures derived from the Highway Performance Monitoring System (HPMS). Although methods to factor the AADT to hourly averages for typical day of week exist, actual volume data is limited to a sparse collection of locations in which volumes are continuously recorded. This paper/poster explores the use of state-of-art machine learning techniques to estimate accurate volume measures that span the highway network providing ubiquitous coverage in space, and point-in-time measures for a specific date and time. Three tree-based ensemble learning models, random forest (RF), gradient boost machine (GBM), and extreme gradient boost (XGBoost), were tested for volume estimation by learning from combined dataset of commercial probe data provided by TomTom, the FHWA's Travel Monitoring Analysis System (TMAS) data, and other infrastructure attributes such as number of lanes, speed limit, and weather. The methods were tested on major corridors and freeways in the metropolitan area of Denver. All three machine learning methods were able to provide hourly volume estimates 24 hours a day, 7 days a week, and 365 days a year with around 18% mean absolute error to true volume and about 5% of error with respect to roadway capacity. The low error measures allow the potential application by transportation agencies.

33 ADVANCED PROPULSION SYSTEMS↗