Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 199 records · Page 11

Chicken Production and Human Clinical Escherichia coli Isolates Differ in Their Carriage of Antimicrobial Resistance and Virulence Factors

Contamination of food animal products by Escherichia coli is a leading cause of foodborne disease outbreaks, hospitalizations, and deaths in humans. Chicken is the most consumed meat both in the United States and across the globe according to the U.S. Department of Agriculture. Although E. coli is a ubiquitous commensal bacterium of the guts of humans and animals, its ability to acquire antimicrobial resistance (AMR) genes and virulence factors (VFs) can lead to the emergence of pathogenic strains that are resistant to critically important antibiotics. Thus, it is important to identify the genetic factors that contribute to the virulence and AMR of E. coli. In this study, we performed in-depth genomic evaluation of AMR genes and VFs of E. coli genomes available through the National Antimicrobial Resistance Monitoring System GenomeTrackr database. Our objective was to determine the genetic relatedness of chicken production isolates and human clinical isolates. To achieve this aim, we first developed a massively parallel analytical pipeline (Reads2Resistome) to accurately characterize the resistome of each E. coli genome, including the AMR genes and VFs harbored. We used random forests and hierarchical clustering to show that AMR genes and VFs are sufficient to classify isolates into different pathogenic phylogroups and host origin. We found that the presence of key type III secretion system and AMR genes differentiated human clinical isolates from chicken production isolates. These results further improve our understanding of the interconnected role AMR genes and VFs play in shaping the evolution of pathogenic E. coli strains.

59 BASIC BIOLOGICAL SCIENCES↗

A spatial-statistical investigation of surface expressions associated with cyclic steaming in the Midway-Sunset Oil Field, California

In the Midway Sunset Oil Field in Central California, operators inject steam into the shallow diatomite formation to enhance heavy oil recovery through imbibition, wettability alteration, and viscosity reduction, among other mechanisms. The injected steam, however, does not always remain in the reservoir or return through the wells. In two zones in the study area, the steam comes out at the surface, creating sinkholes, seeps, and steam outlets. These phenomena, called “surface expressions,” pose safety and environmental hazards. Even though these surface expressions are a widespread problem in Central California, they are not well documented and understood. Possible causes of the surface expressions include: high injection pressure, structurally controlled flow patterns, leakage of steam through old improperly abandoned wells, high injection volumes, or flow along naturally occurring faults, among other possible factors. This work examines attributes of the zones with surface expressions in order to determine factors that may contribute to their occurrence. Spatial statistical analysis using logistic regression, random forests, and classification trees is used to explore the relationship between the surface expressions and geological and production-related attributes. The results point to a significant spatial correlation between the surface expressions and two predictors: concentration of plugged wells and geologic seal thickness. The results guide follow-up studies to further investigate the role of well abandonment and seal thickness in the occurrence of surface expressions.

02 PETROLEUM↗

Voltage Estimation in Low-Voltage Distribution Grids with Distributed Energy Resources

Present distribution grids generally have limited sensing capabilities and are therefore characterized by low observability. Improved observability is a prerequisite for increasing the hosting capacity of distributed energy resources such as solar photovoltaics (PV) in distribution grids. In this context, this paper presents learning-aided low-voltage estimation using untapped but readily available and widely distributed sensors from cable television (CATV) networks. The cable broadband sensors offer timely local voltage magnitude sensing with 5-minute resolution and can provide an order of magnitude more data on the time-varying state of a secondary distribution system than currently deployed utility sensors. The proposed solution incorporates voltage readings from neighboring CATV sensors, taking into account spatio-temporal aspects of the observations, and estimates single-phase voltage magnitudes at all non-monitored low-voltage buses using random forests. The effectiveness of the proposed approach was demonstrated using a multi-phase 1572-bus feeder from the SMART-DS data set for two case studies passive distribution feeder (without PV) and active distribution feeder (with PV). The analysis was conducted on simulated data, and the results show voltage estimates with a high degree of accuracy, even at extremely low percentages of observable nodes.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A comparative study of machine learning models for predicting the state of reactive mixing

Mixing phenomena are important mechanisms controlling flow, species transport, and reaction processes in fluids and porous media. Accurate predictions of reactive mixing are critical for many Earth and environmental science problems such as contaminant fate and remediation, macroalgae growth, and plankton biomass evolution. Here, to investigate the evolution of mixing dynamics under different scenarios (e.g., anisotropy, fluctuating velocity fields), a finite-element-based numerical model was built to solve the fast, irreversible bimolecular reaction-diffusion equations to simulate a range of reactive-mixing scenarios. A total of 2,315 simulations were performed using different sets of model input parameters comprising various spatial scales of vortex structures in the velocity field, time-scales associated with velocity oscillations, the perturbation parameter for the vortex-based velocity, anisotropic dispersion contrast (i.e., ratio of longitudinal-to-transverse dispersion), and molecular diffusion. The outputs comprised concentration profiles of reactants and products. The inputs to and outputs from these simulations were concatenated into feature and label matrices, respectively, to train 20 different machine learning (ML) models intended to emulate system behavior. These 20 ML emulators, based on linear methods, Bayesian methods, ensemble learning methods, and multilayer perceptrons (MLPs), were trained to classify the state of mixing and predict three quantities of interest (QoIs) characterizing species production, decay (i.e., average concentration, square of average concentration), and degree of mixing (i.e., variances of species concentration). Unsurprisingly, linear classifiers and regressors failed to reproduce the QoIs; however, ensemble methods (classifiers and regressors) and the MLP model accurately classified the state of reactive mixing and the QoIs. Among ensemble methods, random forest and decision-tree-based AdaBoost faithfully predicted the QoIs. At run time, trained ML emulators produced results times faster than the finite-element simulations. Due to their low computational expense and high accuracy, ensemble and MLP models are excellent emulators for these numerical simulations and great utilities in uncertainty quantification exercises, which can require 1,000s of forward model runs.

97 MATHEMATICS AND COMPUTING↗

A learning-augmented approach for AC optimal power flow

Because of the high nonlinearity of AC optimal power flow (OPF), numerous efforts have been made in recent decades to find efficient methods. Machine learning (ML) has proven to significantly reduce the computational costs in many real-world problems. Thus, this paper develops a learning-augmented method for solving AC OPF, which integrates both power network equations and ML to yield near-optimal solutions. More specifically, ML models are developed to first predict bus voltage magnitudes and angles. Then, physics-based network equations are employed to calculate the power injection at different buses. Three ML algorithms, i.e., random forest, multi-target decision tree, and extreme learning machine, are explored and compared. To evaluate the efficiency of the proposed learning-augmented AC OPF solver, the MATPOWER Interior Point Solver is adopted as a baseline. Case studies on both 500-bus and 4918-bus test networks show that the proposed learning-augmented method has reduced the computational time by 15–100 times depending on the network size with a minimal loss in optimality.

42 ENGINEERING↗

A Machine Learning–Based Tire Life Prediction Framework for Increasing Life of Commercial Vehicle Tires

In the commercial freight industry, tire retreading decisions are often conservative due to limited knowledge of a tire’s remaining service life. This practice leads to increased costs and material waste. This paper proposes a machine learning–based approach for estimating tire casing life and retreadability, focusing on usage data rather than wear information. This approach could extend the tire’s lifespan and reduce landfill waste. Data integration from diverse tire casing measurement sources presents challenges, including imbalanced removal data. Our methodology addresses these challenges by using historical inspection, telematics, and finite element modeling (FEM) datasets. We introduce “Tire Casing Energy” as a comprehensive usage input and apply a Variance-Reduction Synthetic Minority Oversampling Technique (VR-SMOTE) for data imbalance rectification. A random forest model is used to estimate the state of the tire casing and the casing removal probability, with Bayesian optimization applied for hyperparameter tuning, enhancing model accuracy. Here, the proposed prediction framework is able to differentiate different truck fleets and tire locations based on their usage parameters. With the aid of this machine learning model, the importance and sensitivity of different tire usage parameters can be obtained, which is beneficial to maximize tire life.

Data balancing↗

Addressing bias in bagging and boosting regression models

As artificial intelligence (AI) becomes widespread, there is increasing attention on investigating bias in machine learning (ML) models. Previous research concentrated on classification problems, with little emphasis on regression models. This paper presents an easy-to-apply and effective methodology for mitigating bias in bagging and boosting regression models, that is also applicable to any model trained through minimizing a differentiable loss function. Our methodology measures bias rigorously and extends the ML model's loss function with a regularization term to penalize high correlations between model errors and protected attributes. We applied our approach to three popular tree-based ensemble models: a random forest model (RF), a gradient-boosted model (GBT), and an extreme gradient boosting model (XGBoost). We implemented our methodology on a case study for predicting road-level traffic volume, where RF, GBT, and XGBoost models were shown to have high accuracy. Despite high accuracy, the ML models were shown to perform poorly on roads in minority-populated areas. Our bias mitigation approach reduced minority-related bias by over 50%.

97 MATHEMATICS AND COMPUTING↗

Identification of mechanisms driving heterogeneous void growth in ductile aluminum

Void growth plays a central role in ductile fracture, yet the specific mechanisms that control this remain obscure. Classical models, such as those proposed by Rice and Tracey in 1969, are able to capture average rates of void growth, but cannot capture the heterogeneity of individual void growth. Building on recent work, the present study employs laboratory-based diffraction contrast tomography and in-situ x-ray computed tomography to investigate the effect of grain structure and other microstructural factors on void growth in an Al-2219 alloy. Crystal plasticity finite element (CP-FE) modeling is used alongside experimental data to evaluate the contributions of local mechanical states, grain orientation, grain size, and neighboring microstructural features. No strong linear relationships are found with any of the considered descriptors and void growth rate. Potential complex nonlinear relationships are explored with the use of a random forest regression model, which identifies initial void volume, void aspect ratio, local normal stress state, local shear stress state, and local equivalent plastic strain (EQPS) as features that most improve void growth rate predictions. The combination of these analyses suggests that these features should be prioritized to improve models of void growth.

Diffraction contrast tomography (DCT)↗

Effects of spatial variability in vegetation phenology, climate, landcover, biodiversity, topography, and soil property on soil respiration across a coastal ecosystem

Coastal terrestrial-aquatic interfaces (TAIs) are crucial contributors to global biogeochemical cycles and carbon exchange. A systematic evaluation of the interaction between coastal catchment properties and carbon dioxide (CO2) emission by soil respiration is significant for assessing carbon dynamics and predicting the future trajectory of atmospheric CO2 concentrations in coastal TAIs. The soil CO2 efflux in these transition zones is however poorly understood due to the high spatiotemporal dynamics of TAIs, as various sub-ecosystems in this region are compressed and expanded by complex influences of tides, changes in river levels, climate, and land use. We focus on the Chesapeake Bay region to (i) investigate the spatial heterogeneity of the coastal ecosystem and identify spatial zones with similar environmental characteristics based on the spatial data layers, including vegetation index (kNDVI), climate, landcover, diversity, topography, soil property, and relative tidal elevation; (ii) understand the primary driving factors affecting soil respiration within sub-ecosystems of the coastal ecosystem. Specifically, we employed hierarchical clustering analysis to identify spatial regions with distinct environmental characteristics, followed by the determination of main driving factors using Random Forest regression and SHapley Additive exPlanations. Maximum and minimum temperature are the main drivers common to all sub-ecosystems, while each region also has additional unique major drivers that differentiate them from one another. Precipitation exerts an influence on vegetated lands, while soil pH value holds importance specifically in forested lands. In croplands characterized by high clay content and low sand content, the significant role is attributed to bulk density. Wetlands demonstrate the importance of both elevation and sand content, with clay content being more relevant in non-inundated wetlands than in inundated wetlands. The topographic wetness index significantly contributes to the mixed vegetation areas, including shrub, grass, pasture, and forest. Additionally, our research reveals that dense vegetation land covers and urban/developed areas exhibit distinct soil property drivers. Overall, there is no one-size-fits-all approach to modeling carbon fluxes in coastal TAIs, and our study highlights the importance of further research and monitoring practices to improve our understanding of carbon dynamics and promote the sustainable management of coastal TAIs.

54 ENVIRONMENTAL SCIENCES↗

Landsat 8 monitoring of multi-depth suspended sediment concentrations in Lake Erie’s Maumee River using machine learning

Satellite remote sensing has been widely used to map suspended sediment concentration (SSC) in waterbodies. However, due to the complexity of sediment-water interactions, it has been difficult to derive linear and non-linear regression equations to reliably predict SSC, especially when trying to estimate depth of integrated sediment. Herein, this study uses Landsat 8 OLI (Operational Land Imager) sensor to map SSC within the Maumee River in Ohio, USA, at multiple depth intervals (15, 61, 91, and 182 cm). Simple linear least squares regression (LLSR), and three common machine learning models: random forest (RF), support vector regression (SVR), and model averaged neural network (MANN) were used to estimate SSC at the depth intervals. All machine learning models significantly outperformed LLSR while RF performed the best. In both RF and MANN, R2 (coefficient of determination) increases with depth with a maximum R2 of 0.89 and 0.83, respectively, at a depth of 0–182 cm. The results show that machine learning models can implement nonlinear relationships that produce better predictions than traditional linear regression methods in estimating depth integrated SSC, especially when samples are limited.

47 OTHER INSTRUMENTATION↗

A Machine Learning Initializer for Newton-Raphson AC Power Flow Convergence

Power flow computations are fundamental to many power system studies. Obtaining a converged power flow case is not a trivial task especially in large power grids due to the non-linear nature of the power flow equations. One key challenge is that the widely used Newton based power flow methods are sensitive to the initial voltage magnitude and angle estimates, and a bad initial estimate would lead to non-convergence. This paper addresses this challenge by developing a random-forest (RF) machine learning model to provide better initial voltage magnitude and angle estimates towards achieving power flow convergence. This method was implemented on a real ERCOT 6102 bus system under various operating conditions. By providing better Newton-Raphson initialization, the RF model precipitated the solution of 2,106 cases out of 3,899 non-converging dispatches. These cases could not be solved from flat start or by initialization with the voltage solution of a reference case. Finally, results obtained from the RF initializer performed better when compared with DC power flow initialization, Linear regression, and Decision Trees.

random forest↗

Data-Driven Security Assessment of Power Grids Based on Machine Learning Approach: Preprint

Data-driven security assessment provides key indicators on power system stability using simulations on scheduling models, as opposed to dynamic simulations that are more time-consuming. This paper investigates data-driven security assessment of power grids based on machine learning. Multivariate random forest regression is used as the machine learning algorithm due to its high robustness to the input data. Three stability issues are analyzed using the proposed machine learning tool, including transient stability, frequency stability and small signal stability. The estimation values from machine learning tool are compared with those from dynamic simulations. Results show that the proposed machine learning tool can effectively predict the stability margins for the three stability metrics.

14 SOLAR ENERGY↗

Machine Learning Based Resilience Testing of an Address Randomization Cyber Defense

Moving target defenses (MTDs) are widely used as an active defense strategy for thwarting cyberattacks on cyber-physical systems by increasing diversity of software and network paths. Recently, machine Learning (ML) and deep Learning (DL) models have been demonstrated to defeat some of the cyber defenses by learning attack detection patterns and defense strategies. It raises concerns about the susceptibility of MTD to ML and DL methods. Here, in this article, we analyze the effectiveness of ML and DL models when it comes to deciphering MTD methods and ultimately evade MTD-based protections in real-time systems. Specifically, we consider a MTD algorithm that periodically randomizes address assignments within the MIL-STD-1553 protocol—a military standard serial data bus. Two ML and DL-based tasks are performed on MIL-STD-1553 protocol to measure the effectiveness of the learning models in deciphering the MTD algorithm: 1) determining whether there is an address assignments change i.e., whether the given system employs a MTD protocol and if it does 2) predicting the future address assignments. The supervised learning models (random forest and k-nearest neighbors) effectively detected the address assignment changes and classified whether the given system is equipped with a specified MTD protocol. On the other hand, the unsupervised learning model (K-means) was significantly less effective. The DL model (long short-term memory) was able to predict the future addresses with varied effectiveness based on MTD algorithm's settings.

45 MILITARY TECHNOLOGY, WEAPONRY, AND NATIONAL DEF↗

Development of a metamodelling framework for building energy models with application to fifth-generation district heating and cooling networks

Fully defined physics-based building energy models can accurately represent building systems; however, generating models based on high-level parameters is time consuming and simulation time of complex models can be slow. This article discusses the development of a Metamodelling Framework to create metamodels from a building energy modelling dataset. The framework generates metamodels using either linear regression, random forests, or support vector regressions. A fifth-generation district heating and cooling system analysis use case was used to motivate the development of the framework. The use case required quick and accurate representations of annual building loads reported hourly. Typical annual building modelling approaches can result in a runtime of 10 min. The metamodels runtime was reduced to less than 10 s to load and run an annual simulation with user-defined covariates. The results of the metamodel performance and an abbreviated topology analysis based on the motivating use case will be presented.

32 ENERGY CONSERVATION, CONSUMPTION, AND UTILIZATI↗

Exploring Sensitivity of ICF Outputs to Design Parameters in Experiments Using Machine Learning

We report building a sustainable burn platform in inertial confinement fusion (ICF) requires an understanding of the complex coupling of physical processes and the effects that key experimental design changes have on implosion performance. While simulation codes are used to model ICF implosions, incomplete physics and the need for approximations deteriorate their predictive capability. Identification of relationships between controllable design inputs and measurable outcomes can help guide the future design of experiments and development of simulation codes, which can potentially improve the accuracy of the computational models used to simulate ICF implosions. In this article, we leverage developments in machine learning (ML) and methods for ML feature importance/sensitivity analysis to identify complex relationships in ways that are difficult to process using expert judgment alone. We present work using random forest (RF) regression for prediction of yield, velocity, and other experimental outcomes given a suite of design parameters, along with an assessment of important relationships and uncertainties in the prediction model. We show that RF models are capable of learning and predicting on ICF experimental data with high accuracy, and we extract feature importance metrics that provide insight into the physical significance of different controllable design inputs for various ICF design configurations. These results can be used to augment expert intuition and simulation results for optimal design of future ICF experiments.

70 PLASMA PHYSICS AND FUSION TECHNOLOGY↗

Predicting Elastic Constants of Refractory Complex Concentrated Alloys Using Machine Learning Approach

Refractory complex concentrated alloys (RCCAs) have drawn increasing attention recently owing to their balanced mechanical properties, including excellent creep resistance, ductility, and oxidation resistance. The mechanical and thermal properties of RCCAs are directly linked with the elastic constants. However, it is time consuming and expensive to obtain the elastic constants of RCCAs with conventional trial-and-error experiments. The elastic constants of RCCAs are predicted using a combination of density functional theory simulation data and machine learning (ML) algorithms in this study. The elastic constants of several RCCAs are predicted using the random forest regressor, gradient boosting regressor (GBR), and XGBoost regression models. Based on performance metrics R-squared, mean average error and root mean square error, the GBR model was found to be most promising in predicting the elastic constant of RCCAs among the three ML models. Additionally, GBR model accuracy was verified using the other four RHEAs dataset which was never seen by the GBR model, and reasonable agreements between ML prediction and available results were found. The present findings show that the GBR model can be used to predict the elastic constant of new RHEAs more accurately without performing any expensive computational and experimental work.

36 MATERIALS SCIENCE↗

Applications of Machine Learning to Predicting Core-collapse Supernova Explosion Outcomes

Most existing criteria derived from progenitor properties of core-collapse supernovae are not very accurate in predicting explosion outcomes. We present a novel look at identifying the explosion outcome of core-collapse supernovae using a machine-learning approach. Informed by a sample of 100 2D axisymmetric supernova simulations evolved with F ornax , we train and evaluate a random forest classifier as an explosion predictor. Furthermore, we examine physics-based feature sets including the compactness parameter, the Ertl condition, and a newly developed set that characterizes the silicon/oxygen interface. With over 1500 supernovae progenitors from 9-27 M ⊙ , we additionally train an autoencoder to extract physics-agnostic features directly from the progenitor density profiles. We find that the density profiles alone contain meaningful information regarding their explodability. Both the silicon/oxygen and autoencoder features predict the explosion outcome with ≈90% accuracy. In anticipation of much larger multidimensional simulation sets, we identify future directions in which machine-learning applications will be useful beyond the explosion outcome prediction.

79 ASTRONOMY AND ASTROPHYSICS↗

Tree-Based Ensemble Learning Models for Wall Temperature Predictions in Post-Critical Heat Flux Flow Regimes at Subcooled and Low-Quality Conditions

Accurately predicting post-critical heat flux (CHF) heat transfer is an important but challenging task in water-cooled reactor design and safety analysis. Although numerous heat transfer correlations have been developed to predict post-CHF heat transfer, these correlations are only applicable to relatively narrow ranges of flow conditions due to the complex physical nature of the post-CHF heat transfer regimes. In this paper, a large quantity of experimental data is collected and summarized from the literature for steady-state subcooled and low-quality film boiling regimes with water as the working fluid in vertical tubular test sections. In addition, a low-quality water film boiling (LWFB) database is consolidated with a total of 22,813 experimental data points, which cover a wide flow range of the system pressure from 0.1 to 9.0 MPa, mass flux from 25 to 2750 kg/m 2 s, and inlet subcooling from 1 to 70 °C. Two machine learning (ML) models, based on random forest (RF) and gradient boosted decision tree (GBDT), are trained and validated to predict wall temperatures in post-CHF flow regimes. The trained ML models demonstrate significantly improved accuracies compared to conventional empirical correlations. To further evaluate the performance of these two ML models from a statistical perspective, three criteria are investigated and three metrics are calculated to quantitatively assess the accuracy of these two ML models. For the full LWFB database, the root-mean-square errors between the measured and predicted wall temperatures by the GBDT and RF models are 5.7% and 6.2%, respectively, confirming the accuracy of the two ML models.

22 GENERAL STUDIES OF NUCLEAR REACTORS↗