Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 217 records · Page 12

A Hierarchical Feature-Based Methodology to Perform Cervical Cancer Classification

Prevention of cervical cancer could be performed using Pap smear image analysis. This test screens pre-neoplastic changes in the cervical epithelial cells; accurate screening can reduce deaths caused by the disease. Pap smear test analysis is exhaustive and repetitive work performed visually by a cytopathologist. This article proposes a workload-reducing algorithm for cervical cancer detection based on analysis of cell nuclei features within Pap smear images. We investigate eight traditional machine learning methods to perform a hierarchical classification. We propose a hierarchical classification methodology for computer-aided screening of cell lesions, which can recommend fields of view from the microscopy image based on the nuclei detection of cervical cells. We evaluate the performance of several algorithms against the Herlev and CRIC databases, using a varying number of classes during image classification. Results indicate that the hierarchical classification performed best when using Random Forest as the key classifier, particularly when compared with decision trees, k-NN, and the Ridge methods.

60 APPLIED LIFE SCIENCES↗

Vehicle Position Detection Based on Machine Learning Algorithms in Dynamic Wireless Charging

Dynamic wireless charging (DWC) has emerged as a viable approach to mitigate range anxiety by ensuring continuous and uninterrupted charging for electric vehicles in motion. DWC systems rely on the length of the transmitter, which can be categorized into long-track transmitters and segmented coil arrays. The segmented coil array, favored for its heightened efficiency and reduced electromagnetic interference, stands out as the preferred option. However, in such DWC systems, the need arises to detect the vehicle’s position, specifically to activate the transmitter coils aligned with the receiver pad and de-energize uncoupled transmitter coils. This paper introduces various machine learning algorithms for precise vehicle position determination, accommodating diverse ground clearances of electric vehicles and various speeds. Through testing eight different machine learning algorithms and comparing the results, the random forest algorithm emerged as superior, displaying the lowest error in predicting the actual position.

47 OTHER INSTRUMENTATION↗

Post-Event Fault Identification with Machine Learning for Protection System Validation

Power system protection devices have transitioned over the past few decades from mechanical to analog devices, then to solid state and finally digital. Relays and their associated critical network of equipment have significantly increased in complexity. Even internally, relays have gained significant intricacy, with relatively simple overcurrent or differential functions now being assisted by a myriad of other functions. This is necessary as the grid becomes more complex, but it brings increased difficulty in monitoring and upkeep. Misoperation caused by accidental improper relay settings or deliberate malicious actions is a constant challenge faced by all utilities. These improper settings can be difficult to identify and may require exhaustive post-mortem analysis, typically after a major outage event has already occurred. A mechanism is needed for monitoring the behavior of protection systems to validate that their performance falls within expectations. Relays that fail to isolate a fault or trip when there is no system disturbance can be flagged for settings review in situations where this behavior may not have been noticed due to manual restoration or backup protection operations. This work presents a concept for a machine learning (ML) system capable of validating the performance of protection systems by identifying fault events and characterizing protection system responses based solely on available current and voltage measurements. As a first step in its development, an experimental dataset is generated, and a random forest model is implemented with high accuracy in distinguishing four power system scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Protection System Validation Using Post-Event Anomaly Classification with Machine Learning

Power system protection devices have transitioned over the past few decades from mechanical to analog devices, then to solid state and finally digital. Relays and their associated critical network of equipment have significantly increased in complexity. Even internally, relays have gained significant intricacy, with relatively simple overcurrent or differential functions now being assisted by a myriad of other functions. This is necessary as the grid becomes more complex, but it brings increased difficulty in monitoring and upkeep. Misoperation caused by improper relay settings or malicious actions is a constant challenge faced by all utilities. These improper settings can be difficult to identify and may require exhaustive post-mortem analysis, typically after a major outage event has already occurred. A mechanism is needed for monitoring the behavior of protection systems to validate that they act and perform as expected. This work presents a concept for a machine learning (ML) system capable of validating the performance of protection systems by classifying anomalous events and characterizing protection system responses based solely on available current and voltage measurements. As a first step in its development, an experimental dataset is generated, and a random forest model is implemented with high accuracy in distinguishing four power system scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

A machine learning-based fast frequency response control for a VSC-HVDC system

An HVDC system can realize a very fast frequency response to the disturbed system under a contingency because its active power control is decoupled from the frequency deviation. However, most of existing HVDC frequency control strategies are coupled with system primary frequency control and secondary frequency control. Since the traditional system frequency control is dominated by the thermal generators, the advantage of the fast response of the HVDC system is not made fully used. The development of a frequency response estimation based on a machine learning algorithm provides another approach to improve the frequency response capability of the HVDC system. Different from other frequency deviation tracking strategies, a machine learning based HVDC frequency response control can directly increase the power flow of a HVDC system by estimation of the system generator or load lost. In this paper, a fast frequency response control using a HVDC system for a large power system disturbance based on the multivariate random forest regression (MRFR) algorithm is proposed. The simulation is carried out with an integrated power system model based on the North American interconnections. The simulation results indicate that the proposed MRFR based frequency response control can significantly improve the frequency low point during an event, while stabilizing the frequency in advance.

42 ENGINEERING↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗

Identifying Transient Candidates in the Dark Energy Survey Using Convolutional Neural Networks

The ability to discover new transient candidates via image differencing without direct human intervention is an important task in observational astronomy. For these kind of image classification problems, machine learning techniques such as Convolutional Neural Networks (CNNs) have shown remarkable success. In this work, we present the results of an automated transient candidate identification on images with CNNs for an extant data set from the Dark Energy Survey Supernova program, whose main focus was on using Type Ia supernovae for cosmology. By performing an architecture search of CNNs, we identify networks that efficiently select non-artifacts (e.g., supernovae, variable stars, AGN, etc.) from artifacts (image defects, mis-subtractions, etc.), achieving the efficiency of previous work performed with random Forests, without the need to expend any effort in feature identification. The CNNs also help us identify a subset of mislabeled images. Performing a relabeling of the images in this subset, the resulting classification with CNNs is significantly better than previous results, lowering the false positive rate by 27% at a fixed missed detection rate of 0.05.

79 ASTRONOMY AND ASTROPHYSICS↗

Improved quality metrics for association and reproducibility in chromatin accessibility data using mutual information

Correlation metrics are widely utilized in genomics analysis and often implemented with little regard to assumptions of normality, homoscedasticity, and independence of values. This is especially true when comparing values between replicated sequencing experiments that probe chromatin accessibility, such as assays for transposase-accessible chromatin via sequencing (ATAC-seq). Such data can possess several regions across the human genome with little to no sequencing depth and are thus non-normal with a large portion of zero values. Despite distributed use in the epigenomics field, few studies have evaluated and benchmarked how correlation and association statistics behave across ATAC-seq experiments with known differences or the effects of removing specific outliers from the data. Here, we developed a computational simulation of ATAC-seq data to elucidate the behavior of correlation statistics and to compare their accuracy under set conditions of reproducibility. Using these simulations, we monitored the behavior of several correlation statistics, including the Pearson’s R and Spearman’s ρ coefficients as well as Kendall’s τ and Top–Down correlation. We also test the behavior of association measures, including the coefficient of determination R 2 , Kendall’s W, and normalized mutual information. Our experiments reveal an insensitivity of most statistics, including Spear man’s ρ, Kendall’s τ, and Kendall’s W, to increasing differences between simulated ATAC-seq replicates. The removal of co-zeros (regions lacking mapped sequenced reads) between simulated experiments greatly improves the estimates of correlation and association. After removing co-zeros, the R 2 coefficient and normalized mutual information display the best performance, having a closer one-to-one relationship with the known portion of shared, enhanced loci between simulated replicates. When comparing values between experimental ATAC-seq data using a random forest model, mutual information best predicts ATAC-seq replicate relationships. Collectively, this study demonstrates how measures of correlation and association can behave in epigenomics experiments. We provide improved strategies for quantifying relationships in these increasingly prevalent and important chromatin accessibility assays.

59 BASIC BIOLOGICAL SCIENCES↗

Spatial Mapping of Riverbed Grain-Size Distribution Using Machine Learning

Recent alluvial sediments in riverbeds play a significant role in controlling hydrologic exchange flows (HEFs) in river systems. The alluvial layer is usually associated with strong heterogeneity in physical properties (e.g., permeability and hydraulic conductivity), which affects local HEFs and therefore biogeochemical processes. The spatial distribution of these physical properties needs to be determined to inform the numerical models used to reveal the realistic hydro-biogeochemical behaviors. Such information can be obtained based on the intrinsic link between sediment grain-size distribution and hydraulic properties where sediment texture information is available. However, grain-size measurements are usually spatially sparse and do not have adequate coverage and resolution, particularly for a relatively large domain such as the Hanford Reach of the Columbia River. In this paper, we adopted machine learning (ML) approaches for categorizing and mapping the spatial distributions of riverbed substrate grain size and filling in missing areas of substrate data using the ML models along the reach. Such ML models for substrate size mapping were trained at 13,372 locations using measured substrate sizes along with observed and simulated attributes, including bathymetric attributes (e.g., elevation, slope, and aspect ratio) from LIDAR and bathymetric surveys, and hydrodynamic properties (e.g., water depth, velocity, shear stress, and their statistical moments). An ensemble bagging-based ML technique, Random Forest, was adopted to identify the most influential factors as predictors to develop the predictive models with over-fitting issues addressed. The models were evaluated with respect to each individual substrate size class and the lumped group, and then used to generate the final substrate size maps covering all the grid cells in the numerical modeling domain.

54 ENVIRONMENTAL SCIENCES↗

Yield strength prediction of high-entropy alloys using machine learning

Yield strength at high temperature is an important parameter in the design and application of high entropy alloys (HEAs). However, the experimental measurement of yield strength at high temperature is quite costly, complicated, and time-consuming. Therefore, it is essential to identify and apply a robust method for the accurate prediction of yield strength at high temperature from the available experimental and simulation data. In this study, for the first time, a machine learning (ML) method based on the regression technique of random forest (RF) regressor is used to predict the yield strength of HEAs at the desired temperature. Further, the yield strengths of MoNbTaTiW and HfMoNbTaTiZr at 800 °C and 1200 °C, are predicted using the RF regressor model. We find that the results are consistent with the experimental reports, showing that the RF regressor model predicts the yield strength of HEAs at the desired temperatures with high accuracy.

36 MATERIALS SCIENCE↗

Soil Origin and Plant Genotype Modulate Switchgrass Aboveground Productivity and Root Microbiome Assembly

Switchgrass (Panicum virgatum) is a model perennial grass for bioenergy production that can be productive in agricultural lands that are not suitable for food production. There is growing interest in whether its associated microbiome may be adaptive in low- or no-input cultivation systems. However, the relative impact of plant genotype and soil factors on plant microbiome and biomass are a challenge to decouple. To address this, a common garden greenhouse experiment was carried out using six common switchgrass genotypes, which were each grown in four different marginal soils collected from long-term bioenergy research sites in Michigan and Wisconsin. We characterized the fungal and bacterial root communities with high-throughput amplicon sequencing of the ITS and 16S rDNA markers, and collected phenological plant traits during plant growth, as well as soil chemical traits. At harvest, we measured the total plant aerial dry biomass. Significant differences in richness and Shannon diversity across soils but not between plant genotypes were found. Generalized linear models showed an interaction between soil and genotype for fungal richness but not for bacterial richness. Community structure was also strongly shaped by soil origin and soil origin × plant genotype interactions. Overall, plant genotype effects were significant but low. Random Forest models indicate that important factors impacting switchgrass biomass included NO 3 – , Ca 2+ , PO 4 3– , and microbial biodiversity. We identified 54 fungal and 52 bacterial predictors of plant aerial biomass, which included several operational taxonomic units belonging to Glomeraceae and Rhizobiaceae, fungal and bacterial lineages that are involved in provisioning nutrients to plants.

plant biomass↗

Moisture availability mediates the relationship between terrestrial gross primary production and solar-induced chlorophyll fluorescence: Insights from global-scale variations

Effective use of solar-induced chlorophyll fluorescence (SIF) to estimate and monitor gross primary production (GPP) in terrestrial ecosystems requires a comprehensive understanding and quantification of the relationship between SIF and GPP. To date, this understanding is incomplete and somewhat controversial in the literature. Here we derived the GPP/SIF ratio from multiple data sources as a diagnostic metric to explore its global-scale patterns of spatial variation and potential climatic dependence. We found that the growing season GPP/SIF ratio varied substantially across global land surfaces, with the highest ratios consistently found in boreal regions. Spatial variation in GPP/SIF was strongly modulated by climate variables. The most striking pattern was a consistent decrease in GPP/SIF from cold-and-wet climates to hot-and-dry climates. We propose that the reduction in GPP/SIF with decreasing moisture availability may be related to stomatal responses to aridity. Furthermore, we show that GPP/SIF can be empirically modeled from climate variables using a machine learning (random forest) framework, which can improve the modeling of ecosystem production and quantify its uncertainty in global terrestrial biosphere models. Finally, our results point to the need for targeted field and experimental studies to better understand the patterns observed and to improve the modeling of the relationship between SIF and GPP over broad scales.

59 BASIC BIOLOGICAL SCIENCES↗

Chicken Production and Human Clinical Escherichia coli Isolates Differ in Their Carriage of Antimicrobial Resistance and Virulence Factors

Contamination of food animal products by Escherichia coli is a leading cause of foodborne disease outbreaks, hospitalizations, and deaths in humans. Chicken is the most consumed meat both in the United States and across the globe according to the U.S. Department of Agriculture. Although E. coli is a ubiquitous commensal bacterium of the guts of humans and animals, its ability to acquire antimicrobial resistance (AMR) genes and virulence factors (VFs) can lead to the emergence of pathogenic strains that are resistant to critically important antibiotics. Thus, it is important to identify the genetic factors that contribute to the virulence and AMR of E. coli. In this study, we performed in-depth genomic evaluation of AMR genes and VFs of E. coli genomes available through the National Antimicrobial Resistance Monitoring System GenomeTrackr database. Our objective was to determine the genetic relatedness of chicken production isolates and human clinical isolates. To achieve this aim, we first developed a massively parallel analytical pipeline (Reads2Resistome) to accurately characterize the resistome of each E. coli genome, including the AMR genes and VFs harbored. We used random forests and hierarchical clustering to show that AMR genes and VFs are sufficient to classify isolates into different pathogenic phylogroups and host origin. We found that the presence of key type III secretion system and AMR genes differentiated human clinical isolates from chicken production isolates. These results further improve our understanding of the interconnected role AMR genes and VFs play in shaping the evolution of pathogenic E. coli strains.

59 BASIC BIOLOGICAL SCIENCES↗

A spatial-statistical investigation of surface expressions associated with cyclic steaming in the Midway-Sunset Oil Field, California

In the Midway Sunset Oil Field in Central California, operators inject steam into the shallow diatomite formation to enhance heavy oil recovery through imbibition, wettability alteration, and viscosity reduction, among other mechanisms. The injected steam, however, does not always remain in the reservoir or return through the wells. In two zones in the study area, the steam comes out at the surface, creating sinkholes, seeps, and steam outlets. These phenomena, called “surface expressions,” pose safety and environmental hazards. Even though these surface expressions are a widespread problem in Central California, they are not well documented and understood. Possible causes of the surface expressions include: high injection pressure, structurally controlled flow patterns, leakage of steam through old improperly abandoned wells, high injection volumes, or flow along naturally occurring faults, among other possible factors. This work examines attributes of the zones with surface expressions in order to determine factors that may contribute to their occurrence. Spatial statistical analysis using logistic regression, random forests, and classification trees is used to explore the relationship between the surface expressions and geological and production-related attributes. The results point to a significant spatial correlation between the surface expressions and two predictors: concentration of plugged wells and geologic seal thickness. The results guide follow-up studies to further investigate the role of well abandonment and seal thickness in the occurrence of surface expressions.

02 PETROLEUM↗

Voltage Estimation in Low-Voltage Distribution Grids with Distributed Energy Resources

Present distribution grids generally have limited sensing capabilities and are therefore characterized by low observability. Improved observability is a prerequisite for increasing the hosting capacity of distributed energy resources such as solar photovoltaics (PV) in distribution grids. In this context, this paper presents learning-aided low-voltage estimation using untapped but readily available and widely distributed sensors from cable television (CATV) networks. The cable broadband sensors offer timely local voltage magnitude sensing with 5-minute resolution and can provide an order of magnitude more data on the time-varying state of a secondary distribution system than currently deployed utility sensors. The proposed solution incorporates voltage readings from neighboring CATV sensors, taking into account spatio-temporal aspects of the observations, and estimates single-phase voltage magnitudes at all non-monitored low-voltage buses using random forests. The effectiveness of the proposed approach was demonstrated using a multi-phase 1572-bus feeder from the SMART-DS data set for two case studies passive distribution feeder (without PV) and active distribution feeder (with PV). The analysis was conducted on simulated data, and the results show voltage estimates with a high degree of accuracy, even at extremely low percentages of observable nodes.

24 POWER TRANSMISSION AND DISTRIBUTION↗

A comparative study of machine learning models for predicting the state of reactive mixing

Mixing phenomena are important mechanisms controlling flow, species transport, and reaction processes in fluids and porous media. Accurate predictions of reactive mixing are critical for many Earth and environmental science problems such as contaminant fate and remediation, macroalgae growth, and plankton biomass evolution. Here, to investigate the evolution of mixing dynamics under different scenarios (e.g., anisotropy, fluctuating velocity fields), a finite-element-based numerical model was built to solve the fast, irreversible bimolecular reaction-diffusion equations to simulate a range of reactive-mixing scenarios. A total of 2,315 simulations were performed using different sets of model input parameters comprising various spatial scales of vortex structures in the velocity field, time-scales associated with velocity oscillations, the perturbation parameter for the vortex-based velocity, anisotropic dispersion contrast (i.e., ratio of longitudinal-to-transverse dispersion), and molecular diffusion. The outputs comprised concentration profiles of reactants and products. The inputs to and outputs from these simulations were concatenated into feature and label matrices, respectively, to train 20 different machine learning (ML) models intended to emulate system behavior. These 20 ML emulators, based on linear methods, Bayesian methods, ensemble learning methods, and multilayer perceptrons (MLPs), were trained to classify the state of mixing and predict three quantities of interest (QoIs) characterizing species production, decay (i.e., average concentration, square of average concentration), and degree of mixing (i.e., variances of species concentration). Unsurprisingly, linear classifiers and regressors failed to reproduce the QoIs; however, ensemble methods (classifiers and regressors) and the MLP model accurately classified the state of reactive mixing and the QoIs. Among ensemble methods, random forest and decision-tree-based AdaBoost faithfully predicted the QoIs. At run time, trained ML emulators produced results times faster than the finite-element simulations. Due to their low computational expense and high accuracy, ensemble and MLP models are excellent emulators for these numerical simulations and great utilities in uncertainty quantification exercises, which can require 1,000s of forward model runs.

97 MATHEMATICS AND COMPUTING↗

A learning-augmented approach for AC optimal power flow

Because of the high nonlinearity of AC optimal power flow (OPF), numerous efforts have been made in recent decades to find efficient methods. Machine learning (ML) has proven to significantly reduce the computational costs in many real-world problems. Thus, this paper develops a learning-augmented method for solving AC OPF, which integrates both power network equations and ML to yield near-optimal solutions. More specifically, ML models are developed to first predict bus voltage magnitudes and angles. Then, physics-based network equations are employed to calculate the power injection at different buses. Three ML algorithms, i.e., random forest, multi-target decision tree, and extreme learning machine, are explored and compared. To evaluate the efficiency of the proposed learning-augmented AC OPF solver, the MATPOWER Interior Point Solver is adopted as a baseline. Case studies on both 500-bus and 4918-bus test networks show that the proposed learning-augmented method has reduced the computational time by 15–100 times depending on the network size with a minimal loss in optimality.

42 ENGINEERING↗