Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 181 records · Page 10

Transcripts and genomic intervals associated with variation in metabolite abundance in maize leaves under field conditions

Abstract Plants exhibit extensive environment-dependent intraspecific metabolic variation, which likely plays a role in determining variation in whole plant phenotypes. However, much of the work seeking to use natural variation to link genes and transcript’s impacts on plant metabolism has employed data from controlled environments. Here, we generated and analyzed data on the variation in the abundance of 26 metabolites across 660 maize inbred lines under field conditions. We employ these data and previously published transcript and whole plant phenotype data reported for the same field experiment to identify both genomic intervals (through genome-wide association studies (GWAS)) and transcripts (using both transcriptome-wide association studies (TWAS) and an explainable artificial intelligence (AI) approach based on random forest (RF)) associated with variation in metabolite abundance. Both genome-wide association and random forest-based methods identified substantial numbers of significant associations including genes with plausible links to the metabolites they are associated with. In contrast, the transcriptome-wide association identified only six significant associations. In three cases, genetic markers associated with metabolic variation in our study colocalized with markers linked to variation in non-metabolic traits scored in the same experiment. We speculate that the poor performance of transcriptome-wide association studies in identifying transcript-metabolite associations may reflect a high prevalence of non-linear interactions between transcripts and metabolites and/or a bias towards rare transcripts playing a large role in determining intraspecific metabolic variation.

Mathivanan, Ramesh Kanna↗

Predicting chatter using machine learning and acoustic signals from low-cost microphones

Machining chatter is a phenomenon resulting from self-oscillation between a machining tool and workpiece. This self-oscillation results in variation on the machined product that reduces the ability to meet desired specifications. Chatter is a widely studied topic as it directly relates to the quality of machined products. Here, this study details the application of a Random Forest (RF) classifier with Recursive Feature Elimination (RFE) to machining audio collected by a single microphone during down-milling operations. This approach allows straightforward feature elimination that results in an easily understood set of analyzed dimensions. Stability is predicted solely based on the classification output of the RF classifier. Our approach proves highly predictive with consistent machining setup and a small sample set. We also review transferability between machining setups and present key findings. Our RF approach demonstrates the ability to analyze and classify chatter through a low-cost approach with limited training data required. The motivation for using a single microphone is to enable detection on machines without other sensors, such as accelerometers, present in the machining setup. The value of the in-process sensor and chatter classifier is highlighted because the machining setup included asymmetric dynamics that reduced the accuracy of the traditional analytical stability solution. We see a natural progression to deploying this audio-only methodology with real-time processing and classification using either a laptop or smartphone. This progression will allow visual indicators during the machining process that can alert machinists of progression into unstable machining processes.

42 ENGINEERING↗

Further adoption of conservation tillage can increase maize yields in the western US Corn Belt

Conservation tillage can reduce soil erosion, increase soil health, and decrease labor and fuel input costs. Despite these benefits, potential yield impacts remain an important concern for farmers considering adoption. Previous research suggests that conservation tillage is likely to have the largest yield benefits in more arid conditions, but a lack of field-level analyses across climatic, management and soil conditions limits confidence in such predictions. Satellite imagery provides the opportunity to monitor agricultural lands at sub-field resolution across large spatial scales and wide environmental gradients. Here we investigate the maize yield impacts of conservation tillage in the semi-arid western US Corn Belt, using sub-field resolution datasets on tillage practices and crop yields derived from satellite data spanning four states (Nebraska, Kansas, South Dakota, and North Dakota) between 2008 and 2020. On these datasets, we estimate heterogenous yield outcomes for several thousand maize fields across gradients in climate, soil quality and irrigation status by using a causal forests analysis, an adaptation of the random forests machine-learning algorithm for causal inference on observational data. We find that long-term adoption of conservation tillage increased rainfed maize yields by an average of 9.9% in the region. Impacts on irrigated yields were small and not statistically significant. These results, along with an analysis of variables related to greater than average yield benefits, indicate that improved water infiltration and retention are the primary reasons for conservation tillage benefits. Despite yield benefits, many fields estimated to see increased yields under long term low till have not adopted the practice. Therefore, we identify specific counties likely to benefit most from increased levels of adoption. Our results strengthen the understanding of the impacts of conservation agriculture on crop yields and help define environments and counties most likely to benefit from conservation tillage.

54 ENVIRONMENTAL SCIENCES↗

Data, scripts, and figures associated with a manuscript studying impact of climate and topography on post-fire vegetation recovery.

This data package is associated with the publication “Impact of Topography and Climate on Post-fire Vegetation Recovery Across Different Burn Severity and Land Cover Types through Machine Learning” submitted to Remote Sensing of Environment (Zahura et al. 2023). In this research, a machine learning algorithm, random forest (RF), was utilized to examine the impact of climate and topography on post-fire vegetation recovery. We used enhanced vegetation index (EVI) to examine varying burn severity and land cover types. The data package includes the input files for RF model training, outputs from model predictions and analysis, and python scripts to run the model, analyze the results to understand model performance and interpretability, and plot manuscript figures. This data package contains three folders (Data, Scripts, and Figures), a file-level metadata (FLMD) csv, and a data dictionary (dd) csv. Please see Postfire_recovery_flmd.csv for a list of all files contained in this data package and descriptions for each. The data dictionary (Postfire_recovery_dd.csv) describes the csv column headers. The “Data” folder provides all the inputs and outputs to train the RF model, evaluate performance, and interpret predictions. The “Scripts” folder contains python scripts and jupyter notebooks for model training and result analysis. The “Figures” folder includes the figures used in the manuscript in “.png” and “.jpg” format.

54 ENVIRONMENTAL SCIENCES↗

Protection System Validation with Machine Learning Anomaly Classification

A poster for the Early Career Poster Session. Power system protection devices have transitioned over the past few decades from mechanical to analog devices, then to solid state and finally digital. Relays and their associated critical network of equipment have significantly increased in complexity. Even internally, relays have gained significant intricacy, with relatively simple overcurrent or differential functions now being assisted by a myriad of other functions. This is necessary as the grid becomes more complex, but it brings increased difficulty in monitoring and upkeep. Misoperation caused by improper relay settings or malicious actions is a constant challenge faced by all utilities. These improper settings can be difficult to identify and may require exhaustive post-mortem analysis, typically after a major outage event has already occurred. A mechanism is needed for monitoring the behavior of protection systems to validate that they act and perform as expected. This work presents a concept for a machine learning (ML) system capable of validating the performance of protection systems by classifying anomalous events and characterizing protection system responses based solely on available current and voltage measurements. As a first step in its development, an experimental dataset is generated, and a random forest model is implemented with high accuracy in distinguishing four power system scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

A Hierarchical Feature-Based Methodology to Perform Cervical Cancer Classification

Prevention of cervical cancer could be performed using Pap smear image analysis. This test screens pre-neoplastic changes in the cervical epithelial cells; accurate screening can reduce deaths caused by the disease. Pap smear test analysis is exhaustive and repetitive work performed visually by a cytopathologist. This article proposes a workload-reducing algorithm for cervical cancer detection based on analysis of cell nuclei features within Pap smear images. We investigate eight traditional machine learning methods to perform a hierarchical classification. We propose a hierarchical classification methodology for computer-aided screening of cell lesions, which can recommend fields of view from the microscopy image based on the nuclei detection of cervical cells. We evaluate the performance of several algorithms against the Herlev and CRIC databases, using a varying number of classes during image classification. Results indicate that the hierarchical classification performed best when using Random Forest as the key classifier, particularly when compared with decision trees, k-NN, and the Ridge methods.

60 APPLIED LIFE SCIENCES↗

Vehicle Position Detection Based on Machine Learning Algorithms in Dynamic Wireless Charging

Dynamic wireless charging (DWC) has emerged as a viable approach to mitigate range anxiety by ensuring continuous and uninterrupted charging for electric vehicles in motion. DWC systems rely on the length of the transmitter, which can be categorized into long-track transmitters and segmented coil arrays. The segmented coil array, favored for its heightened efficiency and reduced electromagnetic interference, stands out as the preferred option. However, in such DWC systems, the need arises to detect the vehicle’s position, specifically to activate the transmitter coils aligned with the receiver pad and de-energize uncoupled transmitter coils. This paper introduces various machine learning algorithms for precise vehicle position determination, accommodating diverse ground clearances of electric vehicles and various speeds. Through testing eight different machine learning algorithms and comparing the results, the random forest algorithm emerged as superior, displaying the lowest error in predicting the actual position.

47 OTHER INSTRUMENTATION↗

Post-Event Fault Identification with Machine Learning for Protection System Validation

Power system protection devices have transitioned over the past few decades from mechanical to analog devices, then to solid state and finally digital. Relays and their associated critical network of equipment have significantly increased in complexity. Even internally, relays have gained significant intricacy, with relatively simple overcurrent or differential functions now being assisted by a myriad of other functions. This is necessary as the grid becomes more complex, but it brings increased difficulty in monitoring and upkeep. Misoperation caused by accidental improper relay settings or deliberate malicious actions is a constant challenge faced by all utilities. These improper settings can be difficult to identify and may require exhaustive post-mortem analysis, typically after a major outage event has already occurred. A mechanism is needed for monitoring the behavior of protection systems to validate that their performance falls within expectations. Relays that fail to isolate a fault or trip when there is no system disturbance can be flagged for settings review in situations where this behavior may not have been noticed due to manual restoration or backup protection operations. This work presents a concept for a machine learning (ML) system capable of validating the performance of protection systems by identifying fault events and characterizing protection system responses based solely on available current and voltage measurements. As a first step in its development, an experimental dataset is generated, and a random forest model is implemented with high accuracy in distinguishing four power system scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

Protection System Validation Using Post-Event Anomaly Classification with Machine Learning

Power system protection devices have transitioned over the past few decades from mechanical to analog devices, then to solid state and finally digital. Relays and their associated critical network of equipment have significantly increased in complexity. Even internally, relays have gained significant intricacy, with relatively simple overcurrent or differential functions now being assisted by a myriad of other functions. This is necessary as the grid becomes more complex, but it brings increased difficulty in monitoring and upkeep. Misoperation caused by improper relay settings or malicious actions is a constant challenge faced by all utilities. These improper settings can be difficult to identify and may require exhaustive post-mortem analysis, typically after a major outage event has already occurred. A mechanism is needed for monitoring the behavior of protection systems to validate that they act and perform as expected. This work presents a concept for a machine learning (ML) system capable of validating the performance of protection systems by classifying anomalous events and characterizing protection system responses based solely on available current and voltage measurements. As a first step in its development, an experimental dataset is generated, and a random forest model is implemented with high accuracy in distinguishing four power system scenarios.

24 - POWER TRANSMISSION AND DISTRIBUTION↗

A machine learning-based fast frequency response control for a VSC-HVDC system

An HVDC system can realize a very fast frequency response to the disturbed system under a contingency because its active power control is decoupled from the frequency deviation. However, most of existing HVDC frequency control strategies are coupled with system primary frequency control and secondary frequency control. Since the traditional system frequency control is dominated by the thermal generators, the advantage of the fast response of the HVDC system is not made fully used. The development of a frequency response estimation based on a machine learning algorithm provides another approach to improve the frequency response capability of the HVDC system. Different from other frequency deviation tracking strategies, a machine learning based HVDC frequency response control can directly increase the power flow of a HVDC system by estimation of the system generator or load lost. In this paper, a fast frequency response control using a HVDC system for a large power system disturbance based on the multivariate random forest regression (MRFR) algorithm is proposed. The simulation is carried out with an integrated power system model based on the North American interconnections. The simulation results indicate that the proposed MRFR based frequency response control can significantly improve the frequency low point during an event, while stabilizing the frequency in advance.

42 ENGINEERING↗

Physics-Infused AI/ML Based Digital-Twin Framework for Flow-Induced-Vibration Damage Prediction in a Nuclear Reactor Heat Exchanger

This report summarizes some of the ongoing work related to the development of an expert-elicitation-digital-twin framework for real time damage state prediction in heat exchanger components of a nuclear reactor. The framework is targeted towards predicting damage associated with coupled low cycle fatigue (associated with regular heat-up, cool-down and power operation transients) and high cycle fatigue (associated with flow induced vibration transients). The overall framework will be based on a NoSQL based database, physics-infused-geometry-dependent virtual-sensor data, different AI/ML techniques-based data-driven-predictive-model applications (Apps) and real-time plant sensor measurements available through few existing sensors. Towards this overall goal, this report updates some of the ongoing work, such as on implementation of a NoSQL Database (such as MongoDB), FE based heat transfer analysis of a heat exchanger (e.g. of a PWR steam generator) for generating geometry-dependent virtual sensor data and evaluation of various AI/ML models such as based on multivariate linear regression, ensembled decision-tree based Random-Forest and Gradient-Boosting regression and high-dimensional-kernel-function-transformation based Support-Vector-Machine regression models. The AI/ML models were evaluated for predicting multi-time-series thermal states at thousands of 3D point-clouds

22 GENERAL STUDIES OF NUCLEAR REACTORS↗

Predictive Modeling of NOx Emissions from Lean Direct Injection of Hydrogen and Hydrogen/Natural Gas Blends Using Flame Imaging and Machine Learning

This research paper explores the use of machine learning to relate images of flame structure and luminosity to measured NOx emissions. Images of reactions produced by 16 aero-engine derived injectors for a ground-based turbine operated on a range of fuel compositions, air pressure drops, preheat temperatures and adiabatic flame temperatures were captured and postprocessed. The experimental investigations were conducted under atmospheric conditions, capturing CO, NO and NOx emissions data and OH* chemiluminescence images from 27 test conditions. The injector geometry and test conditions were based on a statistically designed test plan. These results were first analyzed using the traditional analysis approach of analysis of variance (ANOVA). The statistically based test plan yielded 432 data points, leading to a correlation for NOx emissions as a function of injector geometry, test conditions and imaging responses, with 70.2% accuracy. As an alternative approach to predicting emissions using imaging diagnostics as well as injector geometry and test conditions, a random forest machine learning algorithm was also applied to the data and was able to achieve an accuracy of 82.6%. This study offers insights into the factors influencing emissions in ground-based turbines while emphasizing the potential of machine learning algorithms in constructing predictive models for complex systems.

08 HYDROGEN↗

Identifying Transient Candidates in the Dark Energy Survey Using Convolutional Neural Networks

The ability to discover new transient candidates via image differencing without direct human intervention is an important task in observational astronomy. For these kind of image classification problems, machine learning techniques such as Convolutional Neural Networks (CNNs) have shown remarkable success. In this work, we present the results of an automated transient candidate identification on images with CNNs for an extant data set from the Dark Energy Survey Supernova program, whose main focus was on using Type Ia supernovae for cosmology. By performing an architecture search of CNNs, we identify networks that efficiently select non-artifacts (e.g., supernovae, variable stars, AGN, etc.) from artifacts (image defects, mis-subtractions, etc.), achieving the efficiency of previous work performed with random Forests, without the need to expend any effort in feature identification. The CNNs also help us identify a subset of mislabeled images. Performing a relabeling of the images in this subset, the resulting classification with CNNs is significantly better than previous results, lowering the false positive rate by 27% at a fixed missed detection rate of 0.05.

79 ASTRONOMY AND ASTROPHYSICS↗

Improved quality metrics for association and reproducibility in chromatin accessibility data using mutual information

Correlation metrics are widely utilized in genomics analysis and often implemented with little regard to assumptions of normality, homoscedasticity, and independence of values. This is especially true when comparing values between replicated sequencing experiments that probe chromatin accessibility, such as assays for transposase-accessible chromatin via sequencing (ATAC-seq). Such data can possess several regions across the human genome with little to no sequencing depth and are thus non-normal with a large portion of zero values. Despite distributed use in the epigenomics field, few studies have evaluated and benchmarked how correlation and association statistics behave across ATAC-seq experiments with known differences or the effects of removing specific outliers from the data. Here, we developed a computational simulation of ATAC-seq data to elucidate the behavior of correlation statistics and to compare their accuracy under set conditions of reproducibility. Using these simulations, we monitored the behavior of several correlation statistics, including the Pearson’s R and Spearman’s ρ coefficients as well as Kendall’s τ and Top–Down correlation. We also test the behavior of association measures, including the coefficient of determination R 2 , Kendall’s W, and normalized mutual information. Our experiments reveal an insensitivity of most statistics, including Spear man’s ρ, Kendall’s τ, and Kendall’s W, to increasing differences between simulated ATAC-seq replicates. The removal of co-zeros (regions lacking mapped sequenced reads) between simulated experiments greatly improves the estimates of correlation and association. After removing co-zeros, the R 2 coefficient and normalized mutual information display the best performance, having a closer one-to-one relationship with the known portion of shared, enhanced loci between simulated replicates. When comparing values between experimental ATAC-seq data using a random forest model, mutual information best predicts ATAC-seq replicate relationships. Collectively, this study demonstrates how measures of correlation and association can behave in epigenomics experiments. We provide improved strategies for quantifying relationships in these increasingly prevalent and important chromatin accessibility assays.

59 BASIC BIOLOGICAL SCIENCES↗

Spatial Mapping of Riverbed Grain-Size Distribution Using Machine Learning

Recent alluvial sediments in riverbeds play a significant role in controlling hydrologic exchange flows (HEFs) in river systems. The alluvial layer is usually associated with strong heterogeneity in physical properties (e.g., permeability and hydraulic conductivity), which affects local HEFs and therefore biogeochemical processes. The spatial distribution of these physical properties needs to be determined to inform the numerical models used to reveal the realistic hydro-biogeochemical behaviors. Such information can be obtained based on the intrinsic link between sediment grain-size distribution and hydraulic properties where sediment texture information is available. However, grain-size measurements are usually spatially sparse and do not have adequate coverage and resolution, particularly for a relatively large domain such as the Hanford Reach of the Columbia River. In this paper, we adopted machine learning (ML) approaches for categorizing and mapping the spatial distributions of riverbed substrate grain size and filling in missing areas of substrate data using the ML models along the reach. Such ML models for substrate size mapping were trained at 13,372 locations using measured substrate sizes along with observed and simulated attributes, including bathymetric attributes (e.g., elevation, slope, and aspect ratio) from LIDAR and bathymetric surveys, and hydrodynamic properties (e.g., water depth, velocity, shear stress, and their statistical moments). An ensemble bagging-based ML technique, Random Forest, was adopted to identify the most influential factors as predictors to develop the predictive models with over-fitting issues addressed. The models were evaluated with respect to each individual substrate size class and the lumped group, and then used to generate the final substrate size maps covering all the grid cells in the numerical modeling domain.

54 ENVIRONMENTAL SCIENCES↗

Yield strength prediction of high-entropy alloys using machine learning

Yield strength at high temperature is an important parameter in the design and application of high entropy alloys (HEAs). However, the experimental measurement of yield strength at high temperature is quite costly, complicated, and time-consuming. Therefore, it is essential to identify and apply a robust method for the accurate prediction of yield strength at high temperature from the available experimental and simulation data. In this study, for the first time, a machine learning (ML) method based on the regression technique of random forest (RF) regressor is used to predict the yield strength of HEAs at the desired temperature. Further, the yield strengths of MoNbTaTiW and HfMoNbTaTiZr at 800 °C and 1200 °C, are predicted using the RF regressor model. We find that the results are consistent with the experimental reports, showing that the RF regressor model predicts the yield strength of HEAs at the desired temperatures with high accuracy.

36 MATERIALS SCIENCE↗

Soil Origin and Plant Genotype Modulate Switchgrass Aboveground Productivity and Root Microbiome Assembly

Switchgrass (Panicum virgatum) is a model perennial grass for bioenergy production that can be productive in agricultural lands that are not suitable for food production. There is growing interest in whether its associated microbiome may be adaptive in low- or no-input cultivation systems. However, the relative impact of plant genotype and soil factors on plant microbiome and biomass are a challenge to decouple. To address this, a common garden greenhouse experiment was carried out using six common switchgrass genotypes, which were each grown in four different marginal soils collected from long-term bioenergy research sites in Michigan and Wisconsin. We characterized the fungal and bacterial root communities with high-throughput amplicon sequencing of the ITS and 16S rDNA markers, and collected phenological plant traits during plant growth, as well as soil chemical traits. At harvest, we measured the total plant aerial dry biomass. Significant differences in richness and Shannon diversity across soils but not between plant genotypes were found. Generalized linear models showed an interaction between soil and genotype for fungal richness but not for bacterial richness. Community structure was also strongly shaped by soil origin and soil origin × plant genotype interactions. Overall, plant genotype effects were significant but low. Random Forest models indicate that important factors impacting switchgrass biomass included NO 3 – , Ca 2+ , PO 4 3– , and microbial biodiversity. We identified 54 fungal and 52 bacterial predictors of plant aerial biomass, which included several operational taxonomic units belonging to Glomeraceae and Rhizobiaceae, fungal and bacterial lineages that are involved in provisioning nutrients to plants.

plant biomass↗

Moisture availability mediates the relationship between terrestrial gross primary production and solar-induced chlorophyll fluorescence: Insights from global-scale variations

Effective use of solar-induced chlorophyll fluorescence (SIF) to estimate and monitor gross primary production (GPP) in terrestrial ecosystems requires a comprehensive understanding and quantification of the relationship between SIF and GPP. To date, this understanding is incomplete and somewhat controversial in the literature. Here we derived the GPP/SIF ratio from multiple data sources as a diagnostic metric to explore its global-scale patterns of spatial variation and potential climatic dependence. We found that the growing season GPP/SIF ratio varied substantially across global land surfaces, with the highest ratios consistently found in boreal regions. Spatial variation in GPP/SIF was strongly modulated by climate variables. The most striking pattern was a consistent decrease in GPP/SIF from cold-and-wet climates to hot-and-dry climates. We propose that the reduction in GPP/SIF with decreasing moisture availability may be related to stomatal responses to aridity. Furthermore, we show that GPP/SIF can be empirically modeled from climate variables using a machine learning (random forest) framework, which can improve the modeling of ecosystem production and quantify its uncertainty in global terrestrial biosphere models. Finally, our results point to the need for targeted field and experimental studies to better understand the patterns observed and to improve the modeling of the relationship between SIF and GPP over broad scales.

59 BASIC BIOLOGICAL SCIENCES↗