Engineering Papers⌕ Search

SEARCH · Engineering Papers

Results for “Random forests”

Search indexed NASA NTRS and DOE OSTI research on propulsion, heat transfer, battery materials and energy systems. Follow report and document links to the original sources.

Quote a phrase for an exact phrase match. Source license links do not imply unrestricted reuse.

At least 253 records · Page 14

Classification of bacterial plasmid and chromosome derived sequences using machine learning

Plasmids are important genetic elements that facilitate horizonal gene transfer between bacteria and contribute to the spread of virulence and antimicrobial resistance. Most bacterial genome sequences in the public archives exist in draft form with many contigs, making it difficult to determine if a contig is of chromosomal or plasmid origin. Using a training set of contigs comprising 10,584 chromosomes and 10,654 plasmids from the PATRIC database, we evaluated several machine learning models including random forest, logistic regression, XGBoost, and a neural network for their ability to classify chromosomal and plasmid sequences using nucleotide k-mers as features. Based on the methods tested, a neural network model that used nucleotide 6-mers as features that was trained on randomly selected chromosomal and plasmid subsequences 5kb in length achieved the best performance, outperforming existing out-of-the-box methods, with an average accuracy of 89.38% ± 2.16% over a 10-fold cross validation. The model accuracy can be improved to 92.08% by using a voting strategy when classifying holdout sequences. In both plasmids and chromosomes, subsequences encoding functions involved in horizontal gene transfer—including hypothetical proteins, transporters, phage, mobile elements, and CRISPR elements—were most likely to be misclassified by the model. This study provides a straightforward approach for identifying plasmid-encoding sequences in short read assemblies without the need for sequence alignment-based tools.

59 BASIC BIOLOGICAL SCIENCES↗

Statistical and Machine Learning Approaches to Analyzing Pipeline Incidents in the United States (2010–2024)

This study applies machine learning methods to analyze natural gas pipeline incidents in the United States using the Pipeline and Hazardous Materials Safety Administration (PHMSA) Gas Distribution Incident Dataset (2010–2024). The dataset includes over 600 variables describing incident characteristics, infrastructure attributes, and contributing factors associated with unintentional gas releases. The objective is to assess whether these features can reliably predict the underlying cause of pipeline failures. Multinomial logistic regression and Random Forest models were developed to classify incident causes, including excavation damage, corrosion, equipment failure, and natural forces. Results show that excavation damage is both the most frequent and most predictable cause, with models achieving strong performance for this category. However, when excavation damage is excluded, model accuracy declines significantly, with some models performing near random levels. Across all approaches, severe class imbalance and limited variability in key predictors constrain predictive performance. Pipeline age and diameter emerge as the most influential variables, but they provide insufficient discriminatory power to distinguish among less frequent failure types. These findings indicate that non-excavation-related incidents are rare, heterogeneous, and weakly represented in the dataset, limiting the effectiveness of machine learning classification. Overall, this study highlights the structural limitations of the PHMSA dataset for predictive modeling and underscores the need for improved data balance and feature enrichment. The results reinforce excavation damage prevention as the most impactful strategy for reducing pipeline incidents.

03 NATURAL GAS↗

Aided Active Learning (AAL) for Enhanced Critical Heat Flux Prediction

Accurate prediction of critical heat flux (CHF) is crucial for the safe and efficient operation of nuclear reactors. Traditional CHF modeling methods often require extensive experimental data, which are hard to obtain. This study introduces the Aided Active Learning (AAL) framework, which strategically minimizes data requirements without sacrificing model accuracy. Unlike conventional Active Learning (AL), AAL introduces an additional step of randomly selecting a subset from the sample pool before applying the query strategy. To evaluate the performance of AAL, two query strategies—uncertainty-based sampling and error-reduction sampling—were evaluated across the following models: random forest (RF), feedforward neural network (FNN), and variational feedforward neural network (vFNN). The proposed framework demonstrated that AAL effectively reduces the number of training samples needed to achieve comparable predictive accuracy. For the RF model, AL required only 710 samples to achieve an R2 score of 0.98, as compared to the 4,785 samples needed by random sampling. Similarly, the FNN model achieved the same R2 score with just 355 samples when using AL, a significant improvement over the 825 samples required by random sampling. In case of uncertainty-based sampling strategy, vFNN attained an R2 of 0.98 with 3,420 samples, reducing the sample requirement by 47% relative to the 6,440 samples needed for random sampling. Its performance suggests that larger training data are required to fully leverage its uncertainty quantification capabilities.

22 - GENERAL STUDIES OF NUCLEAR REACTORS↗

Measuring photometric redshifts for high-redshift radio source surveys

With the advent of deep, all-sky radio surveys, the need for ancillary data to make the most of the new, high-quality radio data from surveys like the Evolutionary Map of the Universe (EMU), GaLactic and Extragalactic All-sky Murchison Widefield Array survey eXtended, Very Large Array Sky Survey, and LOFAR Two-metre Sky Survey is growing rapidly. Radio surveys produce significant numbers of Active Galactic Nuclei (AGNs) and have a significantly higher average redshift when compared with optical and infrared all-sky surveys. Thus, traditional methods of estimating redshift are challenged, with spectroscopic surveys not reaching the redshift depth of radio surveys, and AGNs making it difficult for template fitting methods to accurately model the source. Machine Learning (ML) methods have been used, but efforts have typically been directed towards optically selected samples, or samples at significantly lower redshift than expected from upcoming radio surveys. This work compiles and homogenises a radio-selected dataset from both the northern hemisphere (making use of Sloan Digital Sky Survey optical photometry) and southern hemisphere (making use of Dark Energy Survey optical photometry). We then test commonly used ML algorithms such as k-Nearest Neighbours (kNN), Random Forest, ANNz, and GPz on this monolithic radio-selected sample. We show that kNN has the lowest percentage of catastrophic outliers, providing the best match for the majority of science cases in the EMU survey. We note that the wider redshift range of the combined dataset used allows for estimation of sources up to z = 3 before random scatter begins to dominate. When binning the data into redshift bins and treating the problem as a classification problem, we are able to correctly identify ≈ 76% of the highest redshift sources—sources at redshift z > 2.51 —as being in either the highest bin (z > 2.51) or second highest (z = 2.25).

79 ASTRONOMY AND ASTROPHYSICS↗

The drivers and predictability of wildfire re-burns in the western United States (US)

Evidence is mounting that the effectiveness of using prescribed burns as a management tactic may be diminishing due to the higher incidence of wildfire re-burns. The development of predictive models of re-burns is thus essential to better understand their primary drivers so that forest management practices can be updated to account for these events. First, we assess the potential for human activity as a driver of re-burns by evaluating re-burn trends both within and outside of the wildland–urban interface (WUI) of the western US. Next, we investigate the predictability of re-burns through the application of both random forest and the explanatory machine learning non-negative matrix factorization using k-means clustering (NMFk) algorithms to predict re-burn occurrence over California based on a number of climate factors. Our findings indicate that while most states showed increasing trends within the WUI when trends were conducted over longer moving windows (e.g. 20 years), California was the only state where the rate of increase was consistently higher in the WUI, indicating a stronger potential for human activity as a driver in that location. Furthermore, we find model performance was found to be robust over most of California (Testing F1 scores = 0.688), although results were highly variable based on EPA level III Ecoregion (F1 scores = 0.0–0.778). Insights provided from this study will lead to a better understanding of climate and human activity drivers of re-burns and how these vary at broad spatial scales so that improvements in forest management practices can be tuned according to the level of change that is expected for a given region.

54 ENVIRONMENTAL SCIENCES↗

Artificial Diversity and Defense Security (ADDSec)

Artificial Diversity and Defense Security (ADDSec) machine learning algorithms are used to classify and cluster threats so that an appropriate response can be initiated as a mitigation strategy. The package includes an ensemble of machine learning algorithms such as Support Vector Machines, naïve bayes, logistic regression, and random forest that evolve with the data to recognize anomalous behavior at the host and network levels. Inputs into the machine learning algorithms include end host system calls, system utilization, packet captures, and syslog messages. The machine learning algorithms can be retrained based on user defined intervals or on the number of packets received. ADDSEC's threat responses include Internet Protocol (IP) Address randomization, application port number randomization, and application library randomization. The IP randomization implementation is built on top of a Software Defined Networking (SDN) framework. The SDN controller installs flows on each of the SDN switches with randomized source and destination IP addresses. The application port numbers are randomized using iptables. The application library randomization is created with a LLVM compiler. All randomization schemes are transparent to the endpoints on the network. Sandia National Laboratories is a multimission laboratory managed and operated by National Technology & Engineering Solutions of Sandia, LLC, a wholly owned subsidiary of Honeywell International Inc., for the U.S. Department of Energy’s National Nuclear Security Administration under contract DE-NA0003525. SAND2021-3379 O

Cox, RebeccaE.↗

Using electronic health record metadata to predict housing instability amongst veterans

Housing instability is considered a significant life stressor and preemptive screening should be applied to identify those at risk for homelessness as early as possible so that they can be targeted for specialized care. We developed models to classify patient outcomes for an established VA Homelessness Screening Clinical Reminder (HSCR), which identifies housing instability, in the two months prior to its administration. Logistic Regression and Random Forest models were fit to classify responses using the last 18 months of document activity. We measure concentration of risk across stratifications of predicted probability and observe an enriched likelihood of finding confirmed false negative responses from veterans with diagnosed housing instability. Positive responses were 34 times more likely to be detected within the top 1 % of patients predicted at risk than from those randomly selected. There is a 1 in 4 chance of detecting false negatives within the top 1 % of predicted risk. Machine learning methods can classify between episodes of housing instability using a data-driven approach that does not rely on variables curated from domain experts. This method has the potential to improve clinicians’ ability to identify veterans who are experiencing housing instability but are not captured by HSCR.

60 APPLIED LIFE SCIENCES↗

Target Selection and Validation of DESI Quasars

The Dark Energy Spectroscopic Instrument (DESI) survey will measure large-scale structures using quasars as direct tracers of dark matter in the redshift range 0.9 < z < 2.1 and using Lyα forests in quasar spectra at z > 2.1. We present several methods to select candidate quasars for DESI, using input photometric imaging in three optical bands (g, r, z) from the DESI Legacy Imaging Surveys and two infrared bands (W1, W2) from the Wide-field Infrared Survey Explorer. These methods were extensively tested during the Survey Validation of DESI. In this paper, we report on the results obtained with the different methods and present the selection we optimized for the DESI main survey. The final quasar target selection is based on a random forest algorithm and selects quasars in the magnitude range of 16.5 < r < 23. Visual selection of ultra-deep observations indicates that the main selection consists of 71% quasars, 16% galaxies, 6% stars, and 7% inconclusive spectra. Using the spectra based on this selection, we build an automated quasar catalog that achieves a fraction of true QSOs higher than 99% for a nominal effective exposure time of ~1000 s. With a 310 deg -2 target density, the main selection allows DESI to select more than 200 deg -2 quasars (including 60 deg -2 quasars with z > 2.1), exceeding the project requirements by 20%. The redshift distribution of the selected quasars is in excellent agreement with quasar luminosity function predictions.

79 ASTRONOMY AND ASTROPHYSICS↗

Mapping tree height in complex terrain of northern China using ultra-high-resolution images

Tree height is a key parameter for estimating forest biomass and carbon sequestration. In recent years, notable progress has been made in mapping tree height using satellite imagery. However, existing tree height products show low accuracy in mountainous and complex terrains, and few studies typically addressed tree height estimations in mountain areas. This study examines the Mentougou district of Beijing, China, characterized by complex terrain and mountainous landscapes. We analyzed two methods for estimating tree height: one using only spectral features and another combining spectral features with topographic factors (elevation, slope, aspect). We used 3-m resolution PlanetScope 8-band multispectral imagery, with 710 field-measured individual tree heights averaged to obtain 471 pixel-level tree height values as ground-truth, to develop tree height prediction models using eXtreme Gradient Boosting (XGBoost), Random Forest (RF), and Gradient Boosting Machine (GBM) models. The results show that the XGBoost model consistently presented the highest accuracy for both methods evaluated. Specifically, the XGBoost model that combined spectral data with elevation and slope variables with an R² of 0.75 and an RMSE of 2.69 m. Using the XGBoost model, we generated the tree height map for the Mentougou area at 3 m resolution, showing tree heights ranging from 0.5 to 30.4 m, and the model’s prediction error standard deviations ranged from 2.50 to 4.71 m, indicating reliable performance across varied terrain. Additionally, we compared and evaluated the global tree height products, identifying limitations in the accuracy within complex terrains. This study demonstrates the potential for accurately predicting tree heights by combining high-resolution multispectral satellites with a terrain factor modeling approach.

Complex terrain↗

Modeling the topographic influence on aboveground biomass using a coupled model of hillslope hydrology and ecosystem dynamics

Abstract. Topographic heterogeneity and lateral subsurface flow at the hillslope scale of ≤1 km may have outsized impacts on tropical forest through their impacts on water available to plants under water-stressed conditions. However, vegetation dynamics and finer-scale hydrologic processes are not concurrently represented in Earth system models. In this study, we integrate the Energy Exascale Earth System Model (E3SM) land model (ELM) that includes the Functionally Assembled Terrestrial Ecosystem Simulator (FATES), with a three-dimensional hydrology model (ParFlow) to explicitly resolve hillslope topography and subsurface flow and perform numerical experiments to understand how hillslope-scale hydrologic processes modulate vegetation along water availability gradients at Barro Colorado Island (BCI), Panama. Our simulations show that groundwater table depth (WTD) can play a large role in governing aboveground biomass (AGB) when drought-induced tree mortality is triggered by hydraulic failure. Analyzing the simulations using random forest (RF) models, we find that the domain-wide simulated AGB and WTD can be well predicted by static topographic attributes, including surface elevation, slope, and convexity, and adding soil moisture or groundwater table depth as predictors further improves the RF models. Different model representations of mortality due to hydraulic failure can change the dominant topographic driver for the simulated AGB. Contrary to the simulations, the observed AGB in the well-drained 50 ha forest census plot within BCI cannot be well predicted by the RF models using topographic attributes and observed soil moisture as predictors, suggesting other factors such as nutrient status may have a larger influence on the observed AGB. The new coupled model may be useful for understanding the diverse impact of local heterogeneity by isolating the water availability and nutrient availability from the other external and internal factors in ecosystem modeling.

54 ENVIRONMENTAL SCIENCES↗

Alert Classification for the ALeRCE Broker System: The Light Curve Classifier

We present the first version of the Automatic Learning for the Rapid Classification of Events (ALeRCE) broker light curve classifier. ALeRCE is currently processing the Zwicky Transient Facility (ZTF) alert stream, in preparation for the Vera C. Rubin Observatory. The ALeRCE light curve classifier uses variability features computed from the ZTF alert stream and colors obtained from AllWISE and ZTF photometry. We apply a balanced random forest algorithm with a two-level scheme where the top level classifies each source as periodic, stochastic, or transient, and the bottom level further resolves each of these hierarchical classes among 15 total classes. This classifier corresponds to the first attempt to classify multiple classes of stochastic variables (including core- and host-dominated active galactic nuclei, blazars, young stellar objects, and cataclysmic variables) in addition to different classes of periodic and transient sources, using real data. We created a labeled set using various public catalogs (such as the Catalina Surveys and Gaia DR2 variable stars catalogs, and the Million Quasars catalog), and we classify all objects with ≥6 g-band or ≥6 r-band detections in ZTF (868,371 sources as of 2020 June 9), providing updated classifications for sources with new alerts every day. For the top level we obtain macro-averaged precision and recall scores of 0.96 and 0.99, respectively, and for the bottom level we obtain macro-averaged precision and recall scores of 0.57 and 0.76, respectively. Updated classifications from the light curve classifier can be found at the ALeRCE Explorer website (http://alerce.online).

47 OTHER INSTRUMENTATION↗

Discovery of Stable Surfaces with Extreme Work Functions by High‐Throughput Density Functional Theory and Machine Learning

Abstract The work function is the key surface property that determines the energy required to extract an electron from the surface of a material. This property is crucial for thermionic energy conversion, band alignment in heterostructures, and electron emission devices. This work presents a high‐throughput workflow using density functional theory (DFT) to calculate the work function and cleavage energy of 33,631 slabs (58,332 work functions) that are created from 3,716 bulk materials. The number of calculated surface properties surpasses the previously largest database by a factor of ≈27. Several surfaces with an ultra‐low (<2 eV) and ultra‐high (>7 eV) work function are identified. Specifically, the (100)‐Ba‐O surface of BaMoO 3 and the (001)‐F surface of Ag 2 F have record‐low (1.25 eV) and record‐high (9.06 eV) steady‐state work functions. Based on this database a physics‐based approach to featurize surfaces is utilized to predict the work function. The random forest model achieves a test mean absolute error (MAE) of 0.09 eV, comparable to the accuracy of DFT. This surrogate model enables rapid predictions of the work function (≈ 10 5 faster than DFT) across a vast chemical space and facilitates the discovery of material surfaces with extreme work functions for energy conversion and electronic device applications.

97 MATHEMATICS AND COMPUTING↗

Performance Prediction of High‐Entropy Perovskites La 0.8 Sr 0.2 Mn x Co y Fe z O 3 with Automated High‐Throughput Characterization of Combinatorial Libraries and Machine Learning

Perovskite oxides form a large family of materials with applications across various fields, owing to their structural and chemical flexibility. Efficient exploration of this extensive compositional space is now achievable through automated high-throughput experimentation combined with machine learning. In this study, we investigate the composition–structure–performance relationships of high-entropy La 0.8 Sr 0.2 Mn x Co y Fe z O 3±𝞭 perovskite oxides (0 < x, y, z <1; x+y+z≈1) for application as oxygen electrodes in Solid Oxide Cells. Following the deposition of a continuous compositional map using thin-film combinatorial pulsed laser deposition, compositional, structural, and performance properties are characterized using six different techniques with mapping capabilities. Random forests effectively model electrochemical performance, consistently identifying Fe-rich oxides as optimal compounds with the lowest area-specific resistance values for oxygen electrodes at 700 °C. Additionally, the models identify a statistical correlation between oxygen sublattice distortion—derived from spectral analysis of Raman-active modes—and enhanced performance.

high entropy oxides↗

Data‐Driven Insights into Rare Earth Mineralization: Machine Learning Applications Using Functional Material Synthesis Data

Understanding rare‐earth element (REE) mineralization mechanisms is essential for developing efficient separation strategies. Although the geochemical pathways that generate REE deposits are qualitatively known, quantitative links between specific conditions and mineralization outcomes remain limited. Herein, the repurpose laboratory REE hydrothermal synthesis data—originally collected for functional‐materials fabrication—as a surrogate for studying mineralization with data‐driven methods. The compiled 1,200+ hydrothermal reaction records and trained three machine‐learning models—K‐nearest neighbors (KNN), random forest (RF), and extreme gradient boosting (XGB)—to predict product elements and phases from precursors, additives, reaction conditions, and engineered features. Validation shows XGB achieves the highest accuracy. Feature importance indicates thermodynamic properties of cations and anions dominate model decisions. Correlations reveal positive relationships among precursor concentration, reaction time, pH, and temperature, consistent with classical crystallization behavior. XGB‐based regressors are built to predict crystallization temperature and pH from precursor/product attributes. Performance is strongest when similar training examples exist, while accuracy declines for underrepresented reactions, notably REE carbonates and heavy‐REE systems. Overall, the study shows that functional‐materials datasets can illuminate REE mineralization and provide priors for exploration and processing. Expanding datasets with less‐studied chemistries and conditions will improve generality and support deposit discovery and more efficient REE recovery.

feature importance analysis↗

Machine learning to predict biomass sorghum yields under future climate scenarios

Crop yield modeling is critical in the design of national strategies for agricultural production, particularly in the context of a changing climate. Forecasting yields of bioenergy crops at fine spatial resolutions can help to evaluate near-term and long-term pathways for scaling up bio-based fuel and chemical production, and for understanding the impacts of abiotic stressors such as severe droughts and temperature extremes on potential biomass supply. In this work we used a large dataset of 28,364 Sorghum bicolor yield samples (uniquely identified by county and year of observation), environmental variables, and multiple approaches to analyze historical trends in sorghum productivity across the USA. We selected the most accurate machine learning approach (a variation of the random forest approach) to predict future trends in sorghum yields under four greenhouse gas (GHG) emission scenarios and two irrigation regimes. We identified irrigation practices, vapor pressure deficit, and time (a proxy for technological improvement) as the most important predictors of sorghum productivity. Our results showed a decreasing trend of sorghum yields over future years (on average 2.7% from 2018 to 2099), with greater decline under a high GHG emissions scenario (3.8%) and in the absence of irrigation (4.6%). Geographically, we observed the steepest predicted declines in the Great Lakes (8.2%), Upper Midwest (7.5%), and Heartland (6.7%) regions. Our study demonstrates the use of machine learning to identify environmental controllers of sorghum biomass yield and predict yields with reasonable accuracy. These results can inform the development of more realistic biomass supply projections for bioenergy if sorghum production is scaled up. (c) 2020 Society of Chemical Industry and John Wiley & Sons, Ltd

09 BIOMASS FUELS↗

ytopt: Autotuning Scientific Applications for Energy Efficiency at Large Scales

As we enter the exascale computing era, efficiently utilizing power and optimizing the performance of scientific applications under power and energy constraints has become critical and challenging. We propose a low-overhead autotuning framework to autotune performance and energy for various hybrid MPI/OpenMP scientific applications at large scales and to explore the tradeoffs between application runtime and power/energy for energy efficient application execution, then use this framework to autotune four ECP proxy applications—XSBench, AMG, SWFFT, and SW4lite. Our approach uses Bayesian optimization with a Random Forest surrogate model to effectively search parameter spaces with up to 6 million different configurations on two large-scale HPC production systems, Theta at Argonne National Laboratory and Summit at Oak Ridge National Laboratory. The experimental results show that our autotuning framework at large scales has low overhead and achieves good scalability. Using the proposed autotuning framework to identify the best configurations, we achieve up to 91.59% performance improvement, up to 21.2% energy savings, and up to 37.84% EDP (energy delay product) improvement on up to 4096 nodes.

Autotuning↗

Identifying hydrologic signatures associated with streamflow depletion caused by groundwater pumping

Abstract Groundwater pumping can reduce streamflow in nearby waterways (‘streamflow depletion’), a process which must be accounted for in integrated management of surface and groundwater resources. However, causal identification of streamflow depletion from hydrographs alone is challenging because pumping impacts are masked by other drivers of hydrologic variability. To identify potential indicators of streamflow depletion, we used synthetic hydrographs and an analytical streamflow depletion model to assess potential pumping impacts on specific hydrograph characteristics (‘hydrologic signatures’) for 215 streamgages spanning the conterminous United States (CONUS). We found that streamflow depletion commonly impacts signatures associated with seasonal and annual low flows and low flow recessions. The largest impacts occurred during dry years, suggesting streamflow depletion may be evident in dry years even where impacts are unmeasurable in wet years. Random forest models indicated that streamflow depletion could significantly impact Annual, Summer, and Fall signatures in most streams. Our finding that multiple hydrologic signatures are consistently responsive to streamflow depletion across CONUS suggests that the underlying hydrological processes linking pumping to streamflow reductions are consistent across diverse settings, information that will aid in identifying indicators of streamflow depletion from streamflow hydrographs.

Lapides, Dana A.↗

Seasonal drivers of dissolved oxygen across a tidal creek–marsh interface revealed by machine learning

Abstract Dissolved oxygen (DO) is a key biogeochemical control in coastal systems, and its concentration and drivers vary markedly through time and space. This makes it difficult to accurately represent coastal DO and associated biogeochemical processes in models, limiting our ability to predict how these systems will respond to global change. We obtained high‐frequency (5‐min) in situ measurements of DO collected at three locations across the interface of a tidal creek and coastal marsh in the Pacific Northwest, USA. Random Forest machine learning models quantified the importance of three categories of environmental drivers (Aquatic, Climatic, and Terrestrial) of DO variability across the creek–marsh interface. We selected two 4‐month datasets representing Summer and Winter seasonal periods to test two hypotheses on the dominant drivers of DO at the coastal interface. We found that the Terrestrial driver—characterized by long periods of anaerobic conditions and episodic pulses in DO after floods—was most important during the Winter, whereas the Aquatic driver—characterized by variability over tidal, diel, and lunar cycles—was most important during the Summer. We explored how future climate change scenarios could alter the drivers of DO variability using a cumulative sums driver–response framework. Our results suggest that under climate change, Aquatic and Climatic drivers may increase in importance during the Summer, potentially linked to changing metabolic regimes and sea level, with Terrestrial driver importance potentially increasing during the Winter. Our approach highlights useful methods for understanding the spatiotemporal complexity of oxygen across coastal interfaces and quantifying the relative importance of distinct environmental drivers.

54 ENVIRONMENTAL SCIENCES↗